Skip to results

Methodology

How the benchmark works

How the calls were made, what each model saw, how scores were assigned, and what the benchmark cannot prove.

Synthetic calls, hidden answer keys, meaning-based judging

Benchmark methodology

How the calls were generated, how models coached them, and how each note was judged.

This benchmark tests sales coaching instincts, not transcript summarization. Each case is a synthetic B2B sales conversation generated from company research, persona design, and a hidden coaching answer key.

Coaching models receive the setup, research, participants, and speaker-labeled transcript. They do not receive the answer key, evaluator notes, or the name of the model that generated the transcript. The judge receives both the coaching note and the answer key, then scores the note for meaning.

The benchmark includes 25 calls generated with GPT and 25 matching calls generated with Claude Sonnet 4.6. Each model is evaluated across all 50 calls.

Generated calls
50
Coaching notes
3100
Models
62
Answer-key points
266
Duration range
18-74m
Avg duration
43m
From scenario to Zoom-like case

Generation pipeline

The pipeline uses Vercel AI Gateway for research, case design, transcript generation, coaching, and judging.

01

Scenario input

Each case begins with a seller, buyer, call type, target duration, turn count, and intended level of seller performance.

02

Research brief

The generator runs web research for both companies and asks an LLM to produce a concise, source-grounded sales-call brief.

03

Hidden answer key

Before the transcript is written, an LLM defines 2 to 6 strengths or flaws, along with the evidence a strong coach should find.

04

Personas

Seller and buyer personas are instructed to reveal or pressure-test those strengths and flaws naturally.

05

Transcript turns

The transcript is generated one turn at a time, with each persona responding only from its own role and context.

06

Call package

Each finished call becomes a speaker-labeled transcript and Zoom-style recording package. Audio is included when available.

07

Coaching

Each model receives the same setup, research, participants, and speaker-labeled transcript. The answer key is excluded.

08

Judging

The judge compares each coaching note with the answer key, credits equivalent ideas, penalizes unsupported claims, and returns an eight-axis scorecard.

What the benchmark covers

Benchmark coverage

Calls are grouped by their intended call type and seller performance.

Transcript generator

GPT-generated25
Sonnet-generated25

Call type

Discovery16
Product demo20
Renewal save4
QBR4
Competitive displacement6

Seller performance

Excellent18
Mixed14
Flawed18
What the judge knows

Ground truth

The answer key records the strengths and flaws a strong sales coach should catch.

Answer-key points
266
Flaws
129
Strengths
137
Discovery47
Next Steps47
Technical Knowledge35
Qualification31
Value Alignment26
Research23
Objection Handling22
Communication Style16
Executive Alignment12
Customer Enablement7
One answer key, one judge

Models and scoring

Each coaching note is judged against its call's answer key.

Claude Fable 5
1
high
Claude Opus 4.7
5
low, medium, high, xhigh, max
Claude Opus 4.8
5
low, medium, high, xhigh, max
Claude Opus 5
5
low, medium, high, xhigh, max
Claude Sonnet 4.6
1
default
Claude Sonnet 5
1
default
DeepSeek V4 Pro
1
default
Gemini 3.1 Pro Preview
1
default
Gemini 3.5 Flash-Lite
4
minimal, low, medium, high
Gemini 3.6 Flash
4
minimal, low, medium, high
GLM 5.2
1
default
GPT-5.4
5
none, low, medium, high, xhigh
GPT-5.5
5
none, low, medium, high, xhigh
GPT-5.6 Luna
6
none, low, medium, high, xhigh, max
GPT-5.6 Sol
6
none, low, medium, high, xhigh, max
GPT-5.6 Terra
6
none, low, medium, high, xhigh, max
Kimi K3
1
max
Muse Spark 1.1
4
minimal, low, medium, high

The benchmark covers 18 model families across 62 model-and-reasoning settings. Each one coaches all 50 calls. GPT-5.5 with high reasoning judges every result.

The leaderboard score weights overall score at 25%, answer-key recall and sales instinct at 20% each, prioritization at 15%, and technical accuracy and false-positive control at 10% each. Raw average and downside floor remain visible for context.

What the scores do and do not mean

Limits

Treat these results as a test of transcript-based coaching on synthetic calls.

  • Each case includes the exact transcript shown to the models, its answer key, and every model score.
  • All calls are synthetic. Half were generated with GPT and half with Claude Sonnet 4.6, using company research and detailed personas.
  • The judge awards partial credit for meaning rather than exact wording and penalizes unsupported claims.
  • Scores cover transcript-based coaching only. Audio and video are not evaluated.