Methodology
How the benchmark works
How the calls were made, what each model saw, how scores were assigned, and what the benchmark cannot prove.
Benchmark methodology
How the calls were generated, how models coached them, and how each note was judged.
This benchmark tests sales coaching instincts, not transcript summarization. Each case is a synthetic B2B sales conversation generated from company research, persona design, and a hidden coaching answer key.
Coaching models receive the setup, research, participants, and speaker-labeled transcript. They do not receive the answer key, evaluator notes, or the name of the model that generated the transcript. The judge receives both the coaching note and the answer key, then scores the note for meaning.
The benchmark includes 25 calls generated with GPT and 25 matching calls generated with Claude Sonnet 4.6. Each model is evaluated across all 50 calls.
- Generated calls
- 50
- Coaching notes
- 3100
- Models
- 62
- Answer-key points
- 266
- Duration range
- 18-74m
- Avg duration
- 43m
Generation pipeline
The pipeline uses Vercel AI Gateway for research, case design, transcript generation, coaching, and judging.
Scenario input
Each case begins with a seller, buyer, call type, target duration, turn count, and intended level of seller performance.
Research brief
The generator runs web research for both companies and asks an LLM to produce a concise, source-grounded sales-call brief.
Hidden answer key
Before the transcript is written, an LLM defines 2 to 6 strengths or flaws, along with the evidence a strong coach should find.
Personas
Seller and buyer personas are instructed to reveal or pressure-test those strengths and flaws naturally.
Transcript turns
The transcript is generated one turn at a time, with each persona responding only from its own role and context.
Call package
Each finished call becomes a speaker-labeled transcript and Zoom-style recording package. Audio is included when available.
Coaching
Each model receives the same setup, research, participants, and speaker-labeled transcript. The answer key is excluded.
Judging
The judge compares each coaching note with the answer key, credits equivalent ideas, penalizes unsupported claims, and returns an eight-axis scorecard.
Benchmark coverage
Calls are grouped by their intended call type and seller performance.
Transcript generator
Call type
Seller performance
Ground truth
The answer key records the strengths and flaws a strong sales coach should catch.
- Answer-key points
- 266
- Flaws
- 129
- Strengths
- 137
Models and scoring
Each coaching note is judged against its call's answer key.
- Claude Fable 5
- 1
- high
- Claude Opus 4.7
- 5
- low, medium, high, xhigh, max
- Claude Opus 4.8
- 5
- low, medium, high, xhigh, max
- Claude Opus 5
- 5
- low, medium, high, xhigh, max
- Claude Sonnet 4.6
- 1
- default
- Claude Sonnet 5
- 1
- default
- DeepSeek V4 Pro
- 1
- default
- Gemini 3.1 Pro Preview
- 1
- default
- Gemini 3.5 Flash-Lite
- 4
- minimal, low, medium, high
- Gemini 3.6 Flash
- 4
- minimal, low, medium, high
- GLM 5.2
- 1
- default
- GPT-5.4
- 5
- none, low, medium, high, xhigh
- GPT-5.5
- 5
- none, low, medium, high, xhigh
- GPT-5.6 Luna
- 6
- none, low, medium, high, xhigh, max
- GPT-5.6 Sol
- 6
- none, low, medium, high, xhigh, max
- GPT-5.6 Terra
- 6
- none, low, medium, high, xhigh, max
- Kimi K3
- 1
- max
- Muse Spark 1.1
- 4
- minimal, low, medium, high
The benchmark covers 18 model families across 62 model-and-reasoning settings. Each one coaches all 50 calls. GPT-5.5 with high reasoning judges every result.
The leaderboard score weights overall score at 25%, answer-key recall and sales instinct at 20% each, prioritization at 15%, and technical accuracy and false-positive control at 10% each. Raw average and downside floor remain visible for context.
Limits
Treat these results as a test of transcript-based coaching on synthetic calls.
- Each case includes the exact transcript shown to the models, its answer key, and every model score.
- All calls are synthetic. Half were generated with GPT and half with Claude Sonnet 4.6, using company research and detailed personas.
- The judge awards partial credit for meaning rather than exact wording and penalizes unsupported claims.
- Scores cover transcript-based coaching only. Audio and video are not evaluated.