Skip to results

Sales coaching benchmark

Compare models

Compare quality, response time, and coach cost across the same 50 calls.

Selected model comparison

6 models selected

Green marks stronger results; red marks weaker results. Lower cost and response time are better.

Selected model comparison. Higher score metrics are better; lower cost and response time are better.
ModelScoreAvgRecallInst.FloorCostTime
GPT-5.6 Solmax90.089.788.891.580.3$0.37-
GPT-5.6 Terralow89.989.589.090.984.5$0.05-
GPT-5.6 Lunaxhigh89.989.789.590.881.7$0.04-
GPT-5.4xhigh89.389.088.190.679.9$0.28-
Claude Fable 5high87.787.586.990.477.0$0.46-
Claude Opus 4.8medium85.685.884.788.172.1$0.14-
All eight scoring dimensionsCompare all eight judging dimensions across 62 models.
Average score by model and scorecard dimension.
ModelOverallAnswer-key recallEvidence groundingFalse-positive controlPrioritizationActionabilitySales instinctTechnical accuracy
gpt-5.6 sol max
GPT-5.6 Sol · max
89.7
88.8
94.3
89.9
89.5
94.3
91.5
91.3
gpt-5.6 terra max
GPT-5.6 Terra · max
89.6
87.9
94.0
91.1
89.9
94.2
91.1
91.6
gpt-5.6 luna max
GPT-5.6 Luna · max
89.7
89.6
93.9
89.2
88.7
93.8
90.8
91.4
gpt-5.6 luna xhigh
GPT-5.6 Luna · xhigh
89.7
89.5
93.7
89.7
88.4
93.7
90.8
91.5
gpt-5.6 terra low
GPT-5.6 Terra · low
89.5
89.0
93.9
89.4
89.2
93.4
90.9
91.7
gpt-5.6 sol low
GPT-5.6 Sol · low
89.5
89.1
93.9
89.3
88.4
93.8
91.2
91.3
gpt-5.6 sol xhigh
GPT-5.6 Sol · xhigh
89.5
88.8
94.2
89.5
88.4
94.1
91.3
91.4
gpt-5.6 terra xhigh
GPT-5.6 Terra · xhigh
89.2
87.5
94.2
91.0
89.3
93.9
91.1
91.7
gpt-5.6 terra high
GPT-5.6 Terra · high
89.3
88.3
93.7
89.6
89.2
93.6
91.0
91.6
gpt-5.6 sol none
GPT-5.6 Sol · none
89.3
88.5
93.8
89.1
88.9
93.8
90.8
91.2
gpt-5.6 sol high
GPT-5.6 Sol · high
89.2
88.9
93.8
89.0
88.1
93.8
90.7
90.9
gpt-5.4 xhigh
GPT-5.4 · xhigh
89.0
88.1
93.9
89.6
88.2
92.9
90.6
91.5
gpt-5.6 sol medium
GPT-5.6 Sol · medium
89.0
88.1
93.4
89.0
88.2
93.9
90.8
91.3
gpt-5.4 high
GPT-5.4 · high
89.0
87.6
93.4
89.9
88.0
92.8
90.6
91.3
gpt-5.5 medium
GPT-5.5 · medium
88.8
88.5
93.7
88.0
87.8
93.3
90.6
91.2
gpt-5.5 xhigh
GPT-5.5 · xhigh
89.0
88.5
93.9
89.2
87.6
93.3
89.9
91.3
gpt-5.6 terra none
GPT-5.6 Terra · none
88.8
87.2
93.7
88.6
88.8
93.7
90.6
91.3
gpt-5.6 terra medium
GPT-5.6 Terra · medium
88.9
87.3
93.7
88.6
88.6
93.2
90.4
91.3
gpt-5.5 high
GPT-5.5 · high
88.6
88.4
93.6
89.0
86.7
93.4
90.2
91.1
gpt-5.6 luna medium
GPT-5.6 Luna · medium
88.6
88.4
92.7
87.1
87.4
93.1
89.8
91.2
gpt-5.6 luna high
GPT-5.6 Luna · high
88.6
87.7
92.9
87.7
87.4
93.4
89.9
91.3
gpt-5.6 luna low
GPT-5.6 Luna · low
88.5
87.5
93.1
88.1
87.5
93.0
89.8
90.8
gpt-5.4 medium
GPT-5.4 · medium
88.3
86.5
93.3
88.8
87.2
92.3
90.0
90.9
gpt-5.5 none
GPT-5.5 · none
88.1
86.8
93.5
87.9
86.9
92.9
90.0
91.1
gpt-5.6 luna none
GPT-5.6 Luna · none
87.7
87.3
92.3
86.7
86.8
92.8
89.2
90.6
gpt-5.5 low
GPT-5.5 · low
87.7
87.0
92.9
87.1
86.3
92.4
89.2
90.4
fable 5 high
Claude Fable 5 · high
87.5
86.9
90.1
83.5
86.7
93.0
90.4
89.6
gpt-5.4 low
GPT-5.4 · low
87.4
86.0
92.1
86.8
86.3
91.8
89.0
90.4
gpt-5.4 none
GPT-5.4 · none
87.4
85.8
92.7
86.7
86.7
91.4
88.7
90.5
opus 4.7 max
Claude Opus 4.7 · max
87.3
86.9
90.6
83.5
85.9
92.7
89.1
89.3
kimi k3 max
Kimi K3 · max
86.6
85.3
90.6
84.0
86.0
93.1
89.9
89.1
opus 5 max
Claude Opus 5 · max
86.6
88.3
89.4
80.3
84.3
93.4
89.6
88.6
opus 5 xhigh
Claude Opus 5 · xhigh
86.6
88.3
88.6
80.2
83.9
93.7
89.6
88.8
opus 4.7 high
Claude Opus 4.7 · high
86.8
86.1
89.0
82.4
85.3
92.1
88.8
88.8
muse spark 1.1 high
Muse Spark 1.1 · high
86.4
86.1
89.4
83.1
85.8
90.4
88.4
88.9
muse spark 1.1 medium
Muse Spark 1.1 · medium
86.2
85.0
89.2
83.0
86.5
90.2
88.9
87.8
opus 5 medium
Claude Opus 5 · medium
86.3
86.5
88.8
80.5
84.7
93.2
89.4
88.3
opus 5 low
Claude Opus 5 · low
85.9
85.5
89.4
81.3
84.2
92.6
89.2
88.1
muse spark 1.1 minimal
Muse Spark 1.1 · minimal
85.7
83.8
88.4
82.9
85.8
88.9
87.9
87.9
opus 5 high
Claude Opus 5 · high
85.5
86.0
88.5
80.3
83.6
92.9
88.8
88.0
muse spark 1.1 low
Muse Spark 1.1 · low
85.4
83.2
89.7
84.0
85.2
90.1
88.1
88.6
opus 4.8 medium
Claude Opus 4.8 · medium
85.8
84.7
89.7
81.8
83.8
90.5
88.1
88.2
opus 4.7 medium
Claude Opus 4.7 · medium
85.6
84.2
89.1
82.6
83.9
91.0
88.2
87.8
opus 4.7 xhigh
Claude Opus 4.7 · xhigh
85.6
84.6
89.0
82.3
83.9
91.6
87.8
88.0
opus 4.7 low
Claude Opus 4.7 · low
85.6
84.1
89.6
82.7
84.1
90.9
87.6
88.5
opus 4.8 max
Claude Opus 4.8 · max
85.4
85.4
88.6
80.9
83.8
91.4
87.6
88.2
opus 4.8 xhigh
Claude Opus 4.8 · xhigh
85.2
84.8
88.7
81.3
83.8
90.4
87.6
88.3
opus 4.8 high
Claude Opus 4.8 · high
84.9
83.6
89.2
81.1
83.7
90.5
87.1
88.6
sonnet 4.6
Claude Sonnet 4.6 · default
84.6
83.8
87.0
79.6
83.2
91.2
87.6
86.6
sonnet 5
Claude Sonnet 5 · default
84.6
83.8
88.5
81.0
82.3
89.3
85.9
87.8
opus 4.8 low
Claude Opus 4.8 · low
84.0
82.8
88.5
80.4
81.9
89.3
85.5
87.5
glm 5.2
GLM 5.2 · default
84.0
82.2
88.4
80.8
81.6
89.4
85.7
87.3
deepseek v4 pro
DeepSeek V4 Pro · default
83.5
81.9
86.9
79.5
81.9
88.5
84.9
86.6
gemini 3.6 flash minimal
Gemini 3.6 Flash · minimal
81.6
79.1
85.6
77.7
79.3
84.9
82.8
85.2
gemini 3.6 flash medium
Gemini 3.6 Flash · medium
79.8
76.6
86.6
78.5
77.6
84.0
81.5
85.4
gemini 3.6 flash high
Gemini 3.6 Flash · high
79.3
75.9
85.0
75.7
76.7
83.6
80.8
84.4
gemini 3.1 pro preview
Gemini 3.1 Pro Preview · default
78.9
74.5
86.2
78.0
76.9
84.1
81.4
84.1
gemini 3.6 flash low
Gemini 3.6 Flash · low
78.5
75.0
84.8
75.2
75.6
81.2
79.7
85.1
gemini 3.5 flash lite high
Gemini 3.5 Flash-Lite · high
77.4
73.1
84.0
75.5
74.1
79.2
78.9
83.9
gemini 3.5 flash lite minimal
Gemini 3.5 Flash-Lite · minimal
74.8
71.2
81.9
72.2
70.0
75.8
75.1
83.0
gemini 3.5 flash lite medium
Gemini 3.5 Flash-Lite · medium
74.6
70.3
82.0
73.3
69.7
75.1
74.9
82.2
gemini 3.5 flash lite low
Gemini 3.5 Flash-Lite · low
72.7
68.2
80.5
70.2
68.0
72.9
72.4
81.7
Mean85.984.790.383.984.590.587.788.9
Coach cost and response timeWhat each coach costs and how long it takes to answer. Call generation and judging are excluded.
Estimated 50-call total
$359.82
Median / call
$0.10
Models
62
gemini 3.5 flash lite low
Gemini 3.5 Flash-Lite · low
Raw provider cost · 50 calls
Score
72.7
Cost / call
$0.0047
Response time
4.9s
50-call total
$0.23
Input / response
4,918 / 1,276
Reasoning
0
Input / output rate
$0.30 / $2.50
deepseek v4 pro
DeepSeek V4 Pro · default
Estimated from saved response
Score
83.5
Cost / call
$0.0047
Response time
50-call total
$0.24
Input / response
4,559 / 3,180
Reasoning
Input / output rate
$0.43 / $0.87
gemini 3.5 flash lite minimal
Gemini 3.5 Flash-Lite · minimal
Raw provider cost · 50 calls
Score
74.8
Cost / call
$0.0053
Response time
5.6s
50-call total
$0.26
Input / response
4,918 / 1,513
Reasoning
0
Input / output rate
$0.30 / $2.50
gemini 3.5 flash lite medium
Gemini 3.5 Flash-Lite · medium
Raw provider cost · 50 calls
Score
74.6
Cost / call
$0.0053
Response time
5.7s
50-call total
$0.27
Input / response
4,918 / 1,332
Reasoning
208
Input / output rate
$0.30 / $2.50
gemini 3.5 flash lite high
Gemini 3.5 Flash-Lite · high
Raw provider cost · 50 calls
Score
77.4
Cost / call
$0.01
Response time
11s
50-call total
$0.53
Input / response
4,918 / 1,421
Reasoning
2,219
Input / output rate
$0.30 / $2.50
gemini 3.6 flash low
Gemini 3.6 Flash · low
Raw provider cost · 50 calls
Score
78.5
Cost / call
$0.02
Response time
8.8s
50-call total
$0.99
Input / response
4,918 / 1,533
Reasoning
129
Input / output rate
$1.50 / $7.50
gpt-5.6 luna low
GPT-5.6 Luna · low
Measured on 1 call
Score
88.5
Cost / call
$0.02
Response time
50-call total
$1.00
Input / response
2,879 / 2,843
Reasoning
21
Input / output rate
$1.00 / $6.00
muse spark 1.1 minimal
Muse Spark 1.1 · minimal
Measured on 1 call
Score
85.7
Cost / call
$0.02
Response time
50-call total
$1.03
Input / response
2,554 / 2,557
Reasoning
1,533
Input / output rate
$1.25 / $4.25
gpt-5.6 luna none
GPT-5.6 Luna · none
Measured on 1 call
Score
87.7
Cost / call
$0.02
Response time
50-call total
$1.03
Input / response
2,879 / 2,970
Reasoning
0
Input / output rate
$1.00 / $6.00
gpt-5.6 luna medium
GPT-5.6 Luna · medium
Measured on 1 call
Score
88.6
Cost / call
$0.02
Response time
50-call total
$1.04
Input / response
2,879 / 2,936
Reasoning
67
Input / output rate
$1.00 / $6.00
gemini 3.6 flash minimal
Gemini 3.6 Flash · minimal
Raw provider cost · 50 calls
Score
81.6
Cost / call
$0.02
Response time
9.7s
50-call total
$1.07
Input / response
4,918 / 1,871
Reasoning
0
Input / output rate
$1.50 / $7.50
muse spark 1.1 low
Muse Spark 1.1 · low
Measured on 1 call
Score
85.4
Cost / call
$0.02
Response time
50-call total
$1.10
Input / response
2,554 / 2,854
Reasoning
1,587
Input / output rate
$1.25 / $4.25
muse spark 1.1 medium
Muse Spark 1.1 · medium
Measured on 1 call
Score
86.2
Cost / call
$0.02
Response time
50-call total
$1.25
Input / response
2,554 / 2,680
Reasoning
2,451
Input / output rate
$1.25 / $4.25
muse spark 1.1 high
Muse Spark 1.1 · high
Measured on 1 call
Score
86.4
Cost / call
$0.03
Response time
50-call total
$1.28
Input / response
2,554 / 2,539
Reasoning
2,728
Input / output rate
$1.25 / $4.25
glm 5.2
GLM 5.2 · default
Estimated from saved response
Score
84.0
Cost / call
$0.03
Response time
50-call total
$1.30
Input / response
4,559 / 4,455
Reasoning
Input / output rate
$1.40 / $4.40
gpt-5.6 luna high
GPT-5.6 Luna · high
Measured on 1 call
Score
88.6
Cost / call
$0.03
Response time
50-call total
$1.42
Input / response
2,879 / 3,560
Reasoning
678
Input / output rate
$1.00 / $6.00
gemini 3.6 flash medium
Gemini 3.6 Flash · medium
Raw provider cost · 50 calls
Score
79.8
Cost / call
$0.03
Response time
15s
50-call total
$1.48
Input / response
4,918 / 1,699
Reasoning
1,264
Input / output rate
$1.50 / $7.50
gemini 3.1 pro preview
Gemini 3.1 Pro Preview · default
Estimated from saved response
Score
78.9
Cost / call
$0.03
Response time
50-call total
$1.60
Input / response
4,559 / 1,900
Reasoning
Input / output rate
$2.00 / $12.00
gpt-5.6 luna xhigh
GPT-5.6 Luna · xhigh
Measured on 1 call
Score
89.7
Cost / call
$0.04
Response time
50-call total
$2.09
Input / response
2,879 / 3,644
Reasoning
2,837
Input / output rate
$1.00 / $6.00
gemini 3.6 flash high
Gemini 3.6 Flash · high
Raw provider cost · 50 calls
Score
79.3
Cost / call
$0.04
Response time
22s
50-call total
$2.11
Input / response
4,918 / 1,849
Reasoning
2,803
Input / output rate
$1.50 / $7.50
gpt-5.4 none
GPT-5.4 · none
Measured on 1 call
Score
87.4
Cost / call
$0.05
Response time
50-call total
$2.40
Input / response
2,879 / 2,726
Reasoning
0
Input / output rate
$2.50 / $15.00
gpt-5.4 low
GPT-5.4 · low
Measured on 1 call
Score
87.4
Cost / call
$0.05
Response time
50-call total
$2.41
Input / response
2,879 / 2,728
Reasoning
10
Input / output rate
$2.50 / $15.00
gpt-5.6 terra low
GPT-5.6 Terra · low
Measured on 1 call
Score
89.5
Cost / call
$0.05
Response time
50-call total
$2.49
Input / response
2,879 / 2,787
Reasoning
57
Input / output rate
$2.50 / $15.00
sonnet 5
Claude Sonnet 5 · default
Estimated from saved response
Score
84.6
Cost / call
$0.05
Response time
50-call total
$2.50
Input / response
4,559 / 4,079
Reasoning
Input / output rate
$2.00 / $10.00
gpt-5.6 terra medium
GPT-5.6 Terra · medium
Measured on 1 call
Score
88.9
Cost / call
$0.05
Response time
50-call total
$2.65
Input / response
2,879 / 3,014
Reasoning
44
Input / output rate
$2.50 / $15.00
gpt-5.6 terra none
GPT-5.6 Terra · none
Measured on 1 call
Score
88.8
Cost / call
$0.05
Response time
50-call total
$2.66
Input / response
2,879 / 3,067
Reasoning
0
Input / output rate
$2.50 / $15.00
gpt-5.4 medium
GPT-5.4 · medium
Measured on 1 call
Score
88.3
Cost / call
$0.06
Response time
50-call total
$2.78
Input / response
2,879 / 2,799
Reasoning
430
Input / output rate
$2.50 / $15.00
gpt-5.6 terra high
GPT-5.6 Terra · high
Measured on 1 call
Score
89.3
Cost / call
$0.06
Response time
50-call total
$2.90
Input / response
2,879 / 3,234
Reasoning
159
Input / output rate
$2.50 / $15.00
gpt-5.6 luna max
GPT-5.6 Luna · max
Measured on 1 call
Score
89.7
Cost / call
$0.06
Response time
50-call total
$3.19
Input / response
2,879 / 3,557
Reasoning
6,588
Input / output rate
$1.00 / $6.00
gpt-5.6 terra xhigh
GPT-5.6 Terra · xhigh
Measured on 1 call
Score
89.2
Cost / call
$0.08
Response time
50-call total
$3.80
Input / response
2,879 / 3,041
Reasoning
1,552
Input / output rate
$2.50 / $15.00
gpt-5.6 terra max
GPT-5.6 Terra · max
Measured on 1 call
Score
89.6
Cost / call
$0.10
Response time
50-call total
$5.02
Input / response
2,879 / 3,620
Reasoning
2,588
Input / output rate
$2.50 / $15.00
gpt-5.4 high
GPT-5.4 · high
Measured on 1 call
Score
89.0
Cost / call
$0.10
Response time
50-call total
$5.19
Input / response
2,879 / 3,045
Reasoning
3,389
Input / output rate
$2.50 / $15.00
sonnet 4.6
Claude Sonnet 4.6 · default
Estimated from saved response
Score
84.6
Cost / call
$0.10
Response time
50-call total
$5.21
Input / response
4,559 / 6,034
Reasoning
Input / output rate
$3.00 / $15.00
gpt-5.6 sol low
GPT-5.6 Sol · low
Measured on 1 call
Score
89.5
Cost / call
$0.12
Response time
50-call total
$5.96
Input / response
2,879 / 3,377
Reasoning
116
Input / output rate
$5.00 / $30.00
gpt-5.5 medium
GPT-5.5 · medium
Measured on 1 call
Score
88.8
Cost / call
$0.12
Response time
50-call total
$6.03
Input / response
2,879 / 3,474
Reasoning
64
Input / output rate
$5.00 / $30.00
opus 4.8 low
Claude Opus 4.8 · low
Measured on 1 call
Score
84.0
Cost / call
$0.12
Response time
50-call total
$6.16
Input / response
5,993 / 3,729
Reasoning
0
Input / output rate
$5.00 / $25.00
gpt-5.6 sol high
GPT-5.6 Sol · high
Measured on 1 call
Score
89.2
Cost / call
$0.12
Response time
50-call total
$6.17
Input / response
2,879 / 3,451
Reasoning
180
Input / output rate
$5.00 / $30.00
gpt-5.6 sol none
GPT-5.6 Sol · none
Measured on 1 call
Score
89.3
Cost / call
$0.12
Response time
50-call total
$6.21
Input / response
2,879 / 3,660
Reasoning
0
Input / output rate
$5.00 / $30.00
gpt-5.5 none
GPT-5.5 · none
Measured on 1 call
Score
88.1
Cost / call
$0.13
Response time
50-call total
$6.26
Input / response
2,879 / 3,694
Reasoning
0
Input / output rate
$5.00 / $30.00
gpt-5.6 sol medium
GPT-5.6 Sol · medium
Measured on 1 call
Score
89.0
Cost / call
$0.13
Response time
50-call total
$6.31
Input / response
2,879 / 3,659
Reasoning
66
Input / output rate
$5.00 / $30.00
opus 4.7 low
Claude Opus 4.7 · low
Measured on 1 call
Score
85.6
Cost / call
$0.13
Response time
50-call total
$6.67
Input / response
5,998 / 4,140
Reasoning
0
Input / output rate
$5.00 / $25.00
gpt-5.5 low
GPT-5.5 · low
Measured on 1 call
Score
87.7
Cost / call
$0.14
Response time
50-call total
$6.80
Input / response
2,879 / 4,051
Reasoning
0
Input / output rate
$5.00 / $30.00
gpt-5.5 high
GPT-5.5 · high
Measured on 1 call
Score
88.6
Cost / call
$0.14
Response time
50-call total
$7.01
Input / response
2,879 / 3,679
Reasoning
516
Input / output rate
$5.00 / $30.00
opus 4.8 medium
Claude Opus 4.8 · medium
Measured on 1 call
Score
85.8
Cost / call
$0.14
Response time
50-call total
$7.18
Input / response
5,993 / 4,548
Reasoning
0
Input / output rate
$5.00 / $25.00
kimi k3 max
Kimi K3 · max
Raw provider cost · 50 calls
Score
86.6
Cost / call
$0.15
Response time
50-call total
$7.32
Input / response
4,485 / 5,322
Reasoning
3,801
Input / output rate
$3.00 / $15.00
opus 4.8 high
Claude Opus 4.8 · high
Measured on 1 call
Score
84.9
Cost / call
$0.15
Response time
50-call total
$7.69
Input / response
5,993 / 4,956
Reasoning
0
Input / output rate
$5.00 / $25.00
opus 4.8 xhigh
Claude Opus 4.8 · xhigh
Measured on 1 call
Score
85.2
Cost / call
$0.16
Response time
50-call total
$8.17
Input / response
5,993 / 5,341
Reasoning
0
Input / output rate
$5.00 / $25.00
gpt-5.6 sol xhigh
GPT-5.6 Sol · xhigh
Measured on 1 call
Score
89.5
Cost / call
$0.17
Response time
50-call total
$8.35
Input / response
2,879 / 3,681
Reasoning
1,407
Input / output rate
$5.00 / $30.00
opus 4.7 high
Claude Opus 4.7 · high
Measured on 1 call
Score
86.8
Cost / call
$0.17
Response time
50-call total
$8.60
Input / response
6,244 / 5,630
Reasoning
0
Input / output rate
$5.00 / $25.00
opus 4.7 xhigh
Claude Opus 4.7 · xhigh
Measured on 1 call
Score
85.6
Cost / call
$0.17
Response time
50-call total
$8.67
Input / response
5,998 / 5,737
Reasoning
0
Input / output rate
$5.00 / $25.00
opus 4.8 max
Claude Opus 4.8 · max
Measured on 1 call
Score
85.4
Cost / call
$0.18
Response time
50-call total
$9.15
Input / response
5,992 / 6,118
Reasoning
0
Input / output rate
$5.00 / $25.00
opus 5 low
Claude Opus 5 · low
Raw provider cost · 50 calls
Score
85.9
Cost / call
$0.20
Response time
1m 30s
50-call total
$10.01
Input / response
7,939 / 6,421
Reasoning
0
Input / output rate
$5.00 / $25.00
gpt-5.5 xhigh
GPT-5.5 · xhigh
Measured on 1 call
Score
89.0
Cost / call
$0.20
Response time
50-call total
$10.16
Input / response
2,879 / 3,704
Reasoning
2,588
Input / output rate
$5.00 / $30.00
opus 4.7 max
Claude Opus 4.7 · max
Measured on 1 call
Score
87.3
Cost / call
$0.21
Response time
50-call total
$10.34
Input / response
5,997 / 7,071
Reasoning
0
Input / output rate
$5.00 / $25.00
opus 5 medium
Claude Opus 5 · medium
Raw provider cost · 50 calls
Score
86.3
Cost / call
$0.25
Response time
1m 59s
50-call total
$12.53
Input / response
7,939 / 8,439
Reasoning
0
Input / output rate
$5.00 / $25.00
gpt-5.4 xhigh
GPT-5.4 · xhigh
Measured on 1 call
Score
89.0
Cost / call
$0.28
Response time
50-call total
$13.90
Input / response
2,879 / 3,690
Reasoning
14,369
Input / output rate
$2.50 / $15.00
opus 5 high
Claude Opus 5 · high
Raw provider cost · 50 calls
Score
85.5
Cost / call
$0.30
Response time
2m 27s
50-call total
$14.86
Input / response
7,939 / 10,300
Reasoning
0
Input / output rate
$5.00 / $25.00
opus 4.7 medium
Claude Opus 4.7 · medium
Measured on 1 call
Score
85.6
Cost / call
$0.33
Response time
50-call total
$16.64
Input / response
12,353 / 10,843
Reasoning
0
Input / output rate
$5.00 / $25.00
opus 5 xhigh
Claude Opus 5 · xhigh
Raw provider cost · 50 calls
Score
86.6
Cost / call
$0.35
Response time
2m 53s
50-call total
$17.36
Input / response
7,939 / 12,298
Reasoning
0
Input / output rate
$5.00 / $25.00
gpt-5.6 sol max
GPT-5.6 Sol · max
Measured on 1 call
Score
89.7
Cost / call
$0.37
Response time
50-call total
$18.72
Input / response
2,879 / 4,041
Reasoning
7,959
Input / output rate
$5.00 / $30.00
opus 5 max
Claude Opus 5 · max
Raw provider cost · 50 calls
Score
86.6
Cost / call
$0.38
Response time
3m 13s
50-call total
$19.20
Input / response
7,939 / 13,772
Reasoning
0
Input / output rate
$5.00 / $25.00
fable 5 high
Claude Fable 5 · high
Measured on 1 call
Score
87.5
Cost / call
$0.46
Response time
50-call total
$22.85
Input / response
5,993 / 7,943
Reasoning
0
Input / output rate
$10.00 / $50.00

Costs for 57 of 62 models use recorded provider inference usage or a controlled same-call measurement, priced at uncached list rates. Gateway markup and reporting charges are excluded. Response time is shown where it was recorded with the coach run. 5 models use an estimate from the saved coaching response because exact token usage was not recorded.

Performance by transcript sourceCompare performance on the 25 GPT-generated calls and 25 Sonnet-generated calls.
Model performance by benchmark segment, shown as average score and delta from each model baseline.
ModelOverall
GPT-generated
n=25
Sonnet-generated
n=25
gpt-5.6 sol max
GPT-5.6 Sol
89.7
93.0+3.3
86.5-3.3
gpt-5.6 terra max
GPT-5.6 Terra
89.6
93.4+3.8
85.8-3.8
gpt-5.6 luna max
GPT-5.6 Luna
89.7
92.4+2.8
86.9-2.8
gpt-5.6 luna xhigh
GPT-5.6 Luna
89.7
92.7+3.1
86.6-3.1
gpt-5.6 terra low
GPT-5.6 Terra
89.5
92.6+3.1
86.4-3.1
gpt-5.6 sol low
GPT-5.6 Sol
89.5
92.9+3.4
86.1-3.4
gpt-5.6 sol xhigh
GPT-5.6 Sol
89.5
93.1+3.6
86.0-3.6
gpt-5.6 terra xhigh
GPT-5.6 Terra
89.2
92.6+3.4
85.8-3.4
gpt-5.6 terra high
GPT-5.6 Terra
89.3
92.6+3.3
86.1-3.3
gpt-5.6 sol none
GPT-5.6 Sol
89.3
92.6+3.2
86.1-3.2
gpt-5.6 sol high
GPT-5.6 Sol
89.2
92.7+3.5
85.7-3.5
gpt-5.4 xhigh
GPT-5.4
89.0
92.0+3.0
86.0-3.0
gpt-5.6 sol medium
GPT-5.6 Sol
89.0
92.4+3.4
85.6-3.4
gpt-5.4 high
GPT-5.4
89.0
92.0+3.1
85.9-3.1
gpt-5.5 medium
GPT-5.5
88.8
91.7+2.8
86.0-2.8
gpt-5.5 xhigh
GPT-5.5
89.0
92.0+3.0
86.0-3.0
gpt-5.6 terra none
GPT-5.6 Terra
88.8
92.1+3.3
85.5-3.3
gpt-5.6 terra medium
GPT-5.6 Terra
88.9
92.4+3.5
85.5-3.5
gpt-5.5 high
GPT-5.5
88.6
91.7+3.2
85.4-3.2
gpt-5.6 luna medium
GPT-5.6 Luna
88.6
91.2+2.6
86.0-2.6
gpt-5.6 luna high
GPT-5.6 Luna
88.6
91.4+2.9
85.7-2.9
gpt-5.6 luna low
GPT-5.6 Luna
88.5
91.1+2.6
86.0-2.6
gpt-5.4 medium
GPT-5.4
88.3
90.9+2.6
85.6-2.6
gpt-5.5 none
GPT-5.5
88.1
92.0+3.9
84.3-3.9
gpt-5.6 luna none
GPT-5.6 Luna
87.7
91.2+3.4
84.3-3.4
gpt-5.5 low
GPT-5.5
87.7
90.8+3.1
84.6-3.1
fable 5 high
Claude Fable 5
87.5
91.0+3.4
84.1-3.4
gpt-5.4 low
GPT-5.4
87.4
90.3+3.0
84.4-3.0
gpt-5.4 none
GPT-5.4
87.4
90.8+3.5
83.9-3.5
opus 4.7 max
Claude Opus 4.7
87.3
90.2+3.0
84.3-3.0
kimi k3 max
Kimi K3
86.6
90.6+4.0
82.6-4.0
opus 5 max
Claude Opus 5
86.6
89.2+2.6
84.1-2.6
opus 5 xhigh
Claude Opus 5
86.6
88.9+2.3
84.2-2.3
opus 4.7 high
Claude Opus 4.7
86.8
89.6+2.7
84.1-2.7
muse spark 1.1 high
Muse Spark 1.1
86.4
88.8+2.4
84.0-2.4
muse spark 1.1 medium
Muse Spark 1.1
86.2
88.1+1.9
84.3-1.9
opus 5 medium
Claude Opus 5
86.3
89.8+3.5
82.8-3.5
opus 5 low
Claude Opus 5
85.9
89.2+3.3
82.6-3.3
muse spark 1.1 minimal
Muse Spark 1.1
85.7
87.3+1.7
84.0-1.7
opus 5 high
Claude Opus 5
85.5
89.3+3.8
81.7-3.8
muse spark 1.1 low
Muse Spark 1.1
85.4
88.6+3.3
82.1-3.3
opus 4.8 medium
Claude Opus 4.8
85.8
89.2+3.4
82.4-3.4
opus 4.7 medium
Claude Opus 4.7
85.6
89.0+3.3
82.3-3.3
opus 4.7 xhigh
Claude Opus 4.7
85.6
89.4+3.8
81.8-3.8
opus 4.7 low
Claude Opus 4.7
85.6
87.6+2.0
83.6-2.0
opus 4.8 max
Claude Opus 4.8
85.4
88.6+3.1
82.3-3.1
opus 4.8 xhigh
Claude Opus 4.8
85.2
88.1+2.9
82.4-2.9
opus 4.8 high
Claude Opus 4.8
84.9
89.4+4.5
80.4-4.5
sonnet 4.6
Claude Sonnet 4.6
84.6
88.8+4.3
80.3-4.3
sonnet 5
Claude Sonnet 5
84.6
87.4+2.8
81.8-2.8
opus 4.8 low
Claude Opus 4.8
84.0
88.0+4.0
80.0-4.0
glm 5.2
GLM 5.2
84.0
86.1+2.1
81.8-2.1
deepseek v4 pro
DeepSeek V4 Pro
83.5
85.8+2.3
81.2-2.3
gemini 3.6 flash minimal
Gemini 3.6 Flash
81.6
84.2+2.6
79.0-2.6
gemini 3.6 flash medium
Gemini 3.6 Flash
79.8
83.7+3.9
76.0-3.9
gemini 3.6 flash high
Gemini 3.6 Flash
79.3
83.2+4.0
75.3-4.0
gemini 3.1 pro preview
Gemini 3.1 Pro Preview
78.9
82.4+3.5
75.4-3.5
gemini 3.6 flash low
Gemini 3.6 Flash
78.5
82.4+3.8
74.7-3.8
gemini 3.5 flash lite high
Gemini 3.5 Flash-Lite
77.4
80.2+2.8
74.6-2.8
gemini 3.5 flash lite minimal
Gemini 3.5 Flash-Lite
74.8
78.5+3.7
71.1-3.7
gemini 3.5 flash lite medium
Gemini 3.5 Flash-Lite
74.6
79.3+4.6
70.0-4.6
gemini 3.5 flash lite low
Gemini 3.5 Flash-Lite
72.7
76.4+3.8
68.9-3.8
Performance by call qualityCompare how models coach excellent, mixed, and flawed calls.
Model performance by benchmark segment, shown as average score and delta from each model baseline.
ModelOverall
Excellent
n=18
Mixed
n=14
Flawed
n=18
gpt-5.6 sol max
GPT-5.6 Sol
89.7
88.7-1.1
88.7-1.0
91.6+1.9
gpt-5.6 terra max
GPT-5.6 Terra
89.6
88.3-1.3
90.2+0.6
90.5+0.9
gpt-5.6 luna max
GPT-5.6 Luna
89.7
88.1-1.6
89.9+0.2
91.2+1.5
gpt-5.6 luna xhigh
GPT-5.6 Luna
89.7
88.1-1.5
89.0-0.7
91.7+2.1
gpt-5.6 terra low
GPT-5.6 Terra
89.5
88.9-0.6
88.4-1.1
90.9+1.4
gpt-5.6 sol low
GPT-5.6 Sol
89.5
89.2-0.3
87.6-1.8
91.2+1.7
gpt-5.6 sol xhigh
GPT-5.6 Sol
89.5
89.0-0.5
88.2-1.3
91.1+1.6
gpt-5.6 terra xhigh
GPT-5.6 Terra
89.2
88.8-0.5
87.4-1.8
91.1+1.9
gpt-5.6 terra high
GPT-5.6 Terra
89.3
89.6+0.2
88.1-1.3
90.1+0.8
gpt-5.6 sol none
GPT-5.6 Sol
89.3
90.0+0.7
87.9-1.4
89.8+0.4
gpt-5.6 sol high
GPT-5.6 Sol
89.2
88.8-0.4
87.8-1.4
90.7+1.5
gpt-5.4 xhigh
GPT-5.4
89.0
88.9-0.1
87.6-1.3
90.1+1.1
gpt-5.6 sol medium
GPT-5.6 Sol
89.0
89.4+0.4
86.1-3.0
90.9+1.9
gpt-5.4 high
GPT-5.4
89.0
88.2-0.7
87.7-1.2
90.7+1.7
gpt-5.5 medium
GPT-5.5
88.8
90.1+1.2
86.7-2.1
89.3+0.4
gpt-5.5 xhigh
GPT-5.5
89.0
89.8+0.8
87.7-1.3
89.2+0.2
gpt-5.6 terra none
GPT-5.6 Terra
88.8
90.0+1.2
87.6-1.1
88.4-0.3
gpt-5.6 terra medium
GPT-5.6 Terra
88.9
89.0+0.1
88.0-0.9
89.6+0.7
gpt-5.5 high
GPT-5.5
88.6
90.1+1.6
84.5-4.1
90.2+1.6
gpt-5.6 luna medium
GPT-5.6 Luna
88.6
88.4-0.1
86.4-2.1
90.3+1.8
gpt-5.6 luna high
GPT-5.6 Luna
88.6
88.3-0.2
87.1-1.5
89.9+1.4
gpt-5.6 luna low
GPT-5.6 Luna
88.5
88.2-0.3
87.1-1.4
89.9+1.4
gpt-5.4 medium
GPT-5.4
88.3
88.1-0.2
85.6-2.7
90.6+2.3
gpt-5.5 none
GPT-5.5
88.1
89.6+1.5
85.0-3.1
89.1+1.0
gpt-5.6 luna none
GPT-5.6 Luna
87.7
87.80.0
86.7-1.0
88.5+0.8
gpt-5.5 low
GPT-5.5
87.7
89.4+1.7
85.6-2.1
87.70.0
fable 5 high
Claude Fable 5
87.5
86.9-0.6
85.9-1.6
89.3+1.8
gpt-5.4 low
GPT-5.4
87.4
87.9+0.6
84.8-2.6
88.8+1.4
gpt-5.4 none
GPT-5.4
87.4
87.9+0.5
84.5-2.9
89.1+1.7
opus 4.7 max
Claude Opus 4.7
87.3
88.8+1.6
83.6-3.6
88.5+1.2
kimi k3 max
Kimi K3
86.6
86.4-0.2
83.7-2.9
89.1+2.4
opus 5 max
Claude Opus 5
86.6
81.5-5.1
86.9+0.2
91.6+5.0
opus 5 xhigh
Claude Opus 5
86.6
82.0-4.6
86.7+0.2
91.0+4.4
opus 4.7 high
Claude Opus 4.7
86.8
87.1+0.3
83.4-3.5
89.2+2.4
muse spark 1.1 high
Muse Spark 1.1
86.4
87.1+0.6
83.6-2.8
87.9+1.5
muse spark 1.1 medium
Muse Spark 1.1
86.2
86.7+0.5
84.2-2.0
87.2+1.0
opus 5 medium
Claude Opus 5
86.3
82.3-4.0
86.30.0
90.3+4.0
opus 5 low
Claude Opus 5
85.9
84.4-1.5
83.1-2.8
89.6+3.7
muse spark 1.1 minimal
Muse Spark 1.1
85.7
85.9+0.2
81.3-4.4
88.8+3.2
opus 5 high
Claude Opus 5
85.5
82.4-3.1
83.0-2.5
90.6+5.1
muse spark 1.1 low
Muse Spark 1.1
85.4
85.8+0.4
83.3-2.1
86.6+1.2
opus 4.8 medium
Claude Opus 4.8
85.8
86.6+0.8
81.1-4.7
88.6+2.9
opus 4.7 medium
Claude Opus 4.7
85.6
87.2+1.6
80.0-5.6
88.4+2.8
opus 4.7 xhigh
Claude Opus 4.7
85.6
86.8+1.2
81.1-4.5
87.9+2.3
opus 4.7 low
Claude Opus 4.7
85.6
86.6+1.0
80.8-4.8
88.4+2.8
opus 4.8 max
Claude Opus 4.8
85.4
85.9+0.5
80.4-5.0
88.8+3.4
opus 4.8 xhigh
Claude Opus 4.8
85.2
88.1+2.8
79.4-5.9
87.0+1.8
opus 4.8 high
Claude Opus 4.8
84.9
87.3+2.4
77.4-7.5
88.4+3.5
sonnet 4.6
Claude Sonnet 4.6
84.6
84.7+0.1
80.1-4.4
87.9+3.3
sonnet 5
Claude Sonnet 5
84.6
84.0-0.6
83.4-1.2
86.2+1.6
opus 4.8 low
Claude Opus 4.8
84.0
86.7+2.7
77.0-7.0
86.7+2.7
glm 5.2
GLM 5.2
84.0
85.8+1.8
79.2-4.8
85.9+1.9
deepseek v4 pro
DeepSeek V4 Pro
83.5
86.2+2.7
76.6-6.9
86.2+2.7
gemini 3.6 flash minimal
Gemini 3.6 Flash
81.6
85.9+4.3
71.7-9.9
84.9+3.4
gemini 3.6 flash medium
Gemini 3.6 Flash
79.8
84.8+4.9
70.1-9.7
82.4+2.6
gemini 3.6 flash high
Gemini 3.6 Flash
79.3
84.1+4.8
68.6-10.6
82.8+3.5
gemini 3.1 pro preview
Gemini 3.1 Pro Preview
78.9
81.6+2.7
69.8-9.1
83.3+4.4
gemini 3.6 flash low
Gemini 3.6 Flash
78.5
84.4+5.9
66.3-12.2
82.1+3.6
gemini 3.5 flash lite high
Gemini 3.5 Flash-Lite
77.4
84.7+7.3
67.0-10.4
78.2+0.8
gemini 3.5 flash lite minimal
Gemini 3.5 Flash-Lite
74.8
83.6+8.8
64.9-10.0
73.8-1.0
gemini 3.5 flash lite medium
Gemini 3.5 Flash-Lite
74.6
82.9+8.2
65.6-9.1
73.4-1.2
gemini 3.5 flash lite low
Gemini 3.5 Flash-Lite
72.7
82.2+9.5
63.3-9.4
70.5-2.2
Performance by call typeCompare model performance across discovery, demos, renewals, QBRs, and competitive calls.
Model performance by benchmark segment, shown as average score and delta from each model baseline.
ModelOverall
Discovery
n=16
Product demo
n=20
Renewal save
n=4
QBR
n=4
Competitive displacement
n=6
gpt-5.6 sol max
GPT-5.6 Sol
89.7
91.0+1.3
88.5-1.3
88.5-1.2
88.5-1.2
92.3+2.6
gpt-5.6 terra max
GPT-5.6 Terra
89.6
90.3+0.7
88.8-0.8
90.5+0.9
87.8-1.9
91.0+1.4
gpt-5.6 luna max
GPT-5.6 Luna
89.7
90.9+1.3
88.6-1.1
88.3-1.4
89.5-0.2
91.0+1.3
gpt-5.6 luna xhigh
GPT-5.6 Luna
89.7
91.0+1.3
88.0-1.7
90.3+0.6
88.3-1.4
92.3+2.7
gpt-5.6 terra low
GPT-5.6 Terra
89.5
91.8+2.3
87.7-1.8
89.8+0.3
88.5-1.0
90.0+0.5
gpt-5.6 sol low
GPT-5.6 Sol
89.5
91.3+1.8
88.2-1.3
86.3-3.2
89.50.0
91.2+1.7
gpt-5.6 sol xhigh
GPT-5.6 Sol
89.5
91.4+1.8
87.8-1.7
90.0+0.5
88.0-1.5
91.0+1.5
gpt-5.6 terra xhigh
GPT-5.6 Terra
89.2
91.0+1.8
87.3-1.9
87.0-2.2
90.0+0.8
91.8+2.6
gpt-5.6 terra high
GPT-5.6 Terra
89.3
90.4+1.0
88.3-1.0
89.8+0.4
88.0-1.3
90.5+1.2
gpt-5.6 sol none
GPT-5.6 Sol
89.3
90.7+1.3
88.2-1.2
88.5-0.8
89.8+0.4
90.0+0.7
gpt-5.6 sol high
GPT-5.6 Sol
89.2
91.7+2.5
86.8-2.5
89.0-0.2
89.5+0.3
90.7+1.5
gpt-5.4 xhigh
GPT-5.4
89.0
89.7+0.7
88.5-0.5
86.1-2.9
88.0-1.0
91.5+2.5
gpt-5.6 sol medium
GPT-5.6 Sol
89.0
91.6+2.6
88.4-0.6
83.3-5.8
87.3-1.8
89.2+0.1
gpt-5.4 high
GPT-5.4
89.0
90.4+1.5
87.6-1.4
89.8+0.8
87.5-1.5
90.0+1.0
gpt-5.5 medium
GPT-5.5
88.8
89.9+1.1
87.9-0.9
87.0-1.8
90.8+1.9
89.0+0.2
gpt-5.5 xhigh
GPT-5.5
89.0
89.3+0.2
88.4-0.6
89.8+0.7
88.8-0.3
90.2+1.1
gpt-5.6 terra none
GPT-5.6 Terra
88.8
89.6+0.8
88.3-0.5
87.0-1.8
89.5+0.7
88.8+0.1
gpt-5.6 terra medium
GPT-5.6 Terra
88.9
90.1+1.2
88.0-0.9
89.8+0.8
87.8-1.2
89.2+0.2
gpt-5.5 high
GPT-5.5
88.6
90.4+1.8
87.8-0.8
80.8-7.8
87.8-0.8
92.0+3.4
gpt-5.6 luna medium
GPT-5.6 Luna
88.6
90.6+2.1
87.4-1.2
87.0-1.6
85.5-3.1
90.0+1.4
gpt-5.6 luna high
GPT-5.6 Luna
88.6
90.4+1.9
87.2-1.4
86.8-1.8
86.5-2.1
90.8+2.3
gpt-5.6 luna low
GPT-5.6 Luna
88.5
90.5+2.0
86.8-1.8
87.8-0.8
88.3-0.3
89.8+1.3
gpt-5.4 medium
GPT-5.4
88.3
90.1+1.8
87.3-0.9
84.3-4.0
86.8-1.5
90.2+1.9
gpt-5.5 none
GPT-5.5
88.1
90.1+2.0
86.5-1.7
86.0-2.1
88.0-0.1
90.0+1.9
gpt-5.6 luna none
GPT-5.6 Luna
87.7
89.4+1.7
86.6-1.1
86.5-1.2
86.8-1.0
88.5+0.8
gpt-5.5 low
GPT-5.5
87.7
89.0+1.3
86.5-1.2
87.0-0.7
86.5-1.2
89.3+1.6
fable 5 high
Claude Fable 5
87.5
89.7+2.2
84.8-2.7
87.3-0.3
88.5+1.0
90.2+2.6
gpt-5.4 low
GPT-5.4
87.4
89.3+1.9
86.5-0.9
86.3-1.1
87.0-0.4
86.2-1.2
gpt-5.4 none
GPT-5.4
87.4
88.9+1.6
86.5-0.9
86.3-1.1
86.5-0.9
87.5+0.1
opus 4.7 max
Claude Opus 4.7
87.3
91.3+4.1
84.5-2.8
89.3+2.0
85.8-1.5
85.5-1.8
kimi k3 max
Kimi K3
86.6
88.6+2.0
84.6-2.0
81.8-4.9
88.3+1.6
90.2+3.5
opus 5 max
Claude Opus 5
86.6
89.3+2.7
83.7-2.9
87.8+1.1
83.3-3.4
90.8+4.2
opus 5 xhigh
Claude Opus 5
86.6
88.8+2.2
84.6-2.0
88.5+1.9
84.0-2.6
87.7+1.1
opus 4.7 high
Claude Opus 4.7
86.8
90.1+3.2
83.7-3.2
86.8-0.1
86.3-0.6
89.2+2.3
muse spark 1.1 high
Muse Spark 1.1
86.4
88.4+2.0
83.9-2.5
86.3-0.2
87.3+0.8
89.0+2.6
muse spark 1.1 medium
Muse Spark 1.1
86.2
88.1+1.9
85.0-1.3
83.0-3.2
86.30.0
87.5+1.3
opus 5 medium
Claude Opus 5
86.3
88.1+1.8
83.7-2.6
88.8+2.5
84.8-1.5
89.8+3.5
opus 5 low
Claude Opus 5
85.9
89.4+3.5
83.6-2.3
79.5-6.4
84.3-1.7
89.7+3.8
muse spark 1.1 minimal
Muse Spark 1.1
85.7
88.3+2.7
83.0-2.7
83.3-2.4
86.0+0.3
89.0+3.3
opus 5 high
Claude Opus 5
85.5
89.3+3.8
84.0-1.5
77.8-7.8
82.5-3.0
87.5+2.0
muse spark 1.1 low
Muse Spark 1.1
85.4
86.4+1.0
83.4-2.0
83.0-2.4
85.8+0.4
90.7+5.3
opus 4.8 medium
Claude Opus 4.8
85.8
89.4+3.6
81.5-4.2
87.3+1.5
86.5+0.7
88.7+2.9
opus 4.7 medium
Claude Opus 4.7
85.6
89.8+4.2
81.6-4.0
86.3+0.6
84.8-0.9
88.0+2.4
opus 4.7 xhigh
Claude Opus 4.7
85.6
87.9+2.3
83.0-2.5
86.3+0.7
84.0-1.6
88.5+2.9
opus 4.7 low
Claude Opus 4.7
85.6
90.2+4.6
82.0-3.6
86.8+1.1
84.5-1.1
85.5-0.1
opus 4.8 max
Claude Opus 4.8
85.4
88.8+3.4
81.3-4.1
84.8-0.7
87.8+2.3
88.8+3.4
opus 4.8 xhigh
Claude Opus 4.8
85.2
88.5+3.3
81.5-3.8
82.8-2.5
86.8+1.5
89.8+4.6
opus 4.8 high
Claude Opus 4.8
84.9
89.9+5.0
81.4-3.5
78.5-6.4
86.0+1.1
86.8+1.9
sonnet 4.6
Claude Sonnet 4.6
84.6
86.9+2.3
83.3-1.2
82.0-2.6
82.0-2.6
85.8+1.3
sonnet 5
Claude Sonnet 5
84.6
86.6+2.0
82.0-2.6
88.0+3.4
82.8-1.8
87.2+2.6
opus 4.8 low
Claude Opus 4.8
84.0
88.1+4.1
80.5-3.5
85.8+1.8
81.8-2.2
85.0+1.0
glm 5.2
GLM 5.2
84.0
86.3+2.3
83.0-1.0
85.5+1.5
79.5-4.5
83.2-0.8
deepseek v4 pro
DeepSeek V4 Pro
83.5
87.4+3.9
81.5-2.0
82.0-1.5
82.3-1.3
81.8-1.7
gemini 3.6 flash minimal
Gemini 3.6 Flash
81.6
86.4+4.9
78.5-3.0
79.0-2.6
78.5-3.1
82.5+0.9
gemini 3.6 flash medium
Gemini 3.6 Flash
79.8
85.2+5.3
77.0-2.9
74.8-5.1
74.8-5.1
82.0+2.2
gemini 3.6 flash high
Gemini 3.6 Flash
79.3
84.7+5.4
74.8-4.4
75.3-4.0
77.5-1.8
83.5+4.2
gemini 3.1 pro preview
Gemini 3.1 Pro Preview
78.9
84.4+5.5
74.7-4.2
74.3-4.7
77.5-1.4
82.3+3.4
gemini 3.6 flash low
Gemini 3.6 Flash
78.5
83.2+4.7
77.0-1.6
69.5-9.0
75.5-3.0
79.3+0.8
gemini 3.5 flash lite high
Gemini 3.5 Flash-Lite
77.4
81.4+4.0
76.0-1.4
71.5-5.9
75.0-2.4
77.0-0.4
gemini 3.5 flash lite minimal
Gemini 3.5 Flash-Lite
74.8
79.3+4.5
72.5-2.3
71.3-3.6
74.5-0.3
73.0-1.8
gemini 3.5 flash lite medium
Gemini 3.5 Flash-Lite
74.6
78.6+3.9
74.8+0.2
65.0-9.6
71.8-2.9
72.0-2.6
gemini 3.5 flash lite low
Gemini 3.5 Flash-Lite
72.7
78.5+5.8
71.3-1.4
66.0-6.7
70.5-2.2
67.8-4.8