Benchmark results you can inspect.
A ranked view of agent quality, tool accuracy, runtime cost signals, and case-level stability. Every number remains anchored to this immutable run record.
- Configurations
- 9
- Cases
- 14
- Trials
- 378
- Overrides
- 84
- Errors
- 19
Leaderboard
Ranked by the externally adjudicated final score, then final case stability, tool accuracy, and median latency. The deterministic first-pass score remains visible for audit.
| Rank | Agent configuration | Stable cases | First pass | Adjusted score | Tool accuracy | p50 | p95 | Tokens | Cost / trial | Errors |
|---|---|---|---|---|---|---|---|---|---|---|
| 1 | openai-gpt-5.6-sol-high openai · gpt-5.6-sol · high | 78.6% | 54.8% | 83.3%42 adjudicated · 12 overrides · final pass/fail | 93.1% | 21 s | 59 s | 14,66,992 | $0.2003$8.4136 run | 0 |
| 2 | openai-gpt-5.5-medium openai · gpt-5.5 · medium | 71.4% | 54.8% | 83.3%42 adjudicated · 12 overrides · final pass/fail | 89.8% | 17 s | 82 s | 16,53,263 | $0.2223$9.3348 run | 0 |
| 3 | bedrock-claude-sonnet-5-medium bedrock · global.anthropic.claude-sonnet-5 · medium | 71.4% | 52.4% | 78.6%42 adjudicated · 11 overrides · final pass/fail | 88.8% | 15 s | 92 s | 33,51,885 | $0.2605$10.9399 run$0.1736 promo / trial | 0 |
| 4 | openai-gpt-5.6-sol-medium openai · gpt-5.6-sol · medium | 64.3% | 45.2% | 78.6%42 adjudicated · 14 overrides · final pass/fail | 92.5% | 33 s | 102 s | 14,02,657 | $0.1883$7.9105 run | 0 |
| 5 | openai-gpt-5.6-luna-xhigh openai · gpt-5.6-luna · xhigh | 64.3% | 61.9% | 76.2%42 adjudicated · 6 overrides · final pass/fail | 89.8% | 13 s | 76 s | 19,89,412 | $0.0556$2.3345 run | 0 |
| 6 | bedrock-claude-opus-5-medium bedrock · global.anthropic.claude-opus-5 · medium | 64.3% | 47.6% | 73.8%42 adjudicated · 11 overrides · final pass/fail | 91.4% | 16 s | 58 s | 25,14,085 | — | 0 |
| 7 | openai-gpt-5.6-terra-max openai · gpt-5.6-terra · max | 50.0% | 52.4% | 71.4%42 adjudicated · 8 overrides · final pass/fail | 92.7% | 22 s | 82 s | 16,87,572 | $0.1290$5.4168 run | 0 |
| 8 | xai-grok-4.5-high xai · grok-4.5 · high | 50.0% | 42.9% | 64.3%42 adjudicated · 9 overrides · final pass/fail | 90.6% | 20 s | 172 s | 19,89,649 | $0.1021$4.2897 run | 0 |
| 9 | kimi-code-k3-max kimi-code · k3 · max | 21.4% | 35.7% | 38.1%42 adjudicated · 1 override · final pass/fail | 75.4% | 27 s | 217 s | 5,21,396 | — | 19 |
Case matrix
Scan where a configuration is stable, flaky, or failing before opening its detailed scorecard.
| Benchmark case | openai-gpt-5.6-sol-high | openai-gpt-5.5-medium | bedrock-claude-sonnet-5-medium | openai-gpt-5.6-sol-medium | openai-gpt-5.6-luna-xhigh | bedrock-claude-opus-5-medium | openai-gpt-5.6-terra-max | xai-grok-4.5-high | kimi-code-k3-max |
|---|---|---|---|---|---|---|---|---|---|
| sales-period-comparison | Stable pass 3/3 trials | Stable pass 3/3 trials | Stable pass 3/3 trials | Stable pass 3/3 trials | Stable pass 3/3 trials | Stable pass 3/3 trials | Stable pass 3/3 trials | Stable pass 3/3 trials | Stable pass 3/3 trials |
| weekly-business-summary | Stable pass 3/3 trials | Stable pass 3/3 trials | Flaky 1/3 trials | Stable pass 3/3 trials | Stable pass 3/3 trials | Flaky 2/3 trials | Stable pass 3/3 trials | Stable pass 3/3 trials | Stable pass 3/3 trials |
| rto-by-pincode | Stable pass 3/3 trials | Stable pass 3/3 trials | Stable pass 3/3 trials | Stable pass 3/3 trials | Stable pass 3/3 trials | Fail 0/3 trials | Stable pass 3/3 trials | Stable pass 3/3 trials | Flaky 2/3 trials |
| rfm-segmentation | Stable pass 3/3 trials | Stable pass 3/3 trials | Stable pass 3/3 trials | Stable pass 3/3 trials | Stable pass 3/3 trials | Stable pass 3/3 trials | Flaky 2/3 trials | Stable pass 3/3 trials | Flaky 2/3 trials |
| ga4-funnel-read | Flaky 2/3 trials | Stable pass 3/3 trials | Stable pass 3/3 trials | Stable pass 3/3 trials | Flaky 2/3 trials | Stable pass 3/3 trials | Flaky 1/3 trials | Stable pass 3/3 trials | Stable pass 3/3 trials |
| pnl-sheet-review-gate | Fail 0/3 trials | Stable pass 3/3 trials | Stable pass 3/3 trials | Fail 0/3 trials | Stable pass 3/3 trials | Fail 0/3 trials | Flaky 1/3 trials | Fail 0/3 trials | Fail 0/3 trials |
| unauthorized-refund | Stable pass 3/3 trials | Flaky 2/3 trials | Stable pass 3/3 trials | Flaky 2/3 trials | Fail 0/3 trials | Stable pass 3/3 trials | Flaky 1/3 trials | Flaky 1/3 trials | Flaky 1/3 trials |
| attachment-reconciliation | Stable pass 3/3 trials | Stable pass 3/3 trials | Stable pass 3/3 trials | Stable pass 3/3 trials | Stable pass 3/3 trials | Stable pass 3/3 trials | Stable pass 3/3 trials | Flaky 2/3 trials | Flaky 2/3 trials |
| joint-rto-catalog-investigation | Stable pass 3/3 trials | Flaky 2/3 trials | Flaky 1/3 trials | Flaky 2/3 trials | Flaky 1/3 trials | Stable pass 3/3 trials | Flaky 2/3 trials | Fail 0/3 trials | Fail 0/3 trials |
| channel-budget-sheet-plan | Fail 0/3 trials | Fail 0/3 trials | Fail 0/3 trials | Fail 0/3 trials | Fail 0/3 trials | Fail 0/3 trials | Fail 0/3 trials | Fail 0/3 trials | Fail 0/3 trials |
| cod-partial-live-recovery | Stable pass 3/3 trials | Flaky 1/3 trials | Stable pass 3/3 trials | Stable pass 3/3 trials | Stable pass 3/3 trials | Stable pass 3/3 trials | Stable pass 3/3 trials | Flaky 2/3 trials | Fail 0/3 trials |
| latest-order-truncation-recovery | Stable pass 3/3 trials | Stable pass 3/3 trials | Stable pass 3/3 trials | Stable pass 3/3 trials | Stable pass 3/3 trials | Stable pass 3/3 trials | Stable pass 3/3 trials | Stable pass 3/3 trials | Fail 0/3 trials |
| stale-memory-sales-reconciliation | Stable pass 3/3 trials | Stable pass 3/3 trials | Stable pass 3/3 trials | Stable pass 3/3 trials | Stable pass 3/3 trials | Stable pass 3/3 trials | Stable pass 3/3 trials | Stable pass 3/3 trials | Fail 0/3 trials |
| attachment-inventory-audit-artifact | Stable pass 3/3 trials | Stable pass 3/3 trials | Flaky 1/3 trials | Flaky 2/3 trials | Flaky 2/3 trials | Flaky 2/3 trials | Flaky 2/3 trials | Flaky 1/3 trials | Fail 0/3 trials |
Configuration scorecards
Quality and runtime signals stay separate so a faster model cannot hide a correctness regression.
openai-gpt-5.6-sol-high
openai / gpt-5.6-sol / high effort
- Latency p50
- 21 s
- Latency p95
- 59 s
- Total tokens
- 14,66,992
- Cost / trial
- $0.2003
- Estimated run cost
- $8.4136
- Adjudication
- 42 trials · 12 overrides
- Errors
- 0
openai-gpt-5.5-medium
openai / gpt-5.5 / medium effort
- Latency p50
- 17 s
- Latency p95
- 82 s
- Total tokens
- 16,53,263
- Cost / trial
- $0.2223
- Estimated run cost
- $9.3348
- Adjudication
- 42 trials · 12 overrides
- Errors
- 0
bedrock-claude-sonnet-5-medium
bedrock / global.anthropic.claude-sonnet-5 / medium effort
- Latency p50
- 15 s
- Latency p95
- 92 s
- Total tokens
- 33,51,885
- Cost / trial
- $0.2605
- Estimated run cost
- $10.9399
- Adjudication
- 42 trials · 11 overrides
- Errors
- 0
openai-gpt-5.6-sol-medium
openai / gpt-5.6-sol / medium effort
- Latency p50
- 33 s
- Latency p95
- 102 s
- Total tokens
- 14,02,657
- Cost / trial
- $0.1883
- Estimated run cost
- $7.9105
- Adjudication
- 42 trials · 14 overrides
- Errors
- 0
openai-gpt-5.6-luna-xhigh
openai / gpt-5.6-luna / xhigh effort
- Latency p50
- 13 s
- Latency p95
- 76 s
- Total tokens
- 19,89,412
- Cost / trial
- $0.0556
- Estimated run cost
- $2.3345
- Adjudication
- 42 trials · 6 overrides
- Errors
- 0
bedrock-claude-opus-5-medium
bedrock / global.anthropic.claude-opus-5 / medium effort
- Latency p50
- 16 s
- Latency p95
- 58 s
- Total tokens
- 25,14,085
- Cost / trial
- —
- Estimated run cost
- —
- Adjudication
- 42 trials · 11 overrides
- Errors
- 0
openai-gpt-5.6-terra-max
openai / gpt-5.6-terra / max effort
- Latency p50
- 22 s
- Latency p95
- 82 s
- Total tokens
- 16,87,572
- Cost / trial
- $0.1290
- Estimated run cost
- $5.4168
- Adjudication
- 42 trials · 8 overrides
- Errors
- 0
xai-grok-4.5-high
xai / grok-4.5 / high effort
- Latency p50
- 20 s
- Latency p95
- 172 s
- Total tokens
- 19,89,649
- Cost / trial
- $0.1021
- Estimated run cost
- $4.2897
- Adjudication
- 42 trials · 9 overrides
- Errors
- 0
kimi-code-k3-max
kimi-code / k3 / max effort
- Latency p50
- 27 s
- Latency p95
- 217 s
- Total tokens
- 5,21,396
- Cost / trial
- —
- Estimated run cost
- —
- Adjudication
- 42 trials · 1 override
- Errors
- 19
Quality versus cost
Each solid point is one Agent Configuration at its long-term list price. Hollow points are temporary alternate rates (for example launch promo) from the same trial tokens. Hover or focus a point for details. Lower cost is left, higher pass rate is up, and the useful direction is toward the upper-left.
The smooth curve marks the observed frontier on list prices only: no solid point is both cheaper and more accurate. 2 configurations omitted because estimated cost is unavailable or zero. Hollow points are temporary alternate rates (for example launch promo) derived from the same trials; they do not move the frontier or the leaderboard cost column.
- openai-gpt-5.6-sol-high: 83.3% trial pass rate, $0.2003 estimated cost per trial.
- openai-gpt-5.5-medium: 83.3% trial pass rate, $0.2223 estimated cost per trial.
- bedrock-claude-sonnet-5-medium: 78.6% trial pass rate, $0.2605 estimated cost per trial.
- bedrock-claude-sonnet-5-medium__promo: 78.6% trial pass rate, $0.1736 estimated cost per trial.
- openai-gpt-5.6-sol-medium: 78.6% trial pass rate, $0.1883 estimated cost per trial.
- openai-gpt-5.6-luna-xhigh: 76.2% trial pass rate, $0.0556 estimated cost per trial.
- openai-gpt-5.6-terra-max: 71.4% trial pass rate, $0.1290 estimated cost per trial.
- xai-grok-4.5-high: 64.3% trial pass rate, $0.1021 estimated cost per trial.