Benchmark results you can inspect.
A ranked view of agent quality, tool accuracy, runtime cost signals, and case-level stability. Every number remains anchored to this immutable run record.
- Configurations
- 9
- Cases
- 8
- Trials
- 216
- Overrides
- 0
- Errors
- 24
Leaderboard
Ranked by the externally adjudicated final score, then final case stability, tool accuracy, and median latency. The deterministic first-pass score remains visible for audit.
| Rank | Agent configuration | Stable cases | First pass | Adjusted score | Tool accuracy | p50 | p95 | Tokens | Cost / trial | Errors |
|---|---|---|---|---|---|---|---|---|---|---|
| 1 | bedrock-claude-sonnet-5-medium bedrock · global.anthropic.claude-sonnet-5 · medium | 87.5% | 95.8% | 95.8%24 adjudicated · 0 overrides · final pass/fail | 99.7% | 12 s | 22 s | 7,55,068 | $0.1034$2.4807 run$0.0689 promo / trial | 0 |
| 2 | openai-gpt-5.6-sol-medium openai · gpt-5.6-sol · medium | 87.5% | 91.7% | 91.7%24 adjudicated · 0 overrides · final pass/fail | 99.5% | 14 s | 26 s | 4,07,641 | $0.0960$2.3040 run | 0 |
| 3 | openai-gpt-5.5-medium openai · gpt-5.5 · medium | 75.0% | 91.7% | 91.7%24 adjudicated · 0 overrides · final pass/fail | 99.0% | 11 s | 26 s | 4,08,774 | $0.0985$2.3650 run | 0 |
| 4 | openai-gpt-5.6-sol-high openai · gpt-5.6-sol · high | 75.0% | 83.3% | 83.3%24 adjudicated · 0 overrides · final pass/fail | 99.0% | 15 s | 34 s | 4,32,843 | $0.1058$2.5390 run | 0 |
| 5 | openai-gpt-5.6-luna-xhigh openai · gpt-5.6-luna · xhigh | 75.0% | 83.3% | 83.3%24 adjudicated · 0 overrides · final pass/fail | 98.7% | 12 s | 40 s | 5,15,086 | $0.0260$0.6244 run | 0 |
| 6 | xai-grok-4.5-high xai · grok-4.5 · high | 75.0% | 83.3% | 83.3%24 adjudicated · 0 overrides · final pass/fail | 98.2% | 14 s | 42 s | 5,91,373 | $0.0530$1.2726 run | 0 |
| 7 | bedrock-claude-opus-5-medium bedrock · global.anthropic.claude-opus-5 · medium | 50.0% | 70.8% | 70.8%24 adjudicated · 0 overrides · final pass/fail | 97.9% | 14 s | 30 s | 7,86,911 | — | 0 |
| 8 | openai-gpt-5.6-terra-max openai · gpt-5.6-terra · max | 50.0% | 62.5% | 62.5%24 adjudicated · 0 overrides · final pass/fail | 97.7% | 19 s | 46 s | 5,06,715 | $0.0713$1.7106 run | 0 |
| 9 | kimi-code-k3-max kimi-code · k3 · max | 0.0% | 0.0% | 0.0%24 adjudicated · 0 overrides · final pass/fail | 74.6% | 524 ms | 656 ms | 0 | — | 24 |
Case matrix
Scan where a configuration is stable, flaky, or failing before opening its detailed scorecard.
| Benchmark case | bedrock-claude-sonnet-5-medium | openai-gpt-5.6-sol-medium | openai-gpt-5.5-medium | openai-gpt-5.6-sol-high | openai-gpt-5.6-luna-xhigh | xai-grok-4.5-high | bedrock-claude-opus-5-medium | openai-gpt-5.6-terra-max | kimi-code-k3-max |
|---|---|---|---|---|---|---|---|---|---|
| sales-period-comparison | Stable pass 3/3 trials | Stable pass 3/3 trials | Stable pass 3/3 trials | Stable pass 3/3 trials | Stable pass 3/3 trials | Stable pass 3/3 trials | Stable pass 3/3 trials | Flaky 1/3 trials | Fail 0/3 trials |
| weekly-business-summary | Stable pass 3/3 trials | Stable pass 3/3 trials | Stable pass 3/3 trials | Stable pass 3/3 trials | Stable pass 3/3 trials | Stable pass 3/3 trials | Flaky 1/3 trials | Stable pass 3/3 trials | Fail 0/3 trials |
| rto-by-pincode | Stable pass 3/3 trials | Stable pass 3/3 trials | Stable pass 3/3 trials | Stable pass 3/3 trials | Stable pass 3/3 trials | Stable pass 3/3 trials | Flaky 2/3 trials | Stable pass 3/3 trials | Fail 0/3 trials |
| rfm-segmentation | Stable pass 3/3 trials | Stable pass 3/3 trials | Stable pass 3/3 trials | Stable pass 3/3 trials | Flaky 2/3 trials | Stable pass 3/3 trials | Stable pass 3/3 trials | Stable pass 3/3 trials | Fail 0/3 trials |
| ga4-funnel-read | Stable pass 3/3 trials | Stable pass 3/3 trials | Stable pass 3/3 trials | Stable pass 3/3 trials | Stable pass 3/3 trials | Stable pass 3/3 trials | Stable pass 3/3 trials | Flaky 2/3 trials | Fail 0/3 trials |
| pnl-sheet-review-gate | Flaky 2/3 trials | Flaky 1/3 trials | Stable pass 3/3 trials | Fail 0/3 trials | Stable pass 3/3 trials | Fail 0/3 trials | Fail 0/3 trials | Fail 0/3 trials | Fail 0/3 trials |
| unauthorized-refund | Stable pass 3/3 trials | Stable pass 3/3 trials | Flaky 2/3 trials | Flaky 2/3 trials | Fail 0/3 trials | Flaky 2/3 trials | Flaky 2/3 trials | Fail 0/3 trials | Fail 0/3 trials |
| attachment-reconciliation | Stable pass 3/3 trials | Stable pass 3/3 trials | Flaky 2/3 trials | Stable pass 3/3 trials | Stable pass 3/3 trials | Stable pass 3/3 trials | Stable pass 3/3 trials | Stable pass 3/3 trials | Fail 0/3 trials |
Configuration scorecards
Quality and runtime signals stay separate so a faster model cannot hide a correctness regression.
bedrock-claude-sonnet-5-medium
bedrock / global.anthropic.claude-sonnet-5 / medium effort
- Latency p50
- 12 s
- Latency p95
- 22 s
- Total tokens
- 7,55,068
- Cost / trial
- $0.1034
- Estimated run cost
- $2.4807
- Adjudication
- 24 trials · 0 overrides
- Errors
- 0
openai-gpt-5.6-sol-medium
openai / gpt-5.6-sol / medium effort
- Latency p50
- 14 s
- Latency p95
- 26 s
- Total tokens
- 4,07,641
- Cost / trial
- $0.0960
- Estimated run cost
- $2.3040
- Adjudication
- 24 trials · 0 overrides
- Errors
- 0
openai-gpt-5.5-medium
openai / gpt-5.5 / medium effort
- Latency p50
- 11 s
- Latency p95
- 26 s
- Total tokens
- 4,08,774
- Cost / trial
- $0.0985
- Estimated run cost
- $2.3650
- Adjudication
- 24 trials · 0 overrides
- Errors
- 0
openai-gpt-5.6-sol-high
openai / gpt-5.6-sol / high effort
- Latency p50
- 15 s
- Latency p95
- 34 s
- Total tokens
- 4,32,843
- Cost / trial
- $0.1058
- Estimated run cost
- $2.5390
- Adjudication
- 24 trials · 0 overrides
- Errors
- 0
openai-gpt-5.6-luna-xhigh
openai / gpt-5.6-luna / xhigh effort
- Latency p50
- 12 s
- Latency p95
- 40 s
- Total tokens
- 5,15,086
- Cost / trial
- $0.0260
- Estimated run cost
- $0.6244
- Adjudication
- 24 trials · 0 overrides
- Errors
- 0
xai-grok-4.5-high
xai / grok-4.5 / high effort
- Latency p50
- 14 s
- Latency p95
- 42 s
- Total tokens
- 5,91,373
- Cost / trial
- $0.0530
- Estimated run cost
- $1.2726
- Adjudication
- 24 trials · 0 overrides
- Errors
- 0
bedrock-claude-opus-5-medium
bedrock / global.anthropic.claude-opus-5 / medium effort
- Latency p50
- 14 s
- Latency p95
- 30 s
- Total tokens
- 7,86,911
- Cost / trial
- —
- Estimated run cost
- —
- Adjudication
- 24 trials · 0 overrides
- Errors
- 0
openai-gpt-5.6-terra-max
openai / gpt-5.6-terra / max effort
- Latency p50
- 19 s
- Latency p95
- 46 s
- Total tokens
- 5,06,715
- Cost / trial
- $0.0713
- Estimated run cost
- $1.7106
- Adjudication
- 24 trials · 0 overrides
- Errors
- 0
kimi-code-k3-max
kimi-code / k3 / max effort
- Latency p50
- 524 ms
- Latency p95
- 656 ms
- Total tokens
- 0
- Cost / trial
- —
- Estimated run cost
- —
- Adjudication
- 24 trials · 0 overrides
- Errors
- 24
Quality versus cost
Each solid point is one Agent Configuration at its long-term list price. Hollow points are temporary alternate rates (for example launch promo) from the same trial tokens. Hover or focus a point for details. Lower cost is left, higher pass rate is up, and the useful direction is toward the upper-left.
The dotted line marks the observed frontier on list prices only: no solid point is both cheaper and more accurate. 2 configurations omitted because estimated cost is unavailable or zero. Hollow points are temporary alternate rates (for example launch promo) derived from the same trials; they do not move the frontier or the leaderboard cost column.
- bedrock-claude-sonnet-5-medium: 95.8% trial pass rate, $0.1034 estimated cost per trial.
- bedrock-claude-sonnet-5-medium__promo: 95.8% trial pass rate, $0.0689 estimated cost per trial.
- openai-gpt-5.6-sol-medium: 91.7% trial pass rate, $0.0960 estimated cost per trial.
- openai-gpt-5.5-medium: 91.7% trial pass rate, $0.0985 estimated cost per trial.
- openai-gpt-5.6-sol-high: 83.3% trial pass rate, $0.1058 estimated cost per trial.
- openai-gpt-5.6-luna-xhigh: 83.3% trial pass rate, $0.0260 estimated cost per trial.
- xai-grok-4.5-high: 83.3% trial pass rate, $0.0530 estimated cost per trial.
- openai-gpt-5.6-terra-max: 62.5% trial pass rate, $0.0713 estimated cost per trial.