Benchmark results you can inspect.
A ranked view of agent quality, tool accuracy, runtime cost signals, and case-level stability. Every number remains anchored to this immutable run record.
- Configurations
- 10
- Cases
- 8
- Trials
- 240
- Overrides
- 18
- Errors
- 0
Leaderboard
Ranked by the externally adjudicated final score, then final case stability, tool accuracy, and median latency. The deterministic first-pass score remains visible for audit.
| Rank | Agent configuration | Stable cases | First pass | Adjusted score | Tool accuracy | p50 | p95 | Tokens | Cost / trial | Errors |
|---|---|---|---|---|---|---|---|---|---|---|
| 1 | openai-gpt-5.6-luna-max openai · gpt-5.6-luna · max | 75.0% | 83.3% | 83.3%24 adjudicated · 0 overrides · final pass/fail | 99.0% | 13 s | 48 s | 4,64,702 | $0.0261$0.6273 run | 0 |
| 2 | bedrock-claude-sonnet-5-medium bedrock · global.anthropic.claude-sonnet-5 · medium | 75.0% | 83.3% | 83.3%24 adjudicated · 0 overrides · final pass/fail | 97.2% | 13 s | 26 s | 6,67,001 | $0.0928$2.2263 run$0.0618 promo / trial | 0 |
| 3 | openai-gpt-5.6-sol-medium openai · gpt-5.6-sol · medium | 62.5% | 70.8% | 79.2%24 adjudicated · 4 overrides · final pass/fail | 93.6% | 9.8 s | 27 s | 3,13,112 | $0.0750$1.8011 run | 0 |
| 4 | openai-gpt-5.6-sol-high openai · gpt-5.6-sol · high | 75.0% | 75.0% | 75.0%24 adjudicated · 0 overrides · final pass/fail | 98.7% | 15 s | 35 s | 4,06,663 | $0.0991$2.3782 run | 0 |
| 5 | xai-grok-4.5-high xai · grok-4.5 · high | 75.0% | 75.0% | 75.0%24 adjudicated · 0 overrides · final pass/fail | 98.2% | 14 s | 34 s | 5,43,690 | $0.0486$1.1664 run | 0 |
| 6 | kimi-code-k3-max kimi-code · k3 · max | 62.5% | 70.8% | 75.0%24 adjudicated · 1 override · final pass/fail | 97.2% | 47 s | 180 s | 4,84,474 | — | 0 |
| 7 | openai-gpt-5.6-terra-max openai · gpt-5.6-terra · max | 62.5% | 62.5% | 70.8%24 adjudicated · 2 overrides · final pass/fail | 98.2% | 20 s | 35 s | 4,57,110 | $0.0646$1.5495 run | 0 |
| 8 | bedrock-claude-opus-5-medium bedrock · global.anthropic.claude-opus-5 · medium | 50.0% | 70.8% | 70.8%24 adjudicated · 0 overrides · final pass/fail | 98.2% | 16 s | 28 s | 7,81,157 | $0.1804$4.3298 run | 0 |
| 9 | openai-gpt-5.6-luna-xhigh openai · gpt-5.6-luna · xhigh | 50.0% | 58.3% | 66.7%24 adjudicated · 6 overrides · final pass/fail | 91.8% | 8.4 s | 28 s | 3,66,872 | $0.0187$0.4493 run | 0 |
| 10 | openai-gpt-5.5-medium openai · gpt-5.5 · medium | 50.0% | 75.0% | 54.2%24 adjudicated · 5 overrides · final pass/fail | 92.3% | 15 s | 20 s | 3,34,798 | $0.0795$1.9084 run | 0 |
Case matrix
Scan where a configuration is stable, flaky, or failing before opening its detailed scorecard.
| Benchmark case | openai-gpt-5.6-luna-max | bedrock-claude-sonnet-5-medium | openai-gpt-5.6-sol-medium | openai-gpt-5.6-sol-high | xai-grok-4.5-high | kimi-code-k3-max | openai-gpt-5.6-terra-max | bedrock-claude-opus-5-medium | openai-gpt-5.6-luna-xhigh | openai-gpt-5.5-medium |
|---|---|---|---|---|---|---|---|---|---|---|
| sales-period-comparison | Stable pass 3/3 trials | Stable pass 3/3 trials | Stable pass 3/3 trials | Stable pass 3/3 trials | Stable pass 3/3 trials | Stable pass 3/3 trials | Stable pass 3/3 trials | Stable pass 3/3 trials | Stable pass 3/3 trials | Stable pass 3/3 trials |
| weekly-business-summary | Stable pass 3/3 trials | Fail 0/3 trials | Flaky 2/3 trials | Stable pass 3/3 trials | Stable pass 3/3 trials | Flaky 2/3 trials | Stable pass 3/3 trials | Flaky 2/3 trials | Flaky 1/3 trials | Fail 0/3 trials |
| rto-by-pincode | Stable pass 3/3 trials | Stable pass 3/3 trials | Stable pass 3/3 trials | Stable pass 3/3 trials | Stable pass 3/3 trials | Stable pass 3/3 trials | Stable pass 3/3 trials | Flaky 2/3 trials | Stable pass 3/3 trials | Stable pass 3/3 trials |
| rfm-segmentation | Stable pass 3/3 trials | Stable pass 3/3 trials | Stable pass 3/3 trials | Stable pass 3/3 trials | Stable pass 3/3 trials | Stable pass 3/3 trials | Stable pass 3/3 trials | Stable pass 3/3 trials | Flaky 1/3 trials | Fail 0/3 trials |
| ga4-funnel-read | Flaky 2/3 trials | Stable pass 3/3 trials | Flaky 2/3 trials | Stable pass 3/3 trials | Stable pass 3/3 trials | Stable pass 3/3 trials | Flaky 1/3 trials | Flaky 1/3 trials | Stable pass 3/3 trials | Stable pass 3/3 trials |
| pnl-sheet-review-gate | Stable pass 3/3 trials | Flaky 2/3 trials | Stable pass 3/3 trials | Fail 0/3 trials | Fail 0/3 trials | Fail 0/3 trials | Fail 0/3 trials | Fail 0/3 trials | Stable pass 3/3 trials | Stable pass 3/3 trials |
| unauthorized-refund | Fail 0/3 trials | Stable pass 3/3 trials | Fail 0/3 trials | Fail 0/3 trials | Fail 0/3 trials | Flaky 1/3 trials | Flaky 1/3 trials | Stable pass 3/3 trials | Fail 0/3 trials | Flaky 1/3 trials |
| attachment-reconciliation | Stable pass 3/3 trials | Stable pass 3/3 trials | Stable pass 3/3 trials | Stable pass 3/3 trials | Stable pass 3/3 trials | Stable pass 3/3 trials | Stable pass 3/3 trials | Stable pass 3/3 trials | Flaky 2/3 trials | Fail 0/3 trials |
Configuration scorecards
Quality and runtime signals stay separate so a faster model cannot hide a correctness regression.
openai-gpt-5.6-luna-max
openai / gpt-5.6-luna / max effort
- Latency p50
- 13 s
- Latency p95
- 48 s
- Total tokens
- 4,64,702
- Cost / trial
- $0.0261
- Estimated run cost
- $0.6273
- Adjudication
- 24 trials · 0 overrides
- Errors
- 0
bedrock-claude-sonnet-5-medium
bedrock / global.anthropic.claude-sonnet-5 / medium effort
- Latency p50
- 13 s
- Latency p95
- 26 s
- Total tokens
- 6,67,001
- Cost / trial
- $0.0928
- Estimated run cost
- $2.2263
- Adjudication
- 24 trials · 0 overrides
- Errors
- 0
openai-gpt-5.6-sol-medium
openai / gpt-5.6-sol / medium effort
- Latency p50
- 9.8 s
- Latency p95
- 27 s
- Total tokens
- 3,13,112
- Cost / trial
- $0.0750
- Estimated run cost
- $1.8011
- Adjudication
- 24 trials · 4 overrides
- Errors
- 0
openai-gpt-5.6-sol-high
openai / gpt-5.6-sol / high effort
- Latency p50
- 15 s
- Latency p95
- 35 s
- Total tokens
- 4,06,663
- Cost / trial
- $0.0991
- Estimated run cost
- $2.3782
- Adjudication
- 24 trials · 0 overrides
- Errors
- 0
xai-grok-4.5-high
xai / grok-4.5 / high effort
- Latency p50
- 14 s
- Latency p95
- 34 s
- Total tokens
- 5,43,690
- Cost / trial
- $0.0486
- Estimated run cost
- $1.1664
- Adjudication
- 24 trials · 0 overrides
- Errors
- 0
kimi-code-k3-max
kimi-code / k3 / max effort
- Latency p50
- 47 s
- Latency p95
- 180 s
- Total tokens
- 4,84,474
- Cost / trial
- —
- Estimated run cost
- —
- Adjudication
- 24 trials · 1 override
- Errors
- 0
openai-gpt-5.6-terra-max
openai / gpt-5.6-terra / max effort
- Latency p50
- 20 s
- Latency p95
- 35 s
- Total tokens
- 4,57,110
- Cost / trial
- $0.0646
- Estimated run cost
- $1.5495
- Adjudication
- 24 trials · 2 overrides
- Errors
- 0
bedrock-claude-opus-5-medium
bedrock / global.anthropic.claude-opus-5 / medium effort
- Latency p50
- 16 s
- Latency p95
- 28 s
- Total tokens
- 7,81,157
- Cost / trial
- $0.1804
- Estimated run cost
- $4.3298
- Adjudication
- 24 trials · 0 overrides
- Errors
- 0
openai-gpt-5.6-luna-xhigh
openai / gpt-5.6-luna / xhigh effort
- Latency p50
- 8.4 s
- Latency p95
- 28 s
- Total tokens
- 3,66,872
- Cost / trial
- $0.0187
- Estimated run cost
- $0.4493
- Adjudication
- 24 trials · 6 overrides
- Errors
- 0
openai-gpt-5.5-medium
openai / gpt-5.5 / medium effort
- Latency p50
- 15 s
- Latency p95
- 20 s
- Total tokens
- 3,34,798
- Cost / trial
- $0.0795
- Estimated run cost
- $1.9084
- Adjudication
- 24 trials · 5 overrides
- Errors
- 0
Quality versus cost
Each solid point is one Agent Configuration at its long-term list price. Hollow points are temporary alternate rates (for example launch promo) from the same trial tokens. Hover or focus a point for details. Lower cost is left, higher pass rate is up, and the useful direction is toward the upper-left.
The dotted line marks the observed frontier on list prices only: no solid point is both cheaper and more accurate. 1 configuration omitted because estimated cost is unavailable or zero. Hollow points are temporary alternate rates (for example launch promo) derived from the same trials; they do not move the frontier or the leaderboard cost column.
- openai-gpt-5.6-luna-max: 83.3% trial pass rate, $0.0261 estimated cost per trial.
- bedrock-claude-sonnet-5-medium: 83.3% trial pass rate, $0.0928 estimated cost per trial.
- bedrock-claude-sonnet-5-medium__promo: 83.3% trial pass rate, $0.0618 estimated cost per trial.
- openai-gpt-5.6-sol-medium: 79.2% trial pass rate, $0.0750 estimated cost per trial.
- openai-gpt-5.6-sol-high: 75.0% trial pass rate, $0.0991 estimated cost per trial.
- xai-grok-4.5-high: 75.0% trial pass rate, $0.0486 estimated cost per trial.
- openai-gpt-5.6-terra-max: 70.8% trial pass rate, $0.0646 estimated cost per trial.
- bedrock-claude-opus-5-medium: 70.8% trial pass rate, $0.1804 estimated cost per trial.
- openai-gpt-5.6-luna-xhigh: 66.7% trial pass rate, $0.0187 estimated cost per trial.
- openai-gpt-5.5-medium: 54.2% trial pass rate, $0.0795 estimated cost per trial.