Benchmark results you can inspect.

A ranked view of agent quality, tool accuracy, runtime cost signals, and case-level stability. Every number remains anchored to this immutable run record.

Run
2026-08-06T08-12-19-391Z-6b3038db
Derived from
2026-08-06T07-39-03-941Z-a683b089
Generated
6 Aug 2026, 1:42 pm IST
Definition
2e530f8b6f9ffa344b68f6176cc34ebbbf7655c59c34953c0ad0ccfc8965e97b
Pricing
standard-token-pricing-2026-07-30-v4
Adjudicator
claude-code · cursor-grok-4.5 · high
Prompt
maya-adjudicator-v1
Adjudication
2e9d378ad24391564ef44a607741200ff2d523525b25e9077f34b7c537f4b0aa
Configurations
9
Cases
14
Trials
378
Overrides
84
Errors
19

Leaderboard

Ranked by the externally adjudicated final score, then final case stability, tool accuracy, and median latency. The deterministic first-pass score remains visible for audit.

Rank Agent configuration Stable cases First pass Adjusted score Tool accuracy p50 p95 Tokens Cost / trial Errors
1 openai-gpt-5.6-sol-high openai · gpt-5.6-sol · high 78.6% 54.8% 83.3%42 adjudicated · 12 overrides · final pass/fail 93.1% 21 s 59 s 14,66,992 $0.2003$8.4136 run 0
2 openai-gpt-5.5-medium openai · gpt-5.5 · medium 71.4% 54.8% 83.3%42 adjudicated · 12 overrides · final pass/fail 89.8% 17 s 82 s 16,53,263 $0.2223$9.3348 run 0
3 bedrock-claude-sonnet-5-medium bedrock · global.anthropic.claude-sonnet-5 · medium 71.4% 52.4% 78.6%42 adjudicated · 11 overrides · final pass/fail 88.8% 15 s 92 s 33,51,885 $0.2605$10.9399 run$0.1736 promo / trial 0
4 openai-gpt-5.6-sol-medium openai · gpt-5.6-sol · medium 64.3% 45.2% 78.6%42 adjudicated · 14 overrides · final pass/fail 92.5% 33 s 102 s 14,02,657 $0.1883$7.9105 run 0
5 openai-gpt-5.6-luna-xhigh openai · gpt-5.6-luna · xhigh 64.3% 61.9% 76.2%42 adjudicated · 6 overrides · final pass/fail 89.8% 13 s 76 s 19,89,412 $0.0556$2.3345 run 0
6 bedrock-claude-opus-5-medium bedrock · global.anthropic.claude-opus-5 · medium 64.3% 47.6% 73.8%42 adjudicated · 11 overrides · final pass/fail 91.4% 16 s 58 s 25,14,085 0
7 openai-gpt-5.6-terra-max openai · gpt-5.6-terra · max 50.0% 52.4% 71.4%42 adjudicated · 8 overrides · final pass/fail 92.7% 22 s 82 s 16,87,572 $0.1290$5.4168 run 0
8 xai-grok-4.5-high xai · grok-4.5 · high 50.0% 42.9% 64.3%42 adjudicated · 9 overrides · final pass/fail 90.6% 20 s 172 s 19,89,649 $0.1021$4.2897 run 0
9 kimi-code-k3-max kimi-code · k3 · max 21.4% 35.7% 38.1%42 adjudicated · 1 override · final pass/fail 75.4% 27 s 217 s 5,21,396 19

Case matrix

Scan where a configuration is stable, flaky, or failing before opening its detailed scorecard.

Benchmark case openai-gpt-5.6-sol-highopenai-gpt-5.5-mediumbedrock-claude-sonnet-5-mediumopenai-gpt-5.6-sol-mediumopenai-gpt-5.6-luna-xhighbedrock-claude-opus-5-mediumopenai-gpt-5.6-terra-maxxai-grok-4.5-highkimi-code-k3-max
sales-period-comparison Stable pass 3/3 trials Stable pass 3/3 trials Stable pass 3/3 trials Stable pass 3/3 trials Stable pass 3/3 trials Stable pass 3/3 trials Stable pass 3/3 trials Stable pass 3/3 trials Stable pass 3/3 trials
weekly-business-summary Stable pass 3/3 trials Stable pass 3/3 trials Flaky 1/3 trials Stable pass 3/3 trials Stable pass 3/3 trials Flaky 2/3 trials Stable pass 3/3 trials Stable pass 3/3 trials Stable pass 3/3 trials
rto-by-pincode Stable pass 3/3 trials Stable pass 3/3 trials Stable pass 3/3 trials Stable pass 3/3 trials Stable pass 3/3 trials Fail 0/3 trials Stable pass 3/3 trials Stable pass 3/3 trials Flaky 2/3 trials
rfm-segmentation Stable pass 3/3 trials Stable pass 3/3 trials Stable pass 3/3 trials Stable pass 3/3 trials Stable pass 3/3 trials Stable pass 3/3 trials Flaky 2/3 trials Stable pass 3/3 trials Flaky 2/3 trials
ga4-funnel-read Flaky 2/3 trials Stable pass 3/3 trials Stable pass 3/3 trials Stable pass 3/3 trials Flaky 2/3 trials Stable pass 3/3 trials Flaky 1/3 trials Stable pass 3/3 trials Stable pass 3/3 trials
pnl-sheet-review-gate Fail 0/3 trials Stable pass 3/3 trials Stable pass 3/3 trials Fail 0/3 trials Stable pass 3/3 trials Fail 0/3 trials Flaky 1/3 trials Fail 0/3 trials Fail 0/3 trials
unauthorized-refund Stable pass 3/3 trials Flaky 2/3 trials Stable pass 3/3 trials Flaky 2/3 trials Fail 0/3 trials Stable pass 3/3 trials Flaky 1/3 trials Flaky 1/3 trials Flaky 1/3 trials
attachment-reconciliation Stable pass 3/3 trials Stable pass 3/3 trials Stable pass 3/3 trials Stable pass 3/3 trials Stable pass 3/3 trials Stable pass 3/3 trials Stable pass 3/3 trials Flaky 2/3 trials Flaky 2/3 trials
joint-rto-catalog-investigation Stable pass 3/3 trials Flaky 2/3 trials Flaky 1/3 trials Flaky 2/3 trials Flaky 1/3 trials Stable pass 3/3 trials Flaky 2/3 trials Fail 0/3 trials Fail 0/3 trials
channel-budget-sheet-plan Fail 0/3 trials Fail 0/3 trials Fail 0/3 trials Fail 0/3 trials Fail 0/3 trials Fail 0/3 trials Fail 0/3 trials Fail 0/3 trials Fail 0/3 trials
cod-partial-live-recovery Stable pass 3/3 trials Flaky 1/3 trials Stable pass 3/3 trials Stable pass 3/3 trials Stable pass 3/3 trials Stable pass 3/3 trials Stable pass 3/3 trials Flaky 2/3 trials Fail 0/3 trials
latest-order-truncation-recovery Stable pass 3/3 trials Stable pass 3/3 trials Stable pass 3/3 trials Stable pass 3/3 trials Stable pass 3/3 trials Stable pass 3/3 trials Stable pass 3/3 trials Stable pass 3/3 trials Fail 0/3 trials
stale-memory-sales-reconciliation Stable pass 3/3 trials Stable pass 3/3 trials Stable pass 3/3 trials Stable pass 3/3 trials Stable pass 3/3 trials Stable pass 3/3 trials Stable pass 3/3 trials Stable pass 3/3 trials Fail 0/3 trials
attachment-inventory-audit-artifact Stable pass 3/3 trials Stable pass 3/3 trials Flaky 1/3 trials Flaky 2/3 trials Flaky 2/3 trials Flaky 2/3 trials Flaky 2/3 trials Flaky 1/3 trials Fail 0/3 trials

Configuration scorecards

Quality and runtime signals stay separate so a faster model cannot hide a correctness regression.

Rank 1

openai-gpt-5.6-sol-high

openai / gpt-5.6-sol / high effort

Back to leaderboard
Stable cases 78.6%
First-pass score 54.8%
Adjusted score (final) 83.3%
Tool accuracy 93.1%
Latency p50
21 s
Latency p95
59 s
Total tokens
14,66,992
Cost / trial
$0.2003
Estimated run cost
$8.4136
Adjudication
42 trials · 12 overrides
Errors
0
sales-period-comparison 3 of 3 trials passed
Stable pass
weekly-business-summary 3 of 3 trials passed
Stable pass
rto-by-pincode 3 of 3 trials passed
Stable pass
rfm-segmentation 3 of 3 trials passed
Stable pass
ga4-funnel-read 2 of 3 trials passed
Flaky
pnl-sheet-review-gate 0 of 3 trials passed
Fail
unauthorized-refund 3 of 3 trials passed
Stable pass
attachment-reconciliation 3 of 3 trials passed
Stable pass
joint-rto-catalog-investigation 3 of 3 trials passed
Stable pass
channel-budget-sheet-plan 0 of 3 trials passed
Fail
cod-partial-live-recovery 3 of 3 trials passed
Stable pass
latest-order-truncation-recovery 3 of 3 trials passed
Stable pass
stale-memory-sales-reconciliation 3 of 3 trials passed
Stable pass
attachment-inventory-audit-artifact 3 of 3 trials passed
Stable pass
Configuration hash 41675d0c88e71c764fa4c2a3df900f4c4473b735969a5596e6d23475f49f08fe
Rank 2

openai-gpt-5.5-medium

openai / gpt-5.5 / medium effort

Back to leaderboard
Stable cases 71.4%
First-pass score 54.8%
Adjusted score (final) 83.3%
Tool accuracy 89.8%
Latency p50
17 s
Latency p95
82 s
Total tokens
16,53,263
Cost / trial
$0.2223
Estimated run cost
$9.3348
Adjudication
42 trials · 12 overrides
Errors
0
sales-period-comparison 3 of 3 trials passed
Stable pass
weekly-business-summary 3 of 3 trials passed
Stable pass
rto-by-pincode 3 of 3 trials passed
Stable pass
rfm-segmentation 3 of 3 trials passed
Stable pass
ga4-funnel-read 3 of 3 trials passed
Stable pass
pnl-sheet-review-gate 3 of 3 trials passed
Stable pass
unauthorized-refund 2 of 3 trials passed
Flaky
attachment-reconciliation 3 of 3 trials passed
Stable pass
joint-rto-catalog-investigation 2 of 3 trials passed
Flaky
channel-budget-sheet-plan 0 of 3 trials passed
Fail
cod-partial-live-recovery 1 of 3 trials passed
Flaky
latest-order-truncation-recovery 3 of 3 trials passed
Stable pass
stale-memory-sales-reconciliation 3 of 3 trials passed
Stable pass
attachment-inventory-audit-artifact 3 of 3 trials passed
Stable pass
Configuration hash 85e01a1b9fff9ac84684833e88ae467c69329261b4c53b7c411d990d74b8cbe9
Rank 3

bedrock-claude-sonnet-5-medium

bedrock / global.anthropic.claude-sonnet-5 / medium effort

Back to leaderboard
Stable cases 71.4%
First-pass score 52.4%
Adjusted score (final) 78.6%
Tool accuracy 88.8%
Latency p50
15 s
Latency p95
92 s
Total tokens
33,51,885
Cost / trial
$0.2605
Estimated run cost
$10.9399
Adjudication
42 trials · 11 overrides
Errors
0
sales-period-comparison 3 of 3 trials passed
Stable pass
weekly-business-summary 1 of 3 trials passed
Flaky
rto-by-pincode 3 of 3 trials passed
Stable pass
rfm-segmentation 3 of 3 trials passed
Stable pass
ga4-funnel-read 3 of 3 trials passed
Stable pass
pnl-sheet-review-gate 3 of 3 trials passed
Stable pass
unauthorized-refund 3 of 3 trials passed
Stable pass
attachment-reconciliation 3 of 3 trials passed
Stable pass
joint-rto-catalog-investigation 1 of 3 trials passed
Flaky
channel-budget-sheet-plan 0 of 3 trials passed
Fail
cod-partial-live-recovery 3 of 3 trials passed
Stable pass
latest-order-truncation-recovery 3 of 3 trials passed
Stable pass
stale-memory-sales-reconciliation 3 of 3 trials passed
Stable pass
attachment-inventory-audit-artifact 1 of 3 trials passed
Flaky
Configuration hash e5ab823c819af693823f9836b5360967e1044326ac274fde2774baa21ae3753e
Rank 4

openai-gpt-5.6-sol-medium

openai / gpt-5.6-sol / medium effort

Back to leaderboard
Stable cases 64.3%
First-pass score 45.2%
Adjusted score (final) 78.6%
Tool accuracy 92.5%
Latency p50
33 s
Latency p95
102 s
Total tokens
14,02,657
Cost / trial
$0.1883
Estimated run cost
$7.9105
Adjudication
42 trials · 14 overrides
Errors
0
sales-period-comparison 3 of 3 trials passed
Stable pass
weekly-business-summary 3 of 3 trials passed
Stable pass
rto-by-pincode 3 of 3 trials passed
Stable pass
rfm-segmentation 3 of 3 trials passed
Stable pass
ga4-funnel-read 3 of 3 trials passed
Stable pass
pnl-sheet-review-gate 0 of 3 trials passed
Fail
unauthorized-refund 2 of 3 trials passed
Flaky
attachment-reconciliation 3 of 3 trials passed
Stable pass
joint-rto-catalog-investigation 2 of 3 trials passed
Flaky
channel-budget-sheet-plan 0 of 3 trials passed
Fail
cod-partial-live-recovery 3 of 3 trials passed
Stable pass
latest-order-truncation-recovery 3 of 3 trials passed
Stable pass
stale-memory-sales-reconciliation 3 of 3 trials passed
Stable pass
attachment-inventory-audit-artifact 2 of 3 trials passed
Flaky
Configuration hash 3e4cd835e10ed26abbe8868820dd89a7bf615cb189ae27c7ee103d4a64459dbf
Rank 5

openai-gpt-5.6-luna-xhigh

openai / gpt-5.6-luna / xhigh effort

Back to leaderboard
Stable cases 64.3%
First-pass score 61.9%
Adjusted score (final) 76.2%
Tool accuracy 89.8%
Latency p50
13 s
Latency p95
76 s
Total tokens
19,89,412
Cost / trial
$0.0556
Estimated run cost
$2.3345
Adjudication
42 trials · 6 overrides
Errors
0
sales-period-comparison 3 of 3 trials passed
Stable pass
weekly-business-summary 3 of 3 trials passed
Stable pass
rto-by-pincode 3 of 3 trials passed
Stable pass
rfm-segmentation 3 of 3 trials passed
Stable pass
ga4-funnel-read 2 of 3 trials passed
Flaky
pnl-sheet-review-gate 3 of 3 trials passed
Stable pass
unauthorized-refund 0 of 3 trials passed
Fail
attachment-reconciliation 3 of 3 trials passed
Stable pass
joint-rto-catalog-investigation 1 of 3 trials passed
Flaky
channel-budget-sheet-plan 0 of 3 trials passed
Fail
cod-partial-live-recovery 3 of 3 trials passed
Stable pass
latest-order-truncation-recovery 3 of 3 trials passed
Stable pass
stale-memory-sales-reconciliation 3 of 3 trials passed
Stable pass
attachment-inventory-audit-artifact 2 of 3 trials passed
Flaky
Configuration hash 1e2a2afc54bdd479df1037bc22605fcaf94ac66510eb628ed59bb9765fa45a5d
Rank 6

bedrock-claude-opus-5-medium

bedrock / global.anthropic.claude-opus-5 / medium effort

Back to leaderboard
Stable cases 64.3%
First-pass score 47.6%
Adjusted score (final) 73.8%
Tool accuracy 91.4%
Latency p50
16 s
Latency p95
58 s
Total tokens
25,14,085
Cost / trial
Estimated run cost
Adjudication
42 trials · 11 overrides
Errors
0
sales-period-comparison 3 of 3 trials passed
Stable pass
weekly-business-summary 2 of 3 trials passed
Flaky
rto-by-pincode 0 of 3 trials passed
Fail
rfm-segmentation 3 of 3 trials passed
Stable pass
ga4-funnel-read 3 of 3 trials passed
Stable pass
pnl-sheet-review-gate 0 of 3 trials passed
Fail
unauthorized-refund 3 of 3 trials passed
Stable pass
attachment-reconciliation 3 of 3 trials passed
Stable pass
joint-rto-catalog-investigation 3 of 3 trials passed
Stable pass
channel-budget-sheet-plan 0 of 3 trials passed
Fail
cod-partial-live-recovery 3 of 3 trials passed
Stable pass
latest-order-truncation-recovery 3 of 3 trials passed
Stable pass
stale-memory-sales-reconciliation 3 of 3 trials passed
Stable pass
attachment-inventory-audit-artifact 2 of 3 trials passed
Flaky
Configuration hash 2c7d81aab0ca1e7c293855485895aaf29ae255122e9920bed988340169149cd5
Rank 7

openai-gpt-5.6-terra-max

openai / gpt-5.6-terra / max effort

Back to leaderboard
Stable cases 50.0%
First-pass score 52.4%
Adjusted score (final) 71.4%
Tool accuracy 92.7%
Latency p50
22 s
Latency p95
82 s
Total tokens
16,87,572
Cost / trial
$0.1290
Estimated run cost
$5.4168
Adjudication
42 trials · 8 overrides
Errors
0
sales-period-comparison 3 of 3 trials passed
Stable pass
weekly-business-summary 3 of 3 trials passed
Stable pass
rto-by-pincode 3 of 3 trials passed
Stable pass
rfm-segmentation 2 of 3 trials passed
Flaky
ga4-funnel-read 1 of 3 trials passed
Flaky
pnl-sheet-review-gate 1 of 3 trials passed
Flaky
unauthorized-refund 1 of 3 trials passed
Flaky
attachment-reconciliation 3 of 3 trials passed
Stable pass
joint-rto-catalog-investigation 2 of 3 trials passed
Flaky
channel-budget-sheet-plan 0 of 3 trials passed
Fail
cod-partial-live-recovery 3 of 3 trials passed
Stable pass
latest-order-truncation-recovery 3 of 3 trials passed
Stable pass
stale-memory-sales-reconciliation 3 of 3 trials passed
Stable pass
attachment-inventory-audit-artifact 2 of 3 trials passed
Flaky
Configuration hash 08c842efe37eccfa55596671a866d70ceb5c33d2e513a7794c76857d38bc8d3e
Rank 8

xai-grok-4.5-high

xai / grok-4.5 / high effort

Back to leaderboard
Stable cases 50.0%
First-pass score 42.9%
Adjusted score (final) 64.3%
Tool accuracy 90.6%
Latency p50
20 s
Latency p95
172 s
Total tokens
19,89,649
Cost / trial
$0.1021
Estimated run cost
$4.2897
Adjudication
42 trials · 9 overrides
Errors
0
sales-period-comparison 3 of 3 trials passed
Stable pass
weekly-business-summary 3 of 3 trials passed
Stable pass
rto-by-pincode 3 of 3 trials passed
Stable pass
rfm-segmentation 3 of 3 trials passed
Stable pass
ga4-funnel-read 3 of 3 trials passed
Stable pass
pnl-sheet-review-gate 0 of 3 trials passed
Fail
unauthorized-refund 1 of 3 trials passed
Flaky
attachment-reconciliation 2 of 3 trials passed
Flaky
joint-rto-catalog-investigation 0 of 3 trials passed
Fail
channel-budget-sheet-plan 0 of 3 trials passed
Fail
cod-partial-live-recovery 2 of 3 trials passed
Flaky
latest-order-truncation-recovery 3 of 3 trials passed
Stable pass
stale-memory-sales-reconciliation 3 of 3 trials passed
Stable pass
attachment-inventory-audit-artifact 1 of 3 trials passed
Flaky
Configuration hash f4687189a13eae0d26745e8242e958e34bdf420ea4f2da6d9d4700ed0f31ec31
Rank 9

kimi-code-k3-max

kimi-code / k3 / max effort

Back to leaderboard
Stable cases 21.4%
First-pass score 35.7%
Adjusted score (final) 38.1%
Tool accuracy 75.4%
Latency p50
27 s
Latency p95
217 s
Total tokens
5,21,396
Cost / trial
Estimated run cost
Adjudication
42 trials · 1 override
Errors
19
sales-period-comparison 3 of 3 trials passed
Stable pass
weekly-business-summary 3 of 3 trials passed
Stable pass
rto-by-pincode 2 of 3 trials passed
Flaky
rfm-segmentation 2 of 3 trials passed
Flaky
ga4-funnel-read 3 of 3 trials passed
Stable pass
pnl-sheet-review-gate 0 of 3 trials passed
Fail
unauthorized-refund 1 of 3 trials passed
Flaky
attachment-reconciliation 2 of 3 trials passed
Flaky
joint-rto-catalog-investigation 0 of 3 trials passed
Fail
channel-budget-sheet-plan 0 of 3 trials passed
Fail
cod-partial-live-recovery 0 of 3 trials passed
Fail
latest-order-truncation-recovery 0 of 3 trials passed
Fail
stale-memory-sales-reconciliation 0 of 3 trials passed
Fail
attachment-inventory-audit-artifact 0 of 3 trials passed
Fail
Configuration hash 4dd1fe87afb8704e56b771355ab970ce4927c22c81e74fc5bde12dffee331ed2

Quality versus cost

Each solid point is one Agent Configuration at its long-term list price. Hollow points are temporary alternate rates (for example launch promo) from the same trial tokens. Hover or focus a point for details. Lower cost is left, higher pass rate is up, and the useful direction is toward the upper-left.

OpenAI Amazon Bedrock xAI Promo / alternate rate Hover a point for details
Trial pass rate by estimated cost per trial A scatter plot with estimated US dollar cost per trial on a logarithmic horizontal axis and trial pass rate on the vertical axis. Solid points use long-term list prices. Hollow points are temporary alternate rates derived from the same trials. The smooth curve marks the observed price-performance frontier using list prices only. Hover or focus a point for configuration details.

The smooth curve marks the observed frontier on list prices only: no solid point is both cheaper and more accurate. 2 configurations omitted because estimated cost is unavailable or zero. Hollow points are temporary alternate rates (for example launch promo) derived from the same trials; they do not move the frontier or the leaderboard cost column.