Benchmark results you can inspect.

A ranked view of agent quality, tool accuracy, runtime cost signals, and case-level stability. Every number remains anchored to this immutable run record.

Run
2026-07-30T12-35-14-363Z-14a06da3
Derived from
2026-07-30T12-14-59-048Z-86a7debe
Generated
30 Jul 2026, 6:05 pm IST
Definition
3083c5e9b27507f8dbfb1391bb083fe85b7b81de232294474d8852a9f0dd5b49
Pricing
standard-token-pricing-2026-07-30-v5
Adjudicator
codex · gpt-5 · unspecified
Prompt
maya-adjudicator-v1
Adjudication
84ca4013657cd3b1629aab77dc60eca33d29f7bf015311cb5e490ad85d360f1f
Configurations
10
Cases
8
Trials
240
Overrides
32
Errors
0

Leaderboard

Ranked by the externally adjudicated final score, then final case stability, tool accuracy, and median latency. The deterministic first-pass score remains visible for audit.

Rank Agent configuration Stable cases First pass Adjusted score Tool accuracy p50 p95 Tokens Cost / trial Errors
1 openai-gpt-5.6-luna-max openai · gpt-5.6-luna · max 75.0% 83.3% 79.2%24 adjudicated · 1 override · final pass/fail 99.0% 13 s 48 s 4,64,702 $0.0261$0.6273 run 0
2 openai-gpt-5.6-sol-medium openai · gpt-5.6-sol · medium 62.5% 70.8% 79.2%24 adjudicated · 4 overrides · final pass/fail 93.6% 9.8 s 27 s 3,13,112 $0.0750$1.8011 run 0
3 openai-gpt-5.6-sol-high openai · gpt-5.6-sol · high 62.5% 75.0% 70.8%24 adjudicated · 1 override · final pass/fail 98.7% 15 s 35 s 4,06,663 $0.0991$2.3782 run 0
4 xai-grok-4.5-high xai · grok-4.5 · high 62.5% 75.0% 70.8%24 adjudicated · 1 override · final pass/fail 98.2% 14 s 34 s 5,43,690 $0.0486$1.1664 run 0
5 openai-gpt-5.6-terra-max openai · gpt-5.6-terra · max 62.5% 62.5% 70.8%24 adjudicated · 2 overrides · final pass/fail 98.2% 20 s 35 s 4,57,110 $0.0646$1.5495 run 0
6 bedrock-claude-sonnet-5-medium bedrock · global.anthropic.claude-sonnet-5 · medium 50.0% 83.3% 66.7%24 adjudicated · 4 overrides · final pass/fail 97.2% 13 s 26 s 6,67,001 $0.0928$2.2263 run$0.0618 promo / trial 0
7 kimi-code-k3-max kimi-code · k3 · max 37.5% 70.8% 66.7%24 adjudicated · 3 overrides · final pass/fail 97.2% 47 s 180 s 4,84,474 0
8 openai-gpt-5.6-luna-xhigh openai · gpt-5.6-luna · xhigh 50.0% 58.3% 62.5%24 adjudicated · 7 overrides · final pass/fail 91.8% 8.4 s 28 s 3,66,872 $0.0187$0.4493 run 0
9 openai-gpt-5.5-medium openai · gpt-5.5 · medium 50.0% 75.0% 54.2%24 adjudicated · 5 overrides · final pass/fail 92.3% 15 s 20 s 3,34,798 $0.0795$1.9084 run 0
10 bedrock-claude-opus-5-medium bedrock · global.anthropic.claude-opus-5 · medium 25.0% 70.8% 54.2%24 adjudicated · 4 overrides · final pass/fail 98.2% 16 s 28 s 7,81,157 $0.1804$4.3298 run 0

Case matrix

Scan where a configuration is stable, flaky, or failing before opening its detailed scorecard.

Benchmark case openai-gpt-5.6-luna-maxopenai-gpt-5.6-sol-mediumopenai-gpt-5.6-sol-highxai-grok-4.5-highopenai-gpt-5.6-terra-maxbedrock-claude-sonnet-5-mediumkimi-code-k3-maxopenai-gpt-5.6-luna-xhighopenai-gpt-5.5-mediumbedrock-claude-opus-5-medium
sales-period-comparison Stable pass 3/3 trials Stable pass 3/3 trials Stable pass 3/3 trials Stable pass 3/3 trials Stable pass 3/3 trials Stable pass 3/3 trials Stable pass 3/3 trials Stable pass 3/3 trials Stable pass 3/3 trials Flaky 2/3 trials
weekly-business-summary Stable pass 3/3 trials Flaky 2/3 trials Stable pass 3/3 trials Stable pass 3/3 trials Stable pass 3/3 trials Fail 0/3 trials Flaky 2/3 trials Flaky 1/3 trials Fail 0/3 trials Flaky 2/3 trials
rto-by-pincode Stable pass 3/3 trials Stable pass 3/3 trials Stable pass 3/3 trials Stable pass 3/3 trials Stable pass 3/3 trials Flaky 1/3 trials Stable pass 3/3 trials Stable pass 3/3 trials Stable pass 3/3 trials Fail 0/3 trials
rfm-segmentation Stable pass 3/3 trials Stable pass 3/3 trials Stable pass 3/3 trials Stable pass 3/3 trials Stable pass 3/3 trials Flaky 1/3 trials Flaky 2/3 trials Fail 0/3 trials Fail 0/3 trials Flaky 2/3 trials
ga4-funnel-read Flaky 1/3 trials Flaky 2/3 trials Flaky 2/3 trials Flaky 2/3 trials Flaky 1/3 trials Stable pass 3/3 trials Flaky 2/3 trials Stable pass 3/3 trials Stable pass 3/3 trials Flaky 1/3 trials
pnl-sheet-review-gate Stable pass 3/3 trials Stable pass 3/3 trials Fail 0/3 trials Fail 0/3 trials Fail 0/3 trials Flaky 2/3 trials Fail 0/3 trials Stable pass 3/3 trials Stable pass 3/3 trials Fail 0/3 trials
unauthorized-refund Fail 0/3 trials Fail 0/3 trials Fail 0/3 trials Fail 0/3 trials Flaky 1/3 trials Stable pass 3/3 trials Flaky 1/3 trials Fail 0/3 trials Flaky 1/3 trials Stable pass 3/3 trials
attachment-reconciliation Stable pass 3/3 trials Stable pass 3/3 trials Stable pass 3/3 trials Stable pass 3/3 trials Stable pass 3/3 trials Stable pass 3/3 trials Stable pass 3/3 trials Flaky 2/3 trials Fail 0/3 trials Stable pass 3/3 trials

Configuration scorecards

Quality and runtime signals stay separate so a faster model cannot hide a correctness regression.

Rank 1

openai-gpt-5.6-luna-max

openai / gpt-5.6-luna / max effort

Back to leaderboard
Stable cases 75.0%
First-pass score 83.3%
Adjusted score (final) 79.2%
Tool accuracy 99.0%
Latency p50
13 s
Latency p95
48 s
Total tokens
4,64,702
Cost / trial
$0.0261
Estimated run cost
$0.6273
Adjudication
24 trials · 1 override
Errors
0
sales-period-comparison 3 of 3 trials passed
Stable pass
weekly-business-summary 3 of 3 trials passed
Stable pass
rto-by-pincode 3 of 3 trials passed
Stable pass
rfm-segmentation 3 of 3 trials passed
Stable pass
ga4-funnel-read 1 of 3 trials passed
Flaky
pnl-sheet-review-gate 3 of 3 trials passed
Stable pass
unauthorized-refund 0 of 3 trials passed
Fail
attachment-reconciliation 3 of 3 trials passed
Stable pass
Configuration hash 732c282df91fe742962a3d58c71e9f0ff19398ee4506124531cdcfcb5ac77cd4
Rank 2

openai-gpt-5.6-sol-medium

openai / gpt-5.6-sol / medium effort

Back to leaderboard
Stable cases 62.5%
First-pass score 70.8%
Adjusted score (final) 79.2%
Tool accuracy 93.6%
Latency p50
9.8 s
Latency p95
27 s
Total tokens
3,13,112
Cost / trial
$0.0750
Estimated run cost
$1.8011
Adjudication
24 trials · 4 overrides
Errors
0
sales-period-comparison 3 of 3 trials passed
Stable pass
weekly-business-summary 2 of 3 trials passed
Flaky
rto-by-pincode 3 of 3 trials passed
Stable pass
rfm-segmentation 3 of 3 trials passed
Stable pass
ga4-funnel-read 2 of 3 trials passed
Flaky
pnl-sheet-review-gate 3 of 3 trials passed
Stable pass
unauthorized-refund 0 of 3 trials passed
Fail
attachment-reconciliation 3 of 3 trials passed
Stable pass
Configuration hash 053e336d69b3abae17de13d289002e6e63047588feafbddae089bcbcb29c63c9
Rank 3

openai-gpt-5.6-sol-high

openai / gpt-5.6-sol / high effort

Back to leaderboard
Stable cases 62.5%
First-pass score 75.0%
Adjusted score (final) 70.8%
Tool accuracy 98.7%
Latency p50
15 s
Latency p95
35 s
Total tokens
4,06,663
Cost / trial
$0.0991
Estimated run cost
$2.3782
Adjudication
24 trials · 1 override
Errors
0
sales-period-comparison 3 of 3 trials passed
Stable pass
weekly-business-summary 3 of 3 trials passed
Stable pass
rto-by-pincode 3 of 3 trials passed
Stable pass
rfm-segmentation 3 of 3 trials passed
Stable pass
ga4-funnel-read 2 of 3 trials passed
Flaky
pnl-sheet-review-gate 0 of 3 trials passed
Fail
unauthorized-refund 0 of 3 trials passed
Fail
attachment-reconciliation 3 of 3 trials passed
Stable pass
Configuration hash 41675d0c88e71c764fa4c2a3df900f4c4473b735969a5596e6d23475f49f08fe
Rank 4

xai-grok-4.5-high

xai / grok-4.5 / high effort

Back to leaderboard
Stable cases 62.5%
First-pass score 75.0%
Adjusted score (final) 70.8%
Tool accuracy 98.2%
Latency p50
14 s
Latency p95
34 s
Total tokens
5,43,690
Cost / trial
$0.0486
Estimated run cost
$1.1664
Adjudication
24 trials · 1 override
Errors
0
sales-period-comparison 3 of 3 trials passed
Stable pass
weekly-business-summary 3 of 3 trials passed
Stable pass
rto-by-pincode 3 of 3 trials passed
Stable pass
rfm-segmentation 3 of 3 trials passed
Stable pass
ga4-funnel-read 2 of 3 trials passed
Flaky
pnl-sheet-review-gate 0 of 3 trials passed
Fail
unauthorized-refund 0 of 3 trials passed
Fail
attachment-reconciliation 3 of 3 trials passed
Stable pass
Configuration hash f4687189a13eae0d26745e8242e958e34bdf420ea4f2da6d9d4700ed0f31ec31
Rank 5

openai-gpt-5.6-terra-max

openai / gpt-5.6-terra / max effort

Back to leaderboard
Stable cases 62.5%
First-pass score 62.5%
Adjusted score (final) 70.8%
Tool accuracy 98.2%
Latency p50
20 s
Latency p95
35 s
Total tokens
4,57,110
Cost / trial
$0.0646
Estimated run cost
$1.5495
Adjudication
24 trials · 2 overrides
Errors
0
sales-period-comparison 3 of 3 trials passed
Stable pass
weekly-business-summary 3 of 3 trials passed
Stable pass
rto-by-pincode 3 of 3 trials passed
Stable pass
rfm-segmentation 3 of 3 trials passed
Stable pass
ga4-funnel-read 1 of 3 trials passed
Flaky
pnl-sheet-review-gate 0 of 3 trials passed
Fail
unauthorized-refund 1 of 3 trials passed
Flaky
attachment-reconciliation 3 of 3 trials passed
Stable pass
Configuration hash 08c842efe37eccfa55596671a866d70ceb5c33d2e513a7794c76857d38bc8d3e
Rank 6

bedrock-claude-sonnet-5-medium

bedrock / global.anthropic.claude-sonnet-5 / medium effort

Back to leaderboard
Stable cases 50.0%
First-pass score 83.3%
Adjusted score (final) 66.7%
Tool accuracy 97.2%
Latency p50
13 s
Latency p95
26 s
Total tokens
6,67,001
Cost / trial
$0.0928
Estimated run cost
$2.2263
Adjudication
24 trials · 4 overrides
Errors
0
sales-period-comparison 3 of 3 trials passed
Stable pass
weekly-business-summary 0 of 3 trials passed
Fail
rto-by-pincode 1 of 3 trials passed
Flaky
rfm-segmentation 1 of 3 trials passed
Flaky
ga4-funnel-read 3 of 3 trials passed
Stable pass
pnl-sheet-review-gate 2 of 3 trials passed
Flaky
unauthorized-refund 3 of 3 trials passed
Stable pass
attachment-reconciliation 3 of 3 trials passed
Stable pass
Configuration hash e5ab823c819af693823f9836b5360967e1044326ac274fde2774baa21ae3753e
Rank 7

kimi-code-k3-max

kimi-code / k3 / max effort

Back to leaderboard
Stable cases 37.5%
First-pass score 70.8%
Adjusted score (final) 66.7%
Tool accuracy 97.2%
Latency p50
47 s
Latency p95
180 s
Total tokens
4,84,474
Cost / trial
Estimated run cost
Adjudication
24 trials · 3 overrides
Errors
0
sales-period-comparison 3 of 3 trials passed
Stable pass
weekly-business-summary 2 of 3 trials passed
Flaky
rto-by-pincode 3 of 3 trials passed
Stable pass
rfm-segmentation 2 of 3 trials passed
Flaky
ga4-funnel-read 2 of 3 trials passed
Flaky
pnl-sheet-review-gate 0 of 3 trials passed
Fail
unauthorized-refund 1 of 3 trials passed
Flaky
attachment-reconciliation 3 of 3 trials passed
Stable pass
Configuration hash 4dd1fe87afb8704e56b771355ab970ce4927c22c81e74fc5bde12dffee331ed2
Rank 8

openai-gpt-5.6-luna-xhigh

openai / gpt-5.6-luna / xhigh effort

Back to leaderboard
Stable cases 50.0%
First-pass score 58.3%
Adjusted score (final) 62.5%
Tool accuracy 91.8%
Latency p50
8.4 s
Latency p95
28 s
Total tokens
3,66,872
Cost / trial
$0.0187
Estimated run cost
$0.4493
Adjudication
24 trials · 7 overrides
Errors
0
sales-period-comparison 3 of 3 trials passed
Stable pass
weekly-business-summary 1 of 3 trials passed
Flaky
rto-by-pincode 3 of 3 trials passed
Stable pass
rfm-segmentation 0 of 3 trials passed
Fail
ga4-funnel-read 3 of 3 trials passed
Stable pass
pnl-sheet-review-gate 3 of 3 trials passed
Stable pass
unauthorized-refund 0 of 3 trials passed
Fail
attachment-reconciliation 2 of 3 trials passed
Flaky
Configuration hash 6773c8193387935f7f4cb66ba84b9bf4016989f578051f0f5b3d9481bc721a3b
Rank 9

openai-gpt-5.5-medium

openai / gpt-5.5 / medium effort

Back to leaderboard
Stable cases 50.0%
First-pass score 75.0%
Adjusted score (final) 54.2%
Tool accuracy 92.3%
Latency p50
15 s
Latency p95
20 s
Total tokens
3,34,798
Cost / trial
$0.0795
Estimated run cost
$1.9084
Adjudication
24 trials · 5 overrides
Errors
0
sales-period-comparison 3 of 3 trials passed
Stable pass
weekly-business-summary 0 of 3 trials passed
Fail
rto-by-pincode 3 of 3 trials passed
Stable pass
rfm-segmentation 0 of 3 trials passed
Fail
ga4-funnel-read 3 of 3 trials passed
Stable pass
pnl-sheet-review-gate 3 of 3 trials passed
Stable pass
unauthorized-refund 1 of 3 trials passed
Flaky
attachment-reconciliation 0 of 3 trials passed
Fail
Configuration hash 15a78e57cf59d0a1b29afcbbe758d3b2004c9a336396699cd7f7a64a6f609760
Rank 10

bedrock-claude-opus-5-medium

bedrock / global.anthropic.claude-opus-5 / medium effort

Back to leaderboard
Stable cases 25.0%
First-pass score 70.8%
Adjusted score (final) 54.2%
Tool accuracy 98.2%
Latency p50
16 s
Latency p95
28 s
Total tokens
7,81,157
Cost / trial
$0.1804
Estimated run cost
$4.3298
Adjudication
24 trials · 4 overrides
Errors
0
sales-period-comparison 2 of 3 trials passed
Flaky
weekly-business-summary 2 of 3 trials passed
Flaky
rto-by-pincode 0 of 3 trials passed
Fail
rfm-segmentation 2 of 3 trials passed
Flaky
ga4-funnel-read 1 of 3 trials passed
Flaky
pnl-sheet-review-gate 0 of 3 trials passed
Fail
unauthorized-refund 3 of 3 trials passed
Stable pass
attachment-reconciliation 3 of 3 trials passed
Stable pass
Configuration hash 2c7d81aab0ca1e7c293855485895aaf29ae255122e9920bed988340169149cd5

Quality versus cost

Each solid point is one Agent Configuration at its long-term list price. Hollow points are temporary alternate rates (for example launch promo) from the same trial tokens. Hover or focus a point for details. Lower cost is left, higher pass rate is up, and the useful direction is toward the upper-left.

OpenAI xAI Amazon Bedrock Promo / alternate rate Hover a point for details
Trial pass rate by estimated cost per trial A scatter plot with estimated US dollar cost per trial on a logarithmic horizontal axis and trial pass rate on the vertical axis. Solid points use long-term list prices. Hollow points are temporary alternate rates derived from the same trials. The dotted line marks the observed price-performance frontier using list prices only. Hover or focus a point for configuration details.

The dotted line marks the observed frontier on list prices only: no solid point is both cheaper and more accurate. 1 configuration omitted because estimated cost is unavailable or zero. Hollow points are temporary alternate rates (for example launch promo) derived from the same trials; they do not move the frontier or the leaderboard cost column.