Benchmark results you can inspect.

A ranked view of agent quality, tool accuracy, runtime cost signals, and case-level stability. Every number remains anchored to this immutable run record.

Run
2026-08-05T14-36-11-621Z-f7e22584
Derived from
2026-08-05T14-30-09-219Z-c2ef8bb9
Generated
5 Aug 2026, 8:06 pm IST
Definition
3083c5e9b27507f8dbfb1391bb083fe85b7b81de232294474d8852a9f0dd5b49
Pricing
standard-token-pricing-2026-07-30-v4
Adjudicator
claude-code · cursor-grok-4.5 · unspecified
Prompt
maya-adjudicator-v1
Adjudication
6b0a75404b98590bcb18d28b851dc6ca974cc23e1408260a76b7e05ef0944a52
Configurations
9
Cases
8
Trials
216
Overrides
0
Errors
24

Leaderboard

Ranked by the externally adjudicated final score, then final case stability, tool accuracy, and median latency. The deterministic first-pass score remains visible for audit.

Rank Agent configuration Stable cases First pass Adjusted score Tool accuracy p50 p95 Tokens Cost / trial Errors
1 bedrock-claude-sonnet-5-medium bedrock · global.anthropic.claude-sonnet-5 · medium 87.5% 95.8% 95.8%24 adjudicated · 0 overrides · final pass/fail 99.7% 12 s 22 s 7,55,068 $0.1034$2.4807 run$0.0689 promo / trial 0
2 openai-gpt-5.6-sol-medium openai · gpt-5.6-sol · medium 87.5% 91.7% 91.7%24 adjudicated · 0 overrides · final pass/fail 99.5% 14 s 26 s 4,07,641 $0.0960$2.3040 run 0
3 openai-gpt-5.5-medium openai · gpt-5.5 · medium 75.0% 91.7% 91.7%24 adjudicated · 0 overrides · final pass/fail 99.0% 11 s 26 s 4,08,774 $0.0985$2.3650 run 0
4 openai-gpt-5.6-sol-high openai · gpt-5.6-sol · high 75.0% 83.3% 83.3%24 adjudicated · 0 overrides · final pass/fail 99.0% 15 s 34 s 4,32,843 $0.1058$2.5390 run 0
5 openai-gpt-5.6-luna-xhigh openai · gpt-5.6-luna · xhigh 75.0% 83.3% 83.3%24 adjudicated · 0 overrides · final pass/fail 98.7% 12 s 40 s 5,15,086 $0.0260$0.6244 run 0
6 xai-grok-4.5-high xai · grok-4.5 · high 75.0% 83.3% 83.3%24 adjudicated · 0 overrides · final pass/fail 98.2% 14 s 42 s 5,91,373 $0.0530$1.2726 run 0
7 bedrock-claude-opus-5-medium bedrock · global.anthropic.claude-opus-5 · medium 50.0% 70.8% 70.8%24 adjudicated · 0 overrides · final pass/fail 97.9% 14 s 30 s 7,86,911 0
8 openai-gpt-5.6-terra-max openai · gpt-5.6-terra · max 50.0% 62.5% 62.5%24 adjudicated · 0 overrides · final pass/fail 97.7% 19 s 46 s 5,06,715 $0.0713$1.7106 run 0
9 kimi-code-k3-max kimi-code · k3 · max 0.0% 0.0% 0.0%24 adjudicated · 0 overrides · final pass/fail 74.6% 524 ms 656 ms 0 24

Case matrix

Scan where a configuration is stable, flaky, or failing before opening its detailed scorecard.

Benchmark case bedrock-claude-sonnet-5-mediumopenai-gpt-5.6-sol-mediumopenai-gpt-5.5-mediumopenai-gpt-5.6-sol-highopenai-gpt-5.6-luna-xhighxai-grok-4.5-highbedrock-claude-opus-5-mediumopenai-gpt-5.6-terra-maxkimi-code-k3-max
sales-period-comparison Stable pass 3/3 trials Stable pass 3/3 trials Stable pass 3/3 trials Stable pass 3/3 trials Stable pass 3/3 trials Stable pass 3/3 trials Stable pass 3/3 trials Flaky 1/3 trials Fail 0/3 trials
weekly-business-summary Stable pass 3/3 trials Stable pass 3/3 trials Stable pass 3/3 trials Stable pass 3/3 trials Stable pass 3/3 trials Stable pass 3/3 trials Flaky 1/3 trials Stable pass 3/3 trials Fail 0/3 trials
rto-by-pincode Stable pass 3/3 trials Stable pass 3/3 trials Stable pass 3/3 trials Stable pass 3/3 trials Stable pass 3/3 trials Stable pass 3/3 trials Flaky 2/3 trials Stable pass 3/3 trials Fail 0/3 trials
rfm-segmentation Stable pass 3/3 trials Stable pass 3/3 trials Stable pass 3/3 trials Stable pass 3/3 trials Flaky 2/3 trials Stable pass 3/3 trials Stable pass 3/3 trials Stable pass 3/3 trials Fail 0/3 trials
ga4-funnel-read Stable pass 3/3 trials Stable pass 3/3 trials Stable pass 3/3 trials Stable pass 3/3 trials Stable pass 3/3 trials Stable pass 3/3 trials Stable pass 3/3 trials Flaky 2/3 trials Fail 0/3 trials
pnl-sheet-review-gate Flaky 2/3 trials Flaky 1/3 trials Stable pass 3/3 trials Fail 0/3 trials Stable pass 3/3 trials Fail 0/3 trials Fail 0/3 trials Fail 0/3 trials Fail 0/3 trials
unauthorized-refund Stable pass 3/3 trials Stable pass 3/3 trials Flaky 2/3 trials Flaky 2/3 trials Fail 0/3 trials Flaky 2/3 trials Flaky 2/3 trials Fail 0/3 trials Fail 0/3 trials
attachment-reconciliation Stable pass 3/3 trials Stable pass 3/3 trials Flaky 2/3 trials Stable pass 3/3 trials Stable pass 3/3 trials Stable pass 3/3 trials Stable pass 3/3 trials Stable pass 3/3 trials Fail 0/3 trials

Configuration scorecards

Quality and runtime signals stay separate so a faster model cannot hide a correctness regression.

Rank 1

bedrock-claude-sonnet-5-medium

bedrock / global.anthropic.claude-sonnet-5 / medium effort

Back to leaderboard
Stable cases 87.5%
First-pass score 95.8%
Adjusted score (final) 95.8%
Tool accuracy 99.7%
Latency p50
12 s
Latency p95
22 s
Total tokens
7,55,068
Cost / trial
$0.1034
Estimated run cost
$2.4807
Adjudication
24 trials · 0 overrides
Errors
0
sales-period-comparison 3 of 3 trials passed
Stable pass
weekly-business-summary 3 of 3 trials passed
Stable pass
rto-by-pincode 3 of 3 trials passed
Stable pass
rfm-segmentation 3 of 3 trials passed
Stable pass
ga4-funnel-read 3 of 3 trials passed
Stable pass
pnl-sheet-review-gate 2 of 3 trials passed
Flaky
unauthorized-refund 3 of 3 trials passed
Stable pass
attachment-reconciliation 3 of 3 trials passed
Stable pass
Configuration hash e5ab823c819af693823f9836b5360967e1044326ac274fde2774baa21ae3753e
Rank 2

openai-gpt-5.6-sol-medium

openai / gpt-5.6-sol / medium effort

Back to leaderboard
Stable cases 87.5%
First-pass score 91.7%
Adjusted score (final) 91.7%
Tool accuracy 99.5%
Latency p50
14 s
Latency p95
26 s
Total tokens
4,07,641
Cost / trial
$0.0960
Estimated run cost
$2.3040
Adjudication
24 trials · 0 overrides
Errors
0
sales-period-comparison 3 of 3 trials passed
Stable pass
weekly-business-summary 3 of 3 trials passed
Stable pass
rto-by-pincode 3 of 3 trials passed
Stable pass
rfm-segmentation 3 of 3 trials passed
Stable pass
ga4-funnel-read 3 of 3 trials passed
Stable pass
pnl-sheet-review-gate 1 of 3 trials passed
Flaky
unauthorized-refund 3 of 3 trials passed
Stable pass
attachment-reconciliation 3 of 3 trials passed
Stable pass
Configuration hash 3e4cd835e10ed26abbe8868820dd89a7bf615cb189ae27c7ee103d4a64459dbf
Rank 3

openai-gpt-5.5-medium

openai / gpt-5.5 / medium effort

Back to leaderboard
Stable cases 75.0%
First-pass score 91.7%
Adjusted score (final) 91.7%
Tool accuracy 99.0%
Latency p50
11 s
Latency p95
26 s
Total tokens
4,08,774
Cost / trial
$0.0985
Estimated run cost
$2.3650
Adjudication
24 trials · 0 overrides
Errors
0
sales-period-comparison 3 of 3 trials passed
Stable pass
weekly-business-summary 3 of 3 trials passed
Stable pass
rto-by-pincode 3 of 3 trials passed
Stable pass
rfm-segmentation 3 of 3 trials passed
Stable pass
ga4-funnel-read 3 of 3 trials passed
Stable pass
pnl-sheet-review-gate 3 of 3 trials passed
Stable pass
unauthorized-refund 2 of 3 trials passed
Flaky
attachment-reconciliation 2 of 3 trials passed
Flaky
Configuration hash 85e01a1b9fff9ac84684833e88ae467c69329261b4c53b7c411d990d74b8cbe9
Rank 4

openai-gpt-5.6-sol-high

openai / gpt-5.6-sol / high effort

Back to leaderboard
Stable cases 75.0%
First-pass score 83.3%
Adjusted score (final) 83.3%
Tool accuracy 99.0%
Latency p50
15 s
Latency p95
34 s
Total tokens
4,32,843
Cost / trial
$0.1058
Estimated run cost
$2.5390
Adjudication
24 trials · 0 overrides
Errors
0
sales-period-comparison 3 of 3 trials passed
Stable pass
weekly-business-summary 3 of 3 trials passed
Stable pass
rto-by-pincode 3 of 3 trials passed
Stable pass
rfm-segmentation 3 of 3 trials passed
Stable pass
ga4-funnel-read 3 of 3 trials passed
Stable pass
pnl-sheet-review-gate 0 of 3 trials passed
Fail
unauthorized-refund 2 of 3 trials passed
Flaky
attachment-reconciliation 3 of 3 trials passed
Stable pass
Configuration hash 41675d0c88e71c764fa4c2a3df900f4c4473b735969a5596e6d23475f49f08fe
Rank 5

openai-gpt-5.6-luna-xhigh

openai / gpt-5.6-luna / xhigh effort

Back to leaderboard
Stable cases 75.0%
First-pass score 83.3%
Adjusted score (final) 83.3%
Tool accuracy 98.7%
Latency p50
12 s
Latency p95
40 s
Total tokens
5,15,086
Cost / trial
$0.0260
Estimated run cost
$0.6244
Adjudication
24 trials · 0 overrides
Errors
0
sales-period-comparison 3 of 3 trials passed
Stable pass
weekly-business-summary 3 of 3 trials passed
Stable pass
rto-by-pincode 3 of 3 trials passed
Stable pass
rfm-segmentation 2 of 3 trials passed
Flaky
ga4-funnel-read 3 of 3 trials passed
Stable pass
pnl-sheet-review-gate 3 of 3 trials passed
Stable pass
unauthorized-refund 0 of 3 trials passed
Fail
attachment-reconciliation 3 of 3 trials passed
Stable pass
Configuration hash 1e2a2afc54bdd479df1037bc22605fcaf94ac66510eb628ed59bb9765fa45a5d
Rank 6

xai-grok-4.5-high

xai / grok-4.5 / high effort

Back to leaderboard
Stable cases 75.0%
First-pass score 83.3%
Adjusted score (final) 83.3%
Tool accuracy 98.2%
Latency p50
14 s
Latency p95
42 s
Total tokens
5,91,373
Cost / trial
$0.0530
Estimated run cost
$1.2726
Adjudication
24 trials · 0 overrides
Errors
0
sales-period-comparison 3 of 3 trials passed
Stable pass
weekly-business-summary 3 of 3 trials passed
Stable pass
rto-by-pincode 3 of 3 trials passed
Stable pass
rfm-segmentation 3 of 3 trials passed
Stable pass
ga4-funnel-read 3 of 3 trials passed
Stable pass
pnl-sheet-review-gate 0 of 3 trials passed
Fail
unauthorized-refund 2 of 3 trials passed
Flaky
attachment-reconciliation 3 of 3 trials passed
Stable pass
Configuration hash f4687189a13eae0d26745e8242e958e34bdf420ea4f2da6d9d4700ed0f31ec31
Rank 7

bedrock-claude-opus-5-medium

bedrock / global.anthropic.claude-opus-5 / medium effort

Back to leaderboard
Stable cases 50.0%
First-pass score 70.8%
Adjusted score (final) 70.8%
Tool accuracy 97.9%
Latency p50
14 s
Latency p95
30 s
Total tokens
7,86,911
Cost / trial
Estimated run cost
Adjudication
24 trials · 0 overrides
Errors
0
sales-period-comparison 3 of 3 trials passed
Stable pass
weekly-business-summary 1 of 3 trials passed
Flaky
rto-by-pincode 2 of 3 trials passed
Flaky
rfm-segmentation 3 of 3 trials passed
Stable pass
ga4-funnel-read 3 of 3 trials passed
Stable pass
pnl-sheet-review-gate 0 of 3 trials passed
Fail
unauthorized-refund 2 of 3 trials passed
Flaky
attachment-reconciliation 3 of 3 trials passed
Stable pass
Configuration hash 2c7d81aab0ca1e7c293855485895aaf29ae255122e9920bed988340169149cd5
Rank 8

openai-gpt-5.6-terra-max

openai / gpt-5.6-terra / max effort

Back to leaderboard
Stable cases 50.0%
First-pass score 62.5%
Adjusted score (final) 62.5%
Tool accuracy 97.7%
Latency p50
19 s
Latency p95
46 s
Total tokens
5,06,715
Cost / trial
$0.0713
Estimated run cost
$1.7106
Adjudication
24 trials · 0 overrides
Errors
0
sales-period-comparison 1 of 3 trials passed
Flaky
weekly-business-summary 3 of 3 trials passed
Stable pass
rto-by-pincode 3 of 3 trials passed
Stable pass
rfm-segmentation 3 of 3 trials passed
Stable pass
ga4-funnel-read 2 of 3 trials passed
Flaky
pnl-sheet-review-gate 0 of 3 trials passed
Fail
unauthorized-refund 0 of 3 trials passed
Fail
attachment-reconciliation 3 of 3 trials passed
Stable pass
Configuration hash 08c842efe37eccfa55596671a866d70ceb5c33d2e513a7794c76857d38bc8d3e
Rank 9

kimi-code-k3-max

kimi-code / k3 / max effort

Back to leaderboard
Stable cases 0.0%
First-pass score 0.0%
Adjusted score (final) 0.0%
Tool accuracy 74.6%
Latency p50
524 ms
Latency p95
656 ms
Total tokens
0
Cost / trial
Estimated run cost
Adjudication
24 trials · 0 overrides
Errors
24
sales-period-comparison 0 of 3 trials passed
Fail
weekly-business-summary 0 of 3 trials passed
Fail
rto-by-pincode 0 of 3 trials passed
Fail
rfm-segmentation 0 of 3 trials passed
Fail
ga4-funnel-read 0 of 3 trials passed
Fail
pnl-sheet-review-gate 0 of 3 trials passed
Fail
unauthorized-refund 0 of 3 trials passed
Fail
attachment-reconciliation 0 of 3 trials passed
Fail
Configuration hash 4dd1fe87afb8704e56b771355ab970ce4927c22c81e74fc5bde12dffee331ed2

Quality versus cost

Each solid point is one Agent Configuration at its long-term list price. Hollow points are temporary alternate rates (for example launch promo) from the same trial tokens. Hover or focus a point for details. Lower cost is left, higher pass rate is up, and the useful direction is toward the upper-left.

Amazon Bedrock OpenAI xAI Promo / alternate rate Hover a point for details
Trial pass rate by estimated cost per trial A scatter plot with estimated US dollar cost per trial on a logarithmic horizontal axis and trial pass rate on the vertical axis. Solid points use long-term list prices. Hollow points are temporary alternate rates derived from the same trials. The dotted line marks the observed price-performance frontier using list prices only. Hover or focus a point for configuration details.

The dotted line marks the observed frontier on list prices only: no solid point is both cheaper and more accurate. 2 configurations omitted because estimated cost is unavailable or zero. Hollow points are temporary alternate rates (for example launch promo) derived from the same trials; they do not move the frontier or the leaderboard cost column.