← leaderboard
Claude Opus 4.8
Claude 4.8Ranks #19 of 43 on reading real contracts into structured billing data — -6.1 pts vs GPT-5.5.
63.2%
Accuracy
#19 of 43
24.3%
Hallucination
HIGH-confidence & wrong
$268
Cost / 1k contracts
$0.27 each · via OpenRouter
40.6s
Median latency
p90 43.8s
±3.8
Run-to-run σ
3 runs
0.24
Value
acc. pts per $/1k
Tokens per contract
Input34,639
Output3,796
Reasoning0
Reliability 100% · valid structured output across 18 calls.
Want the full picture? How we score →