← leaderboard

Claude Opus 4.8

Claude 4.8

Ranks #19 of 43 on reading real contracts into structured billing data-6.1 pts vs GPT-5.5.

63.2%
Accuracy
#19 of 43
24.3%
Hallucination
HIGH-confidence & wrong
$268
Cost / 1k contracts
$0.27 each · via OpenRouter
40.6s
Median latency
p90 43.8s
±3.8
Run-to-run σ
3 runs
0.24
Value
acc. pts per $/1k
Tokens per contract
Input34,639
Output3,796
Reasoning0

Reliability 100% · valid structured output across 18 calls.

Want the full picture? How we score →