← leaderboard
Claude Opus 5
NewClaude 5Ranks #8 of 43 on reading real contracts into structured billing data — -0.9 pts vs GPT-5.5.
68.4%
Accuracy
#8 of 43
10.9%
Hallucination
HIGH-confidence & wrong
$346
Cost / 1k contracts
$0.35 each · via OpenRouter
67.1s
Median latency
p90 94.9s
±2.1
Run-to-run σ
3 runs
0.2
Value
acc. pts per $/1k
Tokens per contract
Input34,733
Output6,886
Reasoning426
Reliability 100% · valid structured output across 18 calls.
Want the full picture? How we score →