Back to leaderboard

Claude Opus 5

NewClaude 5

Ranks #2 of 35 on reading real contracts into structured billing data, +1.8 pts vs GPT-5.5.

77.8%
Accuracy
#2 of 35
3.6%
Hallucination
HIGH-confidence & wrong
$370
Cost / 1k contracts
$0.37 each · via OpenRouter
71.2s
Median latency
p90 89.6s
±0
Run-to-run σ
1 runs
0.21
Value
acc. pts per $/1k
Tokens per contract
Input36,515
Output7,478
Reasoning0

Reliability 100% · valid structured output across 30 calls.

Want the full picture? How we score