Back to leaderboard

Claude Sonnet 5

NewClaude 5

Ranks #8 of 35 on reading real contracts into structured billing data, -0.6 pts vs GPT-5.5.

75.4%
Accuracy
#8 of 35
6.2%
Hallucination
HIGH-confidence & wrong
$200
Cost / 1k contracts
$0.20 each · via OpenRouter
106.7s
Median latency
p90 154.2s
±0
Run-to-run σ
1 runs
0.38
Value
acc. pts per $/1k
Tokens per contract
Input36,515
Output12,714
Reasoning0

Reliability 100% · valid structured output across 30 calls.

Want the full picture? How we score