← leaderboard
Claude Sonnet 5
NewClaude 5Ranks #13 of 43 on reading real contracts into structured billing data — -3.4 pts vs GPT-5.5.
65.9%
Accuracy
#13 of 43
14%
Hallucination
HIGH-confidence & wrong
$185
Cost / 1k contracts
$0.18 each · via OpenRouter
92.6s
Median latency
p90 173.4s
±1.6
Run-to-run σ
3 runs
0.36
Value
acc. pts per $/1k
Tokens per contract
Input34,733
Output11,523
Reasoning1,722
Reliability 100% · valid structured output across 18 calls.
Want the full picture? How we score →