← leaderboard
DeepSeek R1
DeepSeek R1Ranks #27 of 43 on reading real contracts into structured billing data — -10 pts vs GPT-5.5.
59.3%
Accuracy
#27 of 43
24.8%
Hallucination
HIGH-confidence & wrong
$29.3
Cost / 1k contracts
$0.03 each · via OpenRouter
251s
Median latency
p90 590.9s
±6.7
Run-to-run σ
3 runs
2.03
Value
acc. pts per $/1k
Tokens per contract
Input22,105
Output8,480
Reasoning5,619
Reliability 83.3% · valid structured output across 18 calls.
Want the full picture? How we score →