← leaderboard

DeepSeek R1

DeepSeek R1

Ranks #27 of 43 on reading real contracts into structured billing data-10 pts vs GPT-5.5.

59.3%
Accuracy
#27 of 43
24.8%
Hallucination
HIGH-confidence & wrong
$29.3
Cost / 1k contracts
$0.03 each · via OpenRouter
251s
Median latency
p90 590.9s
±6.7
Run-to-run σ
3 runs
2.03
Value
acc. pts per $/1k
Tokens per contract
Input22,105
Output8,480
Reasoning5,619

Reliability 83.3% · valid structured output across 18 calls.

Want the full picture? How we score →