← leaderboard
DeepSeek V4 Pro
NewDeepSeek V4Ranks #26 of 43 on reading real contracts into structured billing data — -9.8 pts vs GPT-5.5.
59.5%
Accuracy
#26 of 43
23.7%
Hallucination
HIGH-confidence & wrong
$45.0
Cost / 1k contracts
$0.04 each · via OpenRouter
143.6s
Median latency
p90 295.4s
±15.1
Run-to-run σ
3 runs
1.32
Value
acc. pts per $/1k
Tokens per contract
Input22,631
Output7,933
Reasoning4,660
Reliability 94.4% · valid structured output across 18 calls.
Want the full picture? How we score →