← leaderboard
DeepSeek V3.2
DeepSeek V3Ranks #17 of 43 on reading real contracts into structured billing data — -5.2 pts vs GPT-5.5.
64.1%
Accuracy
#17 of 43
28.3%
Hallucination
HIGH-confidence & wrong
$7.80
Cost / 1k contracts
$0.01 each · via OpenRouter
51.5s
Median latency
p90 57.4s
±5.2
Run-to-run σ
3 runs
8.26
Value
acc. pts per $/1k
Tokens per contract
Input23,075
Output3,893
Reasoning0
Reliability 100% · valid structured output across 18 calls.
Want the full picture? How we score →