← leaderboard

DeepSeek V3.2

DeepSeek V3

Ranks #17 of 43 on reading real contracts into structured billing data-5.2 pts vs GPT-5.5.

64.1%
Accuracy
#17 of 43
28.3%
Hallucination
HIGH-confidence & wrong
$7.80
Cost / 1k contracts
$0.01 each · via OpenRouter
51.5s
Median latency
p90 57.4s
±5.2
Run-to-run σ
3 runs
8.26
Value
acc. pts per $/1k
Tokens per contract
Input23,075
Output3,893
Reasoning0

Reliability 100% · valid structured output across 18 calls.

Want the full picture? How we score →