← leaderboard

DeepSeek V4 Flash

NewDeepSeek V4

Ranks #11 of 43 on reading real contracts into structured billing data-1.7 pts vs GPT-5.5.

67.6%
Accuracy
#11 of 43
23%
Hallucination
HIGH-confidence & wrong
$5.10
Cost / 1k contracts
$0.01 each · via OpenRouter
102.9s
Median latency
p90 144.6s
±7.4
Run-to-run σ
3 runs
13.15
Value
acc. pts per $/1k
Tokens per contract
Input22,882
Output6,917
Reasoning3,890

Reliability 83.3% · valid structured output across 18 calls.

Want the full picture? How we score →