← leaderboard
Qwen3.7 Flash
NewQwen3Ranks #34 of 43 on reading real contracts into structured billing data — -19.3 pts vs GPT-5.5.
50%
Accuracy
#34 of 43
38%
Hallucination
HIGH-confidence & wrong
$1.80
Cost / 1k contracts
$0.00 each · via OpenRouter
118.5s
Median latency
p90 130.3s
±0
Run-to-run σ
3 runs
27.82
Value
acc. pts per $/1k
Tokens per contract
Input21,404
Output8,885
Reasoning6,313
Reliability 83.3% · valid structured output across 6 calls.
Want the full picture? How we score →