← leaderboard
Qwen3.7 Plus
Qwen3Ranks #18 of 43 on reading real contracts into structured billing data — -5.5 pts vs GPT-5.5.
63.8%
Accuracy
#18 of 43
22.8%
Hallucination
HIGH-confidence & wrong
$20.3
Cost / 1k contracts
$0.02 each · via OpenRouter
234.9s
Median latency
p90 1274.2s
±10.3
Run-to-run σ
3 runs
3.15
Value
acc. pts per $/1k
Tokens per contract
Input23,383
Output9,986
Reasoning5,694
Reliability 83.3% · valid structured output across 18 calls.
Want the full picture? How we score →