← leaderboard
Qwen3.7 Max
Qwen3Ranks #36 of 43 on reading real contracts into structured billing data — -24 pts vs GPT-5.5.
45.3%
Accuracy
#36 of 43
19.6%
Hallucination
HIGH-confidence & wrong
$68.2
Cost / 1k contracts
$0.07 each · via OpenRouter
200.2s
Median latency
p90 349.9s
±13.3
Run-to-run σ
3 runs
0.66
Value
acc. pts per $/1k
Tokens per contract
Input22,889
Output7,785
Reasoning5,175
Reliability 94.4% · valid structured output across 18 calls.
Want the full picture? How we score →