← leaderboard
Qwen3.8 Max
NewQwen3Ranks #31 of 43 on reading real contracts into structured billing data — -15.5 pts vs GPT-5.5.
53.8%
Accuracy
#31 of 43
26%
Hallucination
HIGH-confidence & wrong
$199
Cost / 1k contracts
$0.20 each · via OpenRouter
503.4s
Median latency
p90 600.4s
±7.3
Run-to-run σ
3 runs
0.27
Value
acc. pts per $/1k
Tokens per contract
Input22,930
Output25,488
Reasoning21,203
Reliability 94.4% · valid structured output across 18 calls.
Want the full picture? How we score →