← leaderboard
Mistral Medium 3.5
NewMistralRanks #21 of 43 on reading real contracts into structured billing data — -6.4 pts vs GPT-5.5.
62.9%
Accuracy
#21 of 43
24.6%
Hallucination
HIGH-confidence & wrong
$63.4
Cost / 1k contracts
$0.06 each · via OpenRouter
62.4s
Median latency
p90 65.9s
±2.6
Run-to-run σ
3 runs
0.99
Value
acc. pts per $/1k
Tokens per contract
Input24,045
Output3,642
Reasoning0
Reliability 100% · valid structured output across 18 calls.
Want the full picture? How we score →