Back to leaderboard
Mistral Medium 3.5
NewMistralRanks #13 of 35 on reading real contracts into structured billing data, -4.9 pts vs GPT-5.5.
71.1%
Accuracy
#13 of 35
10.8%
Hallucination
HIGH-confidence & wrong
$66.1
Cost / 1k contracts
$0.07 each · via OpenRouter
26.6s
Median latency
p90 45.7s
±0
Run-to-run σ
1 runs
1.08
Value
acc. pts per $/1k
Tokens per contract
Input24,977
Output3,816
Reasoning0
Reliability 100% · valid structured output across 30 calls.
Want the full picture? How we score