← leaderboard
Mistral Large
NewMistralRanks #22 of 43 on reading real contracts into structured billing data — -7.4 pts vs GPT-5.5.
61.9%
Accuracy
#22 of 43
28.8%
Hallucination
HIGH-confidence & wrong
$18.0
Cost / 1k contracts
$0.02 each · via OpenRouter
73.9s
Median latency
p90 98.8s
±5.2
Run-to-run σ
3 runs
3.43
Value
acc. pts per $/1k
Tokens per contract
Input24,045
Output4,000
Reasoning0
Reliability 100% · valid structured output across 18 calls.
Want the full picture? How we score →