Back to leaderboard

Mistral Large

NewMistral

Ranks #14 of 35 on reading real contracts into structured billing data, -5 pts vs GPT-5.5.

71%
Accuracy
#14 of 35
19%
Hallucination
HIGH-confidence & wrong
$18.7
Cost / 1k contracts
$0.02 each · via OpenRouter
43.3s
Median latency
p90 49.1s
±0
Run-to-run σ
1 runs
3.79
Value
acc. pts per $/1k
Tokens per contract
Input24,977
Output4,171
Reasoning0

Reliability 100% · valid structured output across 30 calls.

Want the full picture? How we score