← leaderboard
MiniMax M3
NewMiniMaxRanks #37 of 43 on reading real contracts into structured billing data — -25.2 pts vs GPT-5.5.
44.1%
Accuracy
#37 of 43
33.6%
Hallucination
HIGH-confidence & wrong
$10.7
Cost / 1k contracts
$0.01 each · via OpenRouter
55.5s
Median latency
p90 390.1s
±11.8
Run-to-run σ
3 runs
4.14
Value
acc. pts per $/1k
Tokens per contract
Input21,125
Output3,594
Reasoning0
Reliability 100% · valid structured output across 18 calls.
Want the full picture? How we score →