← leaderboard
Grok 4.5
NewGrok 4Ranks #5 of 43 on reading real contracts into structured billing data — +1.9 pts vs GPT-5.5.
71.2%
Accuracy
#5 of 43
15.4%
Hallucination
HIGH-confidence & wrong
$93.3
Cost / 1k contracts
$0.09 each · via OpenRouter
99.1s
Median latency
p90 167s
±2.6
Run-to-run σ
3 runs
0.76
Value
acc. pts per $/1k
Tokens per contract
Input22,786
Output7,951
Reasoning3,709
Reliability 100% · valid structured output across 18 calls.
Want the full picture? How we score →