← leaderboard
Grok 4.20
Grok 4Ranks #20 of 43 on reading real contracts into structured billing data — -6.2 pts vs GPT-5.5.
63.1%
Accuracy
#20 of 43
27%
Hallucination
HIGH-confidence & wrong
$34.0
Cost / 1k contracts
$0.03 each · via OpenRouter
30.2s
Median latency
p90 40.9s
±5.6
Run-to-run σ
3 runs
1.86
Value
acc. pts per $/1k
Tokens per contract
Input20,776
Output3,211
Reasoning0
Reliability 100% · valid structured output across 18 calls.
Want the full picture? How we score →