← leaderboard

Grok 4.6

NewGrok 4.6

Ranks #15 of 43 on reading real contracts into structured billing data-4 pts vs GPT-5.5.

65.3%
Accuracy
#15 of 43
11.8%
Hallucination
HIGH-confidence & wrong
$108
Cost / 1k contracts
$0.11 each · via OpenRouter
140.2s
Median latency
p90 194.6s
±3
Run-to-run σ
3 runs
0.6
Value
acc. pts per $/1k
Tokens per contract
Input21,246
Output10,957
Reasoning7,151

Reliability 100% · valid structured output across 18 calls.

Want the full picture? How we score →