← leaderboard
Grok 4.3
Grok 4Ranks #25 of 43 on reading real contracts into structured billing data — -9.5 pts vs GPT-5.5.
59.8%
Accuracy
#25 of 43
24.7%
Hallucination
HIGH-confidence & wrong
$33.7
Cost / 1k contracts
$0.03 each · via OpenRouter
24.6s
Median latency
p90 34.8s
±9.1
Run-to-run σ
3 runs
1.78
Value
acc. pts per $/1k
Tokens per contract
Input20,784
Output3,083
Reasoning982
Reliability 100% · valid structured output across 18 calls.
Want the full picture? How we score →