← leaderboard

Llama 4 Maverick

NewLlama 4

Ranks #28 of 43 on reading real contracts into structured billing data-11.3 pts vs GPT-5.5.

58%
Accuracy
#28 of 43
27%
Hallucination
HIGH-confidence & wrong
$5.50
Cost / 1k contracts
$0.01 each · via OpenRouter
39.7s
Median latency
p90 1205.8s
±5
Run-to-run σ
3 runs
10.53
Value
acc. pts per $/1k
Tokens per contract
Input20,678
Output1,969
Reasoning0

Reliability 100% · valid structured output across 18 calls.

Want the full picture? How we score →