← leaderboard
Llama 4 Scout
NewLlama 4Ranks #42 of 43 on reading real contracts into structured billing data — -36.5 pts vs GPT-5.5.
32.8%
Accuracy
#42 of 43
37.4%
Hallucination
HIGH-confidence & wrong
$2.60
Cost / 1k contracts
$0.00 each · via OpenRouter
42.2s
Median latency
p90 61.5s
±11.2
Run-to-run σ
3 runs
12.67
Value
acc. pts per $/1k
Tokens per contract
Input20,659
Output1,746
Reasoning0
Reliability 100% · valid structured output across 18 calls.
Want the full picture? How we score →