← leaderboard

Llama 3.3 70B

Llama 3

Ranks #43 of 43 on reading real contracts into structured billing data-39.1 pts vs GPT-5.5.

30.2%
Accuracy
#43 of 43
25.4%
Hallucination
HIGH-confidence & wrong
$2.80
Cost / 1k contracts
$0.00 each · via OpenRouter
40.1s
Median latency
p90 66.7s
±16.7
Run-to-run σ
3 runs
10.84
Value
acc. pts per $/1k
Tokens per contract
Input21,235
Output2,073
Reasoning0

Reliability 88.9% · valid structured output across 18 calls.

Want the full picture? How we score →