← leaderboard
GPT-5.6 Luna
NewGPT-5.6Ranks #12 of 43 on reading real contracts into structured billing data — -3.1 pts vs GPT-5.5.
66.2%
Accuracy
#12 of 43
17%
Hallucination
HIGH-confidence & wrong
$4.30
Cost / 1k contracts
$0.00 each · via OpenRouter
26.9s
Median latency
p90 147.5s
±1.8
Run-to-run σ
3 runs
15.28
Value
acc. pts per $/1k
Tokens per contract
Input21,576
Output3,627
Reasoning987
Reliability 100% · valid structured output across 18 calls.
Want the full picture? How we score →