← leaderboard
Command A
NewCommandRanks #30 of 43 on reading real contracts into structured billing data — -13.6 pts vs GPT-5.5.
55.7%
Accuracy
#30 of 43
31.6%
Hallucination
HIGH-confidence & wrong
$90.3
Cost / 1k contracts
$0.09 each · via OpenRouter
52.1s
Median latency
p90 61s
±6.5
Run-to-run σ
3 runs
0.62
Value
acc. pts per $/1k
Tokens per contract
Input24,315
Output2,949
Reasoning0
Reliability 100% · valid structured output across 18 calls.
Want the full picture? How we score →