← leaderboard

Kimi K2.6

NewKimi

Ranks #38 of 43 on reading real contracts into structured billing data-26 pts vs GPT-5.5.

43.3%
Accuracy
#38 of 43
30.8%
Hallucination
HIGH-confidence & wrong
$90.2
Cost / 1k contracts
$0.09 each · via OpenRouter
437.7s
Median latency
p90 1362.9s
±13.9
Run-to-run σ
3 runs
0.48
Value
acc. pts per $/1k
Tokens per contract
Input21,028
Output17,560
Reasoning15,066

Reliability 94.4% · valid structured output across 18 calls.

Want the full picture? How we score →