← leaderboard
GLM-5.1
GLMRanks #35 of 43 on reading real contracts into structured billing data — -23.4 pts vs GPT-5.5.
45.9%
Accuracy
#35 of 43
28.1%
Hallucination
HIGH-confidence & wrong
$85.0
Cost / 1k contracts
$0.09 each · via OpenRouter
374.9s
Median latency
p90 1959.3s
±17
Run-to-run σ
3 runs
0.54
Value
acc. pts per $/1k
Tokens per contract
Input21,279
Output12,547
Reasoning9,918
Reliability 94.4% · valid structured output across 18 calls.
Want the full picture? How we score →