Back to leaderboard

Muse Spark 1.3 Contributor

NewMuse Spark 1.3 Contributor

Ranks #9 of 35 on reading real contracts into structured billing data, -1.5 pts vs GPT-5.5.

74.5%
Accuracy
#9 of 35
5.2%
Hallucination
HIGH-confidence & wrong
$3.20
Cost / 1k contracts
$0.00 each · via OpenRouter
29s
Median latency
p90 45.7s
±0
Run-to-run σ
1 runs
23
Value
acc. pts per $/1k
Tokens per contract
Input21,982
Output5,137
Reasoning2,240

Reliability 100% · valid structured output across 30 calls.

Want the full picture? How we score