Back to leaderboard
Muse Spark 1.3
NewMuse Spark 1.3Ranks #11 of 35 on reading real contracts into structured billing data, -3.2 pts vs GPT-5.5.
72.8%
Accuracy
#11 of 35
5.2%
Hallucination
HIGH-confidence & wrong
$48.9
Cost / 1k contracts
$0.05 each · via OpenRouter
32.6s
Median latency
p90 44.1s
±0
Run-to-run σ
1 runs
1.49
Value
acc. pts per $/1k
Tokens per contract
Input21,982
Output5,045
Reasoning2,106
Reliability 100% · valid structured output across 30 calls.
Want the full picture? How we score