Compare

Put models head to head.

Pick up to four models and read them side by side. The best value in every row is called out. The selection lives in the URL, so any comparison is a link you can share.

2/4 selected
MetricClaude Fable 5.1#1 · Claude Fable 5.1GPT-5.5#7 · GPT-5.5
Accuracyhigher is better79.4%76%
Hallucinationlower is better2.6%6.2%
Cost / 1klower is better$674$316
Valueacc. pts per $/1k0.120.24
p50 latencylower is better61.2s90s
Run-to-run σlower is better±0±0
Reliabilityhigher is better100%100%

Cost × accuracy

selected in focus
39%52%65%78%$2$5$10$20$50$100$200$500Claude Fable 5.15.5cost per 1,000 contracts (log)