Compare

Put models head to head.

Pick up to four models and read them side by side — the best value in every row is called out. The selection lives in the URL, so any comparison is a link you can share.

2/4 selected
MetricGemini 3.6 Flash#1 · Gemini 3GPT-5.5#6 · GPT-5.5
Accuracyhigher is better78.9%69.3%
Hallucinationlower is better14.5%14.8%
Cost / 1klower is better$116$303
Valueacc. pts per $/1k0.680.23
p50 latencylower is better52.8s97.5s
Run-to-run σlower is better±5.6±2.1
Reliabilityhigher is better100%100%

Cost × accuracy

selected in focus
28%42%56%70%$1$2$5$10$20$50$100$200$500Gemini 3.6 Flash5.5cost per 1,000 contracts (log) →