
Can your favorite model read the fine print?
Signed contracts bury their billing terms in pages of prose. FinePrint runs every new model over real ones and checks each field it returns against a human’s answer.
python -m fineprint.eval <model>Run any model you care about. Open source on GitHub.
Top models tested
What brought FinePrint to life
Every invoice begins as a sentence in a contract. Read the fee, the cycle or the currency wrong and the bill goes out wrong. The best model still misses a fifth of them.
How we measure
How a model gets a score.
- Contract in
Real PDF, numbered OCR lines.
01 - Model reads it
Structured billing schema.
02 - Private key check
Each field vs the answer key.
03 - On the leaderboard
Accuracy, cost, and speed.
04
Signed contracts from public filings: order forms, MSAs, renewals. Scanned pages and redlines included. Nothing synthetic.
Every field is hand-labeled. Those labels never ship; only the scores do. That keeps the benchmark honest.
For scoring details, see our methodology.
The results
Every model, ranked.
- 79.4%
- accuracy
- 2.6%
- hallucination
- $674
- cost per 1k
- 61.2s
- p50 latency
| # | Model | Accuracy | Extract | Conv. | Halluc. | $/1k | Value | p50 | Valid |
|---|---|---|---|---|---|---|---|---|---|
| 1 | Claude Fable 5.1NewClaude Fable 5.1 | 79.4% | 75.6% | 84.2% | 2.6% | $674 | 0.12 | 61.2s | 100% |
| 2 | Claude Opus 5NewClaude 5 | 77.8% | 74.9% | 81.8% | 3.6% | $370 | 0.21 | 71.2s | 100% |
| 3 | Claude Opus 4.8Claude 4.8 | 76.9% | 75.3% | 78.9% | 4.6% | $288 | 0.27 | 41.4s | 100% |
| 4 | GPT-5.6 SolNewGPT-5.6 | 76.7% | 76.3% | 77.1% | 4.7% | $265 | 0.29 | 55.9s | 100% |
| 5 | Claude Fable 5NewClaude 5 | 76.3% | 74% | 79.6% | 2.3% | $740 | 0.1 | 71.8s | 100% |
| 6 | Gemini 3.5 Flash LiteNewGemini 3 | 76.1% | 70.9% | 82.7% | 21.8% | $17.2 | 4.43 | 9.1s | 100% |
| 7 | GPT-5.5GPT-5.5 | 76% | 74.3% | 78.1% | 6.2% | $316 | 0.24 | 90s | 100% |
| 8 | Claude Sonnet 5NewClaude 5 | 75.4% | 73% | 78.5% | 6.2% | $200 | 0.38 | 106.7s | 100% |
| 9 | Muse Spark 1.3 ContributorNewMuse Spark 1.3 Contributor | 74.5% | 75.4% | 73.4% | 5.2% | $3.20 | 23 | 29s | 100% |
| 10 | GPT-5.6 TerraNewGPT-5.6 | 72.8% | 76.6% | 68.1% | 2.4% | $44.3 | 1.64 | 33.2s | 100% |
| 11 | Muse Spark 1.3NewMuse Spark 1.3 | 72.8% | 70.5% | 75.7% | 5.2% | $48.9 | 1.49 | 32.6s | 100% |
| 12 | GPT-6 AstraNewGPT-6 Astra | 72% | 72.5% | 71.4% | 3.5% | $647 | 0.11 | 149.6s | 100% |
| 13 | Mistral Medium 3.5NewMistral | 71.1% | 70.4% | 71.9% | 10.8% | $66.1 | 1.08 | 26.6s | 100% |
| 14 | Mistral LargeNewMistral | 71% | 61.8% | 84.4% | 19% | $18.7 | 3.79 | 43.3s | 100% |
| 15 | GPT-5.6 LunaNewGPT-5.6 | 70.9% | 73% | 68.3% | 5.3% | $4.60 | 16 | 30.6s | 100% |
| 16 | GPT-5.4 MiniGPT-5.4 | 70% | 74% | 64.8% | 4.7% | $66.2 | 1.06 | 65.2s | 100% |
Try it
Read a contract with any model.
Try one of the top models live on our samples, or upload your own contract. We run extraction in real time (it can take a minute depending on the model) and show you every field it pulls out.
Read the contract on the left, then Run extraction to see the structured billing schema — every field cited back to the page.
Cost
What accuracy actually costs.
We ran 35 priced models through the same contracts. Reading 1,000 of them costs between $1.90 and $740, depending on which model you pick and how verbose its output is.
Going deeper
Ten other ways to compare them.
Beyond the headline rank, we score speed, value, and hallucinations separately. Here is how the field looks on each axis.
35 models · 12 labs
Accuracy
Fields read correctly. Equivalent answers count as correct. Top 15.
Best value
Accuracy per dollar spent. Cheap and accurate rises. Top 12.
Speed × accuracy
Median latency against accuracy. Up and to the left wins.
Lowest hallucination
Share of HIGH-confidence answers that were wrong. Lower is safer. Top 12.
Economic facts
Accuracy on the money itself: dates, fee amounts, credits and overrides. Top 12.
Document difficulty
Accuracy per anonymized contract. Some documents break everyone.
Accuracy by lab
Average and best accuracy per provider.
Latency tail
p50 (filled) to p90 (hollow). Wider means less predictable.
Price spread
Cost to read 1,000 contracts on a log scale, a ~100× range across the field.
Reliability
Share of calls that returned valid structured output. Only models that dropped any.
A Note From The Team
Flexprice runs usage-based billing. Before any invoice exists, someone has to turn a signed contract into structured billing terms: the fees, the cadence, the currency, the commitments. We do that step in production.
That is why we can score it. We already know which fields decide an invoice, what counts as the same answer when a price is written two different ways, and what it costs when a model gets one wrong and sounds certain about it.
So we wrote the rubric down, labeled a set of real contracts by hand, and ran every model we could reach. We put the whole thing in the open so nobody else picking a model for this has to start from scratch, and so the numbers keep updating as new models ship. If your contracts look nothing like ours, point the harness at your own.
Flexprice team
Fix your billing today with Flexprice
FinePrint finds which models can read the contract.
Flexprice already turns those terms into accurate invoices.