
The document-extraction benchmark · by Flexprice
Can it read the fine print?
Every new model, put through real contracts — the messy PDFs businesses actually run on — and scored on whether it turns them into correct, structured data. A private test set nobody can game.
- 43
- models tested
- 6
- contracts
- 13
- fields / contract
- 9,593
- field judgments
Quality × cost
Accuracy vs. price on real contracts. Up and to the left wins.
Ranks #1 of 43 on contract extraction — +9.6 pts vs GPT-5.5.
The task
Watch a model read a real contract.

Every box is the model’s own citation — hover a field to see where it read it.
Leaderboard
Every model, ranked.
| # | Model | Accuracy ↓ | Halluc. | σ (5 runs) | $/1k | Value | p50 | OK |
|---|---|---|---|---|---|---|---|---|
| 1 | Gemini 3.6 FlashnewGemini 3 | 78.9% | 14.5% | ±5.6 | $116 | 0.68 | 52.8s | 100% |
| 2 | Gemini 3.5 FlashGemini 3 | 78.6% | 13.8% | ±2.9 | $132 | 0.6 | 49.9s | 100% |
| 3 | Kimi K3newKimi | 73.7% | 8.7% | ±3 | $278 | 0.26 | 180.6s | 100% |
| 4 | Claude Fable 5newClaude 5 | 72.5% | 7.8% | ±4.8 | $732 | 0.1 | 70.2s | 100% |
| 5 | Grok 4.5newGrok 4 | 71.2% | 15.4% | ±2.6 | $93.3 | 0.76 | 99.1s | 100% |
| 6 | GPT-5.5GPT-5.5 | 69.3% | 14.8% | ±2.1 | $303 | 0.23 | 97.5s | 100% |
| 7 | Gemini 3.5 Flash LitenewGemini 3 | 69.3% | 27.3% | ±4.7 | $15.7 | 4.43 | 8.5s | 100% |
| 8 | Claude Opus 5newClaude 5 | 68.4% | 10.9% | ±2.1 | $346 | 0.2 | 67.1s | 100% |
| 9 | GPT-5.6 SolnewGPT-5.6 | 68% | 8.7% | ±1.7 | $242 | 0.28 | 78.7s | 100% |
| 10 | Sonar Reasoning ProSonar | 67.8% | 14.9% | ±3.1 | $78.8 | 0.86 | 146.5s | 100% |
| 11 | DeepSeek V4 FlashnewDeepSeek V4 | 67.6% | 23% | ±7.4 | $5.10 | 13 | 102.9s | 83.3% |
| 12 | GPT-5.6 LunanewGPT-5.6 | 66.2% | 17% | ±1.8 | $4.30 | 15 | 26.9s | 100% |
| 13 | Claude Sonnet 5newClaude 5 | 65.9% | 14% | ±1.6 | $185 | 0.36 | 92.6s | 100% |
| 14 | GPT-5.4 MiniGPT-5.4 | 65.5% | 18.8% | ±3.6 | $60.7 | 1.08 | 68.1s | 100% |
| 15 | Grok 4.6newGrok 4.6 | 65.3% | 11.8% | ±3 | $108 | 0.6 | 140.2s | 100% |
| 16 | GPT-5.6 TerranewGPT-5.6 | 64.6% | 16.3% | ±0.2 | $42.0 | 1.54 | 28.8s | 100% |
| 17 | DeepSeek V3.2DeepSeek V3 | 64.1% | 28.3% | ±5.2 | $7.80 | 8.26 | 51.5s | 100% |
| 18 | Qwen3.7 PlusQwen3 | 63.8% | 22.8% | ±10.3 | $20.3 | 3.15 | 234.9s | 83.3% |
| 19 | Claude Opus 4.8Claude 4.8 | 63.2% | 24.3% | ±3.8 | $268 | 0.24 | 40.6s | 100% |
| 20 | Grok 4.20Grok 4 | 63.1% | 27% | ±5.6 | $34.0 | 1.86 | 30.2s | 100% |
| 21 | Mistral Medium 3.5newMistral | 62.9% | 24.6% | ±2.6 | $63.4 | 0.99 | 62.4s | 100% |
| 22 | Mistral LargenewMistral | 61.9% | 28.8% | ±5.2 | $18.0 | 3.43 | 73.9s | 100% |
| 23 | Sonar PronewSonar | 60.8% | 31.8% | ±5.7 | $108 | 0.56 | 16.1s | 100% |
| 24 | GPT-5.4 NanoGPT-5.4 | 60.4% | 14.7% | ±8 | $12.8 | 4.73 | 47.7s | 100% |
| 25 | Grok 4.3Grok 4 | 59.8% | 24.7% | ±9.1 | $33.7 | 1.78 | 24.6s | 100% |
| 26 | DeepSeek V4 PronewDeepSeek V4 | 59.5% | 23.7% | ±15.1 | $45.0 | 1.32 | 143.6s | 94.4% |
| 27 | DeepSeek R1DeepSeek R1 | 59.3% | 24.8% | ±6.7 | $29.3 | 2.03 | 251s | 83.3% |
| 28 | Llama 4 MavericknewLlama 4 | 58% | 27% | ±5 | $5.50 | 11 | 39.7s | 100% |
| 29 | Ministral 14BMistral | 56.5% | 9.8% | ±6.3 | $5.50 | 10 | 60.9s | 100% |
| 30 | Command AnewCommand | 55.7% | 31.6% | ±6.5 | $90.3 | 0.62 | 52.1s | 100% |
| 31 | Qwen3.8 MaxnewQwen3 | 53.8% | 26% | ±7.3 | $199 | 0.27 | 503.4s | 94.4% |
| 32 | Mistral SmallMistral | 52.5% | 42.9% | ±3.5 | $5.50 | 9.61 | 17.1s | 100% |
| 33 | Nova 2 LitenewNova | 50.6% | 36% | ±9.6 | $11.2 | 4.5 | 13.3s | 100% |
| 34 | Qwen3.7 FlashnewQwen3 | 50% | 38% | ±0 | $1.80 | 28 | 118.5s | 83.3% |
| 35 | GLM-5.1GLM | 45.9% | 28.1% | ±17 | $85.0 | 0.54 | 374.9s | 94.4% |
| 36 | Qwen3.7 MaxQwen3 | 45.3% | 19.6% | ±13.3 | $68.2 | 0.66 | 200.2s | 94.4% |
| 37 | MiniMax M3newMiniMax | 44.1% | 33.6% | ±11.8 | $10.7 | 4.14 | 55.5s | 100% |
| 38 | Kimi K2.6newKimi | 43.3% | 30.8% | ±13.9 | $90.2 | 0.48 | 437.7s | 94.4% |
| 39 | Command R+Command | 41.3% | 32.8% | ±17.6 | $87.4 | 0.47 | 811.4s | 100% |
| 40 | GLM-5.2newGLM | 39.6% | 48.6% | ±9.9 | $29.1 | 1.36 | 83.2s | 83.3% |
| 41 | Nova PremiernewNova | 33.3% | 25% | ±4.1 | $81.5 | 0.41 | 92.8s | 100% |
| 42 | Llama 4 ScoutnewLlama 4 | 32.8% | 37.4% | ±11.2 | $2.60 | 13 | 42.2s | 100% |
| 43 | Llama 3.3 70BLlama 3 | 30.2% | 25.4% | ±16.7 | $2.80 | 11 | 40.1s | 88.9% |
Value = accuracy points per $/1k. Pricing from OpenRouter, updated continuously.
The numbers
Ten ways to read the field.
43 models · 14 labs · the same private contracts. Blue marks this generation's new releases.
Accuracy
Fields read correctly, economic-equivalence aware. Top 15.
Best value
Accuracy points per $/1k. Cheap and good rises. Top 12.
Speed × accuracy
Median latency vs accuracy. Up and to the left wins.
Lowest hallucination
Share of HIGH-confidence answers that were wrong. Lower is safer. Top 12.
Most consistent
Run-to-run σ across repeated runs. Lower is more reliable. Top 12.
Document difficulty
Accuracy per anonymized contract. Some documents break everyone.
Accuracy by lab
Average and best accuracy per provider.
Latency tail
p50 (filled) to p90 (hollow). Wider = less predictable. Fastest 12.
Price spread
Cost to read 1,000 contracts, log scale — a ~100× range across the field.
Reliability
Share of calls that returned valid structured output. Only models that dropped any.
What’s inside
A private test set of real contracts.
Most benchmarks test trivia. FinePrint tests whether a model can take a real, messy contract and return correct, structured billing data — the task that actually breaks in production.
Public, license-clear contracts from the web — the messy PDFs businesses actually run on. Never synthetic.
Ground-truth labels stay internal, so no model can train on the answers. We publish the volume, never the data.
Every fee, cadence, currency, entitlement and party is checked. Economic-equivalence aware: $10k/qtr = $40k/yr.
Every model runs each contract 3×. We report mean accuracy and run-to-run σ — one shot hides nondeterminism.
Seed benchmark shown on a labeled subset · scaling to ~200 web-sourced contracts across 6 industries and 4 currencies.
Rigorous, transparent, un-gameable.
Read exactly how we score — the private holdout, the field-level rubric, and why we repeat every run.