A vivid vermilion temple archive of rolled contracts, a robed scholar reading a scroll of fine print cascading down the steps, in a lush orange garden under a cobalt sky

The document-extraction benchmark · by Flexprice

Can it read the fine print?

Every new model, put through real contracts — the messy PDFs businesses actually run on — and scored on whether it turns them into correct, structured data. A private test set nobody can game.

43
models tested
6
contracts
13
fields / contract
9,593
field judgments

Quality × cost

Accuracy vs. price on real contracts. Up and to the left wins.

new prior
x-axis
↖ better value30%45%60%75%$1$2$5$10$20$50$100$200$500$1000Gemini 3.6 FlashGrok 4.5Gemini 3.5 Flash LiteDeepSeek V4 FlashGPT-5.6 LunaQwen3.7 Flashcost per 1,000 contracts (log) →accuracy →
NewGemini 3.6 FlashGemini 3

Ranks #1 of 43 on contract extraction+9.6 pts vs GPT-5.5.

78.9% accuracy14.5% hallucination$116/1k52.8s p50

The task

Watch a model read a real contract.

Real contract · public SEC exhibitWeb Hosting Agreement
Contract page
Extracted → billing schema8 fields

Every box is the model’s own citation — hover a field to see where it read it.

Leaderboard

Every model, ranked.

#ModelAccuracyHalluc.σ (5 runs)$/1kValuep50OK
1Gemini 3.6 FlashnewGemini 3
78.9%
14.5%±5.6$1160.6852.8s100%
2Gemini 3.5 FlashGemini 3
78.6%
13.8%±2.9$1320.649.9s100%
3Kimi K3newKimi
73.7%
8.7%±3$2780.26180.6s100%
4Claude Fable 5newClaude 5
72.5%
7.8%±4.8$7320.170.2s100%
5Grok 4.5newGrok 4
71.2%
15.4%±2.6$93.30.7699.1s100%
6GPT-5.5GPT-5.5
69.3%
14.8%±2.1$3030.2397.5s100%
7Gemini 3.5 Flash LitenewGemini 3
69.3%
27.3%±4.7$15.74.438.5s100%
8Claude Opus 5newClaude 5
68.4%
10.9%±2.1$3460.267.1s100%
9GPT-5.6 SolnewGPT-5.6
68%
8.7%±1.7$2420.2878.7s100%
10Sonar Reasoning ProSonar
67.8%
14.9%±3.1$78.80.86146.5s100%
11DeepSeek V4 FlashnewDeepSeek V4
67.6%
23%±7.4$5.1013102.9s83.3%
12GPT-5.6 LunanewGPT-5.6
66.2%
17%±1.8$4.301526.9s100%
13Claude Sonnet 5newClaude 5
65.9%
14%±1.6$1850.3692.6s100%
14GPT-5.4 MiniGPT-5.4
65.5%
18.8%±3.6$60.71.0868.1s100%
15Grok 4.6newGrok 4.6
65.3%
11.8%±3$1080.6140.2s100%
16GPT-5.6 TerranewGPT-5.6
64.6%
16.3%±0.2$42.01.5428.8s100%
17DeepSeek V3.2DeepSeek V3
64.1%
28.3%±5.2$7.808.2651.5s100%
18Qwen3.7 PlusQwen3
63.8%
22.8%±10.3$20.33.15234.9s83.3%
19Claude Opus 4.8Claude 4.8
63.2%
24.3%±3.8$2680.2440.6s100%
20Grok 4.20Grok 4
63.1%
27%±5.6$34.01.8630.2s100%
21Mistral Medium 3.5newMistral
62.9%
24.6%±2.6$63.40.9962.4s100%
22Mistral LargenewMistral
61.9%
28.8%±5.2$18.03.4373.9s100%
23Sonar PronewSonar
60.8%
31.8%±5.7$1080.5616.1s100%
24GPT-5.4 NanoGPT-5.4
60.4%
14.7%±8$12.84.7347.7s100%
25Grok 4.3Grok 4
59.8%
24.7%±9.1$33.71.7824.6s100%
26DeepSeek V4 PronewDeepSeek V4
59.5%
23.7%±15.1$45.01.32143.6s94.4%
27DeepSeek R1DeepSeek R1
59.3%
24.8%±6.7$29.32.03251s83.3%
28Llama 4 MavericknewLlama 4
58%
27%±5$5.501139.7s100%
29Ministral 14BMistral
56.5%
9.8%±6.3$5.501060.9s100%
30Command AnewCommand
55.7%
31.6%±6.5$90.30.6252.1s100%
31Qwen3.8 MaxnewQwen3
53.8%
26%±7.3$1990.27503.4s94.4%
32Mistral SmallMistral
52.5%
42.9%±3.5$5.509.6117.1s100%
33Nova 2 LitenewNova
50.6%
36%±9.6$11.24.513.3s100%
34Qwen3.7 FlashnewQwen3
50%
38%±0$1.8028118.5s83.3%
35GLM-5.1GLM
45.9%
28.1%±17$85.00.54374.9s94.4%
36Qwen3.7 MaxQwen3
45.3%
19.6%±13.3$68.20.66200.2s94.4%
37MiniMax M3newMiniMax
44.1%
33.6%±11.8$10.74.1455.5s100%
38Kimi K2.6newKimi
43.3%
30.8%±13.9$90.20.48437.7s94.4%
39Command R+Command
41.3%
32.8%±17.6$87.40.47811.4s100%
40GLM-5.2newGLM
39.6%
48.6%±9.9$29.11.3683.2s83.3%
41Nova PremiernewNova
33.3%
25%±4.1$81.50.4192.8s100%
42Llama 4 ScoutnewLlama 4
32.8%
37.4%±11.2$2.601342.2s100%
43Llama 3.3 70BLlama 3
30.2%
25.4%±16.7$2.801140.1s88.9%

Value = accuracy points per $/1k. Pricing from OpenRouter, updated continuously.

The numbers

Ten ways to read the field.

43 models · 14 labs · the same private contracts. Blue marks this generation's new releases.

Accuracy

Fields read correctly, economic-equivalence aware. Top 15.

Gemini 3.6 Flash
78.9%
Gemini 3.5 Flash
78.6%
Kimi K3
73.7%
Claude Fable 5
72.5%
Grok 4.5
71.2%
GPT-5.5
69.3%
Gemini 3.5 Flash Lite
69.3%
Claude Opus 5
68.4%
GPT-5.6 Sol
68%
Sonar Reasoning Pro
67.8%
DeepSeek V4 Flash
67.6%
GPT-5.6 Luna
66.2%
Claude Sonnet 5
65.9%
GPT-5.4 Mini
65.5%
Grok 4.6
65.3%

Best value

Accuracy points per $/1k. Cheap and good rises. Top 12.

Qwen3.7 Flash
28
GPT-5.6 Luna
15
DeepSeek V4 Flash
13
Llama 4 Scout
13
Llama 3.3 70B
11
Llama 4 Maverick
11
Ministral 14B
10
Mistral Small
9.6
DeepSeek V3.2
8.3
GPT-5.4 Nano
4.7
Nova 2 Lite
4.5
Gemini 3.5 Flash Lite
4.4

Speed × accuracy

Median latency vs accuracy. Up and to the left wins.

27%55%82%10s20s50s100s200s500s← faster · median latency (log)Gemini 3.6 Flash — 78.9% · 52.8sGemini 3.5 Flash — 78.6% · 49.9sKimi K3 — 73.7% · 180.6sClaude Fable 5 — 72.5% · 70.2sGrok 4.5 — 71.2% · 99.1sGPT-5.5 — 69.3% · 97.5sGemini 3.5 Flash Lite — 69.3% · 8.5sClaude Opus 5 — 68.4% · 67.1sGPT-5.6 Sol — 68% · 78.7sSonar Reasoning Pro — 67.8% · 146.5sDeepSeek V4 Flash — 67.6% · 102.9sGPT-5.6 Luna — 66.2% · 26.9sClaude Sonnet 5 — 65.9% · 92.6sGPT-5.4 Mini — 65.5% · 68.1sGrok 4.6 — 65.3% · 140.2sGPT-5.6 Terra — 64.6% · 28.8sDeepSeek V3.2 — 64.1% · 51.5sQwen3.7 Plus — 63.8% · 234.9sClaude Opus 4.8 — 63.2% · 40.6sGrok 4.20 — 63.1% · 30.2sMistral Medium 3.5 — 62.9% · 62.4sMistral Large — 61.9% · 73.9sSonar Pro — 60.8% · 16.1sGPT-5.4 Nano — 60.4% · 47.7sGrok 4.3 — 59.8% · 24.6sDeepSeek V4 Pro — 59.5% · 143.6sDeepSeek R1 — 59.3% · 251sLlama 4 Maverick — 58% · 39.7sMinistral 14B — 56.5% · 60.9sCommand A — 55.7% · 52.1sQwen3.8 Max — 53.8% · 503.4sMistral Small — 52.5% · 17.1sNova 2 Lite — 50.6% · 13.3sQwen3.7 Flash — 50% · 118.5sGLM-5.1 — 45.9% · 374.9sQwen3.7 Max — 45.3% · 200.2sMiniMax M3 — 44.1% · 55.5sKimi K2.6 — 43.3% · 437.7sCommand R+ — 41.3% · 811.4sGLM-5.2 — 39.6% · 83.2sNova Premier — 33.3% · 92.8sLlama 4 Scout — 32.8% · 42.2sLlama 3.3 70B — 30.2% · 40.1s

Lowest hallucination

Share of HIGH-confidence answers that were wrong. Lower is safer. Top 12.

Claude Fable 5
7.8%
Kimi K3
8.7%
GPT-5.6 Sol
8.7%
Ministral 14B
9.8%
Claude Opus 5
10.9%
Grok 4.6
11.8%
Gemini 3.5 Flash
13.8%
Claude Sonnet 5
14%
Gemini 3.6 Flash
14.5%
GPT-5.4 Nano
14.7%
GPT-5.5
14.8%
Sonar Reasoning Pro
14.9%

Most consistent

Run-to-run σ across repeated runs. Lower is more reliable. Top 12.

Qwen3.7 Flash
±0
GPT-5.6 Terra
±0.2
Claude Sonnet 5
±1.6
GPT-5.6 Sol
±1.7
GPT-5.6 Luna
±1.8
GPT-5.5
±2.1
Claude Opus 5
±2.1
Grok 4.5
±2.6
Mistral Medium 3.5
±2.6
Gemini 3.5 Flash
±2.9
Kimi K3
±3
Grok 4.6
±3

Document difficulty

Accuracy per anonymized contract. Some documents break everyone.

A
31%
B
57%
C
44%
D
59%
E
80%
F
79%
Gemini 3.6 Flash
52
71
62
80
100
100
Gemini 3.5 Flash
35
71
73
80
100
100
Kimi K3
64
71
49
62
100
95
Claude Fable 5
43
71
51
64
100
100
Grok 4.5
29
71
56
64
100
100
GPT-5.5
27
71
51
64
98
96
Gemini 3.5 Flash Lite
37
68
57
64
100
89
Claude Opus 5
27
71
42
64
100
100
GPT-5.6 Sol
27
71
47
64
98
95
Sonar Reasoning Pro
27
68
47
64
100
98
DeepSeek V4 Flash
19
62
66
58
93
82
GPT-5.6 Luna
28
67
47
64
100
84
Claude Sonnet 5
27
71
38
64
98
93
GPT-5.4 Mini
25
62
42
64
100
95
Grok 4.6
30
76
56
78
77
79
GPT-5.6 Terra
29
71
33
64
100
86
DeepSeek V3.2
36
71
49
64
84
84
Qwen3.7 Plus
35
69
47
61
87
73
Claude Opus 4.8
27
64
33
60
93
98
Grok 4.20
27
52
46
78
87
88
Mistral Medium 3.5
27
64
36
64
100
91
Mistral Large
34
60
51
64
84
80
Sonar Pro
27
51
41
69
92
76
GPT-5.4 Nano
43
61
40
57
86
74
Grok 4.3
27
70
45
67
83
65
DeepSeek V4 Pro
33
47
37
72
83
81
DeepSeek R1
34
57
42
57
93
93
Llama 4 Maverick
21
45
47
69
88
81
Ministral 14B
41
40
43
64
60
77
Command A
21
53
36
49
86
82
Qwen3.8 Max
19
52
47
54
81
68
Mistral Small
25
53
19
64
74
76
Nova 2 Lite
26
46
24
55
80
70
Qwen3.7 Flash
23
43
60
43
79
GLM-5.1
29
52
67
35
55
47
Qwen3.7 Max
44
24
34
66
39
56
MiniMax M3
23
42
30
54
34
76
Kimi K2.6
49
28
47
38
47
46
Command R+
29
37
42
21
55
60
GLM-5.2
20
48
36
41
45
48
Nova Premier
24
27
29
17
50
50
Llama 4 Scout
17
40
14
26
50
48
Llama 3.3 70B
24
20
39
36
20
38

Accuracy by lab

Average and best accuracy per provider.

google×3
75.6% avg · 79 best
anthropic×4
67.5% avg · 73 best
openai×6
65.7% avg · 69 best
xai×4
64.8% avg · 71 best
perplexity×2
64.3% avg · 68 best
deepseek×4
62.6% avg · 68 best
moonshot×2
58.5% avg · 74 best
mistral×4
58.5% avg · 63 best
qwen×4
53.2% avg · 64 best
cohere×2
48.5% avg · 56 best
minimax×1
44.1% avg · 44 best
zhipu×2
42.8% avg · 46 best
amazon×2
42.0% avg · 51 best
meta×3
40.3% avg · 58 best

Latency tail

p50 (filled) to p90 (hollow). Wider = less predictable. Fastest 12.

Gemini 3.5 Flash Lite
p50 8.5s9s
Nova 2 Lite
p50 13.3s17.1s
Sonar Pro
p50 16.1s27.4s
Mistral Small
p50 17.1s18.5s
Grok 4.3
p50 24.6s34.8s
GPT-5.6 Luna
p50 26.9s147.5s
GPT-5.6 Terra
p50 28.8s37.9s
Grok 4.20
p50 30.2s40.9s
Llama 4 Maverick
p50 39.7s1205.8s
Llama 3.3 70B
p50 40.1s66.7s
Claude Opus 4.8
p50 40.6s43.8s
Llama 4 Scout
p50 42.2s61.5s

Price spread

Cost to read 1,000 contracts, log scale — a ~100× range across the field.

$5$10$50$100$300Gemini 3.6 Flash — $116/1kGemini 3.5 Flash — $132/1kKimi K3 — $278/1kClaude Fable 5 — $732/1kGrok 4.5 — $93.3/1kGPT-5.5 — $303/1kGemini 3.5 Flash Lite — $15.7/1kClaude Opus 5 — $346/1kGPT-5.6 Sol — $242/1kSonar Reasoning Pro — $78.8/1kDeepSeek V4 Flash — $5.10/1kGPT-5.6 Luna — $4.30/1kClaude Sonnet 5 — $185/1kGPT-5.4 Mini — $60.7/1kGrok 4.6 — $108/1kGPT-5.6 Terra — $42.0/1kDeepSeek V3.2 — $7.80/1kQwen3.7 Plus — $20.3/1kClaude Opus 4.8 — $268/1kGrok 4.20 — $34.0/1kMistral Medium 3.5 — $63.4/1kMistral Large — $18.0/1kSonar Pro — $108/1kGPT-5.4 Nano — $12.8/1kGrok 4.3 — $33.7/1kDeepSeek V4 Pro — $45.0/1kDeepSeek R1 — $29.3/1kLlama 4 Maverick — $5.50/1kMinistral 14B — $5.50/1kCommand A — $90.3/1kQwen3.8 Max — $199/1kMistral Small — $5.50/1kNova 2 Lite — $11.2/1kQwen3.7 Flash — $1.80/1kGLM-5.1 — $85.0/1kQwen3.7 Max — $68.2/1kMiniMax M3 — $10.7/1kKimi K2.6 — $90.2/1kCommand R+ — $87.4/1kGLM-5.2 — $29.1/1kNova Premier — $81.5/1kLlama 4 Scout — $2.60/1kLlama 3.3 70B — $2.80/1kcost per 1,000 contracts (log) →

Reliability

Share of calls that returned valid structured output. Only models that dropped any.

DeepSeek V4 Flash
83.3%
Qwen3.7 Plus
83.3%
DeepSeek R1
83.3%
Qwen3.7 Flash
83.3%
GLM-5.2
83.3%
Llama 3.3 70B
88.9%
DeepSeek V4 Pro
94.4%
Qwen3.8 Max
94.4%
GLM-5.1
94.4%
Qwen3.7 Max
94.4%

What’s inside

A private test set of real contracts.

Most benchmarks test trivia. FinePrint tests whether a model can take a real, messy contract and return correct, structured billing data — the task that actually breaks in production.

Real documents

Public, license-clear contracts from the web — the messy PDFs businesses actually run on. Never synthetic.

Private test set

Ground-truth labels stay internal, so no model can train on the answers. We publish the volume, never the data.

Field-level scoring

Every fee, cadence, currency, entitlement and party is checked. Economic-equivalence aware: $10k/qtr = $40k/yr.

Repeated runs

Every model runs each contract 3×. We report mean accuracy and run-to-run σ — one shot hides nondeterminism.

6
contracts in this seed
13
labeled fields / contract
9,593
field judgments scored
3
runs per contract
Order formsMaster agreementsRenewals & amendmentsScanned / OCR-noisyRedlined draftsMulti-currency

Seed benchmark shown on a labeled subset · scaling to ~200 web-sourced contracts across 6 industries and 4 currencies.

Rigorous, transparent, un-gameable.

Read exactly how we score — the private holdout, the field-level rubric, and why we repeat every run.

How we score