A vivid vermilion temple archive of rolled contracts, a robed scholar reading a scroll of fine print cascading down the steps, in a lush orange garden under a cobalt sky

Can your favorite model read the fine print?

Signed contracts bury their billing terms in pages of prose. FinePrint runs every new model over real ones and checks each field it returns against a human’s answer.

python -m fineprint.eval <model>

Run any model you care about. Open source on GitHub.

Top models tested

OpenAIGoogleAnthropicxAIMetaMistralDeepSeekQwenKimiZhipuAmazonCoherePerplexityMiniMaxNVIDIAOpenAIGoogleAnthropicxAIMetaMistralDeepSeekQwenKimiZhipuAmazonCoherePerplexityMiniMaxNVIDIA

What brought FinePrint to life

Launch film

Every invoice begins as a sentence in a contract. Read the fee, the cycle or the currency wrong and the bill goes out wrong. The best model still misses a fifth of them.

How we measure

How a model gets a score.

  1. Contract in

    Real PDF, numbered OCR lines.

    01
  2. Model reads it

    Structured billing schema.

    02
  3. Private key check

    Each field vs the answer key.

    03
  4. On the leaderboard

    Accuracy, cost, and speed.

    04
Real documents

Signed contracts from public filings: order forms, MSAs, renewals. Scanned pages and redlines included. Nothing synthetic.

Answers stay private

Every field is hand-labeled. Those labels never ship; only the scores do. That keeps the benchmark honest.

For scoring details, see our methodology.

The results

Every model, ranked.

Latest testedClaude Fable 5.1Claude Fable 5.1

Ranks #1 of 35.

79.4%
accuracy
2.6%
hallucination
$674
cost per 1k
61.2s
p50 latency
#ModelAccuracyExtractConv.Halluc.$/1kValuep50Valid
1Claude Fable 5.1NewClaude Fable 5.1
79.4%
75.6%84.2%2.6%$6740.1261.2s100%
2Claude Opus 5NewClaude 5
77.8%
74.9%81.8%3.6%$3700.2171.2s100%
3Claude Opus 4.8Claude 4.8
76.9%
75.3%78.9%4.6%$2880.2741.4s100%
4GPT-5.6 SolNewGPT-5.6
76.7%
76.3%77.1%4.7%$2650.2955.9s100%
5Claude Fable 5NewClaude 5
76.3%
74%79.6%2.3%$7400.171.8s100%
6Gemini 3.5 Flash LiteNewGemini 3
76.1%
70.9%82.7%21.8%$17.24.439.1s100%
7GPT-5.5GPT-5.5
76%
74.3%78.1%6.2%$3160.2490s100%
8Claude Sonnet 5NewClaude 5
75.4%
73%78.5%6.2%$2000.38106.7s100%
9Muse Spark 1.3 ContributorNewMuse Spark 1.3 Contributor
74.5%
75.4%73.4%5.2%$3.202329s100%
10GPT-5.6 TerraNewGPT-5.6
72.8%
76.6%68.1%2.4%$44.31.6433.2s100%
11Muse Spark 1.3NewMuse Spark 1.3
72.8%
70.5%75.7%5.2%$48.91.4932.6s100%
12GPT-6 AstraNewGPT-6 Astra
72%
72.5%71.4%3.5%$6470.11149.6s100%
13Mistral Medium 3.5NewMistral
71.1%
70.4%71.9%10.8%$66.11.0826.6s100%
14Mistral LargeNewMistral
71%
61.8%84.4%19%$18.73.7943.3s100%
15GPT-5.6 LunaNewGPT-5.6
70.9%
73%68.3%5.3%$4.601630.6s100%
16GPT-5.4 MiniGPT-5.4
70%
74%64.8%4.7%$66.21.0665.2s100%

Value = accuracy points per $/1k. Pricing from OpenRouter, updated continuously.

Try it

Read a contract with any model.

Try one of the top models live on our samples, or upload your own contract. We run extraction in real time (it can take a minute depending on the model) and show you every field it pulls out.

Master services · Jazz Pharmaceuticals · SEC EX-10.2
Extracted schema
⚙

Read the contract on the left, then Run extraction to see the structured billing schema — every field cited back to the page.

Cost

What accuracy actually costs.

We ran 35 priced models through the same contracts. Reading 1,000 of them costs between $1.90 and $740, depending on which model you pick and how verbose its output is.

new prior
x-axis
↖ better value45%60%75%$2$5$10$20$50$100$200$500$1000Claude Fable 5.1Claude Opus 5Claude Opus 4.8GPT-5.6 SolGemini 3.5 Flash LiteMuse Spark 1.3 ContributorQwen3.7 Flashcost per 1,000 contracts (log)accuracy

Going deeper

Ten other ways to compare them.

Beyond the headline rank, we score speed, value, and hallucinations separately. Here is how the field looks on each axis.

35 models · 12 labs

Accuracy

Fields read correctly. Equivalent answers count as correct. Top 15.

Claude Fable 5.1
79.4%
Claude Opus 5
77.8%
Claude Opus 4.8
76.9%
GPT-5.6 Sol
76.7%
Claude Fable 5
76.3%
Gemini 3.5 Flash Lite
76.1%
GPT-5.5
76%
Claude Sonnet 5
75.4%
Muse Spark 1.3 Contributor
74.5%
GPT-5.6 Terra
72.8%
Muse Spark 1.3
72.8%
GPT-6 Astra
72%
Mistral Medium 3.5
71.1%
Mistral Large
71%
GPT-5.6 Luna
70.9%

Best value

Accuracy per dollar spent. Cheap and accurate rises. Top 12.

Muse Spark 1.3 Contributor
23
Qwen3.7 Flash
19
GPT-5.6 Luna
16
Llama 4 Scout
13
Llama 3.3 70B
12
Mistral Small
11
DeepSeek V4 Flash
10
Ministral 14B
9.01
Nemotron 3 Super
8
Llama 4 Maverick
7.94
DeepSeek V3.2
7.89
GPT-5.4 Nano
5.28

Speed × accuracy

Median latency against accuracy. Up and to the left wins.

31%57%82%10s20s50s100s200s500slower is faster · median latency (log)

Lowest hallucination

Share of HIGH-confidence answers that were wrong. Lower is safer. Top 12.

Claude Fable 5
2.3%
GPT-5.6 Terra
2.4%
Claude Fable 5.1
2.6%
GPT-6 Astra
3.5%
Claude Opus 5
3.6%
GPT-5.4 Nano
3.9%
Claude Opus 4.8
4.6%
GPT-5.6 Sol
4.7%
GPT-5.4 Mini
4.7%
Muse Spark 1.3 Contributor
5.2%
Muse Spark 1.3
5.2%
GPT-5.6 Luna
5.3%

Economic facts

Accuracy on the money itself: dates, fee amounts, credits and overrides. Top 12.

GPT-5.6 Terra
76.6%
GPT-5.6 Sol
76.3%
Claude Fable 5.1
75.6%
Muse Spark 1.3 Contributor
75.4%
Claude Opus 4.8
75.3%
Claude Opus 5
74.9%
GPT-5.5
74.3%
Claude Fable 5
74%
GPT-5.4 Mini
74%
Claude Sonnet 5
73%
GPT-5.6 Luna
73%
GPT-6 Astra
72.5%

Document difficulty

Accuracy per anonymized contract. Some documents break everyone.

01
71%
02
70%
03
59%
04
75%
05
43%
06
71%
07
55%
08
72%
09
71%
10
64%
11
52%
12
67%
13
52%
14
61%
15
67%
16
57%
17
66%
18
76%
19
70%
20
59%
21
47%
22
68%
23
68%
24
68%
25
54%
26
79%
27
57%
28
53%
29
50%
30
42%
Claude Fable 5.1
83
92
78
92
46
83
78
92
92
85
55
100
92
54
82
62
92
83
83
71
64
91
85
92
60
100
64
71
58
82
Claude Opus 5
100
83
62
85
75
83
88
83
83
77
55
85
92
88
82
62
83
83
91
79
85
82
85
85
60
91
64
55
58
40
Claude Opus 4.8
91
75
78
85
46
83
88
83
83
85
55
85
92
88
82
62
83
83
91
64
85
80
85
85
83
91
71
44
58
40
GPT-5.6 Sol
83
83
78
92
75
83
63
92
83
77
83
85
36
88
90
78
92
83
75
79
46
80
75
92
60
91
63
71
58
40
Claude Fable 5
91
83
78
85
75
83
88
75
83
77
55
85
92
88
82
62
75
83
91
79
64
90
75
85
60
82
64
55
58
40
Gemini 3.5 Flash Lite
91
92
89
85
46
83
58
75
77
77
60
100
69
58
80
62
85
83
83
77
77
70
85
85
86
100
70
50
62
60
GPT-5.5
91
83
78
92
50
92
88
82
92
85
60
85
46
88
90
78
83
83
82
79
46
73
75
92
60
83
63
56
58
50
Claude Sonnet 5
83
83
78
85
33
83
88
83
83
77
44
85
83
88
100
62
83
83
91
79
46
82
85
85
60
91
64
71
50
40
Muse Spark 1.3 Contributor
91
83
78
77
75
83
88
83
75
77
67
77
46
88
90
78
75
83
82
71
46
70
85
77
83
82
63
71
58
40
GPT-5.6 Terra
91
83
78
85
86
83
63
73
75
69
83
77
46
88
40
78
67
83
82
79
46
70
77
77
83
91
71
71
58
40
Muse Spark 1.3
91
83
62
85
46
83
88
83
75
77
55
69
73
88
82
67
83
83
82
71
30
70
85
77
67
91
63
71
58
40
GPT-6 Astra
83
92
67
85
75
83
50
82
83
77
57
85
46
88
90
78
75
83
75
71
30
70
67
83
60
75
63
57
58
40
Mistral Medium 3.5
91
83
78
85
22
83
63
83
83
77
33
77
62
88
50
88
75
83
60
62
58
80
67
85
83
91
63
38
58
50
Mistral Large
80
73
80
92
54
83
70
92
82
85
40
75
69
64
80
40
64
83
91
62
55
80
83
85
67
91
50
40
62
40
GPT-5.6 Luna
91
92
78
77
75
83
63
83
75
62
67
77
39
88
40
78
67
83
82
79
46
70
75
85
60
91
50
56
58
40
GPT-5.4 Mini
91
83
78
85
63
75
63
83
67
69
67
69
46
75
36
78
75
83
82
50
77
82
58
85
50
91
60
71
58
36
GPT-5.4 Nano
91
83
88
77
86
83
71
83
83
69
43
62
18
100
50
70
67
83
58
62
46
46
58
69
63
91
63
83
50
40
Grok 4.20
73
75
46
77
54
92
39
62
73
69
40
77
54
54
73
62
58
75
75
62
58
40
83
77
44
83
64
40
62
50
Grok 4.3
70
42
78
62
25
67
44
73
55
69
67
69
46
60
67
70
58
83
67
58
30
75
67
69
71
91
78
86
30
60
DeepSeek V3.2
50
55
50
85
0
75
25
77
75
77
40
58
62
54
78
39
92
83
75
62
36
70
85
85
57
73
70
33
58
80
Qwen3.7 Max
55
73
54
77
25
85
33
75
85
62
50
64
69
42
64
64
75
36
46
64
60
73
0
69
71
73
55
56
77
73
Mistral Small
58
46
55
69
46
83
71
36
67
69
71
62
46
50
36
55
75
79
50
77
39
83
77
73
56
77
33
43
46
40
MiniMax M3
20
36
58
77
0
36
33
85
64
58
50
79
73
36
78
0
75
85
69
54
54
55
58
62
38
82
50
46
75
22
DeepSeek V4 Flash
82
75
50
69
39
75
42
64
64
62
38
69
9
29
73
50
75
75
73
77
50
70
58
25
40
82
55
38
42
0
Qwen3.7 Plus
27
36
64
77
39
75
50
67
36
69
46
33
69
46
67
17
75
85
82
54
36
82
67
54
38
64
55
30
69
42
DeepSeek R1
82
75
42
69
25
75
40
67
67
69
56
62
31
33
44
58
67
75
82
33
27
70
58
62
17
73
55
71
42
22
DeepSeek V4 Pro
73
67
33
69
25
75
18
75
75
69
14
8
64
20
50
40
75
75
73
8
67
82
69
62
50
82
50
40
55
30
Nova 2 Lite
64
50
43
77
36
50
71
77
58
62
50
69
18
71
27
55
18
83
60
46
50
40
64
79
20
82
43
29
42
22
Ministral 14B
82
75
50
33
25
83
50
36
55
62
60
25
27
43
67
50
36
36
27
30
50
70
69
33
60
82
50
67
30
60
Nemotron 3 Super
82
75
0
39
9
50
58
67
8
56
69
9
86
70
67
58
75
46
43
27
70
54
54
13
82
50
10
Llama 4 Maverick
50
55
25
62
40
75
13
75
46
62
40
50
27
33
82
50
46
43
30
42
30
73
69
33
43
50
50
30
58
67
GLM-5.2
10
75
54
25
39
9
17
27
83
8
50
71
77
17
82
54
9
75
75
40
36
20
69
33
40
64
67
50
10
11
Qwen3.7 Flash
70
18
14
69
14
18
30
75
64
17
50
8
18
20
30
42
18
75
20
50
30
20
58
17
20
20
60
67
20
22
Llama 4 Scout
0
42
60
46
0
64
33
42
42
39
50
57
18
22
55
33
36
36
46
27
10
36
29
23
25
40
25
33
10
33
Llama 3.3 70B
10
36
0
69
22
9
0
54
27
8
25
42
46
33
46
17
36
50
64
30
0
50
25
33
40
62
0
0
46
46

Accuracy by lab

Average and best accuracy per provider.

anthropic×5
77.2% avg · 79 best
google×1
76.1% avg · 76 best
openai×7
72.3% avg · 77 best
mistral×4
63.1% avg · 71 best
xai×2
63.1% avg · 63 best
deepseek×4
57.4% avg · 63 best
minimax×1
57.0% avg · 57 best
amazon×1
53.6% avg · 54 best
meta×5
53.1% avg · 75 best
qwen×3
50.9% avg · 60 best
nvidia×1
49.8% avg · 50 best
zhipu×1
45.0% avg · 45 best

Latency tail

p50 (filled) to p90 (hollow). Wider means less predictable.

Gemini 3.5 Flash Lite
p50 9.1s12.5s
Nova 2 Lite
p50 9.9s13.4s
Mistral Small
p50 16.2s23.1s
Grok 4.3
p50 20.7s35.2s
Mistral Medium 3.5
p50 26.6s45.7s
Grok 4.20
p50 27.8s35.9s
Muse Spark 1.3 Contributor
p50 29s45.7s
GPT-5.6 Luna
p50 30.6s36.2s
Muse Spark 1.3
p50 32.6s44.1s
GPT-5.6 Terra
p50 33.2s40.2s
GLM-5.2
p50 33.3s92.3s
MiniMax M3
p50 36.5s67.2s

Price spread

Cost to read 1,000 contracts on a log scale, a ~100× range across the field.

$5$10$50$100$300cost per 1,000 contracts (log)

Reliability

Share of calls that returned valid structured output. Only models that dropped any.

Nemotron 3 Super
90%

A Note From The Team

Flexprice runs usage-based billing. Before any invoice exists, someone has to turn a signed contract into structured billing terms: the fees, the cadence, the currency, the commitments. We do that step in production.

That is why we can score it. We already know which fields decide an invoice, what counts as the same answer when a price is written two different ways, and what it costs when a model gets one wrong and sounds certain about it.

So we wrote the rubric down, labeled a set of real contracts by hand, and ran every model we could reach. We put the whole thing in the open so nobody else picking a model for this has to start from scratch, and so the numbers keep updating as new models ship. If your contracts look nothing like ours, point the harness at your own.

Flexprice team

Fix your billing today with Flexprice

FinePrint finds which models can read the contract.
Flexprice already turns those terms into accurate invoices.

Book a demo