Identity
When the deal starts and how big it is.
start_datecurrencyusage_plan_classcontract_valueMethodology
We score one thing: can a model read a contract and get the billing terms right? Here is how we test it, judge each field, and keep improving.
Finance teams still turn contracts into data by hand. We built FinePrint because that work is too costly to get wrong.
Companies still turn contracts into structured data by hand. Poor contract management costs roughly 9% of revenue. High-stakes text in, reliable records out.
Under ASC 606 / IFRS 15, the contract drives recognized revenue. A misread fee or cadence becomes a billing error or audit finding. FinePrint measures who is ready.
We hand the model a real contract. It has to find every billing term we score.
A real PDF, read line by line.
01Every billing term, structured.
02Compared against our answer key.
03Ranked on accuracy, cost, and speed.
04These are the billing fields we expect on every scan. Miss one, and the invoice breaks.
When the deal starts and how big it is.
start_datecurrencyusage_plan_classcontract_valueThe core recurring charge for the product.
amountfrequencytimingInfrastructure or environment charges.
amountfrequencytimingMetered or consumption-based pricing.
amountfrequencytimingPrepaid credits or promotional balances.
amounttypeWhat the customer is allowed to use.
quantityunitperiodMinimum spend and true-up rules.
amountperiodoverage_factortrue_upCustom rates that break the standard schedule.
per-unit ratesotherWho the invoice goes to.
nameemailaddressNothing synthetic in our set. Every document was signed by real parties.
Executed commercial agreements under confidentiality, plus material-agreement exhibits from SEC EDGAR. No synthetic docs.
Order forms, MSAs, renewals, scanned pages, redlines, multiple currencies. The mix mirrors what a RevOps team sees, not the cleanest examples.
What makes them hard
Real contracts are messy. These eight patterns show up in almost every run.
Fees split across columns, not clean prose.
$240k over four quarters. Models must split it.
$10k/qtr and $40k/yr must both score.
Term, coverage, and billing get conflated.
Only the edited figure counts.
USD, EUR, GBP, INR from context.
Scanned PDFs with imperfect text.
MSA plus order form; order form wins.
We hold the answer key back on purpose to keep the board honest.
Ground truth is hand-checked and never published — we only share counts, not the data.
The leaderboard shows aggregates only, with contracts appearing as Doc A–F.
Filings from a long tail, rotated as the set grows.
One accuracy score is never enough. We publish six metrics so the trade-offs stay visible.
Fields the model got right, across every contract and run.
Of HIGH-confidence answers, how many were wrong.
Run-to-run spread. Low σ means repeatable.
Projected cost to read 1,000 contracts.
Median and tail wall-clock time per extraction.
Accuracy points per dollar spent.
We normalize both sides before we compare. Quarterly and annual fees can both score if the math matches.
Example
One contract, a few fields from a single extraction run.
| Field | Expected | Predicted | Result |
|---|---|---|---|
recurring_fee.amount | 25000 | 25000 | Match |
recurring_fee.frequency | quarterly | quarterly | Match |
usage_fee.amount | 180000 / yr | 45000 / qtr | Annualized |
recurring_fee.timing | advanced | n/a | Wrong |
fixed_fee.amount | 0 | 0 | Not scored |
scope_notes | “$250k capacity…” | “annual usage…” | Soft match |
Leaderboard accuracy is correct ÷ scored across all fields and 1 runs per model.
Every contract runs multiple times per model. Humans stay in the loop on every label.
Each contract runs 1× per model. We keep the full distribution.
Run-to-run σ sits next to accuracy on the board.
Ground truth is hand-labeled. Disagreements go to a human adjudicator.
Most of the stack is open. The contracts and labels stay private, by design.
Harness, adapters, scorer, and aggregation are open source. Pricing comes from OpenRouter. Only the corpus and labels stay private.
Gold set is still small while the corpus grows, so σ matters. Extraction depends on OCR quality. Schema targets commercial billing terms, English-first today.