Methodology

How we score models.

We score one thing: can a model read a contract and get the billing terms right? Here is how we test it, judge each field, and keep improving.

01

Why contracts

Finance teams still turn contracts into data by hand. We built FinePrint because that work is too costly to get wrong.

Enterprise

Companies still turn contracts into structured data by hand. Poor contract management costs roughly 9% of revenue. High-stakes text in, reliable records out.

Finance

Under ASC 606 / IFRS 15, the contract drives recognized revenue. A misread fee or cadence becomes a billing error or audit finding. FinePrint measures who is ready.

02

The task

We hand the model a real contract. It has to find every billing term we score.

  1. Contract in

    A real PDF, read line by line.

    01
  2. Model reads it

    Every billing term, structured.

    02
  3. Private key check

    Compared against our answer key.

    03
  4. On the leaderboard

    Ranked on accuracy, cost, and speed.

    04

What we look for in every contract

These are the billing fields we expect on every scan. Miss one, and the invoice breaks.

01

Identity

When the deal starts and how big it is.

start_datecurrencyusage_plan_classcontract_value
02

Platform fee

The core recurring charge for the product.

amountfrequencytiming
03

Hosting fee

Infrastructure or environment charges.

amountfrequencytiming
04

Usage fee

Metered or consumption-based pricing.

amountfrequencytiming
05

Credit grant

Prepaid credits or promotional balances.

amounttype
06

Entitlement

What the customer is allowed to use.

quantityunitperiod
07

Commitment

Minimum spend and true-up rules.

amountperiodoverage_factortrue_up
08

Overrides

Custom rates that break the standard schedule.

per-unit ratesother
09

Customer

Who the invoice goes to.

nameemailaddress
03

The dataset

Nothing synthetic in our set. Every document was signed by real parties.

Where they come from

Executed commercial agreements under confidentiality, plus material-agreement exhibits from SEC EDGAR. No synthetic docs.

What is in the mix

Order forms, MSAs, renewals, scanned pages, redlines, multiple currencies. The mix mirrors what a RevOps team sees, not the cleanest examples.

What makes them hard

Real contracts are messy. These eight patterns show up in almost every run.

Messy tables

Fees split across columns, not clean prose.

Installment math

$240k over four quarters. Models must split it.

Same price, two shapes

$10k/qtr and $40k/yr must both score.

Mixed-up cadence

Term, coverage, and billing get conflated.

Redlines

Only the edited figure counts.

Multi-currency

USD, EUR, GBP, INR from context.

OCR noise

Scanned PDFs with imperfect text.

Cross-document

MSA plus order form; order form wins.

04

Kept honest

We hold the answer key back on purpose to keep the board honest.

Private labels

Ground truth is hand-checked and never published — we only share counts, not the data.

Anonymized results

The leaderboard shows aggregates only, with contracts appearing as Doc A–F.

Fresh documents

Filings from a long tail, rotated as the set grows.

05

Metrics

One accuracy score is never enough. We publish six metrics so the trade-offs stay visible.

Accuracy

Fields the model got right, across every contract and run.

correct ÷ scored
Hallucination

Of HIGH-confidence answers, how many were wrong.

wrong ÷ high-conf
Consistency

Run-to-run spread. Low σ means repeatable.

σ across runs
Cost / 1k

Projected cost to read 1,000 contracts.

tokens × price
Latency

Median and tail wall-clock time per extraction.

p50 / p90
Value

Accuracy points per dollar spent.

accuracy ÷ cost
06

Field scoring

We normalize both sides before we compare. Quarterly and annual fees can both score if the math matches.

Example

One contract, a few fields from a single extraction run.

field-level rubric
FieldExpectedPredictedResult
recurring_fee.amount2500025000Match
recurring_fee.frequencyquarterlyquarterlyMatch
usage_fee.amount180000 / yr45000 / qtrAnnualized
recurring_fee.timingadvancedn/aWrong
fixed_fee.amount00Not scored
scope_notes“$250k capacity…”“annual usage…”Soft match

Leaderboard accuracy is correct ÷ scored across all fields and 1 runs per model.

07

Runs & labels

Every contract runs multiple times per model. Humans stay in the loop on every label.

Repeated runs

Each contract runs 1× per model. We keep the full distribution.

Report the spread

Run-to-run σ sits next to accuracy on the board.

Human labels

Ground truth is hand-labeled. Disagreements go to a human adjudicator.

08

Open & limited

Most of the stack is open. The contracts and labels stay private, by design.

Open

Harness, adapters, scorer, and aggregation are open source. Pricing comes from OpenRouter. Only the corpus and labels stay private.

Limits

Gold set is still small while the corpus grows, so σ matters. Extraction depends on OCR quality. Schema targets commercial billing terms, English-first today.