Data Reliability
Data

Reliability

Calibration data for the confidence scores the API returns: reliability diagrams, ECE, Brier, AUROC and risk–coverage per document class, versioned by calibrator.

calibrator cal-2026.08-4model gemini-2.5-flashgenerated August 17, 2026

Scope

Published
CORD v2, VRDU Registration and VRDU Ad-Buy: 1,846 scored fields on frozen, held-out test splits. Real published reliability curves for the cal-2026.08-4 calibrator. deepform is evaluated internally but not published this cycle; the three published classes (cord-v2, vrdu-registration, vrdu-ad-buy) are the launch set.
Variant
All numbers are the logprob-free variant: the pipeline as it runs when the model exposes no token logprobs. Every document gets this variant.
Served calibrator
Every response carries meta.calibrator_version. These curves describe cal-2026.08-4. Responses stamped identity-0 carry raw pass-through scores; no published curve describes them.
Calibration vs discrimination
ECE measures the gap between claimed confidence and the observed hit rate: 0.068–0.094 here. AUROC measures how the score ranks correct fields above incorrect ones: 0.742–0.844, where 0.5 is chance.
Transfer
These numbers describe these corpora, this model and this calibrator version. To measure your own documents, score a labeled set of your own.
Artifacts
The curve files behind this page are a stable contract: /curves/INDEX.json lists the classes and the served calibrator version, with the per-class curve and split files beside it. When a new calibrator ships, prior versions stay fetchable at their published paths.

CORD v2

Scanned retail receipts: line items and totals.

190 fields · frozen test split
ECE
0.085
Brier
0.110
AUROC
0.797
Fields scored
190

Claimed confidence and observed hit rate differ by 0.085 on average across bins. AUROC 0.797: the probability a correct field outranks an incorrect one, where 0.5 is chance.

Reliability claimed confidence vs observed accuracy, 15 equal-mass bins
observed accuracy000.50.511confidence, mean per binperfect calibration
000.50.511perfect calibration

y observed accuracy x confidence, mean per bin

Risk–coverage error among kept fields, keeping the highest-confidence share first
error among kept fields0%50%100%0%50%100%share of fields kept, highest confidence firstkeep all → 15.3% wrong
0%50%100%0%50%100%

keep all → 15.3% wrong

y error among kept fields x share of fields kept, highest confidence first

Logprob ablation: not published — no scored row carried token logprobs, so the full variant coincides with the logprob-free floor.

model
gemini-2.5-flash
calibrator
cal-2026.08-4
pipeline
caae8699b528d817ff278d1d5eb1b7abdd3d36ed2b02f1ce168dae145bdfda62
split
883d44e1cb60ee3b1feecaffcdba377c7146ba63d8c4306b9c1be3207aed31b7 (salted commitment; salt revealed at launch)
corpus
see corpora/manifests
View the raw bins
BinClaimedObservedn
10.5920.58312
20.7990.53813
30.8420.61513
40.8580.66712
50.8660.84613
60.8781.00013
70.8980.83312
80.9060.92313
90.9100.92313
100.9131.00012
110.9230.84613
120.9321.00013
130.9410.91712
140.9571.00013
150.9571.00013
Coverage keptError among kept
14%0.0%
21%2.5%
31%5.2%
37%4.3%
49%4.3%
60%6.1%
70%5.3%
80%8.6%
89%11.2%
100%15.3%

VRDU Registration

U.S. FARA registration forms: flat fields.

288 fields · frozen test split
ECE
0.068
Brier
0.123
AUROC
0.844
Fields scored
288

Claimed confidence and observed hit rate differ by 0.068 on average across bins. AUROC 0.844: the probability a correct field outranks an incorrect one, where 0.5 is chance.

Reliability claimed confidence vs observed accuracy, 15 equal-mass bins
observed accuracy000.50.511confidence, mean per binperfect calibration
000.50.511perfect calibration

y observed accuracy x confidence, mean per bin

Risk–coverage error among kept fields, keeping the highest-confidence share first
error among kept fields0%50%100%0%50%100%share of fields kept, highest confidence firstkeep all → 27.1% wrong
0%50%100%0%50%100%

keep all → 27.1% wrong

y error among kept fields x share of fields kept, highest confidence first

Logprob ablation: not published — no scored row carried token logprobs, so the full variant coincides with the logprob-free floor.

model
gemini-2.5-flash
calibrator
cal-2026.08-4
pipeline
caae8699b528d817ff278d1d5eb1b7abdd3d36ed2b02f1ce168dae145bdfda62
split
883d44e1cb60ee3b1feecaffcdba377c7146ba63d8c4306b9c1be3207aed31b7 (salted commitment; salt revealed at launch)
corpus
see corpora/manifests
View the raw bins
BinClaimedObservedn
10.1220.00019
20.3120.31619
30.5120.42119
40.6620.63219
50.7080.55020
60.7930.84219
70.8280.94719
80.8410.89519
90.8930.84219
100.9030.95020
110.9070.78919
120.9150.84219
130.9230.94719
140.9271.00019
150.9360.95020
Coverage keptError among kept
12%2.9%
20%3.4%
31%7.9%
41%8.5%
50%9.0%
59%9.5%
70%10.4%
80%14.4%
90%19.6%
100%27.1%

VRDU Ad-Buy

U.S. TV political ad-buy contracts: dense line-item tables.

1,368 fields · frozen test split
ECE
0.094
Brier
0.185
AUROC
0.742
Fields scored
1,368

Claimed confidence and observed hit rate differ by 0.094 on average across bins. AUROC 0.742: the probability a correct field outranks an incorrect one, where 0.5 is chance.

Reliability claimed confidence vs observed accuracy, 15 equal-mass bins
observed accuracy000.50.511confidence, mean per binperfect calibration
000.50.511perfect calibration

y observed accuracy x confidence, mean per bin

Risk–coverage error among kept fields, keeping the highest-confidence share first
error among kept fields0%50%100%0%50%100%share of fields kept, highest confidence firstkeep all → 38.0% wrong
0%50%100%0%50%100%

keep all → 38.0% wrong

y error among kept fields x share of fields kept, highest confidence first

Logprob ablation: not published — no scored row carried token logprobs, so the full variant coincides with the logprob-free floor.

model
gemini-2.5-flash
calibrator
cal-2026.08-4
pipeline
caae8699b528d817ff278d1d5eb1b7abdd3d36ed2b02f1ce168dae145bdfda62
split
883d44e1cb60ee3b1feecaffcdba377c7146ba63d8c4306b9c1be3207aed31b7 (salted commitment; salt revealed at launch)
corpus
see corpora/manifests
View the raw bins
BinClaimedObservedn
10.0540.02291
20.1420.25391
30.4030.37491
40.5780.44091
50.6280.64192
60.6640.83591
70.6840.58291
80.6850.84691
90.7120.68191
100.7480.76192
110.7670.65991
120.8160.69291
130.8460.60491
140.8740.98991
150.9400.91392
Coverage keptError among kept
10%6.9%
19%14.1%
30%19.1%
39%22.9%
48%24.1%
61%24.8%
70%25.6%
80%28.0%
90%31.6%
100%38.0%

How these are made

  1. 1

    An open corpus with human labels. Each class links its source and license in the provenance block.

  2. 2

    A frozen test split. Document ids are salted hashes. The commitment (the SHA-256 of the salt) is published per class; the salt is revealed at launch, so the split cannot be re-drawn after the fact. The test split is never fitted on.

  3. 3

    The production pipeline scores every field. The same code path the API serves. meta.calibrator_version names it; the pipeline hash in each provenance block pins the exact source.

  4. 4

    Metrics. Fields are grouped into 15 equal-mass bins by claimed confidence to form the reliability diagram. ECE is the bin-weighted gap between claimed and observed. Brier is the mean squared error of the claim. AUROC is the probability a correct field outranks an incorrect one. Risk–coverage is the error among kept fields as you keep only the highest-confidence share.

† Array-index scoring. CORD v2 and VRDU Ad-Buy label repeated line items by position, so a value extracted correctly but placed at a neighboring index scores as wrong. Both corpora list line items in document order, so alignment is mostly right; some of the published error on those two classes is indexing noise, not extraction error. It is published as measured, not adjusted.

This page recommends no thresholds. The right cut-off depends on your documents and your risk tolerance. Measure it on a labeled set of your own; the risk–coverage tables show the shape of the trade, not your answer.