Reliability
Calibration data for the confidence scores the API returns: reliability diagrams, ECE, Brier, AUROC and risk–coverage per document class, versioned by calibrator.
Scope
- Published
- CORD v2, VRDU Registration and VRDU Ad-Buy: 1,846 scored fields on frozen, held-out test splits. Real published reliability curves for the cal-2026.08-4 calibrator. deepform is evaluated internally but not published this cycle; the three published classes (cord-v2, vrdu-registration, vrdu-ad-buy) are the launch set.
- Variant
- All numbers are the logprob-free variant: the pipeline as it runs when the model exposes no token logprobs. Every document gets this variant.
- Served calibrator
- Every response carries
meta.calibrator_version. These curves describecal-2026.08-4. Responses stampedidentity-0carry raw pass-through scores; no published curve describes them. - Calibration vs discrimination
- ECE measures the gap between claimed confidence and the observed hit rate: 0.068–0.094 here. AUROC measures how the score ranks correct fields above incorrect ones: 0.742–0.844, where 0.5 is chance.
- Transfer
- These numbers describe these corpora, this model and this calibrator version. To measure your own documents, score a labeled set of your own.
- Artifacts
- The curve files behind this page are a stable contract: /curves/INDEX.json lists the classes and the served calibrator version, with the per-class curve and split files beside it. When a new calibrator ships, prior versions stay fetchable at their published paths.
CORD v2†
Scanned retail receipts: line items and totals.
- ECE
- 0.085
- Brier
- 0.110
- AUROC
- 0.797
- Fields scored
- 190
Claimed confidence and observed hit rate differ by 0.085 on average across bins. AUROC 0.797: the probability a correct field outranks an incorrect one, where 0.5 is chance.
y observed accuracy x confidence, mean per bin
keep all → 15.3% wrong
y error among kept fields x share of fields kept, highest confidence first
Logprob ablation: not published — no scored row carried token logprobs, so the full variant coincides with the logprob-free floor.
- model
- gemini-2.5-flash
- calibrator
- cal-2026.08-4
- pipeline
- caae8699b528d817ff278d1d5eb1b7abdd3d36ed2b02f1ce168dae145bdfda62
- split
- 883d44e1cb60ee3b1feecaffcdba377c7146ba63d8c4306b9c1be3207aed31b7 (salted commitment; salt revealed at launch)
- corpus
- see corpora/manifests
View the raw bins
| Bin | Claimed | Observed | n |
|---|---|---|---|
| 1 | 0.592 | 0.583 | 12 |
| 2 | 0.799 | 0.538 | 13 |
| 3 | 0.842 | 0.615 | 13 |
| 4 | 0.858 | 0.667 | 12 |
| 5 | 0.866 | 0.846 | 13 |
| 6 | 0.878 | 1.000 | 13 |
| 7 | 0.898 | 0.833 | 12 |
| 8 | 0.906 | 0.923 | 13 |
| 9 | 0.910 | 0.923 | 13 |
| 10 | 0.913 | 1.000 | 12 |
| 11 | 0.923 | 0.846 | 13 |
| 12 | 0.932 | 1.000 | 13 |
| 13 | 0.941 | 0.917 | 12 |
| 14 | 0.957 | 1.000 | 13 |
| 15 | 0.957 | 1.000 | 13 |
| Coverage kept | Error among kept |
|---|---|
| 14% | 0.0% |
| 21% | 2.5% |
| 31% | 5.2% |
| 37% | 4.3% |
| 49% | 4.3% |
| 60% | 6.1% |
| 70% | 5.3% |
| 80% | 8.6% |
| 89% | 11.2% |
| 100% | 15.3% |
VRDU Registration
U.S. FARA registration forms: flat fields.
- ECE
- 0.068
- Brier
- 0.123
- AUROC
- 0.844
- Fields scored
- 288
Claimed confidence and observed hit rate differ by 0.068 on average across bins. AUROC 0.844: the probability a correct field outranks an incorrect one, where 0.5 is chance.
y observed accuracy x confidence, mean per bin
keep all → 27.1% wrong
y error among kept fields x share of fields kept, highest confidence first
Logprob ablation: not published — no scored row carried token logprobs, so the full variant coincides with the logprob-free floor.
- model
- gemini-2.5-flash
- calibrator
- cal-2026.08-4
- pipeline
- caae8699b528d817ff278d1d5eb1b7abdd3d36ed2b02f1ce168dae145bdfda62
- split
- 883d44e1cb60ee3b1feecaffcdba377c7146ba63d8c4306b9c1be3207aed31b7 (salted commitment; salt revealed at launch)
- corpus
- see corpora/manifests
View the raw bins
| Bin | Claimed | Observed | n |
|---|---|---|---|
| 1 | 0.122 | 0.000 | 19 |
| 2 | 0.312 | 0.316 | 19 |
| 3 | 0.512 | 0.421 | 19 |
| 4 | 0.662 | 0.632 | 19 |
| 5 | 0.708 | 0.550 | 20 |
| 6 | 0.793 | 0.842 | 19 |
| 7 | 0.828 | 0.947 | 19 |
| 8 | 0.841 | 0.895 | 19 |
| 9 | 0.893 | 0.842 | 19 |
| 10 | 0.903 | 0.950 | 20 |
| 11 | 0.907 | 0.789 | 19 |
| 12 | 0.915 | 0.842 | 19 |
| 13 | 0.923 | 0.947 | 19 |
| 14 | 0.927 | 1.000 | 19 |
| 15 | 0.936 | 0.950 | 20 |
| Coverage kept | Error among kept |
|---|---|
| 12% | 2.9% |
| 20% | 3.4% |
| 31% | 7.9% |
| 41% | 8.5% |
| 50% | 9.0% |
| 59% | 9.5% |
| 70% | 10.4% |
| 80% | 14.4% |
| 90% | 19.6% |
| 100% | 27.1% |
VRDU Ad-Buy†
U.S. TV political ad-buy contracts: dense line-item tables.
- ECE
- 0.094
- Brier
- 0.185
- AUROC
- 0.742
- Fields scored
- 1,368
Claimed confidence and observed hit rate differ by 0.094 on average across bins. AUROC 0.742: the probability a correct field outranks an incorrect one, where 0.5 is chance.
y observed accuracy x confidence, mean per bin
keep all → 38.0% wrong
y error among kept fields x share of fields kept, highest confidence first
Logprob ablation: not published — no scored row carried token logprobs, so the full variant coincides with the logprob-free floor.
- model
- gemini-2.5-flash
- calibrator
- cal-2026.08-4
- pipeline
- caae8699b528d817ff278d1d5eb1b7abdd3d36ed2b02f1ce168dae145bdfda62
- split
- 883d44e1cb60ee3b1feecaffcdba377c7146ba63d8c4306b9c1be3207aed31b7 (salted commitment; salt revealed at launch)
- corpus
- see corpora/manifests
View the raw bins
| Bin | Claimed | Observed | n |
|---|---|---|---|
| 1 | 0.054 | 0.022 | 91 |
| 2 | 0.142 | 0.253 | 91 |
| 3 | 0.403 | 0.374 | 91 |
| 4 | 0.578 | 0.440 | 91 |
| 5 | 0.628 | 0.641 | 92 |
| 6 | 0.664 | 0.835 | 91 |
| 7 | 0.684 | 0.582 | 91 |
| 8 | 0.685 | 0.846 | 91 |
| 9 | 0.712 | 0.681 | 91 |
| 10 | 0.748 | 0.761 | 92 |
| 11 | 0.767 | 0.659 | 91 |
| 12 | 0.816 | 0.692 | 91 |
| 13 | 0.846 | 0.604 | 91 |
| 14 | 0.874 | 0.989 | 91 |
| 15 | 0.940 | 0.913 | 92 |
| Coverage kept | Error among kept |
|---|---|
| 10% | 6.9% |
| 19% | 14.1% |
| 30% | 19.1% |
| 39% | 22.9% |
| 48% | 24.1% |
| 61% | 24.8% |
| 70% | 25.6% |
| 80% | 28.0% |
| 90% | 31.6% |
| 100% | 38.0% |
How these are made
- 1
An open corpus with human labels. Each class links its source and license in the provenance block.
- 2
A frozen test split. Document ids are salted hashes. The commitment (the SHA-256 of the salt) is published per class; the salt is revealed at launch, so the split cannot be re-drawn after the fact. The test split is never fitted on.
- 3
The production pipeline scores every field. The same code path the API serves.
meta.calibrator_versionnames it; the pipeline hash in each provenance block pins the exact source. - 4
Metrics. Fields are grouped into 15 equal-mass bins by claimed confidence to form the reliability diagram. ECE is the bin-weighted gap between claimed and observed. Brier is the mean squared error of the claim. AUROC is the probability a correct field outranks an incorrect one. Risk–coverage is the error among kept fields as you keep only the highest-confidence share.
† Array-index scoring. CORD v2 and VRDU Ad-Buy label repeated line items by position, so a value extracted correctly but placed at a neighboring index scores as wrong. Both corpora list line items in document order, so alignment is mostly right; some of the published error on those two classes is indexing noise, not extraction error. It is published as measured, not adjusted.
This page recommends no thresholds. The right cut-off depends on your documents and your risk tolerance. Measure it on a labeled set of your own; the risk–coverage tables show the shape of the trade, not your answer.