Compare
Document extraction API pricing, 2026: list prices per 1,000 pages from ten vendors
Ten ways to turn a PDF into fields, priced per 1,000 pages. Each vendor figure below is a quotation from that vendor's pricing page with the day it was taken, and the table shows the arithmetic where the vendor bills in credits, tokens or document counts. The two token-priced models carry the cost measured in Velrim's published comparison of six extraction setups (September 2026), because a token price alone does not say what a page costs.
The table
| Vendor | Product | Billing unit | List price per 1,000 pages | Free tier |
|---|---|---|---|---|
| Velrim | Extraction with grounding and confidence | Extracted page | $20 ($0.02 a page) | About 500 pages, no card, 90 days |
| Reducto | Extract / Deep Extract / Parse, Standard tier | Page | $20 / $40 / $10 | $150 in free usage |
| LlamaExtract (LlamaCloud) | Cost-effective / Agentic / Agentic Plus / Turbo | Credit, $1.25 per 1,000 | $10.00 / $31.25 / $75.00 / $43.75 (8, 25, 60 and 35 credits a page) | Free plan, 10K credits |
| Mistral | OCR 4 / Document AI | Page | $4 / $5, OCR $2 on the batch API | |
| Unstructured | Pay-As-You-Go | Page | $15 ($0.015 a page) | 10,000 pages, no card |
| Azure AI Document Intelligence, East US | Read / Layout / Prebuilt / Custom | Page | $1.50 / $10 / $10 / $30 | 500 pages a month |
| Amazon Textract, US West (Oregon) | Detect Document Text / Tables / Forms / Analyze ID | Page, per feature | $1.50 / $15 / $50 / $25 | 1,000 text pages a month for three months, 100 with forms and tables |
| Google Document AI | Enterprise Document OCR / Layout Parser / Form Parser / Custom extractor / Invoice parser | Page, or a count of up to 10 pages for the invoice parser | $1.50 / $10 / $30 / $30 / $0.10 per document of up to 10 pages | 1,000 OCR pages a month |
| Gemini 2.5 Flash, called directly | Model API | Token, 258 a page | $0.08 in page tokens at list, ~$2.2 (measured, token-priced) | Free tier, inputs used to improve Google's products |
| gpt-5.4-mini, called directly | Model API | Token | $0.75 in and $4.50 out per million tokens, ~$1.0 (measured, prompt-cached) |
How each vendor counts a page
Velrim bills per extracted page at $0.02, with grounding and the confidence score included and a prepaid balance from $5 (pricing). Reducto prices each endpoint per 1,000 pages on its Standard tier, with Growth and Enterprise through sales (Reducto pricing, read 2026-09-03).
All features are priced using credits, which are billed per page (or minute for audio). Price per 1,000 Credits: North America $1.25. Europe $1.25.
Extract tier, Extract credits/page, Default parse tier, Parse credits/page, Default total: Agentic Plus, 50, Agentic, 10, 60. Agentic, 15, Agentic, 10, 25. Cost-effective, 5, Cost-effective, 3, 8. Turbo, 35.
A LlamaExtract page costs the extract tier plus the parse tier it runs on: 8 credits on Cost-effective, 25 on Agentic, 60 on Agentic Plus and 35 on Turbo, which at $1.25 per 1,000 credits is $10.00, $31.25, $75.00 and $43.75 per 1,000 pages. The plans start at a free tier with 10K credits, then $50 a month for 40K and $500 a month for 400K (LlamaIndex pricing, read 2026-09-03).
Mistral OCR 4 through the API is priced at $4 per 1,000 pages, with a 50% Batch-API discount, reducing the cost to $2 per 1,000 pages. Document AI is priced at $5 per 1,000 pages. June 23, 2026.
Pay only $0.015 a page after your first 10,000 free pages.
10,000 free pages to start. No card required. All features included.
S0 Read Pages, 1K, 1.5 (0.6 from 1000 units). S0 Batch Layout Pages, 1K, 10.0. S0 Batch Pre-built Pages, 1K, 10.0. S0 Custom Pages, 1K, 30.0. S0 Custom Generative Pages, 1K, 30.0. S0 Add-on for Pages, 1K, 6.0. Free Transactions, 1K, 0.0. Currency USD.
Azure lists 500 free pages a month (Azure AI Document Intelligence pricing, read 2026-09-03). The rows above are the East US consumption meters from the retail prices API, in dollars per 1,000 pages: Read $1.50, Layout $10, Prebuilt $10, Custom $30, and $6 for an add-on.
The pricing per page in the US West (Oregon) region for the first one million pages is $0.0015, and pages after one million are $0.0006
The pricing per page in the US West (Oregon) region for one million pages with tables is $0.015, and $0.01 per page after one million pages. Pages with forms is $0.05 for one million pages, and $0.04 per page after one million.
The Free Tier lasts for three months, and new AWS customers can analyze up to: Detect Document Text API: 1,000 pages per month. 100 Pages per month when using Forms, Tables, and Layout features.
Analyze ID is $0.025 a page for the first 100,000 (Amazon Textract pricing, read 2026-09-03). Textract bills per page per feature, so a page run through Forms and Tables together costs $0.065, or $65 per 1,000.
Enterprise Document OCR Processor: 0 count to 1,000 count $0.00 (Free). 1,000 count to 5,000,000 count $1.50. 5,000,000 count and above $0.60.
Custom extractor: 0 count to 1,000,000 count $30.00. 1,000,000 count and above $20.00. Form Parser: 0 count to 1,000,000 count $30.00. 1,000,000 count and above $20.00. Layout Parser (Includes initial chunking) $10.00.
Invoice parser* $0.10 / 1 count. Expense parser (formerly receipt parser)* $0.10 / 1 count. * 1 count equals up to 10 pages in a document.
Google bills the general processors per page and the specialized parsers per count, where one count is a document of up to 10 pages. An invoice parser count is $0.10, so 1,000 one-page invoices cost $100 and 1,000 ten-page invoices cost the same $100.
Token-priced models
Gemini 2.5 Flash, Standard, Paid Tier, per 1M tokens in USD. Input price: $0.30 (text / image / video), $1.00 (audio). Output price: $2.50. Context caching price: $0.03 (text / image / video), $0.1 (audio), $1.00 / 1,000,000 tokens per hour. Last updated 2026-09-02 UTC.
Each document page is equivalent to 258 tokens.
Model | Short context input | Short context cached input | Short context cache writes | Short context output. gpt-5.4-mini | $0.75 | $0.075 | - | $4.50
At 258 tokens a page, the pages sent to Gemini 2.5 Flash cost $0.08 per 1,000 at the paid tier. The prompt, the schema and the answer are on top, and the answer is billed at $2.50 per million. OpenAI's pricing page has no line on how PDF inputs are billed as of 2026-09-03. Velrim's published comparison of six extraction setups (September 2026) measured both models on 319 pages: the table below carries the result in its last column.
| Arm | macro-F1 (norm) | fabrication (absent fields) | wrong among the top 90% by confidence | $/1k pages |
|---|---|---|---|---|
| Velrim | 0.691 | 11.5% [6.1, 19.8] | 36.1% [29.9, 44.1] | $20 (list; measured matches) |
| Gemini free-decode | 0.709 | 17.0% [10.1, 26.7] | not requested | ~$2.2 (measured, token-priced) |
| Gemini constrained | 0.697 | 12.8% [7.6, 20.5] | not requested | ~$2.2 (measured, token-priced) |
| gpt-5.4-mini free | 0.735 | 10.8% [5.7, 18.4] | none surfaced | ~$1.0 (measured, prompt-cached) |
| gpt-5.4-mini structured | 0.718 | 10.4% [5.3, 17.6] | none surfaced | ~$1.0 (measured, prompt-cached) |
| Mistral OCR 4 | 0.712 | 40.3% [31.0, 52.5] | none surfaced | $5 (list; measured $5.3) |
Reading the table
Text out of a page costs $1.50 per 1,000 at all three clouds. Typed fields out of a page run from $10.00 to $75 per 1,000 at the vendors that bill per page, and a few dollars at the model APIs. The differences at the same price are in what comes back with the value: a page location, a score per field, and evidence for that score. Velrim publishes its error curves and the comparison that put it in the table above at $20, the most expensive of the six rows.