API · Documents
Submit a document
/api/v1/documents
$0.025/page (layout) or $0.004/page (read), + tokens for what you switch on
Send multipart/form-data with file (or several files[]) and an optional options field containing a JSON DocumentOptions object, or a JSON body with a url to fetch. Poll poll_url until status is completed or failed.
With no options the document is read (text, tables, key/value pairs and layout for every page) and its figures and picture-heavy pages are turned into images. That is all: no AI model runs, result is null, and you pay for the pages. Everything else is opt-in: embeddings makes vectors, extraction has the model fill result.data (built-in fields and/or your own schema), and queries has it answer questions.
Body
Send the document one of 2 ways. Pick one: each has its own fields, and "required" means required for that way.
multipart/form-data
The file comes from your computer or server. Send it as file (or several as files[]); there is no url. Options go in their own options part, as a JSON string.
-
filestring optionalThe document: a PDF, image or Office file. Send
fileorfiles[]. -
filesarray of string optionalUp to 10 documents at once, as repeated
files[]parts. Each becomes its own document. -
optionsstring optionalProcessing options, as a JSON string in its own form part. Every option is optional.
8 properties
-
analysislayout | read optional default layoutHow the pages are read.
layout($0.025/page) also finds tables, key/value pairs, figures and reading order.read($0.004/page) is OCR only, about six times cheaper: pages come back without tables, key/value pairs or figures, andmarkdown: truewith it is a 422. Example:"analysis": "read". -
embeddingstrue | EmbeddingOptions optionalMake a vector for every page, for your own search or model. Off unless you send it:
truefor the defaults (one 1,024-number vector per page), or an object to choose the settings. Billed ascompass.input. Read them fromGET /documents/{id}/pages?include=embeddings. Example:"embeddings": true.2 properties
-
dimensions1024 | 512 | 256 optional default 1024Vector length. Compass is trained so shorter vectors stay usable: 512 and 256 cut storage and index size at a small accuracy cost. Pick one per corpus — you cannot compare vectors of different lengths. Example:
"dimensions": 256. -
granularitypage optional default pagepageembeds each page as one vector (from up to its first 20,000 bytes of text), returned byGET /documents/{id}/pages?include=embeddings.
-
-
extractionExtractionOptions optionalHave the model fill
result.data: any of Northdoc's built-infields(summary, parties, dates and more), your own JSONschema, or both. Off unless you send it. Name at least one field or give a schema. Example:"extraction": {"fields": ["summary", "parties"]}.4 properties
-
fieldsarray of title | document_type | summary | dates | parties | amounts | key_facts optionalBuilt-in fields, each with a fixed shape in
result.data:title— the title or heading, a string.document_type— what kind of document it is (invoice, contract, letter…), a string.summary— two or three sentences on what it says.dates—[{label, date}], dates in ISO 8601 where possible.parties—[{name, role}], the people and organisations involved.amounts—[{label, amount, currency}], the money in it.key_facts— the most important facts, one short sentence each.
Only the fields you name come back. A value the document does not confirm is left out, never guessed. Example:
"fields": ["summary", "parties", "dates"]. -
instructionsstring optionalFree-text guidance for ambiguous fields — which of two totals to take, what to do when a field is missing, how to normalise dates. Example:
"instructions": "Use the invoice date, not the print date."at most 4000 characters
-
per_pageboolean optional default falseRun the extraction separately on every page instead of once over the whole document. Each page's result appears on that page (
resultinGET /documents/{id}/pages) as well as underresult.pages. Costs roughly one extraction per page — use it for documents that are a stack of independent records, not for one contract spanning pages. Up to 200 pages; a longer document fails witherror.type: document_too_large. -
schemaobject optionalYour own JSON Schema (
type: object) forresult.datato follow. Nested objects and arrays work, so line items and tables come back structured. Mark the fields you depend on asrequired.
-
-
imagesboolean optional default trueTake pictures of the document: figures (charts, diagrams, photos) are cropped from their pages, and a whole page is kept when pictures cover at least half of it or it has fewer than 25 words (scans, signature and stamp pages, handwriting). Tiny figures (under 1.5% of the page, like logos) and blank pages are skipped. Download them with
GET /documents/{id}/images, or with each page fromGET /documents/{id}/pages?include=image_data. Covered by the page price.When the model runs it reads the pictures where they sit in the text. They are billed as input tokens then (about width × height / 750 each, at most ~1,600); one request carries at most 20 images and 10 MB, and past that the largest are kept and
result.warningssays how many were left out.Works on PDFs, and on PNG and JPEG uploads (under 3.75 MB and 8000px, kept as the page image). TIFF, BMP, HEIF and Office files have no pictures taken.
falsetakes none. Example:"images": false. -
markdownboolean optional default falseAlso produce a markdown rendering of the whole document, read with
GET /documents/{id}?include=markdown. Needsanalysis: layout. Same page price. Example:"markdown": true. -
modelauto | swift | summit optional default autoWhich Northdoc model reads the document, when one runs (for an
extractionorqueries).swift— fast and economical, for everyday documents up to about 180k tokens (roughly 300 pages of dense text).summit— the most capable, for long or dense legal and financial documents, up to about 900k tokens.auto(the default) —swiftup to about 150k tokens,summitbeyond that.
Pin
summitwhen accuracy on hard documents matters most; pinswiftto cap cost. A document too long forswiftis accepted, read and charged for its pages, then fails atextractingwitherror.type: document_too_large;automoves tosummitinstead.model.usedon the document says which one ran (northdoc-swift-1ornorthdoc-summit-1).Swift and Summit are kept current: when a better model becomes available the name moves to it and the version in
model.usedgoes up (northdoc-swift-2), with no change to the API. Example:"model": "summit". -
queriesarray of string optionalQuestions the model answers while the document is processed, returned in order under
result.answers. Each question once: the same question twice is a 422. Cheaper than onePOST /documents/{id}/queriesper question because the document is read once for all of them. Use the queries endpoint for follow-ups you only think of later. Example:"queries": ["What is the total?", "When is it due?"].at most 50 items
-
retention_secondsinteger optional default 86400How long results are kept once the document completes:
expires_atiscompleted_atplus this many seconds. A result you have never fetched is kept for at least 24 hours after it completes, however short this is, so a short retention never deletes a result nobody has read. Reading a result does not moveexpires_at.0means keep indefinitely, which only plans withlimits.max_retention_seconds == 0may do; on capped plans0is clamped to the plan limit and anything larger is a 422. Read the ceiling fromGET /account. After it lapses the document returns410 document_expired. Example:"retention_seconds": 3600.at least 0
-
application/json
The file is already online. Send its link as url and Northdoc downloads it; there is no file. Options go in options, as a normal JSON object.
-
urlstring requiredhttp(s) URL of a PDF, image or Office file. Pre-signed object-storage links are fine.
at most 8192 characters
-
optionsDocumentOptions optionalProcessing options, as a JSON object. Every option is optional.
8 properties
-
analysislayout | read optional default layoutHow the pages are read.
layout($0.025/page) also finds tables, key/value pairs, figures and reading order.read($0.004/page) is OCR only, about six times cheaper: pages come back without tables, key/value pairs or figures, andmarkdown: truewith it is a 422. Example:"analysis": "read". -
embeddingstrue | EmbeddingOptions optionalMake a vector for every page, for your own search or model. Off unless you send it:
truefor the defaults (one 1,024-number vector per page), or an object to choose the settings. Billed ascompass.input. Read them fromGET /documents/{id}/pages?include=embeddings. Example:"embeddings": true.2 properties
-
dimensions1024 | 512 | 256 optional default 1024Vector length. Compass is trained so shorter vectors stay usable: 512 and 256 cut storage and index size at a small accuracy cost. Pick one per corpus — you cannot compare vectors of different lengths. Example:
"dimensions": 256. -
granularitypage optional default pagepageembeds each page as one vector (from up to its first 20,000 bytes of text), returned byGET /documents/{id}/pages?include=embeddings.
-
-
extractionExtractionOptions optionalHave the model fill
result.data: any of Northdoc's built-infields(summary, parties, dates and more), your own JSONschema, or both. Off unless you send it. Name at least one field or give a schema. Example:"extraction": {"fields": ["summary", "parties"]}.4 properties
-
fieldsarray of title | document_type | summary | dates | parties | amounts | key_facts optionalBuilt-in fields, each with a fixed shape in
result.data:title— the title or heading, a string.document_type— what kind of document it is (invoice, contract, letter…), a string.summary— two or three sentences on what it says.dates—[{label, date}], dates in ISO 8601 where possible.parties—[{name, role}], the people and organisations involved.amounts—[{label, amount, currency}], the money in it.key_facts— the most important facts, one short sentence each.
Only the fields you name come back. A value the document does not confirm is left out, never guessed. Example:
"fields": ["summary", "parties", "dates"]. -
instructionsstring optionalFree-text guidance for ambiguous fields — which of two totals to take, what to do when a field is missing, how to normalise dates. Example:
"instructions": "Use the invoice date, not the print date."at most 4000 characters
-
per_pageboolean optional default falseRun the extraction separately on every page instead of once over the whole document. Each page's result appears on that page (
resultinGET /documents/{id}/pages) as well as underresult.pages. Costs roughly one extraction per page — use it for documents that are a stack of independent records, not for one contract spanning pages. Up to 200 pages; a longer document fails witherror.type: document_too_large. -
schemaobject optionalYour own JSON Schema (
type: object) forresult.datato follow. Nested objects and arrays work, so line items and tables come back structured. Mark the fields you depend on asrequired.
-
-
imagesboolean optional default trueTake pictures of the document: figures (charts, diagrams, photos) are cropped from their pages, and a whole page is kept when pictures cover at least half of it or it has fewer than 25 words (scans, signature and stamp pages, handwriting). Tiny figures (under 1.5% of the page, like logos) and blank pages are skipped. Download them with
GET /documents/{id}/images, or with each page fromGET /documents/{id}/pages?include=image_data. Covered by the page price.When the model runs it reads the pictures where they sit in the text. They are billed as input tokens then (about width × height / 750 each, at most ~1,600); one request carries at most 20 images and 10 MB, and past that the largest are kept and
result.warningssays how many were left out.Works on PDFs, and on PNG and JPEG uploads (under 3.75 MB and 8000px, kept as the page image). TIFF, BMP, HEIF and Office files have no pictures taken.
falsetakes none. Example:"images": false. -
markdownboolean optional default falseAlso produce a markdown rendering of the whole document, read with
GET /documents/{id}?include=markdown. Needsanalysis: layout. Same page price. Example:"markdown": true. -
modelauto | swift | summit optional default autoWhich Northdoc model reads the document, when one runs (for an
extractionorqueries).swift— fast and economical, for everyday documents up to about 180k tokens (roughly 300 pages of dense text).summit— the most capable, for long or dense legal and financial documents, up to about 900k tokens.auto(the default) —swiftup to about 150k tokens,summitbeyond that.
Pin
summitwhen accuracy on hard documents matters most; pinswiftto cap cost. A document too long forswiftis accepted, read and charged for its pages, then fails atextractingwitherror.type: document_too_large;automoves tosummitinstead.model.usedon the document says which one ran (northdoc-swift-1ornorthdoc-summit-1).Swift and Summit are kept current: when a better model becomes available the name moves to it and the version in
model.usedgoes up (northdoc-swift-2), with no change to the API. Example:"model": "summit". -
queriesarray of string optionalQuestions the model answers while the document is processed, returned in order under
result.answers. Each question once: the same question twice is a 422. Cheaper than onePOST /documents/{id}/queriesper question because the document is read once for all of them. Use the queries endpoint for follow-ups you only think of later. Example:"queries": ["What is the total?", "When is it due?"].at most 50 items
-
retention_secondsinteger optional default 86400How long results are kept once the document completes:
expires_atiscompleted_atplus this many seconds. A result you have never fetched is kept for at least 24 hours after it completes, however short this is, so a short retention never deletes a result nobody has read. Reading a result does not moveexpires_at.0means keep indefinitely, which only plans withlimits.max_retention_seconds == 0may do; on capped plans0is clamped to the plan limit and anything larger is a 422. Read the ceiling fromGET /account. After it lapses the document returns410 document_expired. Example:"retention_seconds": 3600.at least 0
-
Returns 202
Accepted for processing. A single file or url returns the document itself; a batch returns {"documents": [...]} with one entry per file, in order. Batch entries are independent — a file that is rejected comes back as {"error": {...}} in its slot while the others are still accepted, so check each one.
-
idstringThe document's id. Keep it: every other call needs it.
-
objectstringAlways
document. -
estimated_cost_microintegerCredit held while it runs, in micro-USD. You are charged the real cost at the end.
nullfor a URL until it has been downloaded. -
poll_urlstringWhere to check on it:
GET /documents/{id}. -
statusstringqueued: it has not started yet.
Errors
-
401Missing, revoked or expired key
-
402No credits left, or (for an upload) not enough for its estimated cost. A URL document or a question whose estimate the credit cannot cover is accepted and then fails with
error.type: insufficient_credits. -
413The file is over the plan's upload limit (or the request body over the server limit)
-
415Not a PDF, PNG, JPEG, TIFF, BMP, HEIF, DOCX, XLSX or PPTX
-
422Body or options failed validation (
detailslists the fields), or the document is too large (document_too_large) -
429Per-second burst limit for the plan exceeded; retry after
Retry-Afterseconds
Every error has the same shape. See Errors.