API · Documents
List documents
/api/v1/documents
Newest first. Page with starting_after=<id>; filter with status.
Query parameters
-
limitinteger optional default 20How many to return.
Example:
limit=50 -
starting_afterstring optionalId of the last document on the previous page; returns the ones after it.
Example:
starting_after=8f1c2d3e-… -
statusqueued | processing | completed | failed | expired optionalOnly documents in this state.
Example:
status=completed
Returns 200
A page of documents.
-
objectstringAlways
list. -
dataarray of DocumentDocuments, newest first.
resultis left out of lists.29 fields
-
idstringThe document's id. Keep it: every other call needs it.
-
objectstringAlways
document. -
analysislayout | readThe
analysisoption it ran with. -
byte_sizeintegerFile size in bytes.
-
completed_atstringWhen it finished.
-
content_typestringThe type detected from the file's bytes, e.g.
application/pdf. -
costobjectestimated_microwas held from your credit when it was accepted;settled_microis what it actually cost, once finished.2 fields
-
estimated_microintegerHeld before processing
-
settled_microintegerCharged after processing
-
-
created_atstringWhen you submitted it.
-
errorobjectPresent when
statusisfailed.typeis one ofanalysis_failed,fetch_failed,unsupported_media_type,payload_too_large,document_too_large,schema_mismatch,insufficient_credits,embedding_failedorprocessing_failed(anything else).stageis the step it failed at. A failed document keeps answering200with itserror; only an expired one answers410. You pay only for the work done before it failed.3 fields
-
messagestring -
stagestring -
typestring
-
-
expires_atstringWhen the result will be deleted (
completed_at+retention_seconds);nullif kept forever. -
failed_atstringWhen it failed.
-
filenamestringThe uploaded file's name, or the last part of the URL.
-
markdownstringWith
include=markdownand themarkdownoutput: the document as markdown. -
modelive | testtestwhen made with anak_test_…key: fake results, no charge. -
modelobjectrequestedis themodeloption you sent;usedis the model and version that actually read it, such asnorthdoc-swift-1.2 fields
-
requestedstring -
usedstring
-
-
optionsDocumentOptionsWith
include=options: the options it was sent with, defaults filled in.8 fields
-
analysislayout | read default layoutHow the pages are read.
layout($0.025/page) also finds tables, key/value pairs, figures and reading order.read($0.004/page) is OCR only, about six times cheaper: pages come back without tables, key/value pairs or figures, andmarkdown: truewith it is a 422. Example:"analysis": "read". -
embeddingstrue | EmbeddingOptionsMake a vector for every page, for your own search or model. Off unless you send it:
truefor the defaults (one 1,024-number vector per page), or an object to choose the settings. Billed ascompass.input. Read them fromGET /documents/{id}/pages?include=embeddings. Example:"embeddings": true.2 fields
-
dimensions1024 | 512 | 256 default 1024Vector length. Compass is trained so shorter vectors stay usable: 512 and 256 cut storage and index size at a small accuracy cost. Pick one per corpus — you cannot compare vectors of different lengths. Example:
"dimensions": 256. -
granularitypage default pagepageembeds each page as one vector (from up to its first 20,000 bytes of text), returned byGET /documents/{id}/pages?include=embeddings.
-
-
extractionExtractionOptionsHave the model fill
result.data: any of Northdoc's built-infields(summary, parties, dates and more), your own JSONschema, or both. Off unless you send it. Name at least one field or give a schema. Example:"extraction": {"fields": ["summary", "parties"]}.4 fields
-
fieldsarray of title | document_type | summary | dates | parties | amounts | key_factsBuilt-in fields, each with a fixed shape in
result.data:title— the title or heading, a string.document_type— what kind of document it is (invoice, contract, letter…), a string.summary— two or three sentences on what it says.dates—[{label, date}], dates in ISO 8601 where possible.parties—[{name, role}], the people and organisations involved.amounts—[{label, amount, currency}], the money in it.key_facts— the most important facts, one short sentence each.
Only the fields you name come back. A value the document does not confirm is left out, never guessed. Example:
"fields": ["summary", "parties", "dates"]. -
instructionsstringFree-text guidance for ambiguous fields — which of two totals to take, what to do when a field is missing, how to normalise dates. Example:
"instructions": "Use the invoice date, not the print date."at most 4000 characters
-
per_pageboolean default falseRun the extraction separately on every page instead of once over the whole document. Each page's result appears on that page (
resultinGET /documents/{id}/pages) as well as underresult.pages. Costs roughly one extraction per page — use it for documents that are a stack of independent records, not for one contract spanning pages. Up to 200 pages; a longer document fails witherror.type: document_too_large. -
schemaobjectYour own JSON Schema (
type: object) forresult.datato follow. Nested objects and arrays work, so line items and tables come back structured. Mark the fields you depend on asrequired.
-
-
imagesboolean default trueTake pictures of the document: figures (charts, diagrams, photos) are cropped from their pages, and a whole page is kept when pictures cover at least half of it or it has fewer than 25 words (scans, signature and stamp pages, handwriting). Tiny figures (under 1.5% of the page, like logos) and blank pages are skipped. Download them with
GET /documents/{id}/images, or with each page fromGET /documents/{id}/pages?include=image_data. Covered by the page price.When the model runs it reads the pictures where they sit in the text. They are billed as input tokens then (about width × height / 750 each, at most ~1,600); one request carries at most 20 images and 10 MB, and past that the largest are kept and
result.warningssays how many were left out.Works on PDFs, and on PNG and JPEG uploads (under 3.75 MB and 8000px, kept as the page image). TIFF, BMP, HEIF and Office files have no pictures taken.
falsetakes none. Example:"images": false. -
markdownboolean default falseAlso produce a markdown rendering of the whole document, read with
GET /documents/{id}?include=markdown. Needsanalysis: layout. Same page price. Example:"markdown": true. -
modelauto | swift | summit default autoWhich Northdoc model reads the document, when one runs (for an
extractionorqueries).swift— fast and economical, for everyday documents up to about 180k tokens (roughly 300 pages of dense text).summit— the most capable, for long or dense legal and financial documents, up to about 900k tokens.auto(the default) —swiftup to about 150k tokens,summitbeyond that.
Pin
summitwhen accuracy on hard documents matters most; pinswiftto cap cost. A document too long forswiftis accepted, read and charged for its pages, then fails atextractingwitherror.type: document_too_large;automoves tosummitinstead.model.usedon the document says which one ran (northdoc-swift-1ornorthdoc-summit-1).Swift and Summit are kept current: when a better model becomes available the name moves to it and the version in
model.usedgoes up (northdoc-swift-2), with no change to the API. Example:"model": "summit". -
queriesarray of stringQuestions the model answers while the document is processed, returned in order under
result.answers. Each question once: the same question twice is a 422. Cheaper than onePOST /documents/{id}/queriesper question because the document is read once for all of them. Use the queries endpoint for follow-ups you only think of later. Example:"queries": ["What is the total?", "When is it due?"].at most 50 items
-
retention_secondsinteger default 86400How long results are kept once the document completes:
expires_atiscompleted_atplus this many seconds. A result you have never fetched is kept for at least 24 hours after it completes, however short this is, so a short retention never deletes a result nobody has read. Reading a result does not moveexpires_at.0means keep indefinitely, which only plans withlimits.max_retention_seconds == 0may do; on capped plans0is clamped to the plan limit and anything larger is a 422. Read the ceiling fromGET /account. After it lapses the document returns410 document_expired. Example:"retention_seconds": 3600.at least 0
-
-
page_countintegerNumber of pages (the real count once analysed).
-
pagesarray of DocumentPageWith
include=pages: one object per page.9 fields
-
embeddingobjectWith
include=embeddings: the page's vector, the model that made it and its length.3 fields
-
dimensionsinteger -
modelstring -
vectorarray of number
-
-
imagesarray of ImageWith
include=imagesorinclude=image_data: the page's pictures, in their order on the page (withimage_data, bytes included).13 fields
-
idstringThe image's id.
-
objectstringAlways
image. -
bboxobjectFor a figure, where it sits on the page:
left,top,widthandheightas fractions of the page (0 to 1).nullfor a whole page. -
byte_sizeintegerFile size in bytes.
-
captionstringThe figure's caption as the OCR read it, or
null. -
datastringWith
include=image_data(pages) orinclude=data(images): the file itself, base64-encoded. -
figure_indexintegerFor a figure, its place among the page's
[FIGURE]markers, from 0;nullfor a whole page. -
heightintegerHeight in pixels. The long side is at most 1568.
-
kindfigure | pagefigurefor a picture cropped from a page,pagefor a whole page. -
media_typestringimage/png(figures) orimage/jpeg(pages). -
pageintegerThe page it is from.
-
urlstringWhere to download the image file.
-
widthintegerWidth in pixels.
-
-
key_valuesobjectWith
include=key_values: labelled fields the OCR found, each with itsvalue,key_confidenceandvalue_confidence(0 to 100), andselection_status(SELECTEDorNOT_SELECTEDfor a checkbox, otherwisenull). -
layoutarray of objectWith
include=layout: raw layout elements (headings, paragraphs, tables, figures) with their positions on the page. -
pageintegerPage number, starting at 1.
-
resultobjectWith
extraction.per_page: this page's owndataandcitations; otherwisenull. -
tablesarray of objectWith
include=tables: each table's size and a markdown copy.3 fields
-
columnsinteger -
markdownstring -
rowsinteger
-
-
textstringThe page's text in reading order. Tables are markdown; each figure is a
[FIGURE]marker. -
word_countintegerWords the OCR found on the page.
-
-
purged_atstringWhen the result was deleted.
-
resultExtractionResultWhat the model produced, once
statusiscompleted:nullunless the document was sent with anextractionorqueries. SeeExtractionResult.6 fields
-
answersarray of objectOne per question in
queries, in the order you asked.not_found: true, with an emptyanswer, when the document does not say.5 fields
-
answerstring -
not_foundboolean -
pageinteger -
questionstring -
quotestring
-
-
citationsarray of objectThe citations: one per value in
data.pathpoints intodata(likedata.amounts[0].amount),pageis where it is,quoteis the exact words it came from (or the figure's caption when only a picture shows it) andconfidenceruns from 0 to 1.4 fields
-
confidencenumber -
pageinteger -
pathstring -
quotestring
-
-
dataobjectThe extracted data: the built-in
extraction.fieldsyou named, and the properties of yourextraction.schema. A value the document does not confirm is left out, never guessed. -
pagesarray of objectWith
extraction.per_page: one{page, data, citations}per page. -
per_pagebooleantruewhen it ran withextraction.per_page. Thenpagesreplacesdataandcitations. -
warningsarray of stringAnything you should know about how it was read, such as images left out because the document had more than one request can carry. Absent when there is nothing to say.
-
-
retention_secondsintegerHow long the result is kept after completing (
0= forever). -
retrieved_atstringThe first time you fetched the completed result.
-
source_urlstringThe URL you sent, for URL uploads; otherwise
null. -
stagereceived | fetching | analyzing | rendering | reserving | extracting | embedding | finalizing | doneThe step it is on, in order:
received,fetching(URL uploads only),analyzing(OCR and layout),rendering(pictures of figures and scanned pages),reserving(credit hold re-sized),extracting(the model reads it),embedding(only with theembeddingsoutput),finalizing, thendone. Steps with nothing to do are skipped. -
started_atstringWhen a worker picked it up.
-
statusqueued | processing | completed | failed | expiredWhere it is overall:
queued(waiting for a worker),processing, thencompleted(readresult) orfailed(readerror).expiredmeans retention ran out and the result was deleted. -
textstringWith
include=text: the text the model read, each page under a=== Page N ===line, after any OCR key/value hints. -
timingsobjectMilliseconds spent on each step, e.g.
analyzing_ms. -
usageobjectWith
include=usage: what this document and its queries cost, in total, per SKU and per record.3 fields
-
amount_microinteger -
by_skuobjectPer SKU: the quantity (pages or tokens) and what it cost.
-
recordsarray of UsageRecord11 fields
-
idstring -
amount_microinteger -
created_atstring -
document_idstring -
modelstring -
quantityinteger -
query_idstring -
skustring -
stageanalysis | extraction | query | embedding -
unit_divisorinteger1 for pages, 1000000 for tokens
-
unit_price_microinteger
-
-
-
-
has_morebooleantruewhen there are more: pass the last id asstarting_after.
Errors
-
401Missing, revoked or expired key
-
402Payment required. Either the plan's monthly request quota is spent (
quota_exceeded, trial plans only — pay as you go is never capped) or the workspace is out of credit (insufficient_credits). -
422Body or options failed validation (
detailslists the fields), or the document is too large (document_too_large) -
429Per-second burst limit for the plan exceeded; retry after
Retry-Afterseconds
Every error has the same shape. See Errors.