Documentation menu

API · Documents

Get a document and its result

GET /api/v1/documents/{id}

Poll this until status is completed or failed. result appears once completed, error once failed.

By default the response is metadata plus result, which is null unless the document was sent with an extraction or queries. The bulky parts are opt-in through include, and it only shows data that was produced: markdown exists only for a document sent with markdown: true, and tables and key_values only with analysis: layout (the default). Embeddings are the one extra that is not available here; read them from GET /documents/{id}/pages?include=embeddings.

The first read of a completed document sets retrieved_at. A result nobody has fetched is kept at least 24 hours after it completes, however short its retention_seconds.

Path parameters

  • id string required

    The document id from the submit response.

Query parameters

  • include string optional

    Comma-separated extras to embed in the response: text (the stitched plain text sent to the model), markdown (needs markdown: true on the document), pages (the per-page array), tables and key_values (added to each page, so pair them with pages), usage (priced records for this document), options (the options the document was submitted with).

    Example: include=text,pages,usage

Returns 200

The document (result is present once completed; error once failed)

  • id string

    The document's id. Keep it: every other call needs it.

  • object string

    Always document.

  • analysis layout | read

    The analysis option it ran with.

  • byte_size integer

    File size in bytes.

  • completed_at string

    When it finished.

  • content_type string

    The type detected from the file's bytes, e.g. application/pdf.

  • cost object

    estimated_micro was held from your credit when it was accepted; settled_micro is what it actually cost, once finished.

    2 fields
    • estimated_micro integer

      Held before processing

    • settled_micro integer

      Charged after processing

  • created_at string

    When you submitted it.

  • error object

    Present when status is failed. type is one of analysis_failed, fetch_failed, unsupported_media_type, payload_too_large, document_too_large, schema_mismatch, insufficient_credits, embedding_failed or processing_failed (anything else). stage is the step it failed at. A failed document keeps answering 200 with its error; only an expired one answers 410. You pay only for the work done before it failed.

    3 fields
    • message string
    • stage string
    • type string
  • expires_at string

    When the result will be deleted (completed_at + retention_seconds); null if kept forever.

  • failed_at string

    When it failed.

  • filename string

    The uploaded file's name, or the last part of the URL.

  • markdown string

    With include=markdown and the markdown output: the document as markdown.

  • mode live | test

    test when made with an ak_test_… key: fake results, no charge.

  • model object

    requested is the model option you sent; used is the model and version that actually read it, such as northdoc-swift-1.

    2 fields
    • requested string
    • used string
  • options DocumentOptions

    With include=options: the options it was sent with, defaults filled in.

    8 fields
    • analysis layout | read default layout

      How the pages are read. layout ($0.025/page) also finds tables, key/value pairs, figures and reading order. read ($0.004/page) is OCR only, about six times cheaper: pages come back without tables, key/value pairs or figures, and markdown: true with it is a 422. Example: "analysis": "read".

    • embeddings true | EmbeddingOptions

      Make a vector for every page, for your own search or model. Off unless you send it: true for the defaults (one 1,024-number vector per page), or an object to choose the settings. Billed as compass.input. Read them from GET /documents/{id}/pages?include=embeddings. Example: "embeddings": true.

      2 fields
      • dimensions 1024 | 512 | 256 default 1024

        Vector length. Compass is trained so shorter vectors stay usable: 512 and 256 cut storage and index size at a small accuracy cost. Pick one per corpus — you cannot compare vectors of different lengths. Example: "dimensions": 256.

      • granularity page default page

        page embeds each page as one vector (from up to its first 20,000 bytes of text), returned by GET /documents/{id}/pages?include=embeddings.

    • extraction ExtractionOptions

      Have the model fill result.data: any of Northdoc's built-in fields (summary, parties, dates and more), your own JSON schema, or both. Off unless you send it. Name at least one field or give a schema. Example: "extraction": {"fields": ["summary", "parties"]}.

      4 fields
      • fields array of title | document_type | summary | dates | parties | amounts | key_facts

        Built-in fields, each with a fixed shape in result.data:

        • title — the title or heading, a string.
        • document_type — what kind of document it is (invoice, contract, letter…), a string.
        • summary — two or three sentences on what it says.
        • dates — [{label, date}], dates in ISO 8601 where possible.
        • parties — [{name, role}], the people and organisations involved.
        • amounts — [{label, amount, currency}], the money in it.
        • key_facts — the most important facts, one short sentence each.

        Only the fields you name come back. A value the document does not confirm is left out, never guessed. Example: "fields": ["summary", "parties", "dates"].

      • instructions string

        Free-text guidance for ambiguous fields — which of two totals to take, what to do when a field is missing, how to normalise dates. Example: "instructions": "Use the invoice date, not the print date."

        at most 4000 characters

      • per_page boolean default false

        Run the extraction separately on every page instead of once over the whole document. Each page's result appears on that page (result in GET /documents/{id}/pages) as well as under result.pages. Costs roughly one extraction per page — use it for documents that are a stack of independent records, not for one contract spanning pages. Up to 200 pages; a longer document fails with error.type: document_too_large.

      • schema object

        Your own JSON Schema (type: object) for result.data to follow. Nested objects and arrays work, so line items and tables come back structured. Mark the fields you depend on as required.

    • images boolean default true

      Take pictures of the document: figures (charts, diagrams, photos) are cropped from their pages, and a whole page is kept when pictures cover at least half of it or it has fewer than 25 words (scans, signature and stamp pages, handwriting). Tiny figures (under 1.5% of the page, like logos) and blank pages are skipped. Download them with GET /documents/{id}/images, or with each page from GET /documents/{id}/pages?include=image_data. Covered by the page price.

      When the model runs it reads the pictures where they sit in the text. They are billed as input tokens then (about width × height / 750 each, at most ~1,600); one request carries at most 20 images and 10 MB, and past that the largest are kept and result.warnings says how many were left out.

      Works on PDFs, and on PNG and JPEG uploads (under 3.75 MB and 8000px, kept as the page image). TIFF, BMP, HEIF and Office files have no pictures taken. false takes none. Example: "images": false.

    • markdown boolean default false

      Also produce a markdown rendering of the whole document, read with GET /documents/{id}?include=markdown. Needs analysis: layout. Same page price. Example: "markdown": true.

    • model auto | swift | summit default auto

      Which Northdoc model reads the document, when one runs (for an extraction or queries).

      • swift — fast and economical, for everyday documents up to about 180k tokens (roughly 300 pages of dense text).
      • summit — the most capable, for long or dense legal and financial documents, up to about 900k tokens.
      • auto (the default) — swift up to about 150k tokens, summit beyond that.

      Pin summit when accuracy on hard documents matters most; pin swift to cap cost. A document too long for swift is accepted, read and charged for its pages, then fails at extracting with error.type: document_too_large; auto moves to summit instead. model.used on the document says which one ran (northdoc-swift-1 or northdoc-summit-1).

      Swift and Summit are kept current: when a better model becomes available the name moves to it and the version in model.used goes up (northdoc-swift-2), with no change to the API. Example: "model": "summit".

    • queries array of string

      Questions the model answers while the document is processed, returned in order under result.answers. Each question once: the same question twice is a 422. Cheaper than one POST /documents/{id}/queries per question because the document is read once for all of them. Use the queries endpoint for follow-ups you only think of later. Example: "queries": ["What is the total?", "When is it due?"].

      at most 50 items

    • retention_seconds integer default 86400

      How long results are kept once the document completes: expires_at is completed_at plus this many seconds. A result you have never fetched is kept for at least 24 hours after it completes, however short this is, so a short retention never deletes a result nobody has read. Reading a result does not move expires_at.

      0 means keep indefinitely, which only plans with limits.max_retention_seconds == 0 may do; on capped plans 0 is clamped to the plan limit and anything larger is a 422. Read the ceiling from GET /account. After it lapses the document returns 410 document_expired. Example: "retention_seconds": 3600.

      at least 0

  • page_count integer

    Number of pages (the real count once analysed).

  • pages array of DocumentPage

    With include=pages: one object per page.

    9 fields
    • embedding object

      With include=embeddings: the page's vector, the model that made it and its length.

      3 fields
      • dimensions integer
      • model string
      • vector array of number
    • images array of Image

      With include=images or include=image_data: the page's pictures, in their order on the page (with image_data, bytes included).

      13 fields
      • id string

        The image's id.

      • object string

        Always image.

      • bbox object

        For a figure, where it sits on the page: left, top, width and height as fractions of the page (0 to 1). null for a whole page.

      • byte_size integer

        File size in bytes.

      • caption string

        The figure's caption as the OCR read it, or null.

      • data string

        With include=image_data (pages) or include=data (images): the file itself, base64-encoded.

      • figure_index integer

        For a figure, its place among the page's [FIGURE] markers, from 0; null for a whole page.

      • height integer

        Height in pixels. The long side is at most 1568.

      • kind figure | page

        figure for a picture cropped from a page, page for a whole page.

      • media_type string

        image/png (figures) or image/jpeg (pages).

      • page integer

        The page it is from.

      • url string

        Where to download the image file.

      • width integer

        Width in pixels.

    • key_values object

      With include=key_values: labelled fields the OCR found, each with its value, key_confidence and value_confidence (0 to 100), and selection_status (SELECTED or NOT_SELECTED for a checkbox, otherwise null).

    • layout array of object

      With include=layout: raw layout elements (headings, paragraphs, tables, figures) with their positions on the page.

    • page integer

      Page number, starting at 1.

    • result object

      With extraction.per_page: this page's own data and citations; otherwise null.

    • tables array of object

      With include=tables: each table's size and a markdown copy.

      3 fields
      • columns integer
      • markdown string
      • rows integer
    • text string

      The page's text in reading order. Tables are markdown; each figure is a [FIGURE] marker.

    • word_count integer

      Words the OCR found on the page.

  • purged_at string

    When the result was deleted.

  • result ExtractionResult

    What the model produced, once status is completed: null unless the document was sent with an extraction or queries. See ExtractionResult.

    6 fields
    • answers array of object

      One per question in queries, in the order you asked. not_found: true, with an empty answer, when the document does not say.

      5 fields
      • answer string
      • not_found boolean
      • page integer
      • question string
      • quote string
    • citations array of object

      The citations: one per value in data. path points into data (like data.amounts[0].amount), page is where it is, quote is the exact words it came from (or the figure's caption when only a picture shows it) and confidence runs from 0 to 1.

      4 fields
      • confidence number
      • page integer
      • path string
      • quote string
    • data object

      The extracted data: the built-in extraction.fields you named, and the properties of your extraction.schema. A value the document does not confirm is left out, never guessed.

    • pages array of object

      With extraction.per_page: one {page, data, citations} per page.

    • per_page boolean

      true when it ran with extraction.per_page. Then pages replaces data and citations.

    • warnings array of string

      Anything you should know about how it was read, such as images left out because the document had more than one request can carry. Absent when there is nothing to say.

  • retention_seconds integer

    How long the result is kept after completing (0 = forever).

  • retrieved_at string

    The first time you fetched the completed result.

  • source_url string

    The URL you sent, for URL uploads; otherwise null.

  • stage received | fetching | analyzing | rendering | reserving | extracting | embedding | finalizing | done

    The step it is on, in order: received, fetching (URL uploads only), analyzing (OCR and layout), rendering (pictures of figures and scanned pages), reserving (credit hold re-sized), extracting (the model reads it), embedding (only with the embeddings output), finalizing, then done. Steps with nothing to do are skipped.

  • started_at string

    When a worker picked it up.

  • status queued | processing | completed | failed | expired

    Where it is overall: queued (waiting for a worker), processing, then completed (read result) or failed (read error). expired means retention ran out and the result was deleted.

  • text string

    With include=text: the text the model read, each page under a === Page N === line, after any OCR key/value hints.

  • timings object

    Milliseconds spent on each step, e.g. analyzing_ms.

  • usage object

    With include=usage: what this document and its queries cost, in total, per SKU and per record.

    3 fields
    • amount_micro integer
    • by_sku object

      Per SKU: the quantity (pages or tokens) and what it cost.

    • records array of UsageRecord
      11 fields
      • id string
      • amount_micro integer
      • created_at string
      • document_id string
      • model string
      • quantity integer
      • query_id string
      • sku string
      • stage analysis | extraction | query | embedding
      • unit_divisor integer

        1 for pages, 1000000 for tokens

      • unit_price_micro integer

Errors

  • 401

    Missing, revoked or expired key

  • 402

    Payment required. Either the plan's monthly request quota is spent (quota_exceeded, trial plans only — pay as you go is never capped) or the workspace is out of credit (insufficient_credits).

  • 404

    No document or query with that id for this key (documents are scoped to the key's workspace and live/test mode)

  • 410

    Retention ran out and the result was purged

  • 429

    Per-second burst limit for the plan exceeded; retry after Retry-After seconds

Every error has the same shape. See Errors.