Documentation menu

API · Documents

List documents

GET /api/v1/documents

Newest first. Page with starting_after=<id>; filter with status.

Query parameters

  • limit integer optional default 20

    How many to return.

    Example: limit=50

  • starting_after string optional

    Id of the last document on the previous page; returns the ones after it.

    Example: starting_after=8f1c2d3e-…

  • status queued | processing | completed | failed | expired optional

    Only documents in this state.

    Example: status=completed

Returns 200

A page of documents.

  • object string

    Always list.

  • data array of Document

    Documents, newest first. result is left out of lists.

    29 fields
    • id string

      The document's id. Keep it: every other call needs it.

    • object string

      Always document.

    • analysis layout | read

      The analysis option it ran with.

    • byte_size integer

      File size in bytes.

    • completed_at string

      When it finished.

    • content_type string

      The type detected from the file's bytes, e.g. application/pdf.

    • cost object

      estimated_micro was held from your credit when it was accepted; settled_micro is what it actually cost, once finished.

      2 fields
      • estimated_micro integer

        Held before processing

      • settled_micro integer

        Charged after processing

    • created_at string

      When you submitted it.

    • error object

      Present when status is failed. type is one of analysis_failed, fetch_failed, unsupported_media_type, payload_too_large, document_too_large, schema_mismatch, insufficient_credits, embedding_failed or processing_failed (anything else). stage is the step it failed at. A failed document keeps answering 200 with its error; only an expired one answers 410. You pay only for the work done before it failed.

      3 fields
      • message string
      • stage string
      • type string
    • expires_at string

      When the result will be deleted (completed_at + retention_seconds); null if kept forever.

    • failed_at string

      When it failed.

    • filename string

      The uploaded file's name, or the last part of the URL.

    • markdown string

      With include=markdown and the markdown output: the document as markdown.

    • mode live | test

      test when made with an ak_test_… key: fake results, no charge.

    • model object

      requested is the model option you sent; used is the model and version that actually read it, such as northdoc-swift-1.

      2 fields
      • requested string
      • used string
    • options DocumentOptions

      With include=options: the options it was sent with, defaults filled in.

      8 fields
      • analysis layout | read default layout

        How the pages are read. layout ($0.025/page) also finds tables, key/value pairs, figures and reading order. read ($0.004/page) is OCR only, about six times cheaper: pages come back without tables, key/value pairs or figures, and markdown: true with it is a 422. Example: "analysis": "read".

      • embeddings true | EmbeddingOptions

        Make a vector for every page, for your own search or model. Off unless you send it: true for the defaults (one 1,024-number vector per page), or an object to choose the settings. Billed as compass.input. Read them from GET /documents/{id}/pages?include=embeddings. Example: "embeddings": true.

        2 fields
        • dimensions 1024 | 512 | 256 default 1024

          Vector length. Compass is trained so shorter vectors stay usable: 512 and 256 cut storage and index size at a small accuracy cost. Pick one per corpus — you cannot compare vectors of different lengths. Example: "dimensions": 256.

        • granularity page default page

          page embeds each page as one vector (from up to its first 20,000 bytes of text), returned by GET /documents/{id}/pages?include=embeddings.

      • extraction ExtractionOptions

        Have the model fill result.data: any of Northdoc's built-in fields (summary, parties, dates and more), your own JSON schema, or both. Off unless you send it. Name at least one field or give a schema. Example: "extraction": {"fields": ["summary", "parties"]}.

        4 fields
        • fields array of title | document_type | summary | dates | parties | amounts | key_facts

          Built-in fields, each with a fixed shape in result.data:

          • title — the title or heading, a string.
          • document_type — what kind of document it is (invoice, contract, letter…), a string.
          • summary — two or three sentences on what it says.
          • dates — [{label, date}], dates in ISO 8601 where possible.
          • parties — [{name, role}], the people and organisations involved.
          • amounts — [{label, amount, currency}], the money in it.
          • key_facts — the most important facts, one short sentence each.

          Only the fields you name come back. A value the document does not confirm is left out, never guessed. Example: "fields": ["summary", "parties", "dates"].

        • instructions string

          Free-text guidance for ambiguous fields — which of two totals to take, what to do when a field is missing, how to normalise dates. Example: "instructions": "Use the invoice date, not the print date."

          at most 4000 characters

        • per_page boolean default false

          Run the extraction separately on every page instead of once over the whole document. Each page's result appears on that page (result in GET /documents/{id}/pages) as well as under result.pages. Costs roughly one extraction per page — use it for documents that are a stack of independent records, not for one contract spanning pages. Up to 200 pages; a longer document fails with error.type: document_too_large.

        • schema object

          Your own JSON Schema (type: object) for result.data to follow. Nested objects and arrays work, so line items and tables come back structured. Mark the fields you depend on as required.

      • images boolean default true

        Take pictures of the document: figures (charts, diagrams, photos) are cropped from their pages, and a whole page is kept when pictures cover at least half of it or it has fewer than 25 words (scans, signature and stamp pages, handwriting). Tiny figures (under 1.5% of the page, like logos) and blank pages are skipped. Download them with GET /documents/{id}/images, or with each page from GET /documents/{id}/pages?include=image_data. Covered by the page price.

        When the model runs it reads the pictures where they sit in the text. They are billed as input tokens then (about width × height / 750 each, at most ~1,600); one request carries at most 20 images and 10 MB, and past that the largest are kept and result.warnings says how many were left out.

        Works on PDFs, and on PNG and JPEG uploads (under 3.75 MB and 8000px, kept as the page image). TIFF, BMP, HEIF and Office files have no pictures taken. false takes none. Example: "images": false.

      • markdown boolean default false

        Also produce a markdown rendering of the whole document, read with GET /documents/{id}?include=markdown. Needs analysis: layout. Same page price. Example: "markdown": true.

      • model auto | swift | summit default auto

        Which Northdoc model reads the document, when one runs (for an extraction or queries).

        • swift — fast and economical, for everyday documents up to about 180k tokens (roughly 300 pages of dense text).
        • summit — the most capable, for long or dense legal and financial documents, up to about 900k tokens.
        • auto (the default) — swift up to about 150k tokens, summit beyond that.

        Pin summit when accuracy on hard documents matters most; pin swift to cap cost. A document too long for swift is accepted, read and charged for its pages, then fails at extracting with error.type: document_too_large; auto moves to summit instead. model.used on the document says which one ran (northdoc-swift-1 or northdoc-summit-1).

        Swift and Summit are kept current: when a better model becomes available the name moves to it and the version in model.used goes up (northdoc-swift-2), with no change to the API. Example: "model": "summit".

      • queries array of string

        Questions the model answers while the document is processed, returned in order under result.answers. Each question once: the same question twice is a 422. Cheaper than one POST /documents/{id}/queries per question because the document is read once for all of them. Use the queries endpoint for follow-ups you only think of later. Example: "queries": ["What is the total?", "When is it due?"].

        at most 50 items

      • retention_seconds integer default 86400

        How long results are kept once the document completes: expires_at is completed_at plus this many seconds. A result you have never fetched is kept for at least 24 hours after it completes, however short this is, so a short retention never deletes a result nobody has read. Reading a result does not move expires_at.

        0 means keep indefinitely, which only plans with limits.max_retention_seconds == 0 may do; on capped plans 0 is clamped to the plan limit and anything larger is a 422. Read the ceiling from GET /account. After it lapses the document returns 410 document_expired. Example: "retention_seconds": 3600.

        at least 0

    • page_count integer

      Number of pages (the real count once analysed).

    • pages array of DocumentPage

      With include=pages: one object per page.

      9 fields
      • embedding object

        With include=embeddings: the page's vector, the model that made it and its length.

        3 fields
        • dimensions integer
        • model string
        • vector array of number
      • images array of Image

        With include=images or include=image_data: the page's pictures, in their order on the page (with image_data, bytes included).

        13 fields
        • id string

          The image's id.

        • object string

          Always image.

        • bbox object

          For a figure, where it sits on the page: left, top, width and height as fractions of the page (0 to 1). null for a whole page.

        • byte_size integer

          File size in bytes.

        • caption string

          The figure's caption as the OCR read it, or null.

        • data string

          With include=image_data (pages) or include=data (images): the file itself, base64-encoded.

        • figure_index integer

          For a figure, its place among the page's [FIGURE] markers, from 0; null for a whole page.

        • height integer

          Height in pixels. The long side is at most 1568.

        • kind figure | page

          figure for a picture cropped from a page, page for a whole page.

        • media_type string

          image/png (figures) or image/jpeg (pages).

        • page integer

          The page it is from.

        • url string

          Where to download the image file.

        • width integer

          Width in pixels.

      • key_values object

        With include=key_values: labelled fields the OCR found, each with its value, key_confidence and value_confidence (0 to 100), and selection_status (SELECTED or NOT_SELECTED for a checkbox, otherwise null).

      • layout array of object

        With include=layout: raw layout elements (headings, paragraphs, tables, figures) with their positions on the page.

      • page integer

        Page number, starting at 1.

      • result object

        With extraction.per_page: this page's own data and citations; otherwise null.

      • tables array of object

        With include=tables: each table's size and a markdown copy.

        3 fields
        • columns integer
        • markdown string
        • rows integer
      • text string

        The page's text in reading order. Tables are markdown; each figure is a [FIGURE] marker.

      • word_count integer

        Words the OCR found on the page.

    • purged_at string

      When the result was deleted.

    • result ExtractionResult

      What the model produced, once status is completed: null unless the document was sent with an extraction or queries. See ExtractionResult.

      6 fields
      • answers array of object

        One per question in queries, in the order you asked. not_found: true, with an empty answer, when the document does not say.

        5 fields
        • answer string
        • not_found boolean
        • page integer
        • question string
        • quote string
      • citations array of object

        The citations: one per value in data. path points into data (like data.amounts[0].amount), page is where it is, quote is the exact words it came from (or the figure's caption when only a picture shows it) and confidence runs from 0 to 1.

        4 fields
        • confidence number
        • page integer
        • path string
        • quote string
      • data object

        The extracted data: the built-in extraction.fields you named, and the properties of your extraction.schema. A value the document does not confirm is left out, never guessed.

      • pages array of object

        With extraction.per_page: one {page, data, citations} per page.

      • per_page boolean

        true when it ran with extraction.per_page. Then pages replaces data and citations.

      • warnings array of string

        Anything you should know about how it was read, such as images left out because the document had more than one request can carry. Absent when there is nothing to say.

    • retention_seconds integer

      How long the result is kept after completing (0 = forever).

    • retrieved_at string

      The first time you fetched the completed result.

    • source_url string

      The URL you sent, for URL uploads; otherwise null.

    • stage received | fetching | analyzing | rendering | reserving | extracting | embedding | finalizing | done

      The step it is on, in order: received, fetching (URL uploads only), analyzing (OCR and layout), rendering (pictures of figures and scanned pages), reserving (credit hold re-sized), extracting (the model reads it), embedding (only with the embeddings output), finalizing, then done. Steps with nothing to do are skipped.

    • started_at string

      When a worker picked it up.

    • status queued | processing | completed | failed | expired

      Where it is overall: queued (waiting for a worker), processing, then completed (read result) or failed (read error). expired means retention ran out and the result was deleted.

    • text string

      With include=text: the text the model read, each page under a === Page N === line, after any OCR key/value hints.

    • timings object

      Milliseconds spent on each step, e.g. analyzing_ms.

    • usage object

      With include=usage: what this document and its queries cost, in total, per SKU and per record.

      3 fields
      • amount_micro integer
      • by_sku object

        Per SKU: the quantity (pages or tokens) and what it cost.

      • records array of UsageRecord
        11 fields
        • id string
        • amount_micro integer
        • created_at string
        • document_id string
        • model string
        • quantity integer
        • query_id string
        • sku string
        • stage analysis | extraction | query | embedding
        • unit_divisor integer

          1 for pages, 1000000 for tokens

        • unit_price_micro integer
  • has_more boolean

    true when there are more: pass the last id as starting_after.

Errors

  • 401

    Missing, revoked or expired key

  • 402

    Payment required. Either the plan's monthly request quota is spent (quota_exceeded, trial plans only — pay as you go is never capped) or the workspace is out of credit (insufficient_credits).

  • 422

    Body or options failed validation (details lists the fields), or the document is too large (document_too_large)

  • 429

    Per-second burst limit for the plan exceeded; retry after Retry-After seconds

Every error has the same shape. See Errors.