Documentation menu

API · Documents

Pages of a completed document

GET /api/v1/documents/{id}/pages

Per-page text, word count and (with extraction.per_page) the per-page extraction result.

This is the only endpoint that returns embedding vectors. Each extra is opt-in through include. Vectors exist only for a document sent with embeddings; tables and key/value pairs only with analysis: layout (the default), so a read document returns them empty.

Path parameters

  • id string required

    The document's id.

Query parameters

  • include string optional

    Comma-separated extras per page: embeddings (the vector, its model and dimensions; needs embeddings on the document), tables and key_values (from analysis: layout, the default), layout (raw layout elements with their bounding boxes), images (the page's pictures, each with a download url), image_data (the same pictures with their bytes as base64 data; at most 25 MB of pictures a response, so narrow it with page for a picture-heavy document, or download each from its url).

    Example: include=tables,key_values

  • page string optional

    One page (3) or an inclusive range (1-3). Anything else is a 422.

    Example: page=1-3

Returns 200

The document's pages.

  • object string

    Always list.

  • data array of DocumentPage

    One per page, in page order.

    9 fields
    • embedding object

      With include=embeddings: the page's vector, the model that made it and its length.

      3 fields
      • dimensions integer
      • model string
      • vector array of number
    • images array of Image

      With include=images or include=image_data: the page's pictures, in their order on the page (with image_data, bytes included).

      13 fields
      • id string

        The image's id.

      • object string

        Always image.

      • bbox object

        For a figure, where it sits on the page: left, top, width and height as fractions of the page (0 to 1). null for a whole page.

      • byte_size integer

        File size in bytes.

      • caption string

        The figure's caption as the OCR read it, or null.

      • data string

        With include=image_data (pages) or include=data (images): the file itself, base64-encoded.

      • figure_index integer

        For a figure, its place among the page's [FIGURE] markers, from 0; null for a whole page.

      • height integer

        Height in pixels. The long side is at most 1568.

      • kind figure | page

        figure for a picture cropped from a page, page for a whole page.

      • media_type string

        image/png (figures) or image/jpeg (pages).

      • page integer

        The page it is from.

      • url string

        Where to download the image file.

      • width integer

        Width in pixels.

    • key_values object

      With include=key_values: labelled fields the OCR found, each with its value, key_confidence and value_confidence (0 to 100), and selection_status (SELECTED or NOT_SELECTED for a checkbox, otherwise null).

    • layout array of object

      With include=layout: raw layout elements (headings, paragraphs, tables, figures) with their positions on the page.

    • page integer

      Page number, starting at 1.

    • result object

      With extraction.per_page: this page's own data and citations; otherwise null.

    • tables array of object

      With include=tables: each table's size and a markdown copy.

      3 fields
      • columns integer
      • markdown string
      • rows integer
    • text string

      The page's text in reading order. Tables are markdown; each figure is a [FIGURE] marker.

    • word_count integer

      Words the OCR found on the page.

  • document_id string

    The document these belong to.

Errors

  • 401

    Missing, revoked or expired key

  • 402

    Payment required. Either the plan's monthly request quota is spent (quota_exceeded, trial plans only — pay as you go is never capped) or the workspace is out of credit (insufficient_credits).

  • 404

    No document or query with that id for this key (documents are scoped to the key's workspace and live/test mode)

  • 409

    The document has not finished processing

  • 410

    Retention ran out and the result was purged

  • 422

    Body or options failed validation (details lists the fields), or the document is too large (document_too_large)

  • 429

    Per-second burst limit for the plan exceeded; retry after Retry-After seconds

Every error has the same shape. See Errors.