API · Documents
Pages of a completed document
/api/v1/documents/{id}/pages
Per-page text, word count and (with extraction.per_page) the per-page extraction result.
This is the only endpoint that returns embedding vectors. Each extra is opt-in through include. Vectors exist only for a document sent with embeddings; tables and key/value pairs only with analysis: layout (the default), so a read document returns them empty.
Path parameters
-
idstring requiredThe document's id.
Query parameters
-
includestring optionalComma-separated extras per page:
embeddings(the vector, its model and dimensions; needsembeddingson the document),tablesandkey_values(fromanalysis: layout, the default),layout(raw layout elements with their bounding boxes),images(the page's pictures, each with a downloadurl),image_data(the same pictures with their bytes as base64data; at most 25 MB of pictures a response, so narrow it withpagefor a picture-heavy document, or download each from itsurl).Example:
include=tables,key_values -
pagestring optionalOne page (
3) or an inclusive range (1-3). Anything else is a 422.Example:
page=1-3
Returns 200
The document's pages.
-
objectstringAlways
list. -
dataarray of DocumentPageOne per page, in page order.
9 fields
-
embeddingobjectWith
include=embeddings: the page's vector, the model that made it and its length.3 fields
-
dimensionsinteger -
modelstring -
vectorarray of number
-
-
imagesarray of ImageWith
include=imagesorinclude=image_data: the page's pictures, in their order on the page (withimage_data, bytes included).13 fields
-
idstringThe image's id.
-
objectstringAlways
image. -
bboxobjectFor a figure, where it sits on the page:
left,top,widthandheightas fractions of the page (0 to 1).nullfor a whole page. -
byte_sizeintegerFile size in bytes.
-
captionstringThe figure's caption as the OCR read it, or
null. -
datastringWith
include=image_data(pages) orinclude=data(images): the file itself, base64-encoded. -
figure_indexintegerFor a figure, its place among the page's
[FIGURE]markers, from 0;nullfor a whole page. -
heightintegerHeight in pixels. The long side is at most 1568.
-
kindfigure | pagefigurefor a picture cropped from a page,pagefor a whole page. -
media_typestringimage/png(figures) orimage/jpeg(pages). -
pageintegerThe page it is from.
-
urlstringWhere to download the image file.
-
widthintegerWidth in pixels.
-
-
key_valuesobjectWith
include=key_values: labelled fields the OCR found, each with itsvalue,key_confidenceandvalue_confidence(0 to 100), andselection_status(SELECTEDorNOT_SELECTEDfor a checkbox, otherwisenull). -
layoutarray of objectWith
include=layout: raw layout elements (headings, paragraphs, tables, figures) with their positions on the page. -
pageintegerPage number, starting at 1.
-
resultobjectWith
extraction.per_page: this page's owndataandcitations; otherwisenull. -
tablesarray of objectWith
include=tables: each table's size and a markdown copy.3 fields
-
columnsinteger -
markdownstring -
rowsinteger
-
-
textstringThe page's text in reading order. Tables are markdown; each figure is a
[FIGURE]marker. -
word_countintegerWords the OCR found on the page.
-
-
document_idstringThe document these belong to.
Errors
-
401Missing, revoked or expired key
-
402Payment required. Either the plan's monthly request quota is spent (
quota_exceeded, trial plans only — pay as you go is never capped) or the workspace is out of credit (insufficient_credits). -
404No document or query with that id for this key (documents are scoped to the key's workspace and live/test mode)
-
409The document has not finished processing
-
410Retention ran out and the result was purged
-
422Body or options failed validation (
detailslists the fields), or the document is too large (document_too_large) -
429Per-second burst limit for the plan exceeded; retry after
Retry-Afterseconds
Every error has the same shape. See Errors.