Documentation menu

Guides

Images: charts, scans and signatures

How Northdoc keeps a document's pictures, shows them to its models, and returns them to you.

In short: OCR only sees words, so Northdoc also keeps the pictures on each page, for every document. You can download them, and when the model reads the document it sees them in the right place, so it can read a chart, a stamp or a signature.

  • Figures are cut out

    A chart, diagram or photo on a page of text is cropped. When the model reads the document, it is put back exactly where it sits, followed by its caption.

  • Scans go in whole

    A page that is mostly pictures (half or more), or has fewer than 25 words, is kept as one whole-page picture. That covers scans, signature pages, stamps and handwriting.

  • Clutter is skipped

    Figures smaller than 1.5% of the page (logos, bullets, lines) and truly blank pages are left out.

  • Words stay the source

    Quotes come from the OCR text. When only a picture shows a value, the quote is the figure's caption or a short note of where it is. You can download every picture too.

Here is what the model is given for a page with a chart on it. The chart's picture replaces the [FIGURE] marker the OCR left in the text:

what the model reads
=== Page 3 ===
Revenue grew in every quarter of FY2026, led by…
[picture: the chart, cropped from the page]
Figure (page 3): Quarterly revenue, FY2026
Operating costs fell 4% as…
Default On for every document (images: true). Send false to take none.
File types PDFs are cropped and rendered page by page. A PNG or JPEG upload (under 3.75 MB and 8000px) is kept whole, in its own format and size, when it is mostly picture or has fewer than 25 words; no figures are cut from it. TIFF, BMP, HEIF and Office files have no pictures taken.
Picture size From a PDF, at most 1568px on the long side, the size the model reads at: figures are PNG and whole pages JPEG.
Getting them GET /documents/{id}/images lists them with a download url each; include=data (or image_data on the pages endpoint) puts the bytes inline, at most 25 MB of them a response, so page through a picture-heavy document with page.
Limits One model request carries at most 20 pictures and 10 MB of them. A bigger document keeps the largest pictures, marks the rest in the text, and says how many were left out in result.warnings.
Cost Taking them is covered by the page price. When the model runs (for an extraction or questions) it reads them as input tokens: about width × height ÷ 750 each, at most about 1,600. A full-page picture costs about US$0.0044 on Swift or US$0.013 on Summit, each time the model reads it.
Citations Quotes still come from the OCR text. A value only a picture shows is cited by its page, with the caption or a note of where it is.
Getting them GET /documents/{id}/pages?include=image_data returns each page with its pictures, bytes included as base64. GET /documents/{id}/images lists them on their own, and each one's url returns the PNG or JPEG file. The first [FIGURE] in a page's text is that page's figure_index: 0.
Afterwards Pictures are kept with the document, so later questions see them too, and deleted with it. The file you sent is deleted once they are made.
If it fails If a file cannot be turned into pictures, it is kept as text alone rather than failing.

See it working: a number only a chart shows and a signed, stamped scan.