How Northdoc handles pictures

It keeps the pictures, not just the words.

OCR turns a chart into a caption and a signature into nothing at all. So Northdoc cuts the pictures out and keeps them with the text, each in its place. Then either you take the lot (text, vectors and pictures) for your own model, or Northdoc's AI reads it for you.

1 The page you send

Page 3 · Annual report

Quarterly revenue, FY2026

2 What Northdoc keeps, in reading order

  1. === Page 3 ===
  2. Revenue grew in every quarter of FY2026, led by…
  3. The chart itself, as a picture
  4. Figure (page 3): Quarterly revenue, FY2026
  5. Operating costs fell 4% as…

3 What you get back, either way

Take it yourself · pages[2]
{
  "page": 3,
  "text": "Revenue grew… [FIGURE] Operating costs…",
  "embedding": { "vector": [0.0123, …] },
  "images": [{
    "kind": "figure", "figure_index": 0,
    "media_type": "image/png",
    "data": "iVBORw0KGgoAAAANSUhEUg…"
  }]
}
Or ask Northdoc's AI · result.answers[0]
{
  "question": "What was revenue in Q3?",
  "answer": "$4.2 million", "page": 3
}
  • Figures are cut out

    A chart, diagram or photo on a page of text is cropped. When the model reads the document, it is put back exactly where it sits, followed by its caption.

  • Scans go in whole

    A page that is mostly pictures (half or more), or has fewer than 25 words, is kept as one whole-page picture. That covers scans, signature pages, stamps and handwriting.

  • Clutter is skipped

    Figures smaller than 1.5% of the page (logos, bullets, lines) and truly blank pages are left out.

  • Words stay the source

    Quotes come from the OCR text. When only a picture shows a value, the quote is the figure's caption or a short note of where it is. You can download every picture too.

On by default for PDFs, PNGs and JPEGs. The first [FIGURE] in a page's text is that page's figure_index: 0, so every picture goes back where it was.

Get them yourself The full image rules

Use cases

Built for the documents that matter.

The same three calls cover all of these. What changes is the schema you send and the questions you ask.

Worked examples

Real requests, real responses.

Every example uses two shell variables. Set them once and paste away:

terminal
export NORTHDOC_API="https://northdoc.northcape.tech/api/v1"
export NORTHDOC_KEY="ak_test_…"   # from your workspace, under API keys

Responses are trimmed to the parts each example is about, and ids are shortened with … Test keys return made-up results for free; live keys read your document for real.

Example 01

Get the text, vectors and pictures, without the AI.

No model at all. Send a document and get back every page's text, a vector per page and the document's pictures, to use in your own system.

The document

site-report.pdf: a 12-page building inspection report with photos of defects, a floor plan and a signed last page.

embeddings: true GET /pages GET /images

What happened

  • No AI reads the document: the model only runs when you ask for an extraction or queries. This one cost US$0.30, which is 12 pages of layout analysis plus a fraction of a cent for the vectors. Leave out embeddings too and you pay for the pages alone.
  • One call returns every page with its text, its vector and its pictures. image_data puts each picture's bytes in data, base64-encoded: decode it and you have the PNG or JPEG file.
  • Each page's text keeps its reading order, with tables as markdown and a [FIGURE] marker where each picture sits. The first [FIGURE] on page 3 is the picture with figure_index: 0, so you can put every picture back exactly where it was.
  • You get the photos, plans and charts cropped from their pages, and pages that are mostly picture or nearly wordless (like the signed page 12) whole. Logos and blank pages are skipped.
  • Base64 makes a response about a third bigger than the pictures themselves, so for long documents fetch a few pages at a time with page, or use include=images (no bytes) and download each picture from its url.
  • Nothing stops you asking the document a question later: POST /documents/{id}/queries still works on it.

You send

request
curl -X POST "$NORTHDOC_API/documents" \
  -H "Authorization: Bearer $NORTHDOC_KEY" \
  -F file=@site-report.pdf \
  -F 'options={"embeddings": true}'

You get back

response
{
  "id": "8f1c2d3e-…",
  "object": "document",
  "status": "queued",
  "estimated_cost_micro": 360864,
  "poll_url": "https://northdoc.northcape.tech/api/v1/documents/8f1c2d3e-…"
}

Then

request
# Once status is "completed": every page's text, vector and pictures, in one call
curl "$NORTHDOC_API/documents/$DOC_ID/pages?include=embeddings,image_data" \
  -H "Authorization: Bearer $NORTHDOC_KEY"

You get back

response
{
  "object": "list",
  "document_id": "8f1c2d3e-…",
  "data": [
    {
      "page": 3,
      "text": "## Exterior\n\nCracking is visible along the north wall.\n\n[FIGURE]\n\nThe crack runs from the window head to the slab.",
      "word_count": 214,
      "result": null,
      "embedding": {"model": "northdoc-compass-1", "dimensions": 1024,
                    "vector": [0.0123, -0.0456, 0.0789, "… 1021 more"]},
      "images": [
        {"id": "5c1e…", "object": "image", "page": 3, "kind": "figure", "figure_index": 0,
         "caption": "Photo 2: Cracking to the north wall",
         "media_type": "image/png", "width": 1240, "height": 930, "byte_size": 412880,
         "bbox": {"left": 0.12, "top": 0.38, "width": 0.76, "height": 0.41},
         "url": "https://northdoc.northcape.tech/api/v1/documents/8f1c2d3e-…/images/5c1e…",
         "data": "iVBORw0KGgoAAAANSUhEUgAABNgAAAOi…"}
      ]
    },
    {
      "page": 12,
      "text": "Inspector's declaration\n\nSigned:\n\nDate:",
      "word_count": 6,
      "result": null,
      "embedding": {"model": "northdoc-compass-1", "dimensions": 1024,
                    "vector": [-0.0311, 0.0082, 0.0457, "… 1021 more"]},
      "images": [
        {"id": "7a90…", "object": "image", "page": 12, "kind": "page", "figure_index": null,
         "caption": null,
         "media_type": "image/jpeg", "width": 1109, "height": 1568, "byte_size": 286411,
         "bbox": null,
         "url": "https://northdoc.northcape.tech/api/v1/documents/8f1c2d3e-…/images/7a90…",
         "data": "/9j/4AAQSkZJRgABAQAAAQABAAD…"}
      ]
    }
  ]
}

Then

request
# Or, for big documents: page through a few pages at a time,
# or list the pictures without their bytes and download each one
curl "$NORTHDOC_API/documents/$DOC_ID/pages?include=embeddings,image_data&page=1-10" \
  -H "Authorization: Bearer $NORTHDOC_KEY"

curl "$NORTHDOC_API/documents/$DOC_ID/images" \
  -H "Authorization: Bearer $NORTHDOC_KEY"

curl -o page-3-figure-0.png "$NORTHDOC_API/documents/$DOC_ID/images/5c1e…" \
  -H "Authorization: Bearer $NORTHDOC_KEY"

You get back

response
{
  "object": "list",
  "document_id": "8f1c2d3e-…",
  "data": [
    {"id": "5c1e…", "object": "image", "page": 3, "kind": "figure", "figure_index": 0,
     "caption": "Photo 2: Cracking to the north wall",
     "media_type": "image/png", "width": 1240, "height": 930, "byte_size": 412880,
     "bbox": {"left": 0.12, "top": 0.38, "width": 0.76, "height": 0.41},
     "url": "https://northdoc.northcape.tech/api/v1/documents/8f1c2d3e-…/images/5c1e…"},
    {"id": "7a90…", "object": "image", "page": 12, "kind": "page", "figure_index": null,
     "caption": null,
     "media_type": "image/jpeg", "width": 1109, "height": 1568, "byte_size": 286411,
     "bbox": null,
     "url": "https://northdoc.northcape.tech/api/v1/documents/8f1c2d3e-…/images/7a90…"}
  ]
}

Example 02

The whole flow in Python.

A complete script: send a PDF, wait for it, get every page's text, vector and pictures, save the pictures, and lay the document out for your own vision model.

The document

Any PDF. The script takes its path as an argument.

embeddings: true include=embeddings,image_data requests

What happened

  • ingest() is the three calls: send the PDF with embeddings: true, poll until it is completed, then fetch the pages with include=embeddings,image_data. No AI model runs.
  • save_pictures() decodes each picture's base64 data into a file: page-3-figure-0.png for a figure, page-12.jpg for a whole page.
  • as_content() rebuilds the document in reading order for a vision model: each page's heading, its whole-page picture if it has one, then its text with every [FIGURE] swapped for the picture itself. The blocks are text and base64 image parts, the shape most vision-model APIs accept; adjust the field names to suit yours.
  • The vectors are page["embedding"]["vector"], one per page with text, ready for your search index. Store each with the document id and page number.
  • We ran this script against the API before publishing it. It needs only requests.

The script

ingest.py
"""Ingest a PDF with Northdoc: every page's text, a vector per page and the
document's pictures, with no AI in between. Then save the pictures, and lay
the whole document out as one message your own model can read."""

import base64
import json
import os
import pathlib
import sys
import time

import requests

API = os.environ["NORTHDOC_API"]  # https://northdoc.northcape.tech/api/v1
AUTH = {"Authorization": f"Bearer {os.environ['NORTHDOC_KEY']}"}


def ingest(path):
    # 1. Send the PDF. Every page is read and pictured anyway; "embeddings"
    #    adds a vector per page. No AI model runs: nothing is asked of one.
    options = {"embeddings": True}
    with open(path, "rb") as f:
        r = requests.post(
            f"{API}/documents",
            headers=AUTH,
            files={"file": f},
            data={"options": json.dumps(options)},
        )
    r.raise_for_status()
    url = r.json()["poll_url"]

    # 2. Check every two seconds until it is done.
    while True:
        doc = requests.get(url, headers=AUTH).json()
        if doc["status"] == "completed":
            break
        if doc["status"] == "failed":
            raise RuntimeError(doc["error"])
        time.sleep(2)

    # 3. Every page with its text, its vector and its pictures, bytes included,
    #    ten pages at a time: one response carries at most 25 MB of pictures.
    pages = []
    for first in range(1, doc["page_count"] + 1, 10):
        params = {"include": "embeddings,image_data", "page": f"{first}-{first + 9}"}
        r = requests.get(f"{url}/pages", headers=AUTH, params=params)
        r.raise_for_status()
        pages += r.json()["data"]
    return pages


def save_pictures(pages, folder="pictures"):
    out = pathlib.Path(folder)
    out.mkdir(exist_ok=True)
    for page in pages:
        for image in page["images"]:
            ext = "png" if image["media_type"] == "image/png" else "jpg"
            if image["kind"] == "page":
                name = f"page-{page['page']}.{ext}"
            else:
                name = f"page-{page['page']}-figure-{image['figure_index']}.{ext}"
            (out / name).write_bytes(base64.b64decode(image["data"]))


def as_content(pages):
    """The document as one list of text and image blocks in reading order: a
    whole-page picture straight after its page heading, each figure where its
    [FIGURE] marker was. Text and base64 image parts are the shape most
    vision-model APIs accept; adjust the field names to suit yours."""
    blocks = []
    for page in pages:
        blocks.append({"type": "text", "text": f"=== Page {page['page']} ==="})
        figures = {}
        for image in page["images"]:
            if image["kind"] == "page":
                blocks.append(image_block(image))
            else:
                figures[image["figure_index"]] = image
        parts = page["text"].split("[FIGURE]")
        for i, part in enumerate(parts):
            if part.strip():
                blocks.append({"type": "text", "text": part.strip()})
            if i < len(parts) - 1 and i in figures:
                blocks.append(image_block(figures[i]))
    return blocks


def image_block(image):
    source = {"type": "base64", "media_type": image["media_type"], "data": image["data"]}
    return {"type": "image", "source": source}


if __name__ == "__main__":
    pages = ingest(sys.argv[1])
    save_pictures(pages)
    vectors = [(page["page"], page["embedding"]["vector"]) for page in pages if page["embedding"]]
    content = as_content(pages)
    pictures = sum(len(page["images"]) for page in pages)
    print(f"{len(pages)} pages, {len(vectors)} vectors, {pictures} pictures, {len(content)} blocks")

What your model gets

as_content(pages)
[
  {"type": "text", "text": "=== Page 3 ==="},
  {"type": "text", "text": "## Exterior\n\nCracking is visible along the north wall."},
  {"type": "image", "source": {"type": "base64", "media_type": "image/png",
                               "data": "iVBORw0KGgoAAAANSUhEUgAABNgAAAOi…"}},
  {"type": "text", "text": "The crack runs from the window head to the slab."},
  {"type": "text", "text": "=== Page 12 ==="},
  {"type": "image", "source": {"type": "base64", "media_type": "image/jpeg",
                               "data": "/9j/4AAQSkZJRgABAQAAAQABAAD…"}},
  {"type": "text", "text": "Inspector's declaration\n\nSigned:\n\nDate:"}
]

Run it

terminal
pip install requests
export NORTHDOC_API="https://northdoc.northcape.tech/api/v1"
export NORTHDOC_KEY="ak_live_…"
python ingest.py site-report.pdf

It prints

output
12 pages, 12 vectors, 9 pictures, 47 blocks

$ ls pictures
page-3-figure-0.png  page-4-figure-0.png  page-4-figure-1.png  page-6-figure-0.png
page-7.jpg  page-8-figure-0.png  page-9-figure-0.png  page-10-figure-0.png  page-12.jpg

Example 03

Read an invoice into your own fields.

Give Northdoc the shape you want and get exactly that back, with a citation for every value.

The document

inv-1042.pdf: a two-page supplier invoice with a table of line items.

extraction.schema extraction.instructions

What happened

  • The first call answers straight away with 202 and an id. The work happens in the background.
  • result.data has exactly the fields in your schema. A value the invoice does not show is left out, never made up.
  • Every value has a line in result.citations: the page it is on and the exact words it came from, so you can check it or highlight it.
  • It cost US$0.061: two pages of layout analysis (US$0.05) plus the model's tokens. The US$0.11 held when it was accepted was settled down to that.

You send

request
curl -X POST "$NORTHDOC_API/documents" \
  -H "Authorization: Bearer $NORTHDOC_KEY" \
  -F file=@inv-1042.pdf \
  -F 'options={
    "extraction": {
      "schema": {
        "type": "object",
        "properties": {
          "invoice_number": {"type": "string"},
          "supplier":       {"type": "string"},
          "due_date":       {"type": "string", "format": "date"},
          "line_items": {
            "type": "array",
            "items": {
              "type": "object",
              "properties": {
                "description": {"type": "string"},
                "quantity":    {"type": "number"},
                "amount":      {"type": "number"}
              }
            }
          },
          "gst":   {"type": "number"},
          "total": {"type": "number"}
        },
        "required": ["invoice_number", "total"]
      },
      "instructions": "total is the amount payable including GST."
    }
  }'

You get back

response
{
  "id": "8f1c2d3e-…",
  "object": "document",
  "status": "queued",
  "estimated_cost_micro": 110843,
  "poll_url": "https://northdoc.northcape.tech/api/v1/documents/8f1c2d3e-…"
}

Then

request
curl "$NORTHDOC_API/documents/8f1c2d3e-…" \
  -H "Authorization: Bearer $NORTHDOC_KEY"

You get back

response
{
  "id": "8f1c2d3e-…",
  "status": "completed",
  "stage": "done",
  "page_count": 2,
  "model": {"requested": "auto", "used": "northdoc-swift-1"},
  "result": {
    "data": {
      "invoice_number": "INV-1042",
      "supplier": "Harbour Office Supplies Pty Ltd",
      "due_date": "2026-10-31",
      "line_items": [
        {"description": "Ergonomic chair", "quantity": 4, "amount": 980.00},
        {"description": "LED desk lamp", "quantity": 4, "amount": 141.82}
      ],
      "gst": 112.18,
      "total": 1234.00
    },
    "citations": [
      {"path": "data.invoice_number", "page": 1, "quote": "Invoice No: INV-1042", "confidence": 0.99},
      {"path": "data.line_items[0].amount", "page": 1, "quote": "Ergonomic chair 4 $245.00 $980.00", "confidence": 0.96},
      {"path": "data.gst", "page": 2, "quote": "GST $112.18", "confidence": 0.97},
      {"path": "data.total", "page": 2, "quote": "Total amount due (inc. GST) $1,234.00", "confidence": 0.97}
    ]
  },
  "cost": {"estimated_micro": 110843, "settled_micro": 61275},
  "completed_at": "2026-10-05T01:12:09Z",
  "expires_at": "2026-10-06T01:12:09Z"
}

Example 04

Get a summary, the parties and the dates.

No schema to write: name the built-in fields you want and get them back, cited.

The document

lease.pdf: a nine-page commercial lease for a shop.

extraction.fields

What happened

  • extraction.fields picks from seven built-in fields: title, document_type, summary, dates, parties, amounts and key_facts. Only the ones you name come back.
  • Each has a fixed shape, so your code can rely on it: parties is always a list of name and role, dates a list of label and date.
  • Need something else too? Add your own schema beside the fields. Its properties land in the same result.data.
  • Without an extraction (or queries) no model runs at all, so you only pay for these when you ask for them.

You send

request
curl -X POST "$NORTHDOC_API/documents" \
  -H "Authorization: Bearer $NORTHDOC_KEY" \
  -F file=@lease.pdf \
  -F 'options={"extraction": {"fields": ["summary", "parties", "dates", "amounts"]}}'

You get back

response
{
  "status": "completed",
  "result": {
    "data": {
      "summary": "A five-year lease of Shop 4, 18 Bay Road to Kettle & Co, at $54,000 a year plus GST, with one five-year option to renew.",
      "parties": [
        {"name": "Bay Road Holdings Pty Ltd", "role": "landlord"},
        {"name": "Kettle & Co Pty Ltd", "role": "tenant"}
      ],
      "dates": [
        {"label": "Commencement date", "date": "2026-11-01"},
        {"label": "Expiry date", "date": "2031-10-31"}
      ],
      "amounts": [
        {"label": "Annual rent", "amount": 54000.0, "currency": "AUD"},
        {"label": "Security deposit", "amount": 13500.0, "currency": "AUD"}
      ]
    },
    "citations": [
      {"path": "data.parties[1].name", "page": 1, "quote": "Kettle & Co Pty Ltd (Tenant)", "confidence": 0.98},
      {"path": "data.dates[0].date", "page": 2, "quote": "Commencement Date: 1 November 2026", "confidence": 0.97},
      {"path": "data.amounts[0].amount", "page": 3, "quote": "Annual Rent: $54,000 plus GST", "confidence": 0.96}
    ]
  }
}

Example 05

Ask a contract questions.

Send your questions with the file and get answers with the page and quote behind each one.

The document

contract-of-sale.pdf: a 38-page contract of sale for a house.

queries POST /queries with a schema

What happened

  • Answers come back under result.answers, in the same order you asked.
  • When the contract does not say, you get not_found: true and an empty answer instead of a guess.
  • A question you think of later goes to the queries endpoint. Give it a schema and the answer is an object you can store, not a sentence you have to parse.
  • Later questions on the same document reuse a cached copy of it, so they cost less than the first.

You send

request
curl -X POST "$NORTHDOC_API/documents" \
  -H "Authorization: Bearer $NORTHDOC_KEY" \
  -F file=@contract-of-sale.pdf \
  -F 'options={"queries": [
        "What is the purchase price?",
        "What is the deposit?",
        "When is settlement?",
        "Is a pool safety certificate attached?"
      ]}'

You get back

response
{
  "status": "completed",
  "result": {
    "answers": [
      {"question": "What is the purchase price?", "answer": "$1,250,000",
       "page": 2, "quote": "Price: $1,250,000", "not_found": false},
      {"question": "What is the deposit?", "answer": "$125,000 (10% of the price)",
       "page": 2, "quote": "Deposit: $125,000 payable on the day of sale", "not_found": false},
      {"question": "When is settlement?", "answer": "14 November 2026",
       "page": 3, "quote": "Settlement is due on 14/11/2026", "not_found": false},
      {"question": "Is a pool safety certificate attached?", "answer": "",
       "page": null, "quote": null, "not_found": true}
    ]
  }
}

Then

request
curl -X POST "$NORTHDOC_API/documents/$DOC_ID/queries" \
  -H "Authorization: Bearer $NORTHDOC_KEY" \
  -H "Content-Type: application/json" \
  -d '{
        "question": "List the special conditions.",
        "schema": {
          "type": "object",
          "properties": {
            "conditions": {
              "type": "array",
              "items": {
                "type": "object",
                "properties": {
                  "number":  {"type": "integer"},
                  "summary": {"type": "string"}
                }
              }
            }
          }
        }
      }'

You get back

response
{
  "id": "c0de1234-…",
  "object": "query",
  "status": "completed",
  "question": "List the special conditions.",
  "answer": {
    "answer": {
      "conditions": [
        {"number": 1, "summary": "Subject to finance approval within 21 days"},
        {"number": 2, "summary": "Vendor to remove the garden shed before settlement"}
      ]
    },
    "citations": [
      {"page": 31, "quote": "Special condition 1: This contract is subject to finance"},
      {"page": 32, "quote": "Special condition 2: The vendor must remove the shed"}
    ],
    "confidence": 0.9,
    "not_found": false
  },
  "cost": {"estimated_micro": 31000, "settled_micro": 9120}
}

Example 06

Get a number that only a chart shows.

The images option lets the model look at charts, so answers that are not in the text are still found.

The document

annual-report.pdf: a 24-page annual report. Page 3 has a bar chart of quarterly revenue, and the numbers appear nowhere in the text.

images (on by default)

What happened

  • OCR turns the chart into a caption and nothing else, so on its own the text has no Q3 figure.
  • With images on (the default), the chart is cut out and placed in the text where it sits, followed by Figure (page 3): Quarterly revenue, FY2026. The model reads the bar.
  • Because the value is in a picture, the quote is the figure's caption. The page number is still exact.
  • With images: false the model only has the text, so it honestly says not_found. It costs a little less, which is fine for documents with no pictures.

You send

request
curl -X POST "$NORTHDOC_API/documents" \
  -H "Authorization: Bearer $NORTHDOC_KEY" \
  -F file=@annual-report.pdf \
  -F 'options={"queries": ["What was revenue in Q3?"]}'

You get back

response
{
  "status": "completed",
  "result": {
    "answers": [
      {"question": "What was revenue in Q3?",
       "answer": "$4.2 million",
       "page": 3,
       "quote": "Figure (page 3): Quarterly revenue, FY2026",
       "not_found": false}
    ]
  }
}

Then

request
# The same document with images turned off
curl -X POST "$NORTHDOC_API/documents" \
  -H "Authorization: Bearer $NORTHDOC_KEY" \
  -F file=@annual-report.pdf \
  -F 'options={"images": false, "queries": ["What was revenue in Q3?"]}'

You get back

response
{
  "status": "completed",
  "result": {
    "answers": [
      {"question": "What was revenue in Q3?", "answer": "",
       "page": null, "quote": null, "not_found": true}
    ]
  }
}

Example 07

Check a scanned form is signed and stamped.

Scanned pages are sent as whole pictures, so signatures, stamps and handwriting are read too.

The document

statutory-declaration.pdf: a three-page scan. Page 3 has a handwritten signature, a date and a JP's stamp, and only 12 words of printed text.

images (whole pages) extraction.schema

What happened

  • Page 3 has fewer than 25 words, so instead of cutting out pieces, the whole page goes to the model as a picture, ahead of its text.
  • The signature and the handwritten date are not in the OCR text at all. The model reads them from the picture.
  • For values read from a picture, quote describes where they are on the page. The stamp's printed words were readable, so that quote is the stamp's own text.
  • Handwriting is harder than print, which shows in a lower confidence. Use it to send uncertain values to a person.

You send

request
curl -X POST "$NORTHDOC_API/documents" \
  -H "Authorization: Bearer $NORTHDOC_KEY" \
  -F file=@statutory-declaration.pdf \
  -F 'options={
    "extraction": {
      "schema": {
        "type": "object",
        "properties": {
          "declarant":       {"type": "string"},
          "signed":          {"type": "boolean"},
          "date_signed":     {"type": "string", "format": "date"},
          "witness_stamped": {"type": "boolean"},
          "witness_name":    {"type": "string"}
        }
      }
    }
  }'

You get back

response
{
  "status": "completed",
  "page_count": 3,
  "result": {
    "data": {
      "declarant": "Priya Raman",
      "signed": true,
      "date_signed": "2026-09-28",
      "witness_stamped": true,
      "witness_name": "Daniel Okafor JP"
    },
    "citations": [
      {"path": "data.declarant", "page": 1, "quote": "I, Priya Raman, of 12 Wattle Street", "confidence": 0.98},
      {"path": "data.signed", "page": 3, "quote": "Handwritten signature on the Signature of declarant line", "confidence": 0.93},
      {"path": "data.date_signed", "page": 3, "quote": "Handwritten date 28/9/2026 beside the signature", "confidence": 0.86},
      {"path": "data.witness_name", "page": 3, "quote": "Stamp: DANIEL OKAFOR JP Reg. No. 18842", "confidence": 0.91}
    ]
  }
}

Example 08

Split a stack of receipts, one per page.

Per-page extraction treats each page as its own document and gives you one result per page.

The document

receipts-september.pdf: 18 scanned receipts, one per page.

extraction.per_page extraction.schema

What happened

  • per_page: true runs the extraction once per page, so 18 receipts give you 18 results under result.pages.
  • Each page's result is also on that page in GET /documents/{id}/pages.
  • It costs about one extraction per page, so use it for stacks of separate records, not for one contract that runs across pages.

You send

request
curl -X POST "$NORTHDOC_API/documents" \
  -H "Authorization: Bearer $NORTHDOC_KEY" \
  -F file=@receipts-september.pdf \
  -F 'options={
    "extraction": {
      "per_page": true,
      "schema": {
        "type": "object",
        "properties": {
          "merchant": {"type": "string"},
          "date":     {"type": "string", "format": "date"},
          "total":    {"type": "number"}
        }
      }
    }
  }'

You get back

response
{
  "status": "completed",
  "page_count": 18,
  "result": {
    "per_page": true,
    "pages": [
      {"page": 1,
       "data": {"merchant": "Bean Counter Cafe", "date": "2026-09-02", "total": 18.50},
       "citations": [{"path": "data.total", "page": 1, "quote": "TOTAL $18.50", "confidence": 0.97}]},
      {"page": 2,
       "data": {"merchant": "Officeworks", "date": "2026-09-03", "total": 64.95},
       "citations": [{"path": "data.total", "page": 2, "quote": "Total AUD 64.95", "confidence": 0.98}]}
    ]
  }
}

Example 09

Get tables and labelled fields as they are.

The raw structure the OCR finds, kept for every document: every table as markdown and every labelled field with its confidence.

The document

rate-schedule.pdf: a one-page fee schedule with a table of rates and a header block of labelled fields.

no options include=tables,key_values

What happened

  • Tables and labelled fields come from the layout analysis (the default). You read them from the pages endpoint with include.
  • Tables come back as markdown, which spreadsheets, LLMs and people can all read.
  • These are the OCR's raw findings, with no AI judgement and no model cost. Add an extraction if you want a cleaned-up version in result.data.

You send

request
curl -X POST "$NORTHDOC_API/documents" \
  -H "Authorization: Bearer $NORTHDOC_KEY" \
  -F file=@rate-schedule.pdf

You get back

response
{
  "id": "8f1c2d3e-…",
  "object": "document",
  "status": "queued",
  "estimated_cost_micro": 30000,
  "poll_url": "https://northdoc.northcape.tech/api/v1/documents/8f1c2d3e-…"
}

Then

request
curl "$NORTHDOC_API/documents/$DOC_ID/pages?include=tables,key_values" \
  -H "Authorization: Bearer $NORTHDOC_KEY"

You get back

response
{
  "object": "list",
  "document_id": "8f1c2d3e-…",
  "data": [
    {
      "page": 1,
      "text": "# Fee schedule 2026\n\nEffective date: 1 July 2026\n\n| Service | Rate |\n| - | - |\n| Standard review | $180.00 |\n| Urgent review | $260.00 |",
      "word_count": 19,
      "result": null,
      "key_values": {
        "Effective date": {"value": "1 July 2026", "key_confidence": 94.0,
                           "value_confidence": 92.0, "selection_status": null}
      },
      "tables": [
        {"rows": 3, "columns": 2,
         "markdown": "| Service | Rate |\n| - | - |\n| Standard review | $180.00 |\n| Urgent review | $260.00 |"}
      ]
    }
  ]
}

Example 10

Make a handbook searchable.

Get one vector per page to store in your own search index, alongside the page text.

The document

staff-handbook.pdf: a 60-page policy handbook.

embeddings.dimensions images: false

What happened

  • embeddings turns vectors on: true for the defaults, or an object to set the size, as here.
  • Vectors are only on the pages endpoint, never on GET /documents/{id}.
  • Store each vector with the document id and page number, so a search hit can point at the page.
  • Choose one dimensions for your whole index: vectors of different lengths cannot be compared. A text-only handbook has no pictures worth keeping, so they are off here.

You send

request
curl -X POST "$NORTHDOC_API/documents" \
  -H "Authorization: Bearer $NORTHDOC_KEY" \
  -F file=@staff-handbook.pdf \
  -F 'options={
    "embeddings": {"granularity": "page", "dimensions": 512},
    "images": false
  }'

You get back

response
{
  "id": "8f1c2d3e-…",
  "object": "document",
  "status": "queued",
  "estimated_cost_micro": 1804320,
  "poll_url": "https://northdoc.northcape.tech/api/v1/documents/8f1c2d3e-…"
}

Then

request
curl "$NORTHDOC_API/documents/$DOC_ID/pages?include=embeddings&page=1-2" \
  -H "Authorization: Bearer $NORTHDOC_KEY"

You get back

response
{
  "object": "list",
  "document_id": "8f1c2d3e-…",
  "data": [
    {"page": 1, "text": "# Staff handbook\n\nWelcome to…", "word_count": 212, "result": null,
     "embedding": {"model": "northdoc-compass-1", "dimensions": 512,
                   "vector": [0.0123, -0.0456, 0.0789, "… 509 more"]}},
    {"page": 2, "text": "## Leave\n\nYou can take…", "word_count": 388, "result": null,
     "embedding": {"model": "northdoc-compass-1", "dimensions": 512,
                   "vector": [-0.0311, 0.0082, 0.0457, "… 509 more"]}}
  ]
}

Example 11

Turn an archive into plain text, cheaply.

OCR-only reading with pictures off is the cheapest way to get the words out of a typed document.

The document

letters-1998.pdf: 40 pages of typed correspondence with no tables or pictures.

analysis: read images: false include=text

What happened

  • analysis: "read" costs US$0.004 a page instead of US$0.025. It finds the words but not tables, labelled fields or figures, and markdown: true is refused with it.
  • include=text gives the whole document's text, each page under a === Page N === line: the same text the model reads.
  • No model runs, so 40 pages cost US$0.16. Add "extraction": {"fields": ["title", "summary", "dates", "parties"]} and each file gets a handy index card in result.data too.

You send

request
curl -X POST "$NORTHDOC_API/documents" \
  -H "Authorization: Bearer $NORTHDOC_KEY" \
  -F file=@letters-1998.pdf \
  -F 'options={"analysis": "read", "images": false}'

You get back

response
{
  "id": "8f1c2d3e-…",
  "object": "document",
  "status": "queued",
  "estimated_cost_micro": 192000,
  "poll_url": "https://northdoc.northcape.tech/api/v1/documents/8f1c2d3e-…"
}

Then

request
curl "$NORTHDOC_API/documents/$DOC_ID?include=text" \
  -H "Authorization: Bearer $NORTHDOC_KEY" | jq -r .text

You get back

response
=== Page 1 ===
14 March 1998

Dear Ms Chen,

Thank you for your letter of 2 March…

=== Page 2 ===
…

Example 12

Send several files, or a link.

Upload up to ten files in one request, or point Northdoc at a URL to fetch.

The document

Three monthly statements as separate PDFs, one of which is a spreadsheet saved as .csv by mistake.

files[] url

What happened

  • Each file becomes its own document with its own id and charge. The answers come back in the order you sent the files.
  • One bad file does not sink the batch: its slot holds an error and the rest go ahead. Check every slot.
  • Write files[] with the brackets. Without them, curl's repeated parts collapse into one file.
  • With a URL, options is a real JSON object, not a string. The cost estimate is made after the download, once the size and page count are known.

You send

request
curl -X POST "$NORTHDOC_API/documents" \
  -H "Authorization: Bearer $NORTHDOC_KEY" \
  -F "files[]=@july.pdf" \
  -F "files[]=@august.pdf" \
  -F "files[]=@september.csv"

You get back

response
{
  "documents": [
    {"id": "8f1c…", "object": "document", "status": "queued", "estimated_cost_micro": 90000,
     "poll_url": "https://northdoc.northcape.tech/api/v1/documents/8f1c…"},
    {"id": "9a2d…", "object": "document", "status": "queued", "estimated_cost_micro": 90000,
     "poll_url": "https://northdoc.northcape.tech/api/v1/documents/9a2d…"},
    {"error": {"type": "unsupported_media_type",
               "message": "Unsupported file type. Send a PDF, PNG, JPEG, TIFF, BMP, HEIF, DOCX, XLSX or PPTX."}}
  ]
}

Then

request
curl -X POST "$NORTHDOC_API/documents" \
  -H "Authorization: Bearer $NORTHDOC_KEY" \
  -H "Content-Type: application/json" \
  -d '{
        "url": "https://files.example.com/statements/october.pdf",
        "options": {"retention_seconds": 3600}
      }'

You get back

response
{
  "id": "b7e4…",
  "object": "document",
  "status": "queued",
  "estimated_cost_micro": null,
  "poll_url": "https://northdoc.northcape.tech/api/v1/documents/b7e4…"
}

Example 13

Handle a document that fails.

A failed document still answers 200, with an error that tells you what went wrong and where.

The document

A URL that answers 404.

status: failed error.type error.stage

What happened

  • Polling a failed document is not an error: you get 200, status: "failed" and an error object.
  • error.type is stable, so branch on it. error.stage says which step it stopped at.
  • You only pay for work done before the failure. Here nothing was done, so nothing was charged.
  • Some problems are caught before anything starts, and answer the submit call itself with a 4xx: a file over your plan's size (413), too many pages (422) or a type we cannot read (415).

You send

request
curl "$NORTHDOC_API/documents/$DOC_ID" \
  -H "Authorization: Bearer $NORTHDOC_KEY"

You get back

response
{
  "id": "b7e4…",
  "object": "document",
  "status": "failed",
  "stage": "fetching",
  "error": {
    "type": "unsupported_media_type",
    "message": "The URL could not be fetched (HTTP 404).",
    "stage": "fetching"
  },
  "cost": {"estimated_micro": null, "settled_micro": 0},
  "failed_at": "2026-10-05T01:20:44Z"
}

Try any of these, free.

Create a workspace, copy a test key and paste an example. When it works, swap in a live key and your own document.