Example 01
Get the text, vectors and pictures, without the AI.
No model at all. Send a document and get back every page's text, a vector per page and the document's pictures, to use in your own system.
The document
site-report.pdf: a 12-page building inspection report with photos of defects, a floor plan and a signed last page.
What happened
-
No AI reads the document: the model only runs when you ask for an
extractionorqueries. This one cost US$0.30, which is 12 pages of layout analysis plus a fraction of a cent for the vectors. Leave outembeddingstoo and you pay for the pages alone. -
One call returns every page with its text, its vector and its pictures.
image_dataputs each picture's bytes indata, base64-encoded: decode it and you have the PNG or JPEG file. -
Each page's text keeps its reading order, with tables as markdown and a
[FIGURE]marker where each picture sits. The first[FIGURE]on page 3 is the picture withfigure_index: 0, so you can put every picture back exactly where it was. - You get the photos, plans and charts cropped from their pages, and pages that are mostly picture or nearly wordless (like the signed page 12) whole. Logos and blank pages are skipped.
-
Base64 makes a response about a third bigger than the pictures themselves, so for long documents fetch a few pages at a time with
page, or useinclude=images(no bytes) and download each picture from itsurl. -
Nothing stops you asking the document a question later:
POST /documents/{id}/queriesstill works on it.
You send
curl -X POST "$NORTHDOC_API/documents" \
-H "Authorization: Bearer $NORTHDOC_KEY" \
-F file=@site-report.pdf \
-F 'options={"embeddings": true}'
You get back
{
"id": "8f1c2d3e-…",
"object": "document",
"status": "queued",
"estimated_cost_micro": 360864,
"poll_url": "https://northdoc.northcape.tech/api/v1/documents/8f1c2d3e-…"
}
Then
# Once status is "completed": every page's text, vector and pictures, in one call
curl "$NORTHDOC_API/documents/$DOC_ID/pages?include=embeddings,image_data" \
-H "Authorization: Bearer $NORTHDOC_KEY"
You get back
{
"object": "list",
"document_id": "8f1c2d3e-…",
"data": [
{
"page": 3,
"text": "## Exterior\n\nCracking is visible along the north wall.\n\n[FIGURE]\n\nThe crack runs from the window head to the slab.",
"word_count": 214,
"result": null,
"embedding": {"model": "northdoc-compass-1", "dimensions": 1024,
"vector": [0.0123, -0.0456, 0.0789, "… 1021 more"]},
"images": [
{"id": "5c1e…", "object": "image", "page": 3, "kind": "figure", "figure_index": 0,
"caption": "Photo 2: Cracking to the north wall",
"media_type": "image/png", "width": 1240, "height": 930, "byte_size": 412880,
"bbox": {"left": 0.12, "top": 0.38, "width": 0.76, "height": 0.41},
"url": "https://northdoc.northcape.tech/api/v1/documents/8f1c2d3e-…/images/5c1e…",
"data": "iVBORw0KGgoAAAANSUhEUgAABNgAAAOi…"}
]
},
{
"page": 12,
"text": "Inspector's declaration\n\nSigned:\n\nDate:",
"word_count": 6,
"result": null,
"embedding": {"model": "northdoc-compass-1", "dimensions": 1024,
"vector": [-0.0311, 0.0082, 0.0457, "… 1021 more"]},
"images": [
{"id": "7a90…", "object": "image", "page": 12, "kind": "page", "figure_index": null,
"caption": null,
"media_type": "image/jpeg", "width": 1109, "height": 1568, "byte_size": 286411,
"bbox": null,
"url": "https://northdoc.northcape.tech/api/v1/documents/8f1c2d3e-…/images/7a90…",
"data": "/9j/4AAQSkZJRgABAQAAAQABAAD…"}
]
}
]
}
Then
# Or, for big documents: page through a few pages at a time,
# or list the pictures without their bytes and download each one
curl "$NORTHDOC_API/documents/$DOC_ID/pages?include=embeddings,image_data&page=1-10" \
-H "Authorization: Bearer $NORTHDOC_KEY"
curl "$NORTHDOC_API/documents/$DOC_ID/images" \
-H "Authorization: Bearer $NORTHDOC_KEY"
curl -o page-3-figure-0.png "$NORTHDOC_API/documents/$DOC_ID/images/5c1e…" \
-H "Authorization: Bearer $NORTHDOC_KEY"
You get back
{
"object": "list",
"document_id": "8f1c2d3e-…",
"data": [
{"id": "5c1e…", "object": "image", "page": 3, "kind": "figure", "figure_index": 0,
"caption": "Photo 2: Cracking to the north wall",
"media_type": "image/png", "width": 1240, "height": 930, "byte_size": 412880,
"bbox": {"left": 0.12, "top": 0.38, "width": 0.76, "height": 0.41},
"url": "https://northdoc.northcape.tech/api/v1/documents/8f1c2d3e-…/images/5c1e…"},
{"id": "7a90…", "object": "image", "page": 12, "kind": "page", "figure_index": null,
"caption": null,
"media_type": "image/jpeg", "width": 1109, "height": 1568, "byte_size": 286411,
"bbox": null,
"url": "https://northdoc.northcape.tech/api/v1/documents/8f1c2d3e-…/images/7a90…"}
]
}