Document intelligence API · Processed and stored in Australia

Every page read.
Every answer cited.

Northdoc reads PDFs, scans and Office files the way a person does, text, tables, charts and signatures. Take every page back as text, pictures and vectors, or ask for structured JSON with the page and the quote behind every value.

  • No card to start
  • Test keys are free
  • Pay per page
Reads PDF Scans Word Excel PowerPoint PNG JPEG TIFF HEIF

How it works

From a file to facts you can check.

One request in, one result out. No templates to train, no zones to draw, and nothing to host.

  1. 01

    Send the document

    Upload a PDF, scan, image or Office file, or pass a URL. Ask for vectors, your JSON Schema or questions in the same request, or nothing more than the pages.

    POST /api/v1/documents

  2. 02

    Northdoc reads every page

    OCR and layout analysis capture the text, tables, key-value pairs and reading order. Charts, diagrams, scanned pages and signatures are kept as pictures, in place, for you and the model.

    analyzing → rendering → extracting

  3. 03

    Get answers you can check

    Every page's text and pictures, and whatever you asked for: structured JSON that matches your schema, with the page, the quote and a confidence score for every value. Ask follow-up questions whenever you like.

    GET /api/v1/documents/{id}

Sees the whole page

OCR reads the words. Northdoc reads the document.

A chart, a site plan, a stamp or a signature is a blank to plain OCR. Northdoc crops each figure where it sits and shows the model any page that is mostly picture, so the answer can come from what the page shows as well as what it says.

  • Figures in reading order, with their captions, exactly where the page puts them.
  • Scans, photos and signature pages included whole, so nothing silent is treated as blank.
  • Taken for every document. Download them, or let the model read them, billed per token only when it does. Turn them off with "images": false.

Capabilities

Everything between the upload and the answer.

Your schema, filled in

Send a JSON Schema, or pick built-in fields like a summary, the parties and the dates, and get back data that validates against it. Northdoc checks the model's output and retries once before it reports a mismatch.

A citation for every value

Each field comes with the page it came from, the exact source quote and a confidence score, so a person can verify it in seconds.

Sees what OCR misses

Figures are cropped where they sit and picture-heavy pages are shown whole, so charts, diagrams, stamps and signatures are read rather than skipped.

Questions, answered

Ask up to 50 questions while the document is processed, or ask more afterwards. Answers can follow their own schema and are cited like everything else.

Embeddings for search

A vector per page at 256, 512 or 1024 dimensions, ready for semantic search and retrieval-augmented generation.

Tables and fields, kept

Tables come back as rows and cells, form fields as key-value pairs, and the full document as clean Markdown when you ask for it.

Processed and stored in Australia

OCR, layout analysis and the AI models all run in Australian data centres, and the app and its database are hosted in Sydney.

Kept only as long as you say

Uploads are deleted as soon as they are read. Results are purged on the retention schedule you set, or the moment you delete the document.

Free to build against

Test keys run the same API against deterministic fakes at no cost. Live keys are pay as you go, from prepaid credit.

For developers

One request. A result you can trust.

A plain REST API with an OpenAPI 3.1 spec. Send the file with your schema and your questions, poll for the result, and ask follow-ups whenever you need them, served from a cached copy of the document.

Stable error types
JSON, every time
Test mode
Free, deterministic
Models
Swift & Summit
Spec
OpenAPI 3.1
terminal
$ curl https://northdoc.northcape.tech/api/v1/documents \
    -H "Authorization: Bearer $NORTHDOC_KEY" \
    -F file=@contract-of-sale.pdf \
    -F 'options={"extraction":{"fields":["amounts","dates"]},
                 "queries":["What is the deposit?"]}'

Use cases

Built for the paperwork that runs a business.

From the team behind Matter First, NorthCape's practice management platform for Australian law firms, where every matter starts with a stack of documents.

Legal & conveyancing

Pull parties, dates, prices and special conditions from contracts of sale, leases and title searches, each cited to its clause.

  • Contracts of sale
  • Leases
  • Title searches
  • Court documents

Finance & accounting

Turn invoices, statements and remittances into ledger-ready rows, with totals and line items checked against the page.

  • Invoices
  • Bank statements
  • Remittances
  • Receipts

Insurance & claims

Read claim forms, medical reports and photos of damage together, and get the facts back in your claim system's own shape.

  • Claim forms
  • Assessments
  • Medical reports
  • Policies

Operations

Onboarding packs, compliance certificates and supplier documents, filed and searchable without anyone retyping them.

  • Onboarding packs
  • Certificates
  • Delivery dockets
  • Forms

Security & data

Kept in Australia. Kept on your terms.

Contracts, statements and medical reports are the documents that matter most. Northdoc handles them as if they were ours: processed and stored in Australia, kept only as long as you choose, and never used to train anyone's models.

How we handle your data
  • Onshore at every step

    OCR, AI models and storage all run in Australian data centres, with the app and database in Sydney.

  • Uploads deleted on read

    The file is removed as soon as it has been analysed. Results follow your retention setting.

  • Never used for training

    Our model providers' terms exclude your content from training their models.

  • Every request logged

    Your workspace keeps an audit log of each API call, its key, status and latency.

Pricing

Pay for the pages you process.

No subscriptions. Your first workspace starts with $5 of trial credit, then tops up and pays per page analysed and per token generated.

Free trial

$5 free, for 30 days

Trial credit to build your integration; small documents, short retention.

  • Up to 20 pages per document
  • Files up to 20 MB
  • Results kept a day by default, up to 7 days
  • 2 requests / second
Start free
For production

Pay as you go

$5+ per top-up

Top up credits and pay per page and token. Large documents, no expiry.

  • Up to 1,000 pages per document
  • Files up to 200 MB
  • Keep results as long as you need
  • 10 requests / second
Start, then top up

US$0.025 per page for full layout, US$0.004 for read-only OCR, plus model tokens. See the full rate card

Questions

Everything you'd want to ask first.

The API reference covers authentication, limits and every error the API returns.

Open the reference
What is Northdoc?

Northdoc is a document intelligence API from NorthCape Technology. You send it a PDF, scan, image or Office file and it returns every page's text, tables and pictures, with vectors for search if you want them. Ask it to, and its AI also returns structured data that matches your JSON Schema, or answers to your questions, with a page citation, a source quote and a confidence score for every value.

Which file types can Northdoc read?

PDF, scanned documents, PNG, JPEG, TIFF, BMP and HEIF images, and Word, Excel and PowerPoint files. Scanned and handwritten pages are read with OCR.

Does Northdoc read charts, diagrams and signatures?

Yes. Northdoc crops the figures it finds and keeps pages that are mostly pictures, such as scans, diagrams and signature pages, as images. You can download them, and when the model reads the document it sees them in their place. The OCR text is still what citations quote.

Where are documents processed and stored?

In Australia. OCR and layout analysis, the AI models and the vector embeddings all run in Australian data centres, and the application and its database, where results are kept, are hosted in Sydney.

How long does Northdoc keep my documents?

The uploaded file is deleted as soon as it has been read. Results are kept for the retention period you choose with `retention_seconds` (a day by default, up to seven days on the free trial) and purged after that, or straight away when you delete the document.

Are my documents used to train AI models?

No. Northdoc never trains models on your documents, and the infrastructure it runs on does not use customer content for training either.

How much does Northdoc cost?

Pay as you go, from prepaid credit: US$0.025 per page for full layout analysis or US$0.004 per page for read-only OCR. Vectors and the AI model's tokens are added only when you ask for them. Your first workspace gets free trial credit (one trial per account), and test keys are always free.

Which AI models does Northdoc use?

Northdoc has three models. Swift is fast and economical, for everyday documents. Summit is the most capable, for long or dense legal and financial documents. Compass makes embeddings for search. By default Northdoc picks Swift or Summit by the size of the document, and you can choose one yourself. Swift and Summit are kept current: when a better model becomes available Northdoc moves to it, so you get the improvement without changing your code.

How accurate is the extraction?

Every value is validated against your schema and comes with the page, the quote it was taken from and a confidence score, so your system can accept confident values automatically and route the rest to a person.

Send Northdoc your hardest document.

A workspace, a key and your first answer in a few minutes, on $5 of free credit. No card until you need more.