Document intelligence API · Processed and stored in Australia
Every page read.
Every answer cited.
Northdoc reads PDFs, scans and Office files the way a person does, text, tables, charts and signatures. Take every page back as text, pictures and vectors, or ask for structured JSON with the page and the quote behind every value.
- No card to start
- Test keys are free
- Pay per page
How it works
From a file to facts you can check.
One request in, one result out. No templates to train, no zones to draw, and nothing to host.
-
01
Send the document
Upload a PDF, scan, image or Office file, or pass a URL. Ask for vectors, your JSON Schema or questions in the same request, or nothing more than the pages.
POST /api/v1/documents
-
02
Northdoc reads every page
OCR and layout analysis capture the text, tables, key-value pairs and reading order. Charts, diagrams, scanned pages and signatures are kept as pictures, in place, for you and the model.
analyzing → rendering → extracting
-
03
Get answers you can check
Every page's text and pictures, and whatever you asked for: structured JSON that matches your schema, with the page, the quote and a confidence score for every value. Ask follow-up questions whenever you like.
GET /api/v1/documents/{id}
Sees the whole page
OCR reads the words. Northdoc reads the document.
A chart, a site plan, a stamp or a signature is a blank to plain OCR. Northdoc crops each figure where it sits and shows the model any page that is mostly picture, so the answer can come from what the page shows as well as what it says.
- Figures in reading order, with their captions, exactly where the page puts them.
- Scans, photos and signature pages included whole, so nothing silent is treated as blank.
-
Taken for every document.
Download them, or let the model read them, billed per token only when it does.
Turn them off with
"images": false.
Capabilities
Everything between the upload and the answer.
Your schema, filled in
Send a JSON Schema, or pick built-in fields like a summary, the parties and the dates, and get back data that validates against it. Northdoc checks the model's output and retries once before it reports a mismatch.
A citation for every value
Each field comes with the page it came from, the exact source quote and a confidence score, so a person can verify it in seconds.
Sees what OCR misses
Figures are cropped where they sit and picture-heavy pages are shown whole, so charts, diagrams, stamps and signatures are read rather than skipped.
Questions, answered
Ask up to 50 questions while the document is processed, or ask more afterwards. Answers can follow their own schema and are cited like everything else.
Embeddings for search
A vector per page at 256, 512 or 1024 dimensions, ready for semantic search and retrieval-augmented generation.
Tables and fields, kept
Tables come back as rows and cells, form fields as key-value pairs, and the full document as clean Markdown when you ask for it.
Processed and stored in Australia
OCR, layout analysis and the AI models all run in Australian data centres, and the app and its database are hosted in Sydney.
Kept only as long as you say
Uploads are deleted as soon as they are read. Results are purged on the retention schedule you set, or the moment you delete the document.
Free to build against
Test keys run the same API against deterministic fakes at no cost. Live keys are pay as you go, from prepaid credit.
For developers
One request. A result you can trust.
A plain REST API with an OpenAPI 3.1 spec. Send the file with your schema and your questions, poll for the result, and ask follow-ups whenever you need them, served from a cached copy of the document.
- Stable error types
- JSON, every time
- Test mode
- Free, deterministic
- Models
- Swift & Summit
- Spec
- OpenAPI 3.1
$ curl https://northdoc.northcape.tech/api/v1/documents \
-H "Authorization: Bearer $NORTHDOC_KEY" \
-F file=@contract-of-sale.pdf \
-F 'options={"extraction":{"fields":["amounts","dates"]},
"queries":["What is the deposit?"]}'
{ "status": "completed", "page_count": 14, "result": { "data": { "amounts": [{ "label": "Deposit", "amount": 125000 }, …], "dates": [{ "label": "Settlement", "date": "2026-11-28" }] }, "citations": [{ "path": "data.amounts[0].amount", "page": 2, "quote": "Deposit: $125,000", "confidence": 0.97 }, …], "answers": [{ "question": "What is the deposit?", "answer": "$125,000, 10% of the price", "page": 2 }] } }
Use cases
Built for the paperwork that runs a business.
From the team behind Matter First, NorthCape's practice management platform for Australian law firms, where every matter starts with a stack of documents.
Legal & conveyancing
Pull parties, dates, prices and special conditions from contracts of sale, leases and title searches, each cited to its clause.
- Contracts of sale
- Leases
- Title searches
- Court documents
Finance & accounting
Turn invoices, statements and remittances into ledger-ready rows, with totals and line items checked against the page.
- Invoices
- Bank statements
- Remittances
- Receipts
Insurance & claims
Read claim forms, medical reports and photos of damage together, and get the facts back in your claim system's own shape.
- Claim forms
- Assessments
- Medical reports
- Policies
Operations
Onboarding packs, compliance certificates and supplier documents, filed and searchable without anyone retyping them.
- Onboarding packs
- Certificates
- Delivery dockets
- Forms
Security & data
Kept in Australia. Kept on your terms.
Contracts, statements and medical reports are the documents that matter most. Northdoc handles them as if they were ours: processed and stored in Australia, kept only as long as you choose, and never used to train anyone's models.
How we handle your data-
Onshore at every step
OCR, AI models and storage all run in Australian data centres, with the app and database in Sydney.
-
Uploads deleted on read
The file is removed as soon as it has been analysed. Results follow your retention setting.
-
Never used for training
Our model providers' terms exclude your content from training their models.
-
Every request logged
Your workspace keeps an audit log of each API call, its key, status and latency.
Pricing
Pay for the pages you process.
No subscriptions. Your first workspace starts with $5 of trial credit, then tops up and pays per page analysed and per token generated.
Free trial
$5 free, for 30 days
Trial credit to build your integration; small documents, short retention.
- Up to 20 pages per document
- Files up to 20 MB
- Results kept a day by default, up to 7 days
- 2 requests / second
Pay as you go
$5+ per top-up
Top up credits and pay per page and token. Large documents, no expiry.
- Up to 1,000 pages per document
- Files up to 200 MB
- Keep results as long as you need
- 10 requests / second
US$0.025 per page for full layout, US$0.004 for read-only OCR, plus model tokens. See the full rate card
Questions
Everything you'd want to ask first.
The API reference covers authentication, limits and every error the API returns.
Open the referenceWhat is Northdoc?
Northdoc is a document intelligence API from NorthCape Technology. You send it a PDF, scan, image or Office file and it returns every page's text, tables and pictures, with vectors for search if you want them. Ask it to, and its AI also returns structured data that matches your JSON Schema, or answers to your questions, with a page citation, a source quote and a confidence score for every value.
Which file types can Northdoc read?
PDF, scanned documents, PNG, JPEG, TIFF, BMP and HEIF images, and Word, Excel and PowerPoint files. Scanned and handwritten pages are read with OCR.
Does Northdoc read charts, diagrams and signatures?
Yes. Northdoc crops the figures it finds and keeps pages that are mostly pictures, such as scans, diagrams and signature pages, as images. You can download them, and when the model reads the document it sees them in their place. The OCR text is still what citations quote.
Where are documents processed and stored?
In Australia. OCR and layout analysis, the AI models and the vector embeddings all run in Australian data centres, and the application and its database, where results are kept, are hosted in Sydney.
How long does Northdoc keep my documents?
The uploaded file is deleted as soon as it has been read. Results are kept for the retention period you choose with `retention_seconds` (a day by default, up to seven days on the free trial) and purged after that, or straight away when you delete the document.
Are my documents used to train AI models?
No. Northdoc never trains models on your documents, and the infrastructure it runs on does not use customer content for training either.
How much does Northdoc cost?
Pay as you go, from prepaid credit: US$0.025 per page for full layout analysis or US$0.004 per page for read-only OCR. Vectors and the AI model's tokens are added only when you ask for them. Your first workspace gets free trial credit (one trial per account), and test keys are always free.
Which AI models does Northdoc use?
Northdoc has three models. Swift is fast and economical, for everyday documents. Summit is the most capable, for long or dense legal and financial documents. Compass makes embeddings for search. By default Northdoc picks Swift or Summit by the size of the document, and you can choose one yourself. Swift and Summit are kept current: when a better model becomes available Northdoc moves to it, so you get the improvement without changing your code.
How accurate is the extraction?
Every value is validated against your schema and comes with the page, the quote it was taken from and a confidence score, so your system can accept confident values automatically and route the rest to a person.
Send Northdoc your hardest document.
A workspace, a key and your first answer in a few minutes, on $5 of free credit. No card until you need more.