Reading time: ~8 min
Getting text off a receipt is the easy part. Every OCR engine does that. The hard part starts afterwards: turning "TOTAL 15.56" into a field your code can rely on, for a Polish VAT invoice, a Japanese convenience-store slip and a taxi receipt with no itemisation — all returning the same shape.
Search for this problem and you will find a dozen vendors showing you their response format, and a Reddit thread telling you not to let the model do arithmetic. Nobody publishes the schema itself, or the decisions behind it. This post does both.
Why "just ask the model for JSON" breaks in production
Handing a vision model an image and a sentence like "return the receipt as JSON" works impressively well in a notebook. It fails in three specific ways once real documents arrive.
Schema drift. The same prompt returns total on one document, total_amount on the next, and grandTotal on a third. Your parser works for a week.
Invented values. A model asked for a field it cannot see will often produce something plausible rather than nothing. A missing tax line becomes 0.00, which is indistinguishable from a genuine zero-rated item.
Arithmetic. This is the one people underestimate. Language models are unreliable at summing decimals, and a receipt where the line items add to 14.90 but the printed total says 15.56 is extremely common — service charges, rounding, deposits, discounts applied at the basket level. If you ask the model to compute the total, you get a number that is neither what the paper says nor what the items add up to.
The fix for all three is the same: constrain the output to a fixed schema, tell the model to copy values rather than compute them, and do the maths yourself afterwards. If you are still choosing between approaches here, the trade-offs between classical OCR, purpose-built APIs and general multimodal models are covered in Receipt OCR in 2026: AI vs Traditional.
The fields a receipt and an invoice actually share — and where they diverge
Most teams start with a receipt schema, then discover invoices need more, then bolt fields on. Design for both from the start — the overlap is larger than it looks.
Shared: vendor, transaction time, currency, country, total, tax total, net total, line items.
Receipt-only in practice: there is rarely a document number, rarely a customer, and almost never per-item tax. A supermarket receipt gives you one basket-level VAT summary, not a rate per line.
Invoice-only: a document number you can key on, a vendor tax identifier, per-line net amounts and per-line tax, and frequently multiple tax rates on one document.
That last point matters for the schema. Per-item tax and priceNet fields are meaningful on an invoice and meaningless on a receipt. Rather than returning them as null everywhere, it is cleaner to omit them for receipts entirely — which is what WiseOCR does: the extractor strips those two keys from every line item when documentType is receipt, so a present key always means a real value.
A reference schema you can copy
This is the full response shape, with both document types represented. It is the contract WiseOCR's POST /file endpoint returns, documented at developers.wiseocr.com — but the structure generalises, and the reasoning behind each choice is in the next section.
{
"data": {
"documentType": "invoice",
"documentNo": "FV/2026/09/114",
"vendor": "Meridian Logistics Sp. z o.o.",
"taxNo": "PL5213003700",
"time": "2026-09-14T13:42",
"currency": "PLN",
"country": "PL",
"items": [
{
"name": "Freight — Gdansk to Krakow",
"quantity": "1",
"unitPrice": "1250.00",
"price": "1250.00",
"priceNet": "1016.26",
"tax": "233.74"
},
{
"name": "Fuel surcharge",
"quantity": "2",
"unitPrice": "43.75",
"unitPriceBeforeDiscount": "50.00",
"priceBeforeDiscount": "100.00",
"discount": "-12.50",
"price": "87.50",
"priceNet": "71.14",
"tax": "16.36"
}
],
"totalNet": "1087.40",
"totalTax": "250.10",
"taxDetails": [
{ "rate": "0.23", "amount": "250.10" }
],
"total": "1337.50"
}
}
A receipt returns the same envelope with documentType: "receipt", no documentNo or taxNo, and line items reduced to name, quantity, unitPrice and price.
Five decisions you have to make before you design the schema
1. Money as strings, integers or floats
JSON numbers are IEEE-754 doubles. 19.99 is not exactly representable in one, and 0.1 + 0.2 famously is not 0.3. For money you have three honest options: decimal strings ("1337.50"), integer minor units (133750), or floats and a tolerance everywhere downstream.
The schema above uses decimal strings, including for tax rates ("0.23", not 23). It costs you a Decimal(...) cast at the boundary and buys exactness — and it survives currencies with zero or three decimal places, where "cents" is the wrong unit. Whatever you pick, pick it once. A schema that mixes strings and numbers for amounts is the worst of the three.
2. One timestamp field, not a date and a time
Receipts usually print both; invoices usually print only a date. Two nullable fields means four states to handle. One field in YYYY-MM-DD or YYYY-MM-DDTHH:mm form means two, and the presence of T tells you which you got.
3. Normalise to standards, not to what the paper says
Currency as ISO 4217 (PLN), country as ISO 3166-1 alpha-2 (PL), dates as ISO 8601. The document might say "zł", "PLN", or nothing at all and rely on the language — the extraction layer's job is to resolve that, so your code never sees a currency symbol.
4. Omit or null — but be consistent
Two defensible conventions: always emit every key and use null, or omit keys that do not apply. The one to avoid is mixing them, because then "tax" in data and data["tax"] is not None mean different things in different places.
5. Decide which total is the total
"Amount due", "amount paid", "grand total", "balance" and "subtotal" all appear on real invoices, sometimes on the same one. Pick the gross payable amount as total and keep the net separately. Do not let the field's meaning depend on the document.
Line items: the part that actually fails
Everything above is solvable with a good prompt and a strict schema. Line items are where extraction genuinely struggles, and where you should concentrate your testing.
Discounts. A line can carry a pre-discount unit price, a discount amount and a post-discount total. If you only model unit_price and amount, you silently lose the discount or, worse, record the pre-discount figure as the charge. The schema above keeps unitPriceBeforeDiscount, priceBeforeDiscount and a negative discount — and emits them only when a discount actually exists.
quantity × unitPrice ≠ price. Weighted goods (0.482 kg of tomatoes), multi-buy offers and per-line rounding all break the identity. Treat a mismatch as a signal to review, never as a reason to overwrite the printed figure.
Tables that span pages. A multi-page invoice repeats its header row on every page. Extraction that works page-by-page will either duplicate that header as an item or lose the continuation rows.
Empty is a valid answer. Taxi, parking and fuel receipts frequently have no itemisation at all. An empty items array is correct output, not a failure — make sure your downstream code agrees.
PDFs: text layers, scans and multi-page invoices
"Invoice PDF to JSON" is a different problem from "receipt photo to JSON", and conflating them is a common source of bad accuracy numbers.
A PDF generated by accounting software has a text layer. You can read it directly, exactly, with no OCR at all. A PDF that is a scan is just an image in a wrapper and needs the full pipeline. Many real inboxes contain both, sometimes in the same file — a digital invoice with a scanned delivery note appended.
Two practical consequences. First, detect the text layer before you spend money on OCR. Second, price and rate-limit per page, not per file: a fifteen-page invoice is fifteen units of work, and an API that charges per upload will surprise you. WiseOCR converts PDF pages to images before OCR and bills per page for exactly this reason.
Validating the output before it reaches your ledger
This is the layer almost nobody ships, and it is about twenty lines of code. The extractor's job is to report what the document says. Your job is to decide whether to trust it.
from decimal import Decimal
def check(doc):
problems = []
total = Decimal(doc["total"])
net, tax = doc.get("totalNet"), doc.get("totalTax")
if net and tax and Decimal(net) + Decimal(tax) != total:
problems.append("net + tax does not equal total")
for tax_line in doc.get("taxDetails") or []:
if Decimal(tax_line["amount"]) < 0:
problems.append(f"negative tax at rate {tax_line['rate']}")
items = doc.get("items") or []
if items:
summed = sum(Decimal(i["price"]) for i in items if i.get("price"))
if abs(summed - total) > total * Decimal("0.02"):
problems.append(f"items sum to {summed}, document says {total}")
return problems
Three rules make this useful rather than noisy. Compare against the printed total, never a recomputed one. Allow a tolerance on the item sum — service charges and deposits legitimately break it. And route failures to a human queue instead of rejecting them, because a receipt that fails this check is usually still 90% correct.
Getting this JSON without building the pipeline
If you want the schema above without assembling OCR, a model, a prompt, retries and normalisation yourself, this is one request:
curl -X POST https://api.wiseocr.com/v1/file \
-H "Authorization: $WISEOCR_API_KEY" \
-F "file=@invoice.pdf"
Add -F "skipItems=true" when you only need header data — it is meaningfully faster. The response is the object shown above, plus a usage block telling you how many credits the call consumed and how many remain.
You can try it on a sample document without an account using the receipt or invoice to JSON converter. If you would rather not write code at all, the Make.com integration routes the same JSON into spreadsheets, accounting systems and 2,000+ other apps.
Whichever route you take, keep the validation layer. It is the cheapest insurance in the whole pipeline.
WiseOCR turns receipts and invoices into structured JSON from $0.03 per page. Try the live demo on a document of your own.
