Extracting invoices and contracts with LLMs: schemas, business-rule checks and source citations
JSON that parses is not the same as JSON that is correct. This guide builds a pipeline that knows when to trust a figure and when to send it to a human reviewer.
In brief
- Structured outputs guarantee the JSON structure, not the correctness of the values, so you still need to check stop_reason, normalise enums and apply business rules.
- A few lines of Pydantic validation can catch a total error that the schema lets through.
- Every extracted field should come with a source quote so a reviewer can verify it in seconds.
The invoice says the total is 3,850,000 dong. The model returns 3,580,000. The JSON is valid and parses without error. The client’s accounting system accepts the figure and nobody questions it.
If you build document extraction for clients, this is the failure to plan for from day one. Anthropic’s docs state that structured outputs use constrained decoding so that output matches the schema and can be parsed by downstream steps. But matching the schema is only the first condition.
A demo that returns clean JSON does not answer the two questions every consumer of financial data will ask: is the number right, and if so, which line of the source document does it come from? The guide below walks through five steps to answer both.
What will you build, and what do you need first?
The end result is a Python script that takes invoice text and returns JSON that has passed five checks: schema, stop_reason, enum normalisation, business rules and source citations. Any record that fails a step goes to a human reviewer rather than straight into the system.
You need Python 3, the anthropic SDK, the pydantic library and an API key. The code below is simplified for readability. The model name is left as MODEL_ID; before running it, check the model name and parameter structure against the Structured outputs page in Anthropic’s docs.
Step 1: the leaner the schema, the longer it lasts
Anthropic’s docs warn that a schema using unsupported features will get a 400 error with details. So do not pack business constraints into the schema, such as minimum or maximum on amounts. The schema only needs to describe the shape of the data; save the business rules for Step 4.
INVOICE_SCHEMA = {
"type": "object",
"properties": {
"invoice_number": {"type": "string"},
"issue_date": {"type": "string"},
"currency": {"type": "string", "enum": ["vnd", "usd"]},
"line_items": {"type": "array", "items": {
"type": "object",
"properties": {
"description": {"type": "string"},
"quantity": {"type": "number"},
"unit_price": {"type": "number"},
"amount": {"type": "number"}
},
"required": ["description", "quantity", "unit_price", "amount"],
"additionalProperties": False
}},
"subtotal": {"type": "number"},
"vat": {"type": "number"},
"total": {"type": "number"},
"evidence": {"type": "array", "items": {
"type": "object",
"properties": {
"field": {"type": "string"},
"quote": {"type": "string"}
},
"required": ["field", "quote"],
"additionalProperties": False
}}
},
"required": ["invoice_number", "currency", "line_items",
"subtotal", "vat", "total", "evidence"],
"additionalProperties": False
}
The evidence field is there on purpose: the model must copy, verbatim, the passage it relied on to fill each field. Step 5 puts it to use. Check: send a test request; if you get a 400 error, read the error details and remove the schema features the API does not accept.
Step 2: call the API and read stop_reason before reading the JSON
According to Anthropic’s docs, JSON outputs are configured through output_config.format and are meant for tasks such as extracting data from images or text. The other mode, strict tool use, is for validating tool parameters. For this guide, you use JSON outputs.
import json, anthropic
client = anthropic.Anthropic()
resp = client.messages.create(
model=MODEL_ID,
max_tokens=4096,
messages=[{"role": "user",
"content": "Trích xuất hoá đơn sau. Với mỗi trường, "
"chép nguyên văn đoạn chứa nó vào evidence.\n\n"
+ invoice_text}],
output_config={"format": {"type": "json_schema",
"schema": INVOICE_SCHEMA}},
)
if resp.stop_reason == "refusal":
route_to_human(invoice_text, reason="refusal")
elif resp.stop_reason == "max_tokens":
route_to_human(invoice_text, reason="truncated")
else:
data = json.loads(resp.content[0].text)
(The prompt reads: “Extract the following invoice. For each field, copy the passage containing it verbatim into evidence.”)
The first two if branches are the part many people skip. The docs are clear that when the model refuses, the refusal message takes priority over the schema constraint, so the output may not match the schema. When output is cut off by max_tokens, it may be incomplete and also fail to match the schema.
Long contracts and invoices with hundreds of line items are where the second branch tends to fire. Check: set max_tokens=50 on a long document and confirm the script takes the truncated branch instead of crashing in json.loads.
Step 3: normalise what the schema does not promise
Anthropic’s docs also note that structured outputs do not guarantee the case of enum and const values. You declare "vnd", but downstream code may still receive "VND". If the client’s system does exact string comparison, a difference that small can corrupt an entire batch.
The simple fix is to normalise with a field_validator inside the Pydantic model in the next step, rather than trusting the enum to always be right.
Step 4: let Pydantic catch the money errors
Back to the invoice at the top. There are 2 service lines at 1,500,000 and 1 line at 500,000, so the subtotal is 3,500,000. VAT at 10% is 350,000, giving a total of 3,850,000. The model misreads it as 3,580,000; the number type is still valid, so the schema lets it through.
from pydantic import BaseModel, field_validator, model_validator
class LineItem(BaseModel):
description: str
quantity: float
unit_price: float
amount: float
class Invoice(BaseModel):
invoice_number: str
currency: str
line_items: list[LineItem]
subtotal: float
vat: float
total: float
@field_validator("currency", mode="before")
@classmethod
def normalize_currency(cls, v):
return v.lower()
@model_validator(mode="after")
def check_totals(self):
s = sum(i.amount for i in self.line_items)
if abs(s - self.subtotal) > 1:
raise ValueError(f"subtotal {self.subtotal} != tổng dòng {s}")
if abs(self.subtotal + self.vat - self.total) > 1:
raise ValueError(f"total {self.total} != subtotal + vat")
return self
With 3,580,000, the second check sees that 3,500,000 + 350,000 does not equal 3,580,000 and raises an error. This is the kind of error constrained decoding will never catch, because it only cares about the shape of the data.
When validation fails, there are two ways to handle it. The first is to route the record to a reviewer. The second is to use the Instructor library: according to its introduction page, Instructor validates output with Pydantic and automatically re-asks the model when validation fails.
For financial data, cap the number of retries and log each one, so you can later explain why a figure was changed.
Step 5: every field must point to its source
The evidence field from Step 1 now has a job. A few lines of code check whether the quote the model supplied actually appears in the source text:
def unverified_fields(data, source_text):
norm = lambda s: " ".join(s.split())
src = norm(source_text)
return [e["field"] for e in data["evidence"]
if norm(e["quote"]) not in src]
This is a simplified approach: it only normalises whitespace, so OCR text with spelling errors will trigger false alarms.
If the client needs stricter verification, Anthropic’s Citations API returns the exact passages that support each claim, so you can verify them and show the sources to users.
The current docs state that all active models support citations.
A cautious design splits the work into two separate calls: the first extracts with structured outputs, the second re-asks about the critical fields (totals, contract expiry dates, penalty clauses) with citations enabled. Before combining both features in a single request, check the docs to see whether they can be used together.
What are the most common mistakes?
The most common mistake is treating a successful json.loads as the job being done. The second is stuffing business rules into the schema, hitting a 400 error, then losing an afternoon stripping it back piece by piece. The third is testing only on short invoices, so the first time the max_tokens branch fires is on the client’s real data.
If the client uses OpenAI, the principles still hold. OpenAI recommends Structured Outputs over JSON mode, because both produce valid JSON but only Structured Outputs guarantees schema adherence. Responses also include a refusal field so code can detect when the model declines.
One small detail to remember before a demo. Ted Sanders of OpenAI has said that the first request with each JSON schema will be slow, because the schema must be preprocessed into a context-free grammar. Make a test call before the client walks into the meeting room.
What does this skill look like on a client site?
Picture a logistics company that receives several thousand supplier invoices a month. The head of accounting’s first question will probably not be about the accuracy percentage.
It will be: when the machine gets it wrong, who finds out, and how? The five-step pipeline above answers that: every record either goes straight into the system or into a review queue with a specific reason.
That queue deserves as much design care as the code. Each item should show the source text, the extracted JSON, the reason it was blocked (refusal, truncated, a total mismatch or an unverified field) and the evidence quote for each field, so a reviewer looking at the 3,580,000 invoice only needs to glance at the total line to fix it.
When reading FDE job descriptions, look for phrases such as “document processing”, “data extraction” or “human-in-the-loop”; they signal work like this. On your CV, do not just write “used LLMs to extract invoices”.
Spell out that you check stop_reason, validate totals and attach source citations to every field, because those details show you have thought about when the system fails, not only when it works.
Anyone can build a demo that returns clean JSON. Clients sign with the person who can show what the system does when it meets 3,580,000.