When the client's source team quietly changes the schema: writing data contracts that keep pipelines intact
At some point the client's source team will rename a field without warning. Whether your pipeline notices straight away depends on the data contract you wrote in week one.

In brief
- A schema is not yet a contract. A contract also needs change rules and a named owner.
- The contract must be checked at the system boundary. Start in warning mode, and only later block writes.
- Renaming or removing a field that consumers still read is a breaking change. The safe way to do it is in two phases.
- 1Before the changeThe source writes only total_amount; the contract treats total_amount as required
- 2Phase 1: write bothThe source team writes total_amount and amount_vnd at the same time
- 3Consumer upgradesReads amount_vnd, falling back to total_amount only when the new field is missing
- 4Update the contractIn contract.yaml, total_amount becomes optional
- 5Phase 2: drop the old fieldOnce no consumer reads total_amount, the source team stops writing it
Consumers upgrade first; the source team drops the old field only once nobody reads it.
Graphic: FDE Times
A schema is not yet a contract. In its analysis of Kafka data contracts, Conduktor defines a contract as a schema plus rules about how it is allowed to change and a person who is responsible for it. Without those last two parts, all you have is a description of the data at one point in time.
As an FDE, you hit this gap early. You build pipelines on the client’s data, while the team that develops the client’s source system does not report to you. They can rename a field on a Friday afternoon, and by Monday morning the dashboard is wrong and nobody knows why.
This guide walks you through building four things on your laptop: a contract file, a consumer-side test, a boundary-check configuration and a two-phase field change process. The running example is hypothetical: a retail chain stores its orders in a collection called orders (“orders” in Vietnamese) in a document database.
Why do document databases drift?
Document databases do not store data in fixed rows and columns; they store flexible documents. MongoDB’s documentation states plainly that documents in the same collection do not need to have the same fields. Wikipedia likewise describes documents in one collection as able to differ in both fields and structure.
So you cannot look at one sample document and conclude that it is the schema. Last year’s records may have total_amount while this week’s do not. To find out, you have to read real data before writing any rules.
What you need: Python 3, an editor, and read access to a sample of the client’s source data, or a few JSON documents you create yourself for practice.
Step 1: put only the fields you actually read into the contract
Open your consumer code and list the fields it accesses, not every field that exists in the source. In the example, the pipeline uses only order_id, total_amount and status, so the contract needs only those three.
MongoDB’s schema validation rules do not have to cover every field in a document. You can start with a small contract that is strict exactly where it matters and leaves the rest flexible.
Check: for each field on the list, you should be able to point to the line of code that reads it. Any field you cannot point to comes off the list.
Step 2: write a contract file with all three parts
The file below uses a home-made format, not any tool’s standard. What matters is that it contains all three parts of the definition at the top of this article. (The keys under evolution are Vietnamese: adding new fields is allowed with notice; removing a field or changing its meaning happens only in two phases; renaming counts as removing the old field plus adding a new one.)
# contract.yaml — a home-grown format, used to agree terms with the source team
dataset: orders
owner: "<name of the person on the client's source team>"
consumer_contact: "<your name>"
fields:
order_id: { type: string, required: true }
total_amount: { type: number, required: true, min: 0 }
status: { type: string, required: true, enum: [new, paid, cancelled] }
evolution:
add_new_field: "allowed, with advance notice"
remove_or_change_meaning_of_field: "only in two phases"
rename_field: "treated as removing the old field + adding a new field"
The owner line is the most important in the file. When the schema changes at midnight, it tells you whom to call. If the client cannot name anyone, record that in the meeting minutes as a project risk; do not fill in your own name.
Check: give the file to the source team and ask them to confirm it in writing. A contract signed only by you carries no weight.
Step 3: the consumer-side test
In consumer-driven contract testing, the consumer writes its own tests stating what it expects from the provider. Pact is a popular framework for this. The snippet below is a simplified plain-Python version meant to convey the idea; it is not Pact.
CONTRACT = {"order_id": str, "total_amount": (int, float), "status": str}
ALLOWED_STATUS = {"new", "paid", "cancelled"}
def check(doc):
errors = []
for field, typ in CONTRACT.items():
if field not in doc:
errors.append(f"missing {field}")
elif not isinstance(doc[field], typ):
errors.append(f"{field} wrong type")
if isinstance(doc.get("total_amount"), (int, float)) and doc["total_amount"] < 0:
errors.append("total_amount negative")
# only check the enum when status is present and has the right type, to avoid duplicate errors
if isinstance(doc.get("status"), str) and doc["status"] not in ALLOWED_STATUS:
errors.append("unknown status")
return errors
print(check({"order_id": "A1", "total_amount": 250000, "status": "paid"}))
print(check({"order_id": "A2", "amount_vnd": 250000, "status": "paid"}))
The first call returns []. The second returns ['missing total_amount'] (“missing total_amount”), which is exactly what happens when the source team renames a field without telling anyone. A document with no status at all gets a single missing status error, not an extra unknown status (“unknown status”).
Check: run the function against a sample of real data. If old documents are missing fields, you have just found where the schema drifted in the past. Record it before turning on blocking in the next step.
Step 4: enforce at the boundary, starting with warnings
Conduktor writes that a contract that cannot be enforced is merely a reminder and binds nobody. The test in step 3 runs on your side. A second line of defence belongs right where the data is written.
With MongoDB, you can turn on schema validation to lock down the schema when needed, using rules on data types and value ranges. By default, any insert or update that produces an invalid document is rejected. You can also configure it so invalid documents are still written, with a warning in the log.
# IDEA SKETCH — not MongoDB syntax.
# Look up "Schema Validation" in the MongoDB Manual for the exact syntax.
collection: orders
rules: order_id string required; total_amount number >= 0; status in enum
on_violation: warn # phase 1: log only
# once the logs have stayed clean for a while: switch back to the default (reject)
This is a sketch of the idea, not MongoDB syntax; look up “Schema Validation” in the MongoDB Manual for the exact form. Start with warnings because the source team may have write paths that neither side knows about. Switch on reject immediately and you may break the client’s operational system, which is the quickest way to lose their trust.
Check: let the warning log run for a while, review each type of violation with the source team, fix either the source or the contract, and only then switch to reject.
Step 5: what to do when the source team wants to rename a field
Suppose the source team wants to rename total_amount to amount_vnd. Conduktor’s principle is clear: do not remove or repurpose a field while any system still reads it. Instead, remove it in two phases: make the field optional first, then delete it.
In the example, phase one is the source team writing both fields. You change the consumer to read amount_vnd and fall back to total_amount only when the new field is missing, and you update contract.yaml so that total_amount becomes optional. Phase two begins once you have confirmed that no consumer still reads the old field. Only then does the source team stop writing total_amount.
The word BACKWARD is often misread in terms of which direction it checks. BACKWARD compatibility means the new schema can read old data, so the consumer side must upgrade first. The order in the example follows exactly that: you finish changing the consumer, and only then may the source team drop the old field.
Common mistakes
The first mistake is copying an entire sample document into the contract. The contract bloats, and every harmless change to a field you do not use becomes an incident. The second is switching on reject from day one, before you know how dirty the old data is.
The third is harder to spot: treating a rename as a minor change. To a consumer, a rename is the removal of a field. The last is leaving owner blank, at which point the contract is just documentation.
How this skill shows up at the client
In the first week of a deployment, the data contract is what opens the conversation with the source team. You do not arrive asking for access to the whole database. You arrive with three fields, one change rule and the question of who will sign.
When reading job descriptions, look for phrases such as “data quality”, “schema evolution” or “integrate with customer systems”. On your CV, do not write “experienced with MongoDB”. Write that you renamed a field on production data in two phases without stopping the pipeline.
Before an FDE interview, prepare a real story rather than a definition of a schema: the last time a client’s data changed, did you spot it first, or did the client?
Was this article useful?
Thanks for the feedback!