FDE PulseFDE jobs open 434New in the last 7 days 27
VI

The newspaper of the Forward Deployed Engineer

Guides

Debugging without access to a customer's production: finding faults with logs, data samples and reproductions

When you are the only person who can see a system running, your most valuable skill is turning what you see into a reproduction that someone far away can run straight away.

Kỹ sư ngồi trước laptop hiển thị log và mã nguồn, làm việc tại văn phòng hoặc phòng máy chủ của khách hàng.
Photo: Rubin Observatory/NSF/AURA / CC BY 4.0

In brief

  • Staging rarely matches production exactly, and monitoring tells you something is wrong without giving you enough data to reproduce it.
  • Go from logs to a hypothesis, then from the hypothesis to a self-contained reproduction that runs with a copy-paste.
  • When you escalate, send the evidence you gathered and the hypotheses you ruled out. Don't just forward the alert.
ShareLinkedInFacebookX
GraphicFrom alert to a reproduction you can take off site
  1. 1Narrow down with metricsDivide and conquer: find the layer where the data starts to go wrong
  2. 2Form a hypothesisState one specific cause at a time that the logs can confirm or rule out
  3. 3Test against the logsRecord rejected hypotheses too, because they are also evidence
  4. 4Record the data's shapeInspect samples on site; copy the structure, never the real values
  5. 5Write a minimal reproGenerate fake data, strip unneeded lines, make it run with a copy-paste
  6. 6Escalate with evidenceSend what you saw, what you ruled out and the command to run the repro

What leaves the customer's environment is the shape of the bug, not the real data.

Graphic: FDE Times

It is three in the afternoon, and the customer reports that their ingestion pipeline has been quietly dropping some of its records since the morning. You are on site with them, but the engineering team at headquarters has no way to look at that system. All they know is what you tell them.

This situation appears word for word in one job description. Twenty’s posting for a Senior/Staff Forward Deployed Software Engineer says the office-based team cannot access this environment, so the FDE is the only engineer present where the software actually runs.

That means fixing bugs is only part of the skill. You also have to get the bug out of the customer’s environment in a form other people can run, read and verify, without taking a single byte of the customer’s data with you.

Why can’t you rely on staging and alerts?

Many engineers start with “let me try it on staging”. An article by Speedscale, a company that sells bug-reproduction tools, notes that even where a staging environment exists, it rarely matches production exactly. Configuration, service versions and above all the shape of real data usually differ by enough to make the bug disappear.

Monitoring will not save you either. The same article points out that monitoring tells you something has happened, but does not give you the full requests and responses you need to rebuild the bug. An alert is only a starting point.

The gap between knowing there is a bug and being able to reproduce it is where the FDE’s work lies. Three tools close it: logs to point the way, data samples to understand the shape of the input, and a reproduction to prove it.

Logs are for ruling things out, not for reading cover to cover

The Effective Troubleshooting chapter of Google’s SRE book describes troubleshooting as a loop: form a hypothesis about the cause, then test it. The logs are where you test it. They are not a novel to read from the first line.

The same chapter calls divide and conquer a very useful general-purpose problem-solving technique. In a multi-stage pipeline, you ask at which stage the records still exist and at which stage they vanish. Each answer cuts the area under suspicion in half.

At Twenty’s sites, FDEs monitor platform health and data flows with the LGTM stack: Grafana, Loki, Tempo and Mimir. If your site has similar tools, use them: metrics to narrow down the time window, traces to follow a request across services, and logs to read the detail at the exact point of failure.

A worked example, start to finish

Back to the pipeline that is losing records. This is a hypothetical scenario, built to show the method. Metrics show that the number of records entering the parse stage equals the number leaving the ingest stage, but fewer records leave the parse stage. The area under suspicion has shrunk to one stage.

First hypothesis: the parse service restarted partway through. You filter the logs around that time and find no restart. The hypothesis is rejected, and you write that down. The SRE book stresses that negative results should not be ignored or discounted, because they tell the next person not to go down the same path.

Second hypothesis: some records have empty fields. You inspect a few of the dropped records on site, without copying them out, and notice they share one trait: the quantity field is an empty string rather than a number. You record only the shape of the data, not the real values.

Only now do you write the reproduction. The scikit-learn documentation advises that a minimal reproduction should rely on a small dataset generated by the code at runtime rather than external data. It also notes that most bugs do not depend on the particular structure of the real data, so synthetic data is usually enough.

import json
from ingest.parse import parse_event  # module in the team's repo

# Synthetic data: keeps the shape only, none of the customer's real values
rows = [
    {"id": "r1", "quantity": "3", "ts": "2026-01-01T00:00:00Z"},
    {"id": "r2", "quantity": "",  "ts": "2026-01-01T00:00:01Z"},
]

for r in rows:
    print(r["id"], parse_event(json.dumps(r)))
# Expected: both return an event
# Actual (hypothetical): r2 returns None and logs no error

Look at what has been cut: no queue connection, no retry configuration, no file reads. According to scikit-learn, removing irrelevant lines and dropping non-default options helps you and others narrow down the cause. If you remove a line and the bug is still there, that line is not the culprit.

Send evidence, not a forwarded alert

When you escalate to the team at headquarters, you need to bring a clear picture of what you found, not just what triggered the alert. A good escalation note is therefore structured very differently from a message saying “the system’s broken, can someone take a look”.

Forwarding the alert

  • Pipeline has been losing records since this morning
  • Screenshot of the dashboard
  • Ask: can someone take a look?

Reporting the evidence

  • Lost at the parse stage; ingest is complete
  • Ruled out: service restart, with the time window checked
  • Dropped records have quantity as an empty string
  • 15-line repro file that runs with one command

The most important part is the repro file. Scikit-learn observes that reproduction steps written in prose are often ambiguous. A snippet the recipient can copy, paste and immediately see fail saves many rounds of back-and-forth, and the SRE book likewise says that a reliable reproducing test case makes debugging much faster.

Your own code must be debuggable from afar

The problem also runs the other way. Code you write for a customer site must be testable, debuggable and maintainable by engineers who cannot get into the environment where it runs, and Twenty’s job description lists this as a requirement. In the example above, the bug was hard to find precisely because parse_event returned None without leaving a single log line.

When writing code, ask yourself: if this function fails, will the logs tell someone far away which record it failed on and why? One log line that records the record ID and the name of the offending field, without any sensitive values, can save an entire afternoon of investigation.

Common mistakes

The most common mistake is trusting staging. A bug that does not appear on staging has not stopped existing; quite possibly the staging data simply lacks the shape that triggers it.

The second is forgetting the hypotheses you ruled out. If you do not record that you checked for a restart, a colleague at headquarters will waste time checking exactly that again.

The third is a reproduction that is too big: one that drags in production configuration, reads sample files nobody else has, or worse, contains real data. If the recipient cannot run the file on their own machine straight away, the reproduction has not done its job.

How to show this skill when job hunting

If you are a developer looking to move into FDE work, watch for phrases in job descriptions such as “on-site”, “air-gapped”, “cleared environment”, or requirements around an observability stack. They signal that the role demands exactly this skill.

On your CV, don’t just write “debugged production issues”. Describe a time you narrowed down a bug with logs or traces, rebuilt it with synthetic data and handed it to another team to fix. If you can, link to a reproduction you once submitted to an open-source issue on GitHub so hiring managers can see it for themselves.

A good FDE is not the person with the widest access. It is the person who walks out of the customer’s server room with only a few lines of code, and those lines are enough for the whole team to see a bug they will never get to touch.

4 sources
Read next on the roadmap · Stage 5: DeploymentSentry for FDEs: fixing bugs on a client site you are not atYou have no SSH access, no logs, and a client who just writes "it's broken". Sentry tells you what failed, where, and in which release.