From solutions engineer to FDE: rebuilding a PoC to production standard
The demo that won the client over will break on the first malformed CSV. As an FDE, you are the one who has to fix it.
In brief
- Moving from solutions engineer to FDE means moving from demo code to production code, and taking responsibility for how the deployment turns out.
- The most effective practice: take an existing PoC, add error handling, tests and architecture documentation, then deploy it to a real environment.
- In an FDE interview you must explain why you chose a deployment approach, not just show that the code runs.
Becoming an FDE means moving from the left column to the right: the code has to survive reality, and you own the running system.
Graphic: FDE Times
You have just demoed an LLM-powered tool that classifies support tickets, and the client is delighted. The following Monday it goes into production. At three in the morning a CSV file arrives with a blank row, the model replies with a greeting instead of JSON, and the whole pipeline stops.
If you are a solutions engineer, your job usually ends when the contract is signed. As an FDE, you are the one who gets up at 3am to fix it.
fde.academy’s career-switching guide reduces the difference to two shifts: you write production code rather than demo code, and you own the outcome of the deployment rather than handing it over once the sale is made.
The good news is that the gap has a clear shape, and you can practise each part of it separately. What follows takes a very short PoC, brings it up to production standard step by step, and then shows how to present that skill on a CV and in interviews.
Owning the design is not owning the system
Aced, comparing the two roles, describes FDEs as spending most of their working week writing production code. A solutions architect, by contrast, owns the design, not the running system. If the system breaks or fails to deliver value, fixing it is the FDE’s direct responsibility.
Sundeep Teki, in his 2026 FDE interview guide, puts it more bluntly. You do not file a ticket or blame another team; you fix it yourself. For people coming from consulting, this is the biggest change, bigger than learning new syntax or frameworks.
fde.academy has a line worth remembering: a PoC only has to prove a point, while production code has to survive reality. “Reality” here means dirty data, flaky networks, APIs that return things you did not anticipate, and someone else reading your code six months from now.
Where does a six-line PoC break?
Picture the demo code for the client meeting above:
import csv, json
from llm import client
for row in csv.DictReader(open("tickets.csv")):
out = client.complete(f"Classify this ticket: {row['text']}")
print(row["id"], json.loads(out)["label"])
(The prompt string is Vietnamese for “Classify ticket:”.) On ten hand-picked sample tickets it runs smoothly. Put it in a real environment, though, and at least five things break. Empty tickets are still sent to the model, a ticket several pages long can exceed the context limit, and a single API timeout kills the whole loop.
If the model returns plain text, json.loads throws. And if the model invents a label that does not exist, that label quietly flows on into downstream systems.
The rewrite does not need to be cleverer. It only needs to treat every possible failure as something that will certainly happen. First, declare the valid labels explicitly, along with a dedicated logger:
import json, logging, time
from llm import client
LABELS = {"billing", "bug", "account", "other"}
log = logging.getLogger("triage")
Next, move a single model call into its own function, truncating the input, setting a timeout and validating the returned label:
def call_model(text: str) -> str | None:
out = client.complete(f"Classify this ticket: {text[:4000]}", timeout=20)
label = json.loads(out).get("label")
if label in LABELS:
return label
log.warning("unknown label: %r", label)
return None
The warning reads “unknown label”. Finally, the classify function handles empty input, retries with backoff, and provides an exit when every attempt fails:
def classify(text: str, retries: int = 3) -> str:
if not text or not text.strip():
return "other"
for attempt in range(retries):
try:
label = call_model(text)
if label:
return label
except (TimeoutError, json.JSONDecodeError) as e:
log.warning("attempt %d failed: %s", attempt + 1, e)
time.sleep(2 ** attempt)
return "needs_review"
Note the last line. The system does not crash; it routes the ticket to a needs_review queue for a human to check.
That is a business decision, and you have to agree it with the client: how many tickets a day are they willing to see land in that queue? Production code is full of questions like this; demo code skips them.
Then come tests for exactly the cases just listed:
def test_empty_ticket_goes_to_other():
assert classify(" ") == "other"
def test_non_json_reply_goes_to_review(monkeypatch):
monkeypatch.setattr(client, "complete", lambda *a, **k: "hello there")
assert classify("I was charged twice", retries=1) == "needs_review"
Here the mocked model replies “hello”, and the test ticket reads “I was charged twice”. These two tests take less than ten minutes to write. But they turn a 3am incident into a warning line in the log that you read at 9am.
Four steps to take a PoC to production
fde.academy’s most practical advice is to take an existing PoC and rebuild it to production standard. They break the work into four steps: add error handling, write tests, write architecture documentation, and deploy to a real environment. The example above has only completed the first two.
The third step is often underrated. Architecture documentation need not be long; it just has to answer where data comes in, where it goes, which external services it calls, and where to look in the logs when it fails. For the ticket classifier, a single page like this is enough to start:
# triage-service
Input: tickets from the customer's system (id, text)
Output: one label from {billing, bug, account, other, needs_review}
External calls: LLM API, 20-second timeout, up to 3 retries
When it breaks: check the "triage" logger, count needs_review tickets per day
Operator: your name, and how to reach you when something goes wrong
In English the fields read: in (tickets from the client’s system, with id and text); out (one label from the set); external calls (the LLM API, 20-second timeout, at most 3 retries); on failure (check the “triage” logger and count needs_review tickets per day); operator (your name and how to reach you during an incident). The readers of this page may well be engineers on the client side who will run the system alongside you. The last line matters most, because it states plainly who owns the system.
The fourth step is the real test of your ability. Coursera’s article on AI engineers lists turning models into APIs that other applications can integrate with, and managing production infrastructure, as core responsibilities. For the example above, wrapping classify as an endpoint in a file called api.py takes only a few lines:
from fastapi import FastAPI
from pydantic import BaseModel
from triage import classify
app = FastAPI()
class Ticket(BaseModel):
id: str
text: str
@app.post("/classify")
def classify_ticket(t: Ticket) -> dict:
return {"id": t.id, "label": classify(t.text)}
@app.get("/health")
def health() -> dict:
return {"ok": True}
The /health endpoint is what you hook up to monitoring to know the service is alive. But code sitting on a laptop does not count as deployed. The next step is to package the service so it runs identically everywhere, for instance with a short Dockerfile:
FROM python:3.12-slim
WORKDIR /app
COPY requirements.txt .
RUN pip install --no-cache-dir -r requirements.txt
COPY . .
CMD ["uvicorn", "api:app", "--host", "0.0.0.0", "--port", "8000"]
Then build it, run it, and test it yourself as an outside user would:
docker build -t triage-service .
docker run -d -p 8000:8000 -e LLM_API_KEY="$LLM_API_KEY" triage-service
curl http://localhost:8000/health
curl -X POST http://localhost:8000/classify \
-H "Content-Type: application/json" \
-d '{"id": "t-1", "text": " "}'
Note the -e flag: the API key comes in through an environment variable and never sits in the code or the image. The final curl sends an empty ticket, exactly the case that brought down the demo pipeline, and this time you should get back the label other.
Once the image runs cleanly on your machine, put it on a server or container service that others can reach over the network, and hook /health up to monitoring. At a client, that is usually infrastructure they already use, so the first question to ask is where the service will live and who is allowed to see the logs.
The interview tests exactly this gap
According to Sundeep Teki’s guide, the FDE coding round is roughly LeetCode medium in difficulty but set in a client context. At Palantir, the problem is framed around something you are building for an end user.
The technical deep dive lasts 60 minutes, involves no coding, and is scored on clarifying questions, root-cause analysis and how you weigh trade-offs.
Coursera notes the same of AI engineer interviews: you will have to explain why you chose a particular way to develop, deploy and scale a solution. The deep dive does not ask you to type code, but it still demands operational experience.
If you have never deployed anything yourself, the trade-off section exposes it at once: you can say “there should be retries”, but not how many retries, or why.
There is one more trap when reading job descriptions. Teki’s guide notes that Anthropic does not use the FDE title; it hires Solutions Architects for its Applied AI team, a pre-sales role whose job is to become the client’s trusted technical adviser. The title does not tell you who owns production.
When reading a JD, look for phrases such as “own deployment”, “on-call” and “ship to production”. If it is all “workshop” and “demo”, it is still a consulting role.
Mistakes career switchers often make
The most common is putting a polished demo on your CV instead of a running system. “Built chatbot demos for 5 clients” says less than “Deployed a ticket-classification API with tests, retries and a manual review queue, currently in operation”.
FDE recruiters want to see that you have owned a running system, not just a design.
The second mistake is using try/except to swallow every error so the code gets through. That hides errors rather than handling them. Every fallback branch should log, and should lead to a state someone can see, like needs_review in the example.
The third is believing you need a degree to switch. Coursera, writing about AI engineering in general rather than FDE specifically, argues that a degree is not essential and that career switchers can get in through projects and certifications.
If FDE is your target, choose a project like the example above: a service you deployed yourself, with tests and documentation to go with it.
The last mistake is subtler: when an incident hits, your reflex is to find the team responsible. As an FDE, that team is you.
The gap between consultant and FDE lies in the question you ask yourself after the demo. One asks whether the client liked it; the other asks what will break this system tonight.
Was this article useful?
Thanks for the feedback!