FDE PulseFDE jobs open 441New in 7 days 29Companies hiring 47Remote-friendly 24%Median US pay $216kTop hirer Databricks 125
VI

The newspaper of the Forward Deployed Engineer

Guides

A one-line prompt edit is a deploy: running a 5-20-50-100% canary at the client

A prompt edit that only changes formatting can shift accuracy by around 5%. At a client, every change like that needs a feature flag to split traffic, a control group to compare against and a kill switch that is always ready.

A one-line prompt edit is a deploy: running a 5-20-50-100% canary at the client
Photo: Luke Chesser / Unsplash

In brief

  • A prompt or model that works well in staging can still make the experience worse in production, so every change needs a flag and a control group.
  • Log the variation and prompt_version when the LLM is called. If you log on page load, your canary numbers will be wrong.
  • A small canary can miss the rare questions the model gets wrong, so move up a step only once you have enough samples, not when the dashboard looks fine.
ShareLinkedInFacebookX
GraphicAt each step: promote or roll back?
Promote to the next stepRoll back or wait
Error rateNo more than 1.5x control (example threshold)Above 1.5x control: roll back immediately
p95 latencyNo more than 1.3x control (example threshold)Above 1.3x control: roll back immediately
Handoff to a human agentNo more than 3 percentage points above controlMore than 3 percentage points above: roll back
Quality scoreNo more than 0.05 below control on sampled conversationsMore than 0.05 below: roll back
Sample sizeAt least 500 conversations and 20 refund questions in the canaryNot enough yet: wait, do not raise the percentage

A canary moves up a step only when it breaches no guardrail relative to control and has collected enough samples, including rare questions.

Graphic: FDE Times

Picture this: a client asks you to tweak the prompt so their support chatbot sounds “more polite”. You add a few greetings, switch bullet points to a numbered list and test it in staging. Everything looks fine. Two days later, the customer service team reports that the bot is getting the refund policy wrong.

This scenario is hypothetical, but not far-fetched. LaunchDarkly warns that a small prompt edit or a model swap can work well in staging and still degrade the user experience in production. An analysis on the tianpan.co blog found that minor changes to prompt formatting shifted accuracy by around 5%.

So switching bullets to numbers is a deploy. This guide walks you through building the whole process yourself: a feature flag that splits traffic between the old and new prompts, logging to compare the two groups, a rollback decision function and a schedule that ramps up through 5-20-50-100%.

The code here is illustrative Python, trimmed down and not tied to any vendor’s SDK. In a real engagement, swap it for whatever flagging tool the client already uses.

A canary is really a comparison

The Google SRE Workbook defines canarying as deploying a change to a portion of a service for a limited time and evaluating that change. The portion that receives the change is the canary; the rest is the control. Without a control group alongside it, you are only running a trial, not a canary.

What you need: Python 3, an existing function that calls your LLM (called call_llm here), a queryable place to write logs, and two clearly named prompt versions, for example v1.3 and v2.0.

Step 1: Bundle prompt and model into variations

Do not edit the prompt string directly in code. Declare each configuration as a named variation, so that rolling back means changing a selection rather than redeploying.

FLAG = {
    "name": "support_bot_prompt_v2",
    "enabled": True,           # emergency kill switch
    "rollout_percent": 5,
    "internal_users": {"qa@client.vn"},
}

VARIATIONS = {
    "control": {"model": "current-model", "prompt_version": "v1.3"},
    "canary":  {"model": "current-model", "prompt_version": "v2.0"},
}

(The comment on enabled marks it as the emergency off switch; current-model means “current model”.) In this example the two variations differ only in the prompt. If you want to change the model, do it under a separate flag. When prompt and model change together, you cannot tell which one caused the result.

Step 2: Split users stably, internal users first

Flagsmith describes its own process for LLM features like this: first enable the feature in production for internal users only, then open it to 5% of users, then 50% to run an A/B test, and only then go to 100%. The code below follows that order.

import hashlib

def bucket(user_id: str, flag_name: str) -> int:
    h = hashlib.sha256(f"{flag_name}:{user_id}".encode()).hexdigest()
    return int(h[:8], 16) % 100

def pick_variation(user_id: str, email: str) -> str:
    if not FLAG["enabled"]:
        return "control"
    if email in FLAG["internal_users"]:
        return "canary"
    if bucket(user_id, FLAG["name"]) < FLAG["rollout_percent"]:
        return "canary"
    return "control"

Check: run the function on 10,000 fake user_ids and count how many land in the canary; the result should be around 500. Calling it again with the same user_id must always return the same variation. Without that stability, one person could get an answer from the old prompt on one turn and the new prompt on the next, and your data becomes noisy.

Step 3: Log at the right moment

FeatBit recommends recording the variation at the moment the LLM actually runs, not when the page loads. It also recommends logging promptVersion alongside it, so that if someone quietly edits the prompt mid-rollout, the change does not contaminate the canary results.

import time

def answer(user_id, email, question):
    variation = pick_variation(user_id, email)
    cfg = VARIATIONS[variation]
    started = time.time()
    reply = call_llm(cfg["model"], load_prompt(cfg["prompt_version"]), question)
    log_event({
        "user_id": user_id,
        "flag": FLAG["name"],
        "variation": variation,
        "prompt_version": cfg["prompt_version"],
        "model": cfg["model"],
        "latency_ms": int((time.time() - started) * 1000),
    })
    return reply

call_llm, load_prompt and log_event are functions you already have. Check: every log line must contain all four fields: variation, prompt_version, model and latency. Missing even one means you cannot compare the two groups reliably. In practice, also log the question type (for example refund) so the next step can count samples per type.

Step 4: Track a few metrics, set thresholds in advance

Google SRE advises choosing only the few most important metrics to evaluate a canary, probably no more than a dozen. For a customer support chatbot, you might start with error rate, latency, the rate at which users have to be handed off to a human agent, and a quality score graded on a sample of conversations.

The function below returns one of three decisions: roll back, wait, or promote. Guardrails are checked first, because a clear failure should stop things immediately. To promote, the canary must also have collected enough samples, both in total and for rare question types.

MIN_TOTAL = 500    # minimum total conversations in the canary
MIN_REFUND = 20    # minimum number of refund questions

def decide(canary: dict, control: dict) -> str:
    if canary["error_rate"] > control["error_rate"] * 1.5:
        return "rollback"
    if canary["p95_latency_ms"] > control["p95_latency_ms"] * 1.3:
        return "rollback"
    if canary["handoff_rate"] > control["handoff_rate"] + 0.03:
        return "rollback"
    if canary["quality_score"] < control["quality_score"] - 0.05:
        return "rollback"
    if canary["n_total"] < MIN_TOTAL or canary["n_refund"] < MIN_REFUND:
        return "wait"
    return "promote"

(MIN_TOTAL is the minimum total number of canary conversations; MIN_REFUND the minimum number of refund questions.) Every threshold above is only an example. Agree the real thresholds with the client before switching the canary on, rather than setting them once the numbers are in. Commercial tools take the same approach: LaunchDarkly’s guarded rollouts for AI Configs let you attach guardrail metrics and automatically roll back to the previous variation.

Step 5: How long should each step run?

Why start at 5% rather than 1%? The tianpan.co piece argues that for LLM features, a 1% canary may never reach the tail of the input distribution, the rare cases where the model performs worse.

The author also notes that an LLM canary may need to run for hours or even days to collect enough labelled or implicitly labelled data.

Go back to the hypothetical scenario at the start. Suppose the client’s system handles 2,000 conversations a day, and refund questions make up 2% of traffic. At 5%, the canary receives 100 conversations a day, of which only about 2 concern refunds.

With MIN_REFUND = 20, the 5% step needs around 10 days before decide() can return “promote”. At 1%, the figure is 0.4 refund questions a day, which is close to seeing nothing.

So the duration of each step should be set by the number of samples you need, not by hours. If 10 days is too long, that is a conversation to have with the client at the outset.

Step Who gets the canary Condition to move on (suggested)
Internal The client’s QA team No serious errors found on manual review
5% A small share of real users Enough samples for rare question types, no guardrail breached
20% Broader, with clearer comparison data Quality metrics no worse than control
50% A/B test, two groups of roughly equal size Client signs off on the results
100% All users Keep the flag for a while longer so you can still roll back

Flagsmith names only the 5%, 50% and 100% steps. The 20% step in the table is an added buffer that gives you clearer comparison data before running the A/B test at 50%.

Step 6: Rehearse the kill switch

Flagsmith writes that as soon as they spot a problem, they can turn the flag off instantly. With the code above, setting FLAG["enabled"] = False sends every user back to control. Try this once in staging with the client’s engineers, and record how many seconds it takes for the change to take effect.

Three common mistakes

The first is changing prompt and model under the same flag, so when results turn bad you do not know which to revert. The second is logging only the canary group, which leaves no control data when it is time to compare.

The third is jumping from 5% to 50% after a single afternoon because the dashboard looks fine, while the rare questions have not appeared even once. The “wait” branch in decide() exists precisely to prevent this.

How this skill shows up on the job

A request like “can you tweak the prompt so the bot is more polite” is exactly the kind of change this process is built for. Before taking the work on, have answers ready to two questions: if it goes wrong, how do you turn it off, and how long does that take?

When reading job descriptions, look for phrases such as progressive delivery, feature flags, A/B testing or LLM evaluation. On your CV, do not just write “used LaunchDarkly”. Be specific: you separated prompt_version into the logs, set guardrail thresholds and minimum sample sizes, took a change to 100% through four steps, and had a kill switch that took effect within a stated number of seconds.

Next time a client asks you to swap the model, reply with a staged rollout schedule, not just a note saying it is done.

Was this article useful?

Use with your AI assistantAsk Claude ↗Ask ChatGPT ↗
5 sources
Read next on the roadmap · Stage 5: DeploymentSizing GPUs for a 70B model: from parameters to KV cache and quantisationWhen a client asks how many GPUs they need to run a model in their own datacentre, you should have the answer before you leave the meeting, not after a failed deployment.