A one-line prompt edit is a deploy: running a 5-20-50-100% canary at the client
A prompt edit that only changes formatting can shift accuracy by around 5%. At a client, every change like that needs a feature flag to split traffic, a control group to compare against and a kill switch that is always ready.
In brief
- A prompt or model that works well in staging can still make the experience worse in production, so every change needs a flag and a control group.
- Log the variation and prompt_version when the LLM is called. If you log on page load, your canary numbers will be wrong.
- A small canary can miss the rare questions the model gets wrong, so move up a step only once you have enough samples, not when the dashboard looks fine.
A canary moves up a step only when it breaches no guardrail relative to control and has collected enough samples, including rare questions.
Graphic: FDE Times
Picture this: a client asks you to tweak the prompt so their support chatbot sounds “more polite”. You add a few greetings, switch bullet points to a numbered list and test it in staging. Everything looks fine. Two days later, the customer service team reports that the bot is getting the refund policy wrong.
This scenario is hypothetical, but not far-fetched. LaunchDarkly warns that a small prompt edit or a model swap can work well in staging and still degrade the user experience in production. An analysis on the tianpan.co blog found that minor changes to prompt formatting shifted accuracy by around 5%.
So switching bullets to numbers is a deploy. This guide walks you through building the whole process yourself: a feature flag that splits traffic between the old and new prompts, logging to compare the two groups, a rollback decision function and a schedule that ramps up through 5-20-50-100%.
The code here is illustrative Python, trimmed down and not tied to any vendor’s SDK. In a real engagement, swap it for whatever flagging tool the client already uses.
A canary is really a comparison
The Google SRE Workbook defines canarying as deploying a change to a portion of a service for a limited time and evaluating that change. The portion that receives the change is the canary; the rest is the control. Without a control group alongside it, you are only running a trial, not a canary.
What you need: Python 3, an existing function that calls your LLM (called call_llm here), a queryable place to write logs, and two clearly named prompt versions, for example v1.3 and v2.0.
Step 1: Bundle prompt and model into variations
Do not edit the prompt string directly in code. Declare each configuration as a named variation, so that rolling back means changing a selection rather than redeploying.
FLAG = {
"name": "support_bot_prompt_v2",
"enabled": True, # emergency kill switch
"rollout_percent": 5,
"internal_users": {"qa@client.vn"},
}
VARIATIONS = {
"control": {"model": "current-model", "prompt_version": "v1.3"},
"canary": {"model": "current-model", "prompt_version": "v2.0"},
}
(The comment on enabled marks it as the emergency off switch; current-model means “current model”.) In this example the two variations differ only in the prompt. If you want to change the model, do it under a separate flag. When prompt and model change together, you cannot tell which one caused the result.
Step 2: Split users stably, internal users first
Flagsmith describes its own process for LLM features like this: first enable the feature in production for internal users only, then open it to 5% of users, then 50% to run an A/B test, and only then go to 100%. The code below follows that order.
import hashlib
def bucket(user_id: str, flag_name: str) -> int:
h = hashlib.sha256(f"{flag_name}:{user_id}".encode()).hexdigest()
return int(h[:8], 16) % 100
def pick_variation(user_id: str, email: str) -> str:
if not FLAG["enabled"]:
return "control"
if email in FLAG["internal_users"]:
return "canary"
if bucket(user_id, FLAG["name"]) < FLAG["rollout_percent"]:
return "canary"
return "control"
Check: run the function on 10,000 fake user_ids and count how many land in the canary; the result should be around 500. Calling it again with the same user_id must always return the same variation. Without that stability, one person could get an answer from the old prompt on one turn and the new prompt on the next, and your data becomes noisy.
Step 3: Log at the right moment
FeatBit recommends recording the variation at the moment the LLM actually runs, not when the page loads. It also recommends logging promptVersion alongside it, so that if someone quietly edits the prompt mid-rollout, the change does not contaminate the canary results.
import time
def answer(user_id, email, question):
variation = pick_variation(user_id, email)
cfg = VARIATIONS[variation]
started = time.time()
reply = call_llm(cfg["model"], load_prompt(cfg["prompt_version"]), question)
log_event({
"user_id": user_id,
"flag": FLAG["name"],
"variation": variation,
"prompt_version": cfg["prompt_version"],
"model": cfg["model"],
"latency_ms": int((time.time() - started) * 1000),
})
return reply
call_llm, load_prompt and log_event are functions you already have. Check: every log line must contain all four fields: variation, prompt_version, model and latency. Missing even one means you cannot compare the two groups reliably. In practice, also log the question type (for example refund) so the next step can count samples per type.
Step 4: Track a few metrics, set thresholds in advance
Google SRE advises choosing only the few most important metrics to evaluate a canary, probably no more than a dozen. For a customer support chatbot, you might start with error rate, latency, the rate at which users have to be handed off to a human agent, and a quality score graded on a sample of conversations.
The function below returns one of three decisions: roll back, wait, or promote. Guardrails are checked first, because a clear failure should stop things immediately. To promote, the canary must also have collected enough samples, both in total and for rare question types.
MIN_TOTAL = 500 # minimum total conversations in the canary
MIN_REFUND = 20 # minimum number of refund questions
def decide(canary: dict, control: dict) -> str:
if canary["error_rate"] > control["error_rate"] * 1.5:
return "rollback"
if canary["p95_latency_ms"] > control["p95_latency_ms"] * 1.3:
return "rollback"
if canary["handoff_rate"] > control["handoff_rate"] + 0.03:
return "rollback"
if canary["quality_score"] < control["quality_score"] - 0.05:
return "rollback"
if canary["n_total"] < MIN_TOTAL or canary["n_refund"] < MIN_REFUND:
return "wait"
return "promote"
(MIN_TOTAL is the minimum total number of canary conversations; MIN_REFUND the minimum number of refund questions.) Every threshold above is only an example. Agree the real thresholds with the client before switching the canary on, rather than setting them once the numbers are in. Commercial tools take the same approach: LaunchDarkly’s guarded rollouts for AI Configs let you attach guardrail metrics and automatically roll back to the previous variation.
Step 5: How long should each step run?
Why start at 5% rather than 1%? The tianpan.co piece argues that for LLM features, a 1% canary may never reach the tail of the input distribution, the rare cases where the model performs worse.
The author also notes that an LLM canary may need to run for hours or even days to collect enough labelled or implicitly labelled data.
Go back to the hypothetical scenario at the start. Suppose the client’s system handles 2,000 conversations a day, and refund questions make up 2% of traffic. At 5%, the canary receives 100 conversations a day, of which only about 2 concern refunds.
With MIN_REFUND = 20, the 5% step needs around 10 days before decide() can return “promote”. At 1%, the figure is 0.4 refund questions a day, which is close to seeing nothing.
So the duration of each step should be set by the number of samples you need, not by hours. If 10 days is too long, that is a conversation to have with the client at the outset.
| Step | Who gets the canary | Condition to move on (suggested) |
|---|---|---|
| Internal | The client’s QA team | No serious errors found on manual review |
| 5% | A small share of real users | Enough samples for rare question types, no guardrail breached |
| 20% | Broader, with clearer comparison data | Quality metrics no worse than control |
| 50% | A/B test, two groups of roughly equal size | Client signs off on the results |
| 100% | All users | Keep the flag for a while longer so you can still roll back |
Flagsmith names only the 5%, 50% and 100% steps. The 20% step in the table is an added buffer that gives you clearer comparison data before running the A/B test at 50%.
Step 6: Rehearse the kill switch
Flagsmith writes that as soon as they spot a problem, they can turn the flag off instantly. With the code above, setting FLAG["enabled"] = False sends every user back to control. Try this once in staging with the client’s engineers, and record how many seconds it takes for the change to take effect.
Three common mistakes
The first is changing prompt and model under the same flag, so when results turn bad you do not know which to revert. The second is logging only the canary group, which leaves no control data when it is time to compare.
The third is jumping from 5% to 50% after a single afternoon because the dashboard looks fine, while the rare questions have not appeared even once. The “wait” branch in decide() exists precisely to prevent this.
How this skill shows up on the job
A request like “can you tweak the prompt so the bot is more polite” is exactly the kind of change this process is built for. Before taking the work on, have answers ready to two questions: if it goes wrong, how do you turn it off, and how long does that take?
When reading job descriptions, look for phrases such as progressive delivery, feature flags, A/B testing or LLM evaluation. On your CV, do not just write “used LaunchDarkly”. Be specific: you separated prompt_version into the logs, set guardrail thresholds and minimum sample sizes, took a change to 100% through four steps, and had a kill switch that took effect within a stated number of seconds.
Next time a client asks you to swap the model, reply with a staged rollout schedule, not just a note saying it is done.
Was this article useful?
Thanks for the feedback!
5 sources
- Chapter 16 - Canarying Releases (Google SRE Workbook)
- Progressive Delivery for Building LLM-Powered Features
- Guarded Rollouts for AI Configs · 2025-07-24
- Why Gradual Rollouts Don't Work for AI Features (And What to Do Instead) · 2026-04-15
- LLM Guardrails With Feature Flags: Route, Compare, and Roll Back Safely · 2026-06-13