Simulated users: how to test a multi-turn conversational agent before the client sees it
An agent can pass every sample test and still fail on a real user's third question. To catch that before the demo, build a fake customer with its own goal, its own personality and a limited supply of patience.

In brief
- Static question–answer pairs cannot test a multi-turn agent, because the user's next message depends on the agent's last reply.
- A simulated user needs a hidden goal, a consistent persona and a patience budget; grade on database state, not on what the agent says it did.
- Run each scenario several times and report pass^k: on τ-bench, GPT-4o succeeded on under 50% of tasks and scored pass^8 below 25% in the retail domain.
Picture the last week before handover. The returns agent you built for a client has passed all two hundred sample question–answer pairs. Then someone on the client’s staff sits down to try it, types “I want to exchange my shoes”, and by the third turn the agent asks for the order number they have just given it.
A bug like this does not live in any single answer. It only shows up as the conversation goes on, when the user answers incompletely, changes their mind or gets impatient. That is why an FDE building conversational agents needs one more skill: building a simulated user to “interview” the agent hundreds of times before a real customer touches it.
Why is a static test set not enough?
The Strands Evals team says plainly that a static dataset of input–output pairs, however large, cannot capture the dynamics of a conversation.
The user’s next message depends on what the agent has just said. If the agent asks the wrong question, a real person answers differently, and every turn after that heads down a branch the sample test set never anticipated.
So you need a conversation partner that reacts. τ-bench, a benchmark published by Sierra, does exactly this: a language model plays the user and talks to an agent that can call business APIs. Sierra describes the simulator as an LLM guided by instructions specific to each scenario.
What does a simulated user need?
The first thing is a goal, and the goal must be hidden. Tian Pan, writing about synthetic users for evaluating multi-turn agents, stresses that the goal must not be revealed to the agent in advance; the agent has to draw it out through conversation. That is precisely the skill you want to test, so do not accidentally put the answer in the agent’s system prompt.
The second is a consistent persona. Strands Evals requires the simulated user to keep the same communication style, level of expertise and personality from start to finish. If an “older customer who rarely uses apps” suddenly writes like an engineer by turn five, the scenario is broken.
The third is knowing when to stop. Strands Evals tracks the simulated user’s goal alongside the conversation to know when it should end. Tian Pan adds something that is often forgotten: real people have finite patience.
An overly patient simulator makes the agent look right on paths that real customers would have abandoned long ago, so you need a “patience budget”.
Example: a returns agent for a retail chain
Suppose your client is a shoe retailer and the agent can look up orders, change sizes and issue refunds. A simulated scenario might look like the one below. The persona is a 50-year-old customer who rarely uses apps, gives short answers and gets impatient easily; the hidden goal is to exchange the shoes in order W1234 for a size 42, or get a refund to the card if size 42 is out of stock; the reveal rules say to give the order number only when asked and never volunteer the old size.
SCENARIO = {
"persona": "50-year-old customer, rarely uses apps, gives short answers, gets impatient easily",
"hidden_goal": "Exchange the shoes in order W1234 for size 42; "
"if size 42 is out of stock, refund to the card",
"reveal_rules": "Only give the order number when the agent asks; never volunteer the old size",
"patience_turns": 6,
"expected_db": {"W1234": {"status": "exchanged", "size": 42}},
}
Only the simulator knows the hidden goal. The reveal rules force the agent to ask the right questions. A six-turn budget means that if the agent beats around the bush, the fake customer will say “forget it, I’ll call the hotline” and end the conversation. The code for a single run takes only a few lines (the inline comments note that the agent may call APIs and write to the database, and that the simulator checks for itself whether the goal has been met):
def run_episode(agent, sim, scenario, db):
msg = sim.start(scenario)
for _ in range(scenario["patience_turns"]):
reply = agent.respond(msg, db) # agent may call APIs and write to db
msg, done = sim.next(reply) # sim checks for itself whether the goal is met
if done:
break
return db.snapshot() == scenario["expected_db"]
The last line matters most. The agent can say, very politely, “I’ve changed the size for you” while nothing in the database has changed. τ-bench grades by comparing the database state after each task with the expected outcome, and you should do exactly the same.
Passing once is not passing
Running the scenario above once and getting it right tells you little. Sierra uses a metric called pass^k to measure reliability: whether the agent completes the same task across multiple runs.
The results in the τ-bench paper are sobering: even a leading function-calling agent such as GPT-4o succeeded on fewer than 50% of tasks, and its pass^8 in the retail domain was below 25%.
A quick calculation shows why the number falls so fast. If a scenario succeeds 60% of the time at random and runs are independent, the probability of passing all 8 runs is 0.6 to the power of 8, about 1.7%. An agent that is “usually right” is still almost certain to fail someone on the first day of go-live.
How to do this on a client site
Start with the five to ten business flows the client cares about most, taken from discovery sessions or call-centre logs. For each flow, write a scenario with a hidden goal, a persona, reveal rules, a patience budget and an expected database state. Add a few difficult personas: someone who changes their mind halfway, someone who gives wrong information, someone who asks about things out of scope.
Then run each scenario k times against a copy of the database that is reset before every run. Report both the success rate and pass^k to the client, along with a few representative failed transcripts. Transcripts help clients understand faster than any chart, and they double as next week’s fix list.
Three common traps
The first trap is a simulator that is too well-behaved, as discussed above: without a patience budget, every detour looks like success. The second trap is subtler.
Tian Pan warns that when the simulator and the agent use the same kind of model, the evaluation turns into a “hall of mirrors”: the two sides converse fluently but sound nothing like real people.
The third trap is assuming the simulator resembles real people without checking. A paper posted on arXiv in May 2026 proposes the realsim framework for comparing simulated users with real conversations from a distributional perspective rather than sentence by sentence. Before each batch of runs, a quick review with this table helps:
| Trap | Check before running |
|---|---|
| Overly patient simulator | Does every scenario have a turn budget and a give-up line? |
| Hall of mirrors | Does the user role use a different model, or at least a prompt very different from the agent’s? |
| Simulator unlike real people | Have you compared simulated transcripts with real ones on message length, number of turns and when users give up? |
Putting this skill on your CV
On a CV, describe this skill in one concrete line rather than a generic summary, for example: “Built a simulated-user evaluation suite for 10 returns flows, graded on database state, raising pass^8 from X to Y”.
For interviews, have a failed transcript ready (anonymised) and explain how you found the bug, what you fixed and how pass^k changed. A story with before-and-after numbers is more convincing than any list of tools.
Clients will not remember what your agent scored on the sample test set. They will remember the first time it asked again for the order number they had just given it, and ideally you will have seen that bug before they did, in a simulated transcript.
Was this article useful?
Thanks for the feedback!
5 sources
- τ-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains (arXiv) · 2024-06-17
- 𝜏-Bench: Benchmarking AI agents for the real-world · 2024-06-20
- Simulate realistic users to evaluate multi-turn AI agents in Strands Evals · 2026-04-02
- Synthetic users for multi-turn agent eval (Tian Pan) · 2026-04-27
- Synthetic Users, Real Differences: an Evaluation Framework for User Simulation in Multi-Turn Conversations · 2026-05-04