FDE PulseFDE jobs open 441New in 7 days 29Companies hiring 47Remote-friendly 24%Median US pay $216kTop hirer Databricks 125
VI

The newspaper of the Forward Deployed Engineer

Guides

Voice agents for call centres: splitting a 500ms budget, handling interruptions and handing off to humans

Twilio's reference STT→LLM→TTS pipeline takes about 1.1 seconds, more than twice the 500ms target. Even when the agent gives the right answer, callers still get annoyed if it talks over them or bungles a transfer.

Voice agents for call centres: splitting a 500ms budget, handling interruptions and handing off to humans
Photo: Petr Macháček / Unsplash

In brief

  • Twilio's reference cascaded pipeline takes about 1.1 seconds mouth-to-ear. Future AGI sets a target for sales and support agents of under 350ms on average and 500ms at P95.
  • With barge-in, a false interruption is worse than a slightly late one. Choose the turn detection mode before tuning any other parameter.
  • A cold transfer closes the caller's session. A warm transfer has the agent brief a supervisor on the context before connecting the call.
ShareLinkedInFacebookX
GraphicThe 1.1-second reference pipeline versus the 500ms target
Twilio reference pipeline500ms budget at P95
STT (two different measures)350ms for the STT stage100–150ms, to first partial only
LLM, first token375ms200–300ms with prefix caching
TTS, first audio100ms80–150ms with a streaming provider
Network, buffers, orchestrationRoughly 290ms remainingNetwork 30–50ms, orchestration and guardrails 50–100ms
Total mouth-to-earAbout 1.1 seconds500ms at P95

Not every stage needs shrinking: Twilio's 100ms TTS is already within budget. The excess sits in the LLM, STT and the plumbing around the model.

Graphic: FDE Times

Picture a demo in a quiet meeting room where everything goes well. A week later the agent goes live on a real call centre line. Callers finish speaking and are met with a long silence, so they ask, “Hello, is anyone there?”

One caller coughs and the agent goes quiet. Another asks to speak to a person, gets transferred, and has to explain everything from the start.

None of these failures has anything to do with the quality of the answers. They are failures of timing and turn-taking, and that is exactly the work an FDE has to do on the customer’s site.

Twilio says it plainly: latency is the constraint that shapes the whole design of a voice agent. The three failures above map to three skills worth practising: splitting the latency budget, handling interruptions, and handing calls to humans.

Half a second has to be split across every stage

The number to remember is mouth-to-ear: the time from the moment the caller stops speaking to the moment they hear the agent’s first sound. A Future AGI guide sets a working target for sales and support agents of under 350ms on average and 500ms at P95.

P95 matters more than the average. An average of 300ms sounds fine, but if one turn in 20 takes 1.2 seconds, that is the turn the caller remembers. So from day one, reports to the customer should include both the average and percentiles.

Adding up one turn: where does 1.1 seconds come from?

Twilio built a reference cascaded pipeline of STT, LLM and TTS that comes to about 1.1 seconds (1,115ms) mouth-to-ear. It splits the budget into STT at 350ms, LLM at 375ms (to first token, TTFT) and TTS at 100ms (to first byte, TTFB), and states clearly that this is only a starting point.

The three stages add up to 825ms. The remaining 290ms or so goes to places few people look: the network, audio re-encoding, buffers. The first lesson is not to measure only the model. A quarter of the latency can sit in the plumbing around it.

Now set that beside the 500ms P95 budget Future AGI proposes. This comes from a vendor, so treat the numbers as reference points only:

Stage 500ms budget (P95) Notes
Network 30–50ms Place servers close to callers
STT, first partial 100–150ms Use interim results; don’t wait for the full sentence
LLM TTFT 200–300ms Use prefix caching
TTS, first audio 80–150ms Use a provider with streaming
Orchestration and guardrails 50–100ms Includes inline checks

Do the arithmetic. Taking the lower bound of every stage gives 460ms, just under 500ms. Taking the upper bound of every stage gives 750ms, well over target.

So you cannot let every stage hit its upper bound. The four stages other than the LLM already cost 260ms at their lower bounds, which leaves the LLM 240ms at most. If the customer’s LLM must run at 300ms for security reasons, the total is 560ms even with every other stage at the best figure in the table.

At that point you have to push one stage below the table’s figure, for instance by moving inline guardrails to run in parallel, or renegotiate the target with the customer. A budget is a trade-off problem, not a checklist. Note too that summing per-stage P95s gives only a conservative estimate, because slow turns rarely hit every stage at once.

The first job on the customer’s site is therefore simple: record a timestamp at every stage boundary. Log the first-partial mark and the final-transcript mark separately, because they measure different things, and you need to know which one the customer’s pipeline passes to the LLM. The pseudocode below is enough to get started:

# each conversation turn records its timestamps (ms)
turn = {
  "user_end": t0,          # VAD reports the caller stopped speaking
  "stt_first_partial": t1,
  "stt_final": t2,         # final transcript
  "llm_first_token": t3,
  "tts_first_byte": t4,
  "audio_played": t5,      # measure on the caller side if possible
}
stages = {
  "stt_partial": t1 - t0,
  "stt_final": t2 - t0,
  "llm_ttft": t3 - t2,     # change the baseline to t1 if the pipeline sends partials to the LLM
  "tts": t4 - t3,
  "network_and_buffer": t5 - t4,
  "total": t5 - t0,
}
# collect 50-100 turns, compute mean and p95 for each stage

(The comments note, in order: each turn records timestamps in ms; VAD signals the caller has stopped speaking; final transcript; measure on the caller’s side if possible; switch the baseline to t1 if the pipeline sends partials to the LLM; collect 50–100 turns and compute mean and P95 per stage. network_and_buffer is network and buffer; total is the total.)

Interruptions: choose how turns are detected first, tune sensitivity second

LiveKit defines interruption handling as deciding what the agent does when the user starts talking over it. A bigger decision comes before that: the turn detection mode, meaning how the system knows the caller has finished speaking. LiveKit’s documentation advises settling this first, because it determines whether the other parameters still apply.

Do it the other way round and you can spend hours tuning VAD thresholds, then switch turn detection modes and find that some of the parameters you tuned no longer have any effect.

For barge-in, another Future AGI guide sets a budget: TTS must stop within 60ms, and the total time until the TTS buffer is fully flushed must be under 150ms. The guide cites no independent benchmark, so this too is only a reference point.

The same author stresses that a false interruption is worse than a slightly late one, and sets a target false barge-in rate below 2%.

LiveKit calls this a false interruption: VAD hears a sound and makes the agent stop, but no transcript appears. The usual culprits are coughs or background noise. By default, the agent waits a moment and then resumes speaking.

For a call centre, this is where testing has to happen in the field. Callers stand on the street, ring from a market or a motorbike, or interject “yes”, “uh-huh” while the agent is reading out a contract number. Record a few dozen calls in genuinely noisy conditions and count the false stops; do not tune sensitivity on the feel of a meeting room.

Handing off to a human: cold or warm?

There will always be calls the agent should not handle itself. The question is how to transfer them.

Cold transfer

  • Uses SIP REFER to transfer the call
  • The caller's LiveKit session is closed
  • The carrier's trunk must be configured to allow call transfer

Agent-assisted warm transfer

  • The caller is placed on hold
  • The supervisor is called into a private consultation room
  • The agent summarises the context before connecting the call

A cold transfer suits cases where the call only needs to reach the right department, such as accounts. When the caller is frustrated or the problem is complex, a warm transfer is almost always worth the extra few seconds of waiting. LiveKit recommends using its prebuilt warm transfer task for most cases rather than writing your own.

What the FDE really has to design is the content of the summary. Imagine the agent speaking to the staff member in the consultation room:

“The caller is Ms Lan, calling for the second time about a late delivery. She has verified her phone number and wants to cancel the order for a refund. She is quite frustrated because she had a long wait last time.”

Those three sentences answer three questions: who is calling, what has been done, and what is needed next.

Steps when you start work

Measure before optimising: put timestamps on every stage, collect at least a few dozen turns, then report both average and P95. Next, compare against the budget table and find the stage that overshoots the most. That is usually the first thing worth fixing.

Then settle the turn detection mode with the customer’s team, and only after that tune barge-in using real audio from the call centre. Finally, sit down with the call centre team lead to map out the situations that must go to a human, decide which use a cold transfer and which a warm one, and write the summary template with them.

Common mistakes

The first mistake is reporting only average latency and measuring only the model, forgetting the network and buffers. The second is making barge-in too sensitive so the agent “stops really fast”, after which it stops every time a motorbike goes past.

The third is using a cold transfer for everything, forcing callers to start their story again, or writing a warm transfer flow by hand when a prebuilt task already exists.

If you are preparing to apply for an FDE role, the three skills in this article can be told in numbers. On a CV, something specific such as “cut P95 mouth-to-ear from X to Y with streaming TTS and prefix caching” is far more convincing than “built a voice agent”.

Exercise: work out the budget for a specific customer

Suppose the customer says their LLM has a P95 TTFT of 220ms. Subtract 220ms from 500ms and you have 280ms left for the network, STT, TTS and orchestration. The lower bounds for those four stages in the table are 30, 100, 80 and 50ms, totalling 260ms.

That leaves just 20ms of slack across all four stages. The question to take to the first meeting: which stage in the customer’s current pipeline overshoots its lower bound the most, and can you win that back by switching provider or moving servers? Redo the calculation with the customer’s real numbers before promising any latency target.

Call centre callers do not score the model’s intelligence. They remember only whether the agent kept them waiting, whether it talked over them, and whether, when they needed it, it passed them to a person who already knew the story.

Was this article useful?

Use with your AI assistantAsk Claude ↗Ask ChatGPT ↗
6 sources
Read next on the roadmap · Stage 5: DeploymentGo-live cutover: moving the data, running in parallel and planning the way backA go-live night at a client can fall apart even when the code is right, because nobody knows who has the authority to say "roll back", or which data to roll back to.