Quick answer: Use one agent until one prompt must hold two incompatible policies (what it may disclose or commit to), a realtime model must change mid-call, or a phase must ship on its own schedule. Each split adds a transition: the six handoff bugs traced below were filed against LiveKit Agents and Pipecat between March and September 2026; one is still open.
What changes when one voice agent hands off to another
LiveKit’s own documentation states the default plainly: “each new agent or task starts with a fresh conversation history for their LLM prompt”, and the new agent “only sees its own instructions unless you explicitly pass conversation history using chat_ctx“. That applies to a handoff returned from a tool and to session.update_agent() alike (LiveKit, Agents and handoffs, read 24 September 2026). So the first thing a split changes is not latency or cost. It is memory. The caller told the first agent their name, their postcode and the reason they called; the second agent, by default, knows none of it.
Pipecat starts from a different default but lands in the same place. Its multi-agent guide says each LLMContextWorker “gets its own context by default, so agents don’t see each other’s history”, and you share a conversation by passing the same context= to each worker (Pipecat, Multiple LLM Agents). A plain LLMWorker, the kind Pipecat’s Agent Handoff walkthrough uses, does not manage conversation context on its own and relies on context coming from elsewhere, such as the main agent’s aggregators. Which worker class you picked decides what the next agent remembers. Pipecat Flows, which moves one worker between conversation nodes rather than between agents, defaults the other way: APPEND keeps the full history and RESET clears it; a third strategy, RESET_WITH_SUMMARY, is deprecated and not available in a flow config (Pipecat, Context Strategies).
| Framework and construct | What the next unit sees by default | How you carry context across |
|---|---|---|
LiveKit Agents: agent handoff (a tool returns an Agent) or update_agent() |
Its own instructions only. Fresh history | Pass chat_ctx: a copy (copy(exclude_instructions=True) drops the old system prompt), a summary, copy().truncate(max_items=6) for the last six items, or key facts kept in userdata and injected as a system message |
LiveKit Agents: AgentTask under a supervisor |
A scoped copy; the supervisor keeps full context | The task returns a typed result to the agent that called it |
LiveKit Agents: TaskGroup |
Context shared within the group | Summarised and returned to the controlling agent when the group finishes |
Pipecat: LLMContextWorker activated by activate_worker(..., deactivate_self=True) |
Its own context. No history from the previous worker | Pass one shared LLMContext via context=, or inject messages with LLMWorkerActivationArgs |
Pipecat: plain LLMWorker behind a bus bridge |
No context of its own; it sees what the main agent’s aggregators send | Share the main agent’s context; inject the handoff reason with LLMWorkerActivationArgs |
| Pipecat Flows: node transition inside one worker | Full history (APPEND is the default) |
Set context_strategy to reset per node to clear it |
The quotable version: in LiveKit Agents, a second agent is a stranger to the caller unless you introduce them in code. The complete session transcript still exists in session.history; it just is not in the new agent’s prompt.
This page is about that in-session switch between AI agents on one live call. Handing a caller to a person is a different problem, with a different payload, covered in the AI-to-human handoff packet.
Single agent vs multi agent: the four questions that decide it
LiveKit’s workflow guide opens its advice with “Start with a single agent and a small set of tools”, and says to split only on a concrete limitation. It names four: instruction bloat, conflicting tool access, multi-turn data collection, and backtracking (LiveKit, Workflows). Read against LiveKit’s own pattern table, multi-turn collection maps to tasks and backtracking to task groups; the table reserves full agent handoffs for distinct roles, model specialisation and permission boundaries. Pipecat’s guide argues the other side for its framework: giving each agent its own LLM “keeps every context small and focused”. Both are vendor positions. These four questions settle it for your call.
1. Would one prompt have to hold two incompatible policies?
This is the question that matters most, and the next section turns it into a rule. A policy here means two things: what the agent may disclose, and what it may commit to. If phase A must refuse something phase B must do, one prompt carries both a refusal and its exception, and the model decides which one applies on every turn.
2. Does the model or the voice have to change mid-call?
On a speech-to-text, LLM, text-to-speech pipeline, often not. LiveKit’s Agent.update_options() replaces the STT, VAD, LLM or TTS on the active agent without a handoff; the LLM and TTS switch at the next generation. That method is Python-only today. A realtime model is different: LiveKit states that “Realtime models can’t be replaced on an active agent”, the call raises a RuntimeError, and the documented route is to hand off to a new agent. If you run OpenAI Realtime or Gemini Live and need a different model for one phase, the handoff is not optional, and a replaced provider session has its own failure signature, covered in detecting a realtime session that was replaced mid-call.
3. Is the prompt, not the tool list, too big?
Tool count on its own is a weak reason to split. LiveKit’s guide says a single agent can handle multi-step flows by changing its instructions or its available tools between phases. The signal is behavioural: the model starts underperforming on its primary task. Measure that before you restructure.
4. Do you need to test and ship one phase without retesting the others?
A separate agent is a separately testable unit with its own prompt, tools and pass criteria. If your verification phase changes under a different review cadence from your sales phase, that is a real reason to split. If every change touches the whole call anyway, the split buys you an extra transition to test and nothing else.
The two-policy rule: when to split an AI agent
The two-policy rule: split a voice agent when a single prompt would have to hold two incompatible policies about what it may disclose or what it may commit to, not when it merely has many tools. LiveKit’s handoffs page lists “different permissions” as one of four reasons to create separate agents, with the example of “An agent with payment API access versus one handling general inquiries”. The rule sharpens that into a test you can run on paper.
The test takes ten minutes. Write each phase’s policy as refusal sentences: “Before identity is verified, never read back an account balance.” “Never quote a price.” “Never confirm a booking until the slot is written to the calendar.” Then read the sets against each other. If any sentence in phase A forbids something phase B is required to do, you have two policies. If the sentences only add up, you have one policy and a long prompt.
| Situation | Policies in conflict? | Verdict |
|---|---|---|
| Intake that must not disclose account details, followed by account servicing that must | Yes: disclosure forbidden, then required | Split at verification. Pass facts, not the transcript |
| An appointment setter that may commit to a calendar slot but never to price, followed by a closer who discusses terms | Yes: a commitment forbidden, then required | Split, or keep terms out of the call entirely |
| One agent with 30 tools, all governed by the same disclosure rules | No | One agent; gate tools by phase |
| A booking flow where callers correct earlier answers | No | One agent plus a task group, which LiveKit documents as handling regression to earlier steps |
| Wanting a different voice for the billing phase on a pipeline stack | No | One agent; swap TTS with update_options() (Python) |
| Needing a different realtime model for one phase | No, but the constraint is mechanical | Handoff is required; the policy test does not apply |
Why the rule is about refusals rather than tools: a tool you gate out of a phase is simply absent, and the model cannot call it. A refusal written into a shared prompt is present on every turn, next to the instruction that overrides it. Splitting the agents takes the forbidden instruction out of the prompt altogether. The cost is a transition, and every transition is a place the call can die.
One agent, tasks or a full handoff: the comparison
LiveKit documents five patterns and says they are not mutually exclusive. All five are listed here in its terms, with the Pipecat construct that does the nearest job. The latency column is LiveKit’s own wording for LiveKit’s patterns.
| Pattern | Who holds the call | Latency cost (LiveKit) | Right for | Wrong for | Nearest Pipecat construct |
|---|---|---|---|---|---|
| Single agent plus tools | One agent throughout | None | One policy, few distinct phases | Two incompatible policies; realtime model swap | One LLMWorker, or Flows nodes in one worker |
| Supervisor pattern (tasks) | One agent; tasks take temporary control and return a typed result | Minimal | Verification, consent capture, structured collection | A phase that must hold control for the rest of the call | Flows node with its own functions |
| Subagent delegation | One agent; a background model works and answers in a later turn | None on the conversational turn | Slow reasoning that would stall the turn | Anything the caller must hear before the next turn | Job coordination between workers |
| Agent handoff | Control transfers fully; the old agent does not take part afterwards | Handoff overhead per transition | Distinct roles, model specialisation, permission boundaries | Tidiness alone; frequent back-and-forth | activate_worker(..., deactivate_self=True) |
| Task group | Ordered sequence of tasks, context shared in the group | Minimal | Ordered collection where callers go back and correct | Open-ended conversation | Flows nodes; returning to an earlier node is a transition you define |
Read as a decision: the handoff is the only one of the five patterns in which control transfers fully and the original agent takes no further part, and the only one LiveKit prices as an overhead on every transition. If you are unsure how much that overhead is on your stack, measure it on the turn after the switch, which is the turn LiveKit Agents did not record in its latency metrics before 1.8.1; see why the first turn after an agent handoff was missing from latency dashboards.
Five documented ways an in-call handoff dies
These are six public issues in LiveKit Agents and Pipecat, two open-source voice agent frameworks, grouped into five failure modes. Every row was read in full, and each state was checked on 24 September 2026. The last column is what reaches the caller, not what reaches the log. Where an issue does not describe the caller’s experience, the cell says so.
| Failure mode | Issue | Reported against | State on 24 Sep 2026 | What the caller experiences |
|---|---|---|---|---|
| 1a. The handoff silently does not happen | livekit/agents #7320, filed 17 Sep 2026 | livekit-agents 1.8.1; also present in 1.8.2 | Closed 18 Sep 2026 by PR 7321. The guard is in the 1.8.3 wheel on PyPI (23 Sep 2026) and absent from 1.8.2 | The same agent keeps answering “as if nothing happened”. No error, no handoff event |
| 1b. Same batch race, pipeline path | livekit/agents #5150, filed 18 Mar 2026 | livekit-agents 1.4.3 | Closed 18 Sep 2026; the maintainer’s closing comment points to PR 7321 | Silence. The new agent never greets; when the caller speaks again the cycle repeats |
| 2. The handoff waits forever | livekit/agents #6858, filed 14 Aug 2026 | livekit-agents 1.6.8 | Closed 1 Sep 2026. Fix (PR 6865) listed in the 1.8.0 release notes, 5 Sep 2026 | The call “sits in silence” until the caller hangs up |
| 3. The code path does not support a handoff | livekit/agents #5936, filed 2 Jun 2026 | livekit-agents 1.5.10 | Open. Community PR 6254 not merged | The issue does not describe the audio; it reports that the session is not forwarded to the new agent after the interim “please hold” update |
| 4. The incoming agent cannot connect | pipecat-ai/pipecat #4926, filed 1 Jul 2026 | pipecat 1.4.0 with GeminiLiveLLMService |
Closed 27 Jul 2026; fix in v1.7.0, released 1 Aug 2026 | Not described in the issue. The log shows the new agent’s connection failing with 1007 three times, then retries stop |
| 5. The outgoing agent keeps talking | pipecat-ai/pipecat #4979, filed 8 Jul 2026 | pipecat 1.5.0, Flows multi-worker handoff | Closed 11 Aug 2026, pointing to v1.6.0. The fix is opt-in: see below | The assistant repeats a near-identical reply two or three times |
What each one actually is
Rows 1a and 1b: a handoff that shares a tool-call batch. When the model emits a handoff and an ordinary tool call in the same turn, LiveKit collected tool outputs in completion order and assigned the handoff target unconditionally, so an ordinary tool that finished last overwrote the handoff with nothing. The #7320 reporter found it in production when a receptionist agent emitted a data-update tool and a routing tool together, and noted that from outside “it just looks like the LLM decided not to route”. That sentence is the reason this failure survives testing: it reads as model behaviour, not as a bug. Their interim workaround was calling session.update_agent() inside the tool, at the price of the next agent starting one message short. #5150 described a batched handoff failing on the pipeline path, although its reporter’s own analysis blamed a follow-up reply on the old agent racing the handoff and said the overwrite had not caused their loop; the maintainer at first could not reproduce it and noted that an old agent speaking a tool reply before handing off is designed behaviour, then closed it against the same fix in September.
Row 2: a parked speculative reply. With preemptive generation on, a caller who speaks at the wrong moment during uninterruptible speech (a transfer hold line, or any awaited AgentTask, which sets itself uninterruptible) leaves a speculative reply neither scheduled nor cancelled. The activity pause that an AgentTask handoff takes then waits for it indefinitely. The tell in the logs is skipping reply to user input, current speech generation cannot be interrupted.
Row 3: handoff after an interim update. Returning a new agent from an async tool after ctx.update() does not hand off. The maintainer’s reply was that this is “not supported right now” and suggested calling session.update_agent() directly, inside ctx.foreground(), taking care not to call it twice when parallel async tools are running.
Row 4: a realtime provider that rejects the seeded context. Seeding a Gemini Live session from a shared context that contained tool calls produced 1007 Request contains an invalid argument. The same code with OpenAIRealtimeLLMService worked. Pipecat 1.7.0 added a GeminiLiveLLMAdapter that summarises function calls and their responses as text, because Gemini Live only accepts text and media content.
Row 5: a deactivated worker that still answers. In Pipecat Flows, the transfer function activated the router and also returned a next node; setting that node queued an LLM run directly on the worker that had just been deactivated, and deactivation only gates frames arriving from the bus. The v1.6.0 fix added a NO_RESPONSE outcome and changed the example to use it. PR 4980, which would have skipped the run on an inactive worker inside FlowManager, was closed unmerged, and in the pipecat-ai 1.11.0 wheel we unpacked, FlowManager still queues the run whenever respond_immediately is true, with no check on whether the worker is active. If you wrote your own Flows handoff, Pipecat’s shipped fix reaches it only if your handoff returns NO_RESPONSE. The issue’s interim workaround, returning the next node with respond_immediately=False, stops the duplicate reply, but the closed PR found that setting a node on a deactivated worker can also overwrite the active worker’s tool list, which NO_RESPONSE avoids by not moving to a new node at all.
Two patterns run through the table. Three of the five modes produce silence or no change rather than an error, so the call looks like a model that chose badly. And four of the six were closed within five weeks of being filed: the risk is less the bug than the version pinned in your lock file.
Check which rows your build is exposed to
The script below reads a pip freeze file and prints the rows above whose fix is not in your installed versions. It checks exposure, meaning the fix is absent, not whether you actually hit the bug; every LiveKit row flags all versions below the fix, because the issues report the version they were found on, not the version the bug first appeared in. The one lower bound, pipecat-ai 1.5.0 for #4979, is the release the issue names as introducing the multi-worker framework. It is plain Python 3 with no dependencies.
import re, sys
def ver(s):
parts = [int(x) for x in s.split(".")[:3]]
return tuple(parts + [0] * (3 - len(parts)))
def installed(freeze_text):
pkgs, unknown = {}, []
for line in freeze_text.splitlines():
line = line.strip()
m = re.match(r"^([A-Za-z0-9_.-]+)==(\S+)$", line)
name = m.group(1).lower().replace("_", "-") if m else ""
if m and re.fullmatch(r"[0-9]+(\.[0-9]+)*", m.group(2)):
pkgs[name] = ver(m.group(2))
elif name in ("livekit-agents", "pipecat-ai") or re.search(r"(livekit[-_]agents|pipecat[-_]ai)\b(?![-_])", line, re.I):
unknown.append(line)
return pkgs, unknown
def exposure(pkgs):
out = []
lk = pkgs.get("livekit-agents")
pc = pkgs.get("pipecat-ai")
if lk:
if lk < (1, 8, 3):
out.append("livekit #7320/#5150: a handoff batched with another tool can be dropped (guard added in 1.8.3)")
if lk < (1, 8, 0):
out.append("livekit #6858: AgentTask can deadlock on a parked preemptive generation (fixed in 1.8.0)")
out.append("livekit #5936: still open - no handoff by return value after ctx.update(); use session.update_agent()")
if pc:
if pc < (1, 7, 0):
out.append("pipecat #4926: GeminiLiveLLMService fails with 1007 when a handoff seeds tool calls (fixed in 1.7.0)")
if pc >= (1, 5, 0):
out.append("pipecat #4979: a deactivated Flows worker can still answer; return NO_RESPONSE (1.6.0+) from the handoff")
return out
if __name__ == "__main__":
pkgs, unknown = installed(open(sys.argv[1]).read())
for line in unknown:
print("version not read, check by hand: " + line)
for row in exposure(pkgs) or ([] if unknown else ["no documented handoff rows match"]):
print(row)
Run it as pip freeze > freeze.txt && python3 handoff_exposure.py freeze.txt. We ran it against three fixture files. A freeze pinning livekit-agents 1.7.1 and pipecat-ai 1.5.0 printed all five lines. A freeze pinning livekit-agents 1.8.3 and pipecat-ai 1.11.0 printed two: the open LiveKit issue, and the Pipecat opt-in, which no version upgrade removes. A freeze with neither framework printed “no documented handoff rows match”. A framework installed from a Git or file reference, or pinned to a pre-release, prints a “version not read, check by hand” line instead of an all-clear. Add rows as new issues land; the table is a snapshot of 24 September 2026, not a guarantee.
Who a single agent is wrong for, and who multi-agent is wrong for
A single agent is wrong for you once one prompt carries two incompatible refusal policies. It is also wrong if you run a realtime model and one phase genuinely needs a different model, because LiveKit will not swap it in place. And it is wrong if one phase has to be tested and released on its own schedule, for example a verification step that changes under compliance review while the sales conversation changes weekly.
Multi-agent is wrong for you when the only motivation is tidiness, because every transition is a place the call can die. It is wrong when the reason is tool count, since both LiveKit and Pipecat Flows let one agent change tools by phase. It is wrong when callers move back and forth between topics, since each move is a transition, and LiveKit documents task groups for exactly the case of going back to correct an earlier answer. And it is wrong if nobody on the team owns framework upgrades: four of the six issues above were fixed by an upgrade alone, which only helps a team that upgrades. If you are still choosing the framework itself, the trade-offs are in choosing between LiveKit Agents and Pipecat.
What running multi-agent costs you to own
The build is the cheap part. The recurring cost scales with transitions, not with agents, and the count grows faster than people expect. This is arithmetic on a design, not a measurement.
| Design | Directed transitions | Minimum transition tests (3 per transition) |
|---|---|---|
| One agent, tools gated by phase | 0 | 0 |
| 2 agents, either can hand to the other | 2 × 1 = 2 | 6 |
| 4 agents, router and 3 specialists that always return to the router | 3 out + 3 back = 6 | 18 |
| 4 agents, any can hand to any | 4 × 3 = 12 | 36 |
The three tests per transition come straight from the failure table: the handoff fires when it shares a turn with another tool call (rows 1a and 1b); it completes when the caller talks over the transfer line (row 2); and the outgoing agent says nothing after it (row 5). Each transition also needs a context decision (full copy, summary, or facts only), an announcement line so the caller knows why the voice changed, and a latency check on the first turn after it.
Volume turns small rates into daily events. Purely as an illustration: at 10,000 calls a day with one transition each, a silent failure rate of one in a thousand is ten broken calls a day that look like an agent choosing not to route. Both inputs are hypothetical, chosen for the arithmetic; neither is a measured figure.
Zian publishes its digital team as separate role agents: Outbound Appointment Setter, Customer Support Agent, Sales Call Closer and Appointment Show-Specialist. Those are separate jobs with separate outcomes. The two-policy rule is the test for whether separate jobs also need separate prompts inside one call, and for most calls one policy is enough. Where scripts change, PrecisionPitch AI™ split-tests the approach against real outcomes, which is the measurement a phase split should earn its place against.
Frequently asked questions
Should I use one AI agent or multiple agents for a phone call?
Start with one agent and a small set of tools. Split when one prompt would have to hold two incompatible policies about what the agent may disclose or commit to, when a realtime model must change for one phase, or when one phase has to be tested and released on its own schedule. Tool count alone is not a reason, because one agent can change its tools between phases.
Does the new agent remember the conversation after a handoff?
Not by default in LiveKit Agents. The LiveKit Agents and handoffs documentation states that each new agent or task starts with a fresh conversation history, and that the new agent only sees its own instructions unless you pass chat_ctx explicitly. In Pipecat, each LLMContextWorker also keeps its own context unless you pass a shared one.
Why does my voice agent not transfer to the other agent even though the tool was called?
One documented cause in LiveKit Agents is a handoff returned in the same tool-call batch as another tool. The other tool finishing last overwrote the handoff, with no error. It was fixed by PR 7321, which is in livekit-agents 1.8.3. On older versions, avoid batching the handoff with other tools or call session.update_agent inside the tool.
Can I change the voice or the LLM mid-call without a second agent?
On a pipeline stack, yes in Python: Agent.update_options replaces the STT, VAD, LLM or TTS on the active agent, and the LLM and TTS change at the next generation. A realtime model cannot be replaced on an active agent; LiveKit raises a RuntimeError and documents a handoff to a new agent instead.
Why does my Pipecat agent repeat itself after a handoff?
Pipecat issue 4979 traced repeated replies after a Flows multi-worker handoff to the deactivated worker still running an LLM completion. Pipecat 1.6.0 added a NO_RESPONSE outcome and changed its example to use it. Your own handoff gets that fix only if it returns NO_RESPONSE, because FlowManager still queues the run whenever respond_immediately is true.
Is a multi-agent voice setup slower?
LiveKit lists handoff overhead per transition as the latency cost of agent handoffs, against none for a single agent with tools and minimal for tasks. Measure the first turn after the handoff specifically, because an aggregate across all turns dilutes it.
Where every figure on this page comes from
| Figure | Who published it | Link | Date read |
|---|---|---|---|
Four reasons to create separate agents, including different permissions; new agents start with a fresh conversation history; exclude_instructions; userdata summary; truncate(max_items=6); realtime models cannot be replaced (RuntimeError); update_options() Python-only |
LiveKit | Agents and handoffs | 24 Sep 2026 |
| Five workflow patterns; four split signals; “handoff overhead per transition” | LiveKit | Workflows | 24 Sep 2026 |
Each LLMContextWorker has its own context by default; shared via context= |
Pipecat | Multiple LLM Agents | 24 Sep 2026 |
Flows context strategy defaults to APPEND; RESET clears; RESET_WITH_SUMMARY deprecated |
Pipecat | Context Strategies | 24 Sep 2026 |
| #7320 filed 17 Sep 2026 on 1.8.1, closed 18 Sep 2026 via PR 7321 | livekit/agents on GitHub | Issue 7320 | 24 Sep 2026 |
| livekit-agents 1.8.3 uploaded 23 Sep 2026; guard present in its wheel, absent from 1.8.2 | PyPI (upload date); our own inspection of both wheels (guard) | livekit-agents 1.8.3 | 24 Sep 2026 |
| #5150 filed 18 Mar 2026 on 1.4.3, closed 18 Sep 2026 | livekit/agents on GitHub | Issue 5150 | 24 Sep 2026 |
| #6858 filed 14 Aug 2026 on 1.6.8, closed 1 Sep 2026; PR 6865 in 1.8.0, released 5 Sep 2026 | livekit/agents on GitHub | Issue 6858; 1.8.0 release notes | 24 Sep 2026 |
| #5936 filed 2 Jun 2026 on 1.5.10, open; PR 6254 unmerged | livekit/agents on GitHub | Issue 5936 | 24 Sep 2026 |
| #4926 filed 1 Jul 2026 on 1.4.0; error 1007; 3 consecutive failures; closed 27 Jul 2026; fix in v1.7.0, released 1 Aug 2026 | pipecat-ai/pipecat on GitHub | Issue 4926; v1.7.0 release notes | 24 Sep 2026 |
#4979 filed 8 Jul 2026 on 1.5.0; reply repeated 2 to 3 times; closed 11 Aug 2026; NO_RESPONSE in v1.6.0, released 21 Jul 2026 |
pipecat-ai/pipecat on GitHub | Issue 4979; v1.6.0 release notes | 24 Sep 2026 |
FlowManager in pipecat-ai 1.11.0 still queues the LLM run with no active-worker check |
Zian AI, first-party inspection of the published 1.11.0 wheel | Fix context: v1.6.0 notes | 24 Sep 2026 |
| More than 10,000 leads a day | Zian AI (first-party claim released by the owner) | Not externally published | 24 Sep 2026 |
Build agents with the policy lines drawn where they belong
Zian AI builds autonomous AI sales agents with live phone, SMS, email and WhatsApp outreach across 30+ languages, with SmartReach AI™ orchestrating message, channel and timing and PrecisionPitch AI™ split-testing the approach. Private model deployment on customer infrastructure is also available. Zian AI is currently in partnership-application beta.