Quick answer: OpenAI’s Realtime API lets one response write to the default conversation at a time; a response.create sent before the active response’s response.done is rejected. Of 17 public reports, 5 were interruptions whose cancel was never sent or never awaited, 4 tool results, 4 extra app messages, 2 server-VAD races and 2 orphaned responses. Wait for response.done, or cancel first.
If your voice agent went quiet mid-call with this error, this page covers what it means, the five ways it happens, a log test that tells them apart, and a small gate for your own event loop. Every case below links to the original report on GitHub (LiveKit Agents, the OpenAI Agents SDK for Python, OpenAI’s realtime agents demo) or the OpenAI Developer Community forum, with its state as at 28 September 2026.
What “Conversation already has an active response” means
The Realtime API keeps one default conversation per session. OpenAI’s Realtime client events reference states the rule in the response.create entry: “Only one Response can write to the default Conversation at a time, but otherwise multiple Responses can be created in parallel.” A response is active from the moment the server emits response.created until it emits response.done, which the server events reference says is “Always emitted, no matter the final state”, whether the response completed, was cancelled, failed or came back incomplete. Any response.create aimed at the default conversation inside that window is refused.
The refusal arrives as an ordinary error server event. You will see one of two message forms, depending on when your logs are from:
| Seen in reports from | error.code |
error.message |
Example report |
|---|---|---|---|
| October 2024 to January 2025 | null |
Conversation already has an active response | LiveKit Agents #1056 (8 November 2024) |
| April 2025 | conversation_already_has_active_response |
Conversation already has an active response | LiveKit Agents #2064 (22 April 2025) |
| August 2025 onwards | conversation_already_has_active_response |
Conversation already has an active response in progress: resp_… Wait until the response is finished before creating a new one. | OpenAI Agents SDK #1466 (14 August 2025) |
Two consequences follow. First, match on the code and on the message prefix, or your handler will miss older logs where the code was null. Second, the error is not fatal to the session: the server events reference says “Most errors are recoverable and the session will stay open.” The rejected response.create never produces a response.created, so nothing speaks: in LiveKit Agents #1056 the agent “just stops talking” until the caller speaks again, and in #6817 the orphaned reply “dies as 10s timeout” inside the framework.
We could not find the error code or message anywhere in OpenAI’s own documentation. On 28 September 2026 we searched the rendered client events and server events references, the Realtime conversations, voice activity detection and error codes guides, the API changelog and the 4.9 MB combined documentation export (llms-full.txt). Each saved page passed a positive control, a term we knew was on it, found in the same file: response.done or create_response on the Realtime pages and the combined export, the 429 rate-limit code on the error codes guide, and gpt-realtime release entries on the changelog. The closest the docs come is a caution in the create_response field description: “If interrupt_response is set to false this may fail to create a response if the model is already responding.”
Seventeen documented cases, sorted by trigger
We collected every public report we could find of a developer hitting this error, read each thread, and sorted it by what sent the second response.create. The method, including exclusions, is below the table.
| # | Report (opened) | What the reporter describes | Trigger | State as at 28 September 2026 |
|---|---|---|---|---|
| 1 | LiveKit Agents #1056 (8 Nov 2024) | Error after function calls with MultimodalAgent; one reporter put it at “about 50% of the function calls” | Tool result | Closed as not planned (7 Oct 2025). PR #1204 (cancel active response for tool call) merged 13 Dec 2024; issue was reopened 6 Jan 2025 after further reports |
| 2 | OpenAI Agents SDK #2971 (20 Apr 2026) | Async tools each trigger a response while one is in progress; agent cuts off and waits | Tool result | Closed by PR #3140, merged 6 May 2026, which added tests only; the maintainer said v0.14.2 already defers response.create until response.done |
| 3 | OpenAI forum 1005582 (5 Nov 2024) | function_call_output then response.create after each call; fails after a few calls, even though earlier responses all reached response.done |
Tool result | 23 posts, no accepted answer |
| 4 | OpenAI forum 1172573 (1 Apr 2025) | Slow tool; caller speaks again; tool result arrives while the model is answering the new speech | Tool result | 5 posts, no accepted answer |
| 5 | LiveKit Agents #2064 (22 Apr 2025) | generate_reply right after a VAD interruption |
Interruption not cancelled, or cancel not awaited | Closed 22 Apr 2025; the maintainer found “the VAD interruption didn’t interrupt the realtime session” and pointed to PR #2072, merged the same day |
| 6 | LiveKit Agents #5642 (5 May 2026) | await session.interrupt() did not cancel in-flight requests before the next one |
Interruption not cancelled, or cancel not awaited | Closed by PR #5703 (cancel realtime generation when speech is interrupted), merged 14 May 2026 |
| 7 | OpenAI Agents SDK #1907 (16 Oct 2025) | Guardrail trip with interrupt_response set; no response.cancel is sent |
Interruption not cancelled, or cancel not awaited | Closed by PR #1968, merged 21 Oct 2025 |
| 8 | OpenAI Agents SDK #1912 (16 Oct 2025) | Manual cancel after an output guardrail, then a safe message; error logged in a follow-up comment | Interruption not cancelled, or cancel not awaited | Closed 30 Jun 2026 by a maintainer comment that cancel and sequencing are now automatic; no merged PR is linked as closing it |
| 9 | OpenAI Agents SDK #2760 (23 Mar 2026) | Guardrail interrupt sends response.cancel, then the follow-up response.create goes out before the cancelled response’s response.done |
Interruption not cancelled, or cancel not awaited | Closed by PR #2763 (wait for response.done before the follow-up response.create), merged 24 Mar 2026 |
| 10 | OpenAI Agents SDK #1466 (14 Aug 2025) | send_message while a response is being generated or not yet played |
Extra app message | Closed as not planned by the stale bot, 25 Aug 2025 |
| 11 | openai-realtime-agents #3 (18 Jan 2025) | Error logged when transitioning between agents | Extra app message | Closed 17 May 2025 by a maintainer who said state handling and cancelling were improved; no PR linked |
| 12 | OpenAI forum 983613 (18 Oct 2024) | Two response.create events back to back |
Extra app message | 2 posts, no accepted answer |
| 13 | OpenAI forum 1371442 (11 Jan 2026) | “Just a moment” line sent when a function_call item appears |
Extra app message | Accepted answer: play a local “thinking” sound instead |
| 14 | LiveKit Agents #6817 (12 Aug 2026) | generate_reply() races a server-VAD auto-response; lost reply “dies as 10s timeout” |
Server-VAD race | Closed by PR #6818, merged 14 Aug 2026, which fails the reply fast rather than preventing the collision |
| 15 | LiveKit Agents #7514 (28 Sep 2026) | Asks for opt-in serialisation of response.create; errors are frequent in production after the fast-fail |
Server-VAD race | Open as at 28 September 2026, no linked PR |
| 16 | LiveKit Agents #6223 (25 Jun 2026) | generate_reply times out locally but the server keeps generating; the next reply collides |
Orphaned response | Closed by PR #6244 (discard orphaned response after interrupt or timeout), merged 27 Jun 2026 |
| 17 | OpenAI forum 1374711 (20 Feb 2026) | A response ran 1 minute 3 seconds then failed with a server error; later responses failed with this code meanwhile | Orphaned response | 1 post, no reply |
The count: 17 reports, 12 GitHub issues across three repositories and 5 forum threads. By trigger: interruption not cancelled or cancel not awaited 5, tool result 4, extra app message 4, server-VAD race 2, orphaned response 2. Of the 12 GitHub issues, 11 are closed and one, LiveKit Agents #7514, is open as at 28 September 2026. None of the 5 forum threads ends in a fix from OpenAI staff; one has a community-accepted answer.
How we counted. We searched the livekit/agents, pipecat-ai/pipecat, openai/openai-agents-python, openai/openai-agents-js and openai/openai-realtime-agents issue trackers for the code and the message, and the OpenAI Developer Community for both. Pipecat and the JavaScript Agents SDK returned zero issues (a control search for “openai realtime” on Pipecat returned 53, so the search itself worked). A report counts once, by the trigger its opening post describes, and only if a developer says they hit the error. Three threads were left out: LiveKit Agents #6205 names the error only as a risk for a proposed retry design, OpenAI Agents SDK #1951 is a code bug report that names the error only by linking to #1907, and forum thread 1115740 repeats the opening post of 1005582 almost word for word. One reporter filed #6817, #7514 and #6205, so the two server-VAD rows share an author. Where an issue shows a closing pull request, we read the pull request’s own state; “closed” is not the same as “fixed”, and rows 1, 2, 8, 10, 11 and 14 show why.
The table maps where the error gets reported, not how often each trigger happens in production: framework users file issues, teams on raw WebSockets post on the forum. What it does show is that every one of the 17 comes down to one sender, the app or the server, issuing response.create while a response was still active.
Which trigger is mine? A log test that tells them apart
You do not need to guess. Log every event in both directions, with a direction tag, and look at what happened between the last response.done and the error. The five triggers leave different fingerprints:
In the log since the last response.done |
Trigger | First fix to try |
|---|---|---|
You sent conversation.item.create with a function_call_output, then response.create |
Tool result | Queue the create until response.done; send one create for all outputs of a turn |
Server sent input_audio_buffer.committed and a response.created you did not ask for |
Server-VAD race | Decide who owns response creation; retry the loser after response.done |
Your own response.created has no response.done yet, and you sent another create |
Interruption not cancelled or cancel not awaited, or an extra app message, or a retry after a local timeout | Send response.cancel and wait for response.done with status cancelled, or wait |
| No response is visibly active and none appears afterwards | Orphaned or stale response | Look back for a local timeout, a failed response or a dropped response.done |
Note the third row: from a log alone, an interruption that never sent response.cancel (or sent one and did not wait for its response.done), a mid-reply message and a retry after a local timeout look identical. They share a fix, so the ambiguity costs you little. This script applies the table to a JSONL event log:
"""Classify every conversation_already_has_active_response in a Realtime event log.
Input: JSONL, one event per line, each with "dir": "out" (you sent it) or "in" (server sent it)."""
import json, sys
ERR_CODE = "conversation_already_has_active_response"
ERR_TEXT = "Conversation already has an active response" # older logs: code is null
AUTO_RESPONSE = True # False if turn_detection.create_response is off or you commit audio yourself
def is_collision(ev):
e = ev.get("error") or {}
return ev.get("type") == "error" and (e.get("code") == ERR_CODE or (e.get("message") or "").startswith(ERR_TEXT))
def classify(events):
results = []
active = {} # response id -> "server" or "client"
unclaimed_creates = 0 # response.create sent, response.created not yet seen
vad_turn = False # input_audio_buffer.committed seen: server VAD may auto-respond
last_out = [] # client events since the last response.done
for i, ev in enumerate(events):
t, d = ev.get("type"), ev.get("dir")
if d == "out":
last_out.append(t if t != "conversation.item.create" else "item:" + ev["item"]["type"])
if t == "response.create":
unclaimed_creates += 1
elif t == "input_audio_buffer.committed":
vad_turn = True
elif t == "response.created":
rid = ev["response"]["id"]
if vad_turn and AUTO_RESPONSE:
vad_turn = False; active[rid] = "server" # auto-response to the committed user turn
elif unclaimed_creates:
unclaimed_creates -= 1; active[rid] = "client"
else:
active[rid] = "server" # nobody asked: server VAD auto-response
elif t == "response.done":
active.pop(ev["response"]["id"], None)
last_out = []
elif is_collision(ev):
unclaimed_creates = max(0, unclaimed_creates - 1) # the rejected create never gets a response.created
later = [e for e in events[i+1:i+6] if e.get("type") == "response.created"]
if "item:function_call_output" in last_out and active:
why = "A: tool result submitted and response.create sent while a response was still active"
elif "server" in active.values():
why = "B: your response.create raced a server-VAD auto-response"
elif "client" in active.values():
why = ("D: second response.create while your own response was still active "
"(also how a barge-in whose response.cancel was never sent or never awaited, or a retry after a local timeout, looks)")
elif later:
why = "B (in transit): server created a response you had not yet seen; retry after its response.done"
else:
why = "E: no active response visible locally; look for an earlier timeout, failed or orphaned response"
results.append((i, why))
return results
if __name__ == "__main__":
evs = [json.loads(l) for l in open(sys.argv[1]) if l.strip()]
hits = classify(evs)
for line, why in hits:
print(f"event {line}: {why}")
print(f"{len(hits)} collision(s) in {len(evs)} events")
We ran it on 28 September 2026 against thirteen event logs: twelve generated by the simulated server described in the one-in-flight rule section below, and one hand-written log copying the forum 1005582 pattern, where every earlier response had finished. It labelled the two direct tool-result logs A (three collisions between them), both server-VAD logs B (the gated run included, because the gate cannot prevent the in-transit race), the direct interruption, extra-message and timeout logs D, and the hand-written log E. It found zero collisions in the other five gated logs. If you have turned create_response off or commit audio yourself, set AUTO_RESPONSE = False, or your own responses will be read as server-VAD ones. It is a triage aid, not proof: a server-created response whose response.created has not reached you yet is invisible to any client-side log.
“I get the error after a tool call”
Symptom: the agent calls a function, you send the result, and the follow-up reply never plays. Cause: if you run the tool as soon as the call’s arguments arrive, the response that contains the call may not have sent its response.done yet, and a fast tool returns inside that window. In LiveKit #1056 the maintainer put it as the error happening “when the function is finished ‘too soon’”. With parallel tools it gets worse: the same maintainer found that “when there are more than one function calls, each of them will trigger a response request, and the later one will raise the error if the response of the first one is not finished.”
Fix: submit each function_call_output as soon as it is ready, but hold the response.create until the current response’s response.done, and send one create for all the tool outputs of a turn, not one per output. The OpenAI Agents SDK for Python now does exactly this: in #2971 its maintainer describes a _ResponseCreateSequencer whose wait condition “requires no ongoing response before sending the actual response.create.” If you queue two creates instead of coalescing them, the error goes away and you get the opposite symptom, two replies to one turn, which is the failure we diagnosed in why an AI voice agent answers twice per turn. If the tool is slow enough that the silence itself is the problem, that is a separate fix, covered in why a voice agent goes silent on tool calls.
“It happens right after the caller interrupts”
Symptom: the caller barges in, your app stops playback and asks for a new reply, and the new reply is rejected. Cause: stopping playback on your side does not stop generation on OpenAI’s side. In three of the five reports in this group, the interruption never reached the server as a cancel: LiveKit #2064’s VAD interruption “didn’t interrupt the realtime session”; in #5642 the maintainer found interrupt() became “a no-op because it’s called before a speech is created”; in Agents SDK #1907 a guardrail trip sent no response.cancel when interrupt_response was set. In the other two, a cancel was sent but the replacement did not wait for it to finish. In #1912 the reporter cancelled by hand and then sent a replacement message; the maintainer later said the SDK now “forces response.cancel” and sequences the replacement “behind completion of the cancelled response.” In Agents SDK #2760 the guardrail path sent response.cancel, then a response.create without waiting for response.done; the reporter “ran into this while testing with turn_detection.interrupt_response enabled on a setup with some network latency.” PR #2763, merged 24 March 2026, fixed it exactly as below: “wait for realtime response.done before follow-up response.create”.
Fix: send response.cancel and wait for its response.done, then create. The client events reference says the server “will respond with a response.done event with a status of response.status=cancelled”, and that cancelling when nothing is in progress returns an error but leaves the session unaffected, so a defensive cancel costs nothing. If you use server VAD with interrupt_response: true, the server cancels for you when the caller starts speaking; the Realtime conversations guide says “The server will automatically cancel any in-progress model response and a response.cancelled event will be emitted.” Either way, key your next create off the confirmation, not off your own playback state. If the interruptions themselves are false, with the agent cutting itself off on noise, see why an AI voice agent keeps interrupting itself.
“I send my own message while the agent is speaking”
Symptom: you inject a line (“just a moment”, a handoff greeting, a reminder) while the model is mid-reply, and it is refused. Cause: it is a second writer to the default conversation. Fix: three options, depending on what the line is for.
- It must be spoken in sequence: queue it behind
response.done, or cancel the current response first if the new line supersedes it. - It is side work, not part of the dialogue (a classification, a summary, a moderation check): make it out-of-band. The conversations guide says out-of-band responses “which are not added to the default conversation state” are created by setting
response.conversationtonone, and the reference says such responses can run “in parallel”. Tag them withmetadataso you can match theresponse.done. - It is a holding line while a tool runs: the accepted answer in forum thread 1371442 was to play a local “thinking” sound instead of generating audio at all. It cannot collide because it never touches the API.
Do not use out-of-band responses for spoken replies: the model will not remember them on the next turn.
“Server VAD is racing my response.create”
Symptom: errors cluster at the end of caller turns, or when you call a reply from an event handler (a DTMF digit, a timer, an on_enter greeting). Cause: with turn_detection.create_response on, the server creates a response itself when the caller stops speaking, so there are two writers. LiveKit #6817 describes it plainly: “two independent writers can start a response: (1) the server itself after VAD end-of-speech, and (2) the client.”
This one cannot be fully prevented from the client. The server’s response.created takes time to reach you; if your create is already on the wire, it will collide no matter how careful your state tracking is. LiveKit’s maintainer said so in December 2024 (the error “cannot be completely avoided in server_vad mode because the OAI server may create the response at any time”), and #7514, open as at 28 September 2026, lists the same in-transit race as a limitation of its own proposal. So the fix has two parts: pick one owner for response creation (turn create_response off and create every response yourself, or stop issuing creates at turn end), and treat the residual collision as retryable. Match the error to your create by error.event_id, which the server events reference defines as “The event_id of the client event that caused the error”, then resend after the winning response’s response.done. Several reports in our set, #1466 and forum 1005582 among them, show event_id as null, so keep a fallback to your most recent create.
“It fails after a timeout or a failed response”
Symptom: the error appears with no visible active response, or after your own code gave up waiting. Cause: a response your app has stopped tracking is still alive on the server. In LiveKit #6223 a local generate_reply timeout abandoned a response the server kept generating. In forum thread 1374711 a response took 1 minute 3 seconds to end in a server error, and the reporter’s other responses failed with this code meanwhile. Fix: when your code gives up on a response, send response.cancel for it rather than just forgetting it (LiveKit’s PR #6244 did this, “discard orphaned response after interrupt or timeout”), and let only response.done clear your active-response flag. Never clear it on a timer.
The one-in-flight rule, as code
Every fix above is one rule. We call it the one-in-flight rule: nothing sends response.create to the default conversation until the last response.done has arrived; if you cannot wait, send response.cancel and wait for that instead. Out-of-band responses are the only exception. Here it is as a transport-agnostic gate you call instead of sending response.create directly:
"""One-in-flight gate: at most one response.create outstanding against the default conversation."""
import itertools
ERR_CODE = "conversation_already_has_active_response"
ERR_TEXT = "Conversation already has an active response"
class ResponseGate:
def __init__(self, send, max_retries=2):
self.send, self.max_retries = send, max_retries
self.active = None # id of the response writing to the default conversation
self.waiting = None # event_id of our create that has no response.created yet
self.queue, self.sent, self.last_eid = [], {}, None
self._ids = itertools.count(1)
self.cancel_on_created = False # a barge-in arrived before our response.created
def request(self, params=None, tries=0):
params = params or {}
if params.get("conversation") == "none": # out-of-band: allowed in parallel
self.send({"type": "response.create", "response": params}); return "sent out-of-band"
if self.active or self.waiting:
if not params and any(not p for p, _ in self.queue):
return "coalesced" # one reply for N tool outputs, not N replies
self.queue.append((params, tries)); return "queued"
eid = f"rc_{next(self._ids)}"
self.sent[eid] = (params, tries); self.waiting = self.last_eid = eid
for old in list(self.sent)[:-8]:
del self.sent[old] # keep the last 8 for error correlation
self.send({"type": "response.create", "event_id": eid, "response": params})
return "sent"
def cancel(self):
if self.active: # the queued create goes out on response.done
self.send({"type": "response.cancel", "response_id": self.active})
elif self.waiting: # not created yet: cancel it the moment it is
self.cancel_on_created = True
def on_event(self, ev):
t = ev.get("type")
if t == "response.created" and ev["response"].get("conversation_id"):
self.active, self.waiting = ev["response"]["id"], None
if self.cancel_on_created:
self.cancel_on_created = False
self.send({"type": "response.cancel", "response_id": self.active})
elif t == "response.done" and ev["response"]["id"] == self.active:
self.active = None
if self.queue:
self.request(*self.queue.pop(0))
elif t == "error" and (ev["error"].get("code") == ERR_CODE or
(ev["error"].get("message") or "").startswith(ERR_TEXT)): # code is null in older logs
eid = ev["error"].get("event_id") or self.last_eid # event_id is null in some reports
params, tries = self.sent.pop(eid, ({}, 0))
if eid == self.waiting:
self.waiting = None
if tries >= self.max_retries:
return "gave up" # surface it; do not loop forever
self.queue.insert(0, (params, tries + 1)) # retry after the active response.done
if not self.active:
self.request(*self.queue.pop(0))
elif t == "error" and self.waiting and ev["error"].get("event_id") == self.waiting:
self.waiting, self.cancel_on_created = None, False # our create failed for another reason: do not wedge the gate
We did not run this against OpenAI’s servers. We ran it on 28 September 2026 against a simulated server that enforces the documented rule (one default-conversation response at a time, out-of-band responses in parallel, response.done always emitted) with 20 ms of latency each way and 800 ms responses, once with the app sending response.create directly and once through the gate:
| Scenario (simulated) | Direct: creates / errors | Through the gate: creates / errors | What the gate did |
|---|---|---|---|
| Tool returns 300 ms into the reply that called it | 2 / 1 | 2 / 0 | Held the create until response.done |
| Two tools return 20 ms apart | 3 / 2 | 2 / 0 | Held, then coalesced two requests into one reply |
| Server VAD auto-responds 5 ms before the app asks | 1 / 1 | 2 / 1 | Could not prevent the in-transit collision; retried after response.done, so the app’s reply still played |
| Caller barges in at 400 ms | 2 / 1 | 2 / 0 | Sent response.cancel, created on the cancelled response.done |
| Side classification sent mid-reply | 2 / 1 | 2 / 0 | Sent it out-of-band, in parallel |
| App times out at 320 ms and asks again | 2 / 1 | 2 / 0 | Kept tracking the abandoned response; queued the retry |
The simulation proves the gate follows the rule, not how the real server times things, which is where forum reports such as 1005582 point. The retry cap exists for that case: after two failed retries the gate stops and surfaces the error rather than looping.
On a framework, check before writing any of this. LiveKit Agents merged fixes for its interrupt path (PR #5703), its orphaned responses (PR #6244) and its lost-reply timeout (PR #6818), but serialisation of response.create against server VAD is still an open request (#7514) as at 28 September 2026. The OpenAI Agents SDK for Python already sequences tool-triggered creates and, since PR #2763, guardrail follow-ups. Upgrading may be the whole fix. If you are still choosing a speech-to-speech stack, response lifecycle handling is one of the differences we compare in Gemini Live vs OpenAI Realtime for phone agents.
Fix it yourself, or hand it over
The DIY fix is small (the gate and the log test are about sixty lines each); the cost sits elsewhere. You need two-way event logging in production; without it a VAD race and a stale response look alike. You need a test that plays barge-ins and slow tools against a real session. And you need someone who owns the event loop when the framework changes underneath it; three of the LiveKit fixes in the table landed between May and August 2026.
| Your situation | Sensible route |
|---|---|
| Framework user, error appears after upgrades or only occasionally | Upgrade first; check the framework’s tracker against the table above before writing code |
| Raw WebSocket or WebRTC client, one developer owns the event loop | Add the gate and the two-way log; this is a day of work, not a project |
| Server VAD plus app-initiated replies (DTMF, timers, greetings) | Pick a single owner for response creation; keep a retry path for the in-transit race |
| Production phone lines where a silent turn loses a sale, and nobody owns the event loop | A managed agent platform moves the event loop to the vendor; ask any vendor, including us at Zian, what happens to a reply that loses this race |
Zian AI builds autonomous sales agents that run on phone, SMS, email and WhatsApp. It is in partnership-application beta; to join, Apply For Partnership.
Where every figure on this page comes from
| Figure | Who published it | Link | Date read |
|---|---|---|---|
“Only one Response can write to the default Conversation at a time”; out-of-band responses in parallel; response.cancel returns response.done with status cancelled |
OpenAI | Realtime client events reference | 28 September 2026 |
response.done “Always emitted, no matter the final state”; “Most errors are recoverable”; error.event_id definition; conversation_id null for out-of-band |
OpenAI | Realtime server events reference | 28 September 2026 |
Out-of-band via response.conversation = none; server auto-cancels on speech start |
OpenAI | Realtime conversations guide | 28 September 2026 |
| 17 reports; by trigger interruption 5, tool result 4, extra app message 4, server-VAD race 2, orphaned response 2; 12 GitHub (11 closed, 1 open), 5 forum; 3 threads excluded | Zian, our count of the reports linked in the table above | This page, method under the table | 28 September 2026 |
| “about 50% of the function calls” | A reporter on LiveKit Agents #1056 | livekit/agents #1056 | 28 September 2026 |
| Lost reply “dies as 10s timeout” | Reporter, LiveKit Agents #6817 | livekit/agents #6817 | 28 September 2026 |
| Response ran 1 minute 3 seconds before failing | Reporter, OpenAI Developer Community | forum thread 1374711 | 28 September 2026 |
| PR numbers and merge dates (#1204, #2072, #5703, #6244, #6818, #1968, #2763, #3140) | GitHub, on each repository | Linked from each issue in the table above | 28 September 2026 |
| 53 Pipecat issues for “openai realtime” (search control) | GitHub issue search, pipecat-ai/pipecat | pipecat-ai/pipecat issues | 28 September 2026 |
| Simulation results (creates / errors per scenario); 20 ms latency, 800 ms responses | Zian, our own simulation run against the code on this page | This page, one-in-flight rule section | 28 September 2026 |
Frequently asked questions
What does “Conversation already has an active response” mean?
It means the Realtime API refused your response.create because another response is still writing to the default conversation. OpenAI’s Realtime client events reference says only one response can write to the default conversation at a time. A response stays active from response.created until response.done, so wait for response.done or send response.cancel before creating the next one.
Is conversation_already_has_active_response a bug in OpenAI’s Realtime API?
Usually not. Each of the 17 reports we sorted fits one of five patterns in which a second response.create meets a response the server still treats as active: a tool result, an interruption whose cancel was never sent or never awaited, an extra app message, a server VAD auto-response or an orphaned response. Some forum posters, such as the opener of thread 1005582, saw it after every earlier response had finished, which points to a response you cannot see yet rather than a random fault.
Why does my agent go silent after this error?
The rejected response.create never produces a response.created, so nothing is generated for it. The session stays open, and reporters on LiveKit Agents #1056 found the agent resumed when the caller spoke again. To avoid the silent turn, resend the lost request after the active response’s response.done.
Can I run two responses at once in the Realtime API?
Only if one of them is out-of-band. Set response.conversation to none and the response runs in parallel without writing to the default conversation, which suits classifications and summaries. It does not suit spoken replies, because the model will not remember an out-of-band answer on the next turn.
Should I just retry response.create when I get this error?
Yes, but after the active response’s response.done, not straight away, or the retry collides again. Match the error to your request with error.event_id, fall back to your most recent create when that field is null, and cap the retries so a stale response cannot keep you in a loop.
Does OpenAI document this error code?
As at 28 September 2026 we found neither the code nor the message in OpenAI’s documentation, after searching the rendered client and server event references, the conversations, voice activity detection and error codes guides, the changelog and the combined export. The rule the error enforces is documented, in the response.create entry of the client events reference.