Quick answer: Pick by where the caller’s audio starts. Browser or mobile app: WebRTC, which OpenAI recommends for clients. Audio already on your server, such as a Twilio Media Streams call: WebSocket. A phone number on a SIP trunk: SIP to sip.api.openai.com on port 5061. On WebRTC or SIP, add a server-side sideband WebSocket for tools.
This page is only about the transport decision for the OpenAI Realtime API: WebRTC, WebSocket or SIP, and the server-side control channel OpenAI calls a sideband. Other questions have their own pages. For OpenAI Realtime compared with Google’s Gemini Live on price and call length, see our Gemini Live vs OpenAI Realtime comparison for phone agents. For the full SIP setup, see the direct SIP guide with its cutover gate and firewall allowlist. For the Twilio side of a WebSocket call, see Twilio ConversationRelay vs Media Streams. Every fact about OpenAI below comes from OpenAI’s own developer documentation (developers.openai.com), read on 28 September 2026, when the current Realtime model in those docs was gpt-realtime-2.1.
Should I use WebRTC, WebSocket or SIP for the OpenAI Realtime API?
You rarely get a free choice between the three. The answer depends on two facts about your deployment, and both are usually settled before anyone opens the OpenAI docs.
Question one: where does the caller’s audio first become internet packets? In a browser tab or a native app, the device’s own microphone produces it. On a phone call, the carrier produces it: the PSTN leg ends at a SIP trunk or at a CPaaS such as Twilio, and one of those hands you the audio. In an agent framework room, a recorder or a contact-centre media fork, it is already on a server you run.
Question two: who has to hold the secrets and run the tools? The standard OpenAI API key, your CRM credentials, and any tool that books a meeting or looks up an account all belong on a server you control. OpenAI’s server-side controls guide states the general case: the Realtime API lets clients connect directly over WebRTC or SIP, “however, you’ll most likely want tool use and other business logic to reside on your application server to keep this logic private and client-agnostic.”
Put the two answers together and you get the rule this page is built on. We call it the Birthplace Rule: connect OpenAI at the point where the caller’s audio is born, and put your secrets and tools on a server-side sideband, never on the audio path. A browser gets WebRTC. A SIP trunk gets SIP. Audio that is already on your server goes over a WebSocket. Only the WebSocket case skips the sideband, because your server is already the one holding the connection.
The decision table: where the audio starts decides the transport
Each row below is a real deployment shape. Every cell is taken from OpenAI’s WebRTC, WebSocket, SIP and server-side controls guides and the Realtime API reference, all read on 28 September 2026. Where a cell is our inference rather than OpenAI’s wording, the text after the table says so.
| Where the caller’s audio starts | Connect OpenAI with | Who holds the standard API key, and what the client end holds | Who absorbs network jitter and plays the audio | Audio format you handle | Where tools and business logic run |
|---|---|---|---|---|---|
| Web browser (click-to-talk widget, web app) | WebRTC to /v1/realtime/calls, either through the unified interface or with an ephemeral key |
Your server holds the key. In the unified interface the browser holds nothing. In the ephemeral flow it holds an ek_ client secret that lasts 600 seconds by default and can be set anywhere from 10 to 7,200 |
The browser’s WebRTC stack: audio travels on media tracks, not in your code | None: the peer connection handles audio, and events arrive on the oai-events data channel |
Your server, through a sideband WebSocket opened with the call ID |
| Native mobile app | WebRTC. OpenAI’s client recommendation names “a web browser or mobile device”, and WARP connection-time optimisations are available through libwebrtc field trials | The same as the browser row | The app’s WebRTC stack | None | Your server, through a sideband WebSocket |
| Phone call on a SIP trunk you control (a carrier trunk or your own SBC) | SIP to sip:PROJECT_ID@sip.api.openai.com;transport=tls, with sip-eu.api.openai.com for European data residency |
Your server holds the key and uses it to accept the call and attach the sideband. The trunk holds no OpenAI key, because the SIP URI addresses your project by its ID | The carrier and trunk, which send SRTP media straight to OpenAI’s media addresses. Your server is not on the media path | None | Your server, through a WebSocket opened with the call_id from the realtime.call.incoming webhook |
| Phone call through a CPaaS that streams call audio to your server (for example Twilio Media Streams) | WebSocket, server to server, at wss://api.openai.com/v1/realtime |
Your server holds the key. The CPaaS never sees it | You: your relay buffers, paces and forwards frames between two sockets | Base64 audio in JSON events. Set audio/pcmu (G.711 μ-law) or audio/pcma (G.711 A-law) to match the carrier’s codec. audio/pcm must be 24 kHz |
Your server, on the same socket |
| Audio already on a server (an agent-framework room such as LiveKit or Pipecat, a recorder, a contact-centre media fork) | WebSocket, server to server | Your server holds the key | You | The same three formats. Resample anything that is not 24 kHz PCM or G.711 | Your server, on the same socket |
Two cells are our reading, not OpenAI’s words. First, in the unified interface your server sends the SDP offer to /v1/realtime/calls itself. OpenAI’s server-side controls guide says the response from that endpoint carries the call ID in its Location header, so in the unified flow your server should receive the call ID directly. In the ephemeral flow, the browser gets the call ID and has to pass it back to your server. Second, “LiveKit or Pipecat” are our examples of server-side audio. OpenAI names Twilio, Telnyx, LiveKit and Daily/Pipecat as partner integrations for its GPT-Live API, not for this table.
The quotable version of the table: for the OpenAI Realtime API, the browser gets WebRTC, the SIP trunk gets SIP, audio already on your server gets a WebSocket, and a sideband WebSocket puts your server back in control of the first two.
What OpenAI’s own docs say about each transport
OpenAI states a preference for each transport in plain words. The wording below is quoted from the pages linked in the source table at the end.
| Transport | OpenAI’s wording | Page |
|---|---|---|
| WebRTC | “When connecting to a Realtime model from the client (like a web browser or mobile device), we recommend using WebRTC rather than WebSockets for more consistent performance.” | Realtime API with WebRTC |
| WebSocket | “a great choice for connecting to the OpenAI Realtime API in server-to-server applications. For browser and mobile clients, we recommend connecting via WebRTC.” | Realtime API with WebSocket |
| WebSocket from a browser | “It is possible to use WebSocket in browsers with an ephemeral API token … but if you are connecting from a client like a browser or mobile app, WebRTC will be a more robust solution in most cases.” | Realtime API with WebSocket |
| Audio playback over WebSocket | “WebRTC will be more robust sending media to client devices over uncertain network conditions.” | Realtime conversations |
| SIP | “If you want to connect a phone number to the Realtime API, use a SIP trunking provider (e.g., Twilio).” | Realtime API with SIP |
| Sideband | “A sideband connection means there are two active connections to the same Realtime session: one from the user’s client and one from your application server.” | Webhooks and server-side controls |
Why WebRTC wins for browsers and mobile apps
The network reason is simple. A WebSocket runs over TCP; the IETF’s RFC 6455 describes the protocol as “layered over TCP”. TCP provides “a reliable, in-order, byte-stream service” (RFC 9293), so one lost packet on a patchy mobile or home connection holds up every audio frame behind it until it is resent. On a live voice call, late audio is as bad as lost audio. WebRTC was designed for media on unreliable networks, and on the Realtime API the peer connection handles the audio for you. OpenAI’s WebRTC guide says that over WebRTC “you don’t have to handle audio events from the model in the same granular way you must with WebSockets.” Microphone audio goes up a media track, the model’s voice comes back on another, and JSON events travel on a data channel labelled oai-events.
The key question has two documented answers:
- Unified interface. The browser posts its SDP offer to your server. Your server combines it with the session configuration in a multipart form and sends it to
/v1/realtime/callsusing your standard key. OpenAI describes this as “simpler setup and faster connections”, with the trade-off that your application server sits in the critical path every time a session starts. The browser never holds a credential. - Ephemeral key. Your server calls
/v1/realtime/client_secretswith the standard key and hands the browser anek_secret. The browser then connects to OpenAI directly. According to the API reference, the secret expires 600 seconds after creation unless you set a value between 10 and 7,200 seconds. Expiry stops new sessions only: “The session itself may continue after that time once started,” and “a secret can be used to create multiple sessions until it expires.”
For native apps, OpenAI documents WebRTC Abridged Roundtrip Protocol (WARP), which cuts the number of network round trips needed to start a session. Native clients built on a current libwebrtc can turn it on with three field trials. Browsers are more limited. Chrome supports DTLS 1.3 without an origin trial. SNAP needs an origin trial in Chrome 151 to 156. OpenAI says no browser origin trial exists for SPED, and that Firefox, Safari and browsers on iOS “can support standard WebRTC without supporting full WARP”. WARP makes the session start faster. It changes nothing about the transport choice.
The one-line version: use WebRTC whenever the caller is on a device you do not control, because TCP’s in-order delivery turns packet loss into audible stalls and WebRTC’s media stack does not.
When a WebSocket is the right answer, including most phone bridges
A WebSocket is right when the audio is already on your server. OpenAI calls it “perhaps the lowest-level interface available to interact with a Realtime model, where you will be responsible for both sending and processing Base64-encoded audio chunks over the socket connection.” Your server connects to wss://api.openai.com/v1/realtime?model=gpt-realtime-2.1 with the standard key in the Authorization header. It appends caller audio with input_audio_buffer.append, at no more than 15 MB per chunk, and reads the model’s voice from response.output_audio.delta events. The response.output_audio.done and response.done events carry transcripts, not audio bytes, so your playback code has to be built on the delta events.
The Realtime API reference lists three audio formats: audio/pcm, where “only a 24kHz sample rate is supported”; audio/pcmu, which is G.711 μ-law; and audio/pcma, which is G.711 A-law. Input and output are configured separately. OpenAI’s own example takes 24 kHz PCM in and sends μ-law out. For a phone bridge, set the format to match the codec the carrier leg already uses and you avoid resampling. How each API treats 8 kHz telephone audio is covered in our OpenAI Realtime vs Gemini Live phone-agent comparison, so this page does not repeat it.
The common phone case is a CPaaS that streams call audio to your server over its own WebSocket. OpenAI’s Realtime conversations guide links Twilio’s launch integration as the example. In that design your server sits between two sockets: the carrier’s and OpenAI’s. It owns buffering, barge-in (clearing queued audio when the caller talks over the agent) and call teardown. That extra hop is the price of a design where no SIP trunk points at OpenAI. The Twilio half of the bridge, and whether you want raw media or Twilio’s text-level relay, is the subject of our ConversationRelay vs Media Streams guide.
A WebSocket is also the fallback for runtimes that have no WebRTC stack. OpenAI’s browser-socket example notes that the standard WebSocket interface works in “browser-like environments like Deno and Cloudflare Workers.”
When the call arrives over SIP
SIP removes your server from the media path. Your trunk sends the call to sip:PROJECT_ID@sip.api.openai.com;transport=tls. OpenAI fires a realtime.call.incoming webhook to your server. Your server accepts it at /v1/realtime/calls/{call_id}/accept with the model, voice, tools and instructions, or rejects it (603 Decline by default). Signalling runs over TLS on port 5061, and SRTP media flows between the trunk and four published OpenAI /28 ranges. The call also has refer and hangup endpoints, and hangup works on WebRTC sessions too. One limit matters for sales teams: the GPT-Live view of OpenAI’s telephony page says outbound SIP calling “is available through the Live API, not the Realtime API call-creation endpoint”, so a Realtime SIP agent answers calls rather than placing them. Certificates, allowlists, the SBC settings and the five checks to run before moving traffic are covered in the direct SIP guide, and we will not repeat them here.
What is the sideband pattern, and how do I set it up?
A sideband is a second connection to a running Realtime session, opened by your server while the audio travels over WebRTC or SIP. It uses the same URL as a normal Realtime WebSocket, with a call ID in place of a model: wss://api.openai.com/v1/realtime?call_id={call_id}. Once it is open, OpenAI says you can “listen for events and configure the session just as you would from a typical Realtime API WebSocket connection”. That covers monitoring, session.update to change instructions, and answering tool calls.
Where the call ID comes from depends on the transport:
- WebRTC. The response from
/v1/realtime/callscarries aLocationheader such as/v1/realtime/calls/rtc_123456, and the last path segment is the call ID. In the ephemeral flow the browser receives it and must send it to your server. Authenticate that hand-off like any other request, because a call ID is the key to a live conversation. - SIP. The
realtime.call.incomingwebhook carriescall_id. Accept the call first, then open the socket. OpenAI says the WebSocket “will live for the life of the SIP call.” The model is fixed when you accept, so themodelargument is ignored on this connection.
The sideband is where the Birthplace Rule pays off. The browser or the carrier carries the audio, your server holds the key and the tools, and neither end has to trust the other with more than it needs. A sideband WebSocket is how an OpenAI Realtime call stays fast at the edge and private at the centre.
What goes wrong with a sideband: documented failures
This is the part most tutorials skip. The rows below come from OpenAI’s developer forum and OpenAI’s own guidance. Each forum thread was read on 28 September 2026, and its status is given as at that date.
| Symptom | What the source shows | Status as at 28 September 2026 | What to do |
|---|---|---|---|
Sideband WebSocket on a WebRTC call is rejected with HTTP 404, or call_id_not_found (“No session found for the provided call_id”) |
Thread 1360198, opened 27 September 2025. The poster took the call ID from the Location header and still got 404. One reply suggested authenticating with the session’s ephemeral key instead of the standard key. Two later replies (17 and 28 October 2025) said that worked for them. OpenAI’s documentation example authenticates with the standard key. |
Closed, with no accepted answer and no OpenAI staff explanation in the thread | Log the exact call ID and key prefix on every attach. Check that the key belongs to the project that created the call. If that is correct, test the ephemeral key before rewriting anything else. If the 404 comes from the SIP accept endpoint instead, see our guide to a SIP accept that returns 404 call_id_not_found |
| Sideband drops after the call goes quiet | Thread 1368689, opened 8 December 2025. The poster traced the drop to their WebSocket library’s ping_timeout with no reply arriving. A reply from an account titled OpenAI Staff said “Nothing on the server side enforces a timeout” and that they would investigate. |
Open, three posts, no fix posted | Treat the sideband as something that can drop without the call ending. Alert when it closes, and reconnect with the same call ID. Our guide to a voice agent that goes silent on a dead socket sets out a last-event clock for detecting this |
| A tool runs twice | OpenAI’s GPT-Live version of the server-side controls page: “If both connections receive a function-call event, execute the function once.” This is guidance, not a reported bug. | Current guidance | Give each tool a single executor (the server), and make bookings idempotent on the call ID |
| Old answers say the server cannot see a WebRTC session | Thread 1125810, February 2025. The server socket received only session.created, and a reply said a relay was the only option. The current server-side controls guide documents the call-ID sideband for both WebRTC and SIP. |
Superseded by the current docs | Check the date on forum answers before you build a relay you no longer need |
Can the browser change my agent’s instructions?
Assume it can. The WebRTC data channel exists so the client can “send and receive other client and server events”, and the Realtime reference says a client may send session.update “at any time to update any field except for voice and model.” The client-secret reference also says the session configuration you attach to an ephemeral key “can also be overridden by the client connection.” So anything that must not change, such as pricing rules, a disclosure script or the list of tools the agent may call, cannot be protected just by putting it in the session configuration. Keep authority on the server: run every tool call there, check each call’s arguments against your own records, and use the sideband to re-apply instructions if you detect drift. The same logic is why a phone call over SIP is, in one respect, simpler: the carrier end has no data channel to send events on.
Who each option is wrong for
- WebRTC is wrong for phone calls that already reach you as a server-side stream. Wrapping carrier audio in a WebRTC client on your own server adds a media stack and gains nothing a WebSocket would not give you.
- WebSocket is wrong for end-user devices. OpenAI recommends against it for browsers and mobile clients, and a direct browser socket puts an ephemeral key and the full event surface in the page.
- WebSocket bridges are wrong for teams that do not want to own audio pacing, barge-in and two-socket teardown. If you have a SIP trunk, SIP moves all of that to the carrier and OpenAI.
- SIP is wrong for outbound dialling on the Realtime API, which OpenAI routes to its Live API. It is also wrong for anyone without a trunk that can do TLS signalling and SRTP media, and for networks that cannot open port 5061 and the published media ranges.
- Skipping the sideband is wrong for any agent that touches a CRM, a calendar or money. Without one, on WebRTC the tool calls arrive at the browser, and on SIP nothing on your server receives them.
What running this yourself involves
Everything above can be built from OpenAI’s published docs, and the transport is the smaller part of the job. For a phone agent on SIP, the checklist includes a signed-webhook endpoint that verifies OpenAI’s signature and accepts or rejects each call, and one sideband per live call with reconnect handling for the silent-drop case. It includes a single tool executor that is idempotent on the call ID, a firewall rule for port 5061 plus the four SRTP ranges, and per-call logs keyed by call ID so a dropped call can be traced across the carrier, OpenAI and your own code. For a WebSocket bridge, add the relay’s buffering, barge-in and teardown on both sockets. None of it is exotic. All of it has to be run, watched and upgraded when OpenAI changes an event name, which the beta-to-GA migration has already done once: new session shapes, new event names and a new /v1/realtime/calls endpoint.
The transport decides whether a call connects. Whether calls turn into meetings depends on what happens after that: which leads get called, when, how often, and with which script. On Zian’s side, the learning engine tracks around 420,000 data points, and the platform has taken accounts from roughly 2% conversion to around 8%. Those are Zian’s own first-party figures, not independent measurements. If you would rather spend your engineering time on that layer than on sockets, that is the case for a managed agent platform. If the transport layer is your product, the table above is the place to start.
Frequently asked questions
Should I use WebRTC or WebSocket for the OpenAI Realtime API?
Use WebRTC when the caller is on a browser or mobile app, and a WebSocket when the audio is already on your server. OpenAI’s WebRTC guide recommends WebRTC rather than WebSockets for client connections for more consistent performance, and calls WebSockets a great choice for server-to-server applications.
Can I connect a phone number to the OpenAI Realtime API?
Yes, for inbound calls. Point a SIP trunk at sip:PROJECT_ID@sip.api.openai.com with TLS on port 5061, then accept each call from your server when the realtime.call.incoming webhook fires, as described in OpenAI’s Realtime SIP guide. The GPT-Live view of the same page says outbound SIP calling is available through the Live API, not the Realtime API call-creation endpoint.
What is a sideband connection in the OpenAI Realtime API?
It is a second WebSocket from your server to a Realtime session whose audio is travelling over WebRTC or SIP. You open wss://api.openai.com/v1/realtime?call_id= with the call ID, then monitor events, update instructions and answer tool calls there. OpenAI documents it in its server-side controls guide.
Why does my sideband WebSocket return 404 call_id_not_found?
It means OpenAI found no session for that call ID. On a WebRTC call, first confirm you took the ID from the Location header and that your key belongs to the project that created the call. In a developer forum thread (closed, no accepted answer, as at 28 September 2026), two users said authenticating with the session’s ephemeral key fixed it.
Is it safe to give the browser an ephemeral key?
It is the documented way to let a browser connect directly, and safer than a standard key: per OpenAI’s client secret reference it expires after 600 seconds by default, configurable from 10 to 7,200. But the same reference says the client connection can override the session configuration, so keep tools and business rules on your server.
Can I use a WebSocket from the browser instead of WebRTC?
You can, with an ephemeral key, but OpenAI’s WebSocket guide says WebRTC will be a more robust solution in most cases for browser or mobile clients. A WebSocket runs over TCP, so one lost packet delays the audio behind it until it is resent.
Where every figure on this page comes from
| Figure | Who published it | Link | Date read |
|---|---|---|---|
OpenAI’s client recommendation: WebRTC “rather than WebSockets for more consistent performance”; unified interface vs ephemeral key; the unified interface puts your server in the critical path; the oai-events data channel |
OpenAI, Realtime API with WebRTC | developers.openai.com | 2026-09-28 |
WebSocket “a great choice” for server-to-server; “more robust solution in most cases” for browser clients; “lowest-level interface”; Deno and Cloudflare Workers note; wss://api.openai.com/v1/realtime URL |
OpenAI, Realtime API with WebSocket | developers.openai.com | 2026-09-28 |
SIP URI sip:PROJECT_ID@sip.api.openai.com;transport=tls and sip-eu; realtime.call.incoming webhook; accept, reject (603 Decline default), refer and hangup; port 5061; four SRTP /28 ranges; outbound SIP “through the Live API, not the Realtime API call-creation endpoint”; Twilio, Telnyx, LiveKit and Daily/Pipecat named as GPT-Live partner integrations |
OpenAI, Telephony and SIP | developers.openai.com | 2026-09-28 |
Sideband definition; Location header call ID (rtc_123456 example); ?call_id= URL; SIP sideband “will live for the life of the SIP call”; “execute the function once” (GPT-Live version of the page) |
OpenAI, Webhooks and server-side controls | developers.openai.com | 2026-09-28 |
15 MB maximum per input_audio_buffer.append chunk; response.output_audio.delta; “more robust sending media to client devices over uncertain network conditions”; example session with 24 kHz PCM input and μ-law output |
OpenAI, Realtime conversations | developers.openai.com | 2026-09-28 |
Audio formats audio/pcm (“Only a 24kHz sample rate is supported”), audio/pcmu, audio/pcma; session.update may change any field except voice and model |
OpenAI, Realtime client events reference | developers.openai.com | 2026-09-28 |
| Client secret expiry: default 600 seconds, range 10 to 7,200 seconds; sessions may continue after expiry; one secret can create multiple sessions; configuration “can also be overridden by the client connection” | OpenAI, Create client secret reference | developers.openai.com | 2026-09-28 |
| WARP: three libwebrtc field trials; DTLS 1.3 in Chrome without an origin trial; SNAP origin trial in Chrome 151 to 156; no browser origin trial for SPED; Firefox, Safari and iOS browsers may lack full WARP | OpenAI, WebRTC with WARP | developers.openai.com | 2026-09-28 |
Current model name gpt-realtime-2.1; beta-to-GA changes (new session shapes, new event names, /v1/realtime/calls) |
OpenAI, Getting started with the Realtime API | developers.openai.com | 2026-09-28 |
| WebSocket is “layered over TCP” | IETF, RFC 6455 | rfc-editor.org | 2026-09-28 |
| TCP provides “a reliable, in-order, byte-stream service” | IETF, RFC 9293 | rfc-editor.org | 2026-09-28 |
Sideband 404 / call_id_not_found; opened 27 September 2025; one suggestion plus two confirmations (17 and 28 October 2025) that the ephemeral key worked; closed with no accepted answer |
OpenAI Developer Community, thread 1360198 (user reports, not OpenAI documentation) | community.openai.com | 2026-09-28 |
| Sideband drop after silence; opened 8 December 2025; “Nothing on the server side enforces a timeout” from an account titled OpenAI Staff; open, three posts | OpenAI Developer Community, thread 1368689 (user reports) | community.openai.com | 2026-09-28 |
February 2025 thread: the server socket received only session.created; a reply said a relay “would be the only option” |
OpenAI Developer Community, thread 1125810 (user reports) | community.openai.com | 2026-09-28 |
| Learning engine tracks around 420,000 data points; roughly 2% conversion to around 8% | Zian AI, first-party figures released by the company | First-party; no external URL | 2026-09-28 |
Want an agent that already handles the call path, so your team can focus on who gets called and what gets said? Apply For Partnership