Mostly no. Bringing your own OpenAI and Deepgram keys changes who invoices you, not what you pay. Vapi charges its US$0.05 per minute hosting fee either way and passes models through at cost, so the model line only moves bills. Deepgram publishes a credit: US$0.010 per minute for your own text-to-speech, from 15 September 2026.
If I bring my own OpenAI and Deepgram keys, do I actually pay less?
That is the question as people actually ask it, and it turns on one thing almost nobody prices: what the platform takes off your bill when you supply the key, against what that component costs at your own rate card.
Vapi states the mechanism plainly: once your API key is validated, you will not be charged when using that provider through Vapi, and you will be charged directly by the provider instead (read 14 September 2026). Its FAQ settles the arithmetic — on model costs, “we add no fee (vendor cost passes-through)”. If the platform adds no margin on the model, removing the model from its invoice cannot save margin that was never charged.
The honest framing is three lines, not one. Call it the BYOK credit test:
- The credit. What comes off the platform bill per minute when you supply the key, from the vendor published tier table.
- Your rate. What that component costs at your rate card, which is rack rate unless you hold committed spend.
- The operating cost. What owning the key costs monthly in rotation, quota alarms and being paged mid-call.
Bringing your own key cuts cost only when line 1 exceeds line 2 plus line 3 divided by your monthly minutes. Most framings of this question skip line 1 or line 3.
What a provider key changes on the invoice, and what it does not
The three platforms that publish pricing take three incompatible positions, which is why buyers cannot compare them without building a model. All figures read at each vendor own pricing page, 14 September 2026.
| Platform | Platform fee | Model costs | Effect of your own key |
|---|---|---|---|
| Vapi, Build plan | US$0.05 per minute hosting; 10 concurrent lines included, US$10 per extra line per month | “At cost” pass-through for speech-to-text, LLM and text-to-speech | Model line becomes US$0 on the Vapi invoice and arrives from the provider instead. Hosting fee unchanged. |
| Retell AI, pay as you go | US$0.055 per minute voice infrastructure, inside a published US$0.07 to US$0.31 band | Itemised menu: platform voices US$0.015 per minute, ElevenLabs voices US$0.040, GPT 4.1 US$0.045, US telephony via Twilio US$0.015 | A Custom LLM option exists in the estimator with no published per-minute price. Custom telephony and SIP trunking carry no Retell charge. |
| Bland | US$0.14 per minute with no platform fee, or US$0.12 plus a US$299 monthly platform fee | Bundled: “No token charges. No model-provider pass-throughs.” | No model key to bring. Bland does support your own telephony carrier. |
Read that as a structure, not a league ladder. A platform fee is a normal commercial fact; the point is that it sits outside the part a key can move. The per-minute anatomy underneath is in voice AI cost per minute, realtime versus pipeline, and the reasons a bill exceeds minutes times rate are in why your AI voice agent bill beats minutes times rate.
Worked example: one call-minute, line by line, at published rates
A pipeline stack priced from vendor pages, every input substitutable. Two assumptions are ours: the agent speaks about 45% of the call at roughly 900 characters per spoken minute, giving 405 characters of synthesis per call-minute; and the LLM is called four times per call-minute at about 2,000 input and 60 output tokens.
| Line | Published rate, 14 September 2026 | Per call-minute | Does a key move it? |
|---|---|---|---|
| Vapi hosting | US$0.05 per minute | US$0.0500 | No |
| Twilio outbound, US and Canada | US$0.0140 per minute | US$0.0140 | No, you pay a carrier either way |
| Deepgram Nova-3 Monolingual streaming, two channels | US$0.0077 per minute regular | US$0.0154 | Yes |
| OpenAI gpt-5-mini | US$0.25 per 1M input, US$2.00 per 1M output | US$0.0025 | Yes |
| Deepgram Aura-2 text-to-speech | US$0.030 per 1,000 characters | US$0.0122 | Yes |
| Total | US$0.0941 | US$0.0301 of it |
Two details there cost money. Vapi cost documentation states that it transcribes audio from both the caller and the assistant, so the estimate includes two audio channels and your speech-to-text rate effectively doubles per call-minute. And model choice moves more than the key does: on the same token assumptions gpt-5.4 at US$2.50 and US$15.00 per million costs US$0.0236 per call-minute against US$0.0025 for gpt-5-mini.
Under Vapi billing that US$0.0301 arrives on the Vapi invoice at cost. Under your key it arrives from Deepgram and OpenAI. The total is the same number unless one side holds a discount the other does not.
Where bringing your own key actually crosses over
Deepgram is the useful case because it publishes the credit explicitly: its Voice Agent API tier table prices the same tier with and without your components, so the value of your key is published rather than inferred. From 15 September 2026, on pay as you go, Standard is US$0.075 per minute, Standard with your own text-to-speech US$0.065, Custom with your own LLM US$0.065, and Custom with your own LLM and text-to-speech US$0.050.
| What you bring | Published credit per minute | Your cost at rack rate | Net |
|---|---|---|---|
| Text-to-speech, Deepgram Aura-2, US$0.030 per 1k characters | US$0.010 | US$0.0122 | Worse by US$0.0022 |
| Text-to-speech, Deepgram Aura-1, US$0.0150 per 1k characters | US$0.010 | US$0.0061 | Better by US$0.0039 |
| LLM, OpenAI gpt-5-mini | US$0.010 | US$0.0025 | Better by US$0.0075 |
| LLM, OpenAI gpt-5.4 | US$0.010 | US$0.0236 | Worse by US$0.0136 |
| Both, gpt-5-mini and Aura-2 | US$0.025 | US$0.0147 | Better by US$0.0103 |
| Both, gpt-5.4 and ElevenLabs Flash, US$0.05 per 1k characters | US$0.025 | US$0.0439 | Worse by US$0.0189 |
The rack-rate column above is computed from our two assumptions, not from vendor data: 405 characters of synthesis and four LLM calls per call-minute.
The crossover is a rate, not a volume. At 405 characters per call-minute, a US$0.010 credit is repaid by any text-to-speech priced at or below US$0.0247 per 1,000 characters; above that your key costs money every minute. Carried to monthly totals on the Deepgram Standard tier from 15 September 2026:
| Monthly minutes | Platform billed, Standard | Own keys, lean models | Own keys, premium models |
|---|---|---|---|
| 2,000 | US$150.00 | US$129.26, saving US$20.74 | US$187.70, US$37.70 worse |
| 10,000 | US$750.00 | US$646.30, saving US$103.70 | US$938.50, US$188.50 worse |
| 50,000 | US$3,750.00 | US$3,231.50, saving US$518.50 | US$4,692.50, US$942.50 worse |
Telephony sits outside all six figures: the Deepgram Voice Agent API is a websocket service and you bring a carrier either way. So does concurrency — Retell includes 20 concurrent calls then charges US$8.00 each per month, Vapi includes 10 lines at US$10 per extra line. Those are monthly commitments a per-minute comparison never surfaces.
At what monthly volume does bringing my own key start to pay?
The volume threshold is real but second-order. What moves it is whether you hold a discount and whether anyone owns the keys.
| Condition | Effect on the break-even | Why |
|---|---|---|
| Platform passes models through at cost | Never arrives on price alone | The credit equals what you would have paid; the delta is zero |
| Platform publishes a bring-your-own tier discount | Becomes a per-minute rate comparison | Published credit against your rack rate, as in the table above |
| You already hold committed spend with the provider | Falls sharply | Your negotiated rate can beat the platform aggregate rate |
| Pay as you go, no commitment | Usually never arrives | Platforms buy in aggregate; you buy at rack rate |
| Considering a prepaid tier to close the gap | Roughly 25,600 call-minutes per month | Deepgram Growth starts at US$4,000 a year of credits redeemed against usage; two-channel Nova-3 Growth at US$0.0130 per call-minute consumes that in about 307,700 minutes |
| Nobody owns key rotation and quota alarms | Never arrives | The failure mode is a dropped call, not a bigger invoice |
| Models are bundled with no itemised menu | Not applicable | There is no key to bring |
| You need the data-processing agreement in your own name | Price stops deciding | See the next section |
Hold the scale of the prize in view. At the best published Deepgram credit of US$0.025 per minute, bringing both keys returns US$25 per 1,000 minutes. Divide your loaded hourly cost by 25 to get the thousands of minutes a month needed before one hour of key operations is paid for.
Does bringing my own key change where my audio is processed?
This is where buyers most often get the wrong answer, and where cost is frequently not the motive at all. A provider key changes the billing relationship and whose contract governs the model call. It does not by itself change the path the audio takes.
Vapi documents this. Its data-flow page states that even with maximum custom configuration certain data still passes through Vapi orchestration: raw audio routed in real time to transcriber and voice, transcribed text through orchestration analysis, and the orchestration layer — endpointing, interruption detection, emotion detection, backchanneling, filler injection — listed as Vapi only for both bring-your-own-key and custom-server configurations. A Deepgram key means Deepgram bills you and your Deepgram terms govern that inference. It does not mean the audio stopped crossing the platform.
What a key does give you:
- A direct contract with the model provider, so retention, training-use and sub-processor terms flow from your own agreement rather than being inherited.
- Access to the provider own regional options. OpenAI publishes that regional processing endpoints for data residency carry a 10% uplift for eligible models released on or after 5 March 2026.
- Audit evidence in your own tenancy, from provider-side usage logs and spend records you can produce yourself.
What it does not remove is the orchestration hop. If that hop is the real concern, a key is the wrong instrument and you want the model running somewhere you control: Zian AI supports private model deployment on customer infrastructure for this class of requirement. The hop-by-hop analysis is in data residency for AI voice agents, and the questions to put to any vendor are in the voice-AI vendor security questionnaire.
What owning the key costs that no pricing page shows
Line 3 of the BYOK credit test decides most real cases, and none of it reaches an invoice:
- Rate limits become yours. A platform sizes aggregate capacity across its customer base; your account has its own tier limits, and hitting one mid-call produces a failed conversational turn rather than a queued request. Alarm on provider 429s before you move a production agent.
- Quota and spend monitoring become yours. A hard spend cap that trips at 2am stops calls; a soft one stops nothing.
- Key rotation becomes yours, across every environment and every agent that references it, with no gap.
- You are the one paged. When a provider degrades, the platform cannot fail over on your behalf for a component it holds no credentials for.
All of it is doable: rotate keys on a schedule, set a soft and a hard spend alert per provider, and keep a second provider configured as fallback. Call it a day to set up and an hour or two a month to maintain, plus on-call exposure. Put your own numbers on those hours, compare them with the US$25 per 1,000 minutes the best published credit returns, and the decision makes itself. If a platform move is on the table too, what actually transfers when you switch AI voice platforms covers the rest.
Frequently asked questions
Does bringing my own API key remove the platform fee?
No. On Vapi the hosting fee of US$0.05 per minute is charged on the Build plan whether or not you supply provider keys, and the pricing page lists it separately from model provider cost. A provider key moves the speech-to-text, LLM and text-to-speech lines onto the provider invoice. The platform line does not move. Read at the Vapi pricing page on 14 September 2026.
Which voice platforms let me bring my own model keys?
Of the three that publish pricing, Vapi supports provider keys for transcription, LLM and voice and lists the supported providers in its data-flow documentation. Retell AI offers a Custom LLM option in its pricing estimator and charges nothing for custom telephony or SIP trunking. Bland bundles the models into one per-minute rate and states that there are no token charges and no model-provider pass-throughs, so there is no model key to bring, though Bland does support your own telephony carrier.
Will I get a better rate than the platform does?
Usually not, unless you already have committed spend or an enterprise agreement with that provider. Platforms buy in aggregate and you buy at rack rate. Vapi frames its own provider-key feature around customers who hold a custom model or an enterprise account with volume pricing, which is the honest scope of the benefit.
Does a provider key give me a direct data-processing agreement?
It gives you a direct contractual relationship with that model provider for the inference your key performs. It does not change the path the audio takes. Vapi documents that even with maximum custom configuration, raw audio and transcribed text still pass through its orchestration layer, and that the layer is not available for bring-your-own-key or custom-server substitution. If removing that hop is the requirement, private deployment answers it and a key does not.
How current are these prices?
Every figure here was read at the owning vendor pricing page on 14 September 2026. Voice AI prices change monthly. Deepgram was showing limited-time promotional streaming rates with no end date stated on its pricing page as at that date, and separately stated that standard Voice Agent API pricing applies from 15 September 2026. Re-read the source pages before you commit.
Work out your own crossover
The arithmetic above is reusable: substitute your own characters per call-minute, token counts and published rates. Apply For Partnership to discuss private model deployment on your own infrastructure.