Scaling AI Calling: Concurrency Limits at 10,000 Calls - Zian AI

Scaling AI Calling: Concurrency Limits at 10,000 Calls

Concurrency, not daily volume, is the binding constraint. Ten thousand dial attempts a day, at an assumed 67.5 seconds of line time per attempt across an 11-hour window, needs about 17 lines on average and roughly 34 at an assumed 2× peak. Vapi includes 10 concurrent call slots by default and Retell AI 20 on pay-as-you-go (both checked 2 September 2026).

A pilot at 100 calls a day tells you almost nothing about 10,000, because at 100 attempts the same arithmetic needs one line and nothing is under load. What decides whether 10,000 works is carrier calls-per-second caps, number reputation, provider rate limits, the calling windows in the Telecommunications (Telemarketing and Research Calls) Industry Standard 2017, scrubbing throughput under the Do Not Call Register Act 2006, and the size of your handoff queue. Limits below come from Twilio, Telnyx, Vapi, Retell AI, ElevenLabs and Bland’s own pages.

Converting a daily target into a concurrency requirement

Line time is consumed by every attempt, not just the answered ones — ringing occupies a channel too.

concurrent lines = (attempts per day × average line-seconds per attempt) ÷ (window hours × 3,600), then × a peak factor

Substitute your own measured numbers; the ones below are placeholders showing the shape of it:

  • Attempts per day: 10,000
  • Connect rate: 25%; answered calls occupy the line 180 seconds
  • Unanswered attempts occupy the line 30 seconds of ring time
  • Average line-seconds per attempt = (0.25 × 180) + (0.75 × 30) = 45 + 22.5 = 67.5
  • Total = 10,000 × 67.5 = 675,000 line-seconds, or 187.5 line-hours
  • Window = 11 hours (9 am to 8 pm on a weekday, the outer limit set by section 8 of the Industry Standard for non-research calls; section 15 notes that state and territory laws may narrow it further) = 39,600 seconds
  • 675,000 ÷ 39,600 = 17.05, so about 17 lines average
  • Dialling is never flat. Applying a 2× peak factor — an assumption, not a measurement; replace it with your own hour-by-hour distribution — gives 34 lines at peak

Two things fall out. The number is smaller than people expect — tens of lines, not hundreds — and it is still above the default allocation on most platforms. It is also very sensitive to the window: if your productive window is six hours rather than eleven, 675,000 ÷ 21,600 = 31.25 lines average, or about 63 at the same 2× peak. Shrinking the window is the fastest way to inflate your concurrency requirement. Redo the sum at 100 calls a day and you get 0.17 — a single line, idle most of the day, which is why pilots surface none of this.

Carrier and platform limits, as published by the vendors

Two ceilings apply. Concurrency is how many calls can be live at once; calls per second (CPS) is how fast you may start new ones, and for outbound it is usually the tighter of the two. Every figure below is from the vendor’s own documentation or pricing page, checked 2 September 2026. “Not published” means the vendor states no number on the page linked.

Platform Default concurrent calls Calls per second Daily / hourly cap Burst
Twilio Programmable Voice Not published for the Voice API; the SIP Trunking limits page states unlimited concurrent calls, but qualifies it as dependent on carrier support and says new accounts without an approved Business Profile have limited concurrent calls 1 CPS by default on calls created through the /Calls endpoint; accounts with an approved Business Profile can raise it to 5 in the Console. Calls above the rate are queued, not rejected Not published Not published
Telnyx 2 at initial setup, 10 after level 2 verification; above 10 channels is by request to Telnyx support 20 CPS per source IP address or SIP username; the CPS surcharge page puts the first 5 CPS at no charge Not published Not published
Vapi 10 concurrent call slots Not published Not published Add-on concurrency lines purchasable in the dashboard; no burst multiplier published
Retell AI 20 on pay-as-you-go workspaces Not published as a number; set per telephony path in the dashboard Not published Lower of 3× your limit or your limit plus 300, at a US$0.10/min surcharge applied to the entire call duration
ElevenLabs Agents 4 (Free) to 40 (Business) concurrent calls by plan Not published Not published 3× your limit or 300, whichever is lower, charged at 2× the standard rate
Bland 10 (Start), 50 (Build), 100 (Scale), custom on Enterprise Not published 100 / 2,000 / 5,000 calls a day and 100 / 1,000 / 1,000 an hour on those three plans; unlimited on Enterprise Not published

Read that against 34 lines at peak. A default Vapi or Telnyx allocation is under a third of it. Bland’s published Scale daily cap of 5,000 calls rules out a 10,000-call day on that plan whatever the concurrency, and its 1,000-an-hour cap bites sooner still, because 10,000 attempts over 11 hours averages about 900 an hour before any peak. Telnyx’s 20 CPS is generous; Twilio’s default 1 CPS is not — in theory it starts 39,600 calls in an 11-hour window, but it absorbs no burst, and a batch release of several thousand records queues behind a one-per-second gate. Ask every vendor for the number in writing, not the adjective.

Number reputation and answer rates

Concurrency you can buy. Answer rate you cannot. Placing 10,000 attempts from a handful of numbers concentrates volume per number to a level carrier analytics treat as bulk traffic, and a spam-labelled number drops your connect rate — which, in the arithmetic above, raises the attempts needed for the same conversations and pushes concurrency up again.

The controls are the number pool, per-number daily caps, rotation, retirement of labelled numbers and correct attestation. We have covered the identity side already: STIR/SHAKEN and AI voice agents in Australia, branded calling and choosing an Australian phone number. What changes at volume is that number reputation stops being a set-up task and becomes a daily metric with its own alarm.

Latency, provider rate limits and retry storms

A voice turn crosses several providers — speech to text, the model, text to speech, telephony — each with its own rate limit, and each returning an error rather than degrading gracefully. Under load the failure is rarely “the agent got slower”; it is a burst of HTTP 429s or SIP 503s in your busiest minute. Telnyx publishes 503 CPS Limit Reached (P05) for traffic over the calls-per-second ceiling in its SIP response code list, and 403 User channel limit exceeded for traffic over the concurrent channel limit; Retell AI queues an inbound call and, if no slot opens after about 40 seconds, ends it as concurrency_limit_reached or routes it to a fallback number.

The dangerous part is what your code does next. A naive retry turns a brief capacity dip into a retry storm: rejects are re-queued immediately, arrive while the limit is still breached, and the queue grows faster than it drains. Three rules bound it — exponential backoff with jitter, a cap on attempts per record per day, and a circuit breaker that stops dialling when the reject rate crosses a threshold. Telnyx’s own guidance is to back off exponentially, alert at 80% of capacity and queue rather than fail at the limit.

Latency budgets are a separate discipline, covered in sub-second response times for voice AI. The point here: a budget measured on an idle system is not a budget. Measure p95 at peak concurrency.

Compliance error rates multiply

This is the strongest argument for treating scale as a distinct engineering problem. Take a 0.5% error rate — an illustrative figure, not a measured one; substitute your own audit result. At 100 calls a day that is invisible: one bad call every two days. The same rate at 10,000 calls a day is 50 a day, 250 in a five-day working week and roughly 13,000 across 260 working days. The rate did not change; the absolute number became a compliance programme. Three throughput problems appear:

  • Calling windows are per recipient, not per campaign. Section 8 of the Industry Standard prohibits non-research calls on a weekday before 9 am or after 8 pm, on a Saturday before 9 am or after 5 pm, and at any time on a Sunday, plus seven named national public holidays and any weekday holiday given in lieu of one of them. Subsection 8(4) fixes the relevant time as the time of day at the place that is the relevant account-holder’s usual residential address. A 9 am AEST start is 7 am in Perth, so a national list dialled on one clock breaches the window on every Western Australian record for the first two hours. On the Federal Register of Legislation record checked 2 September 2026 the Standard is in force and due for repeal on 1 April 2027 under section 50 of the Legislation Act 2003 — see what the 2027 sunset means for AI callers.
  • Scrubbing has a shelf life. The exception in subsection 11(3) of the Do Not Call Register Act 2006 depends on the number having been submitted in a list and reported unlisted during the 30-day period ending at the end of the day the call was made. At 10,000 attempts a day, with records recycled through retries, scrub freshness has to be a per-record timestamp you can query and block on — not a monthly job.
  • Consent records must be retrievable per call. The exception in subsection 8(5) of the Standard only helps if you can produce, for a specific call, evidence that consent covered that day and time. At volume, “we have consent” is not an answer; a record ID is.

The penalty structure makes the multiplication concrete. Under subsection 25(3) of the Do Not Call Register Act 2006, the maximum civil penalty payable by a body corporate with no prior record is 100 penalty units for a contravention of subsection 11(1), and where a court finds two or more contraventions on a particular day, the total for that day must not exceed 2,000 penalty units. The Act states those caps in penalty units, not dollars. The Crimes (Amount of a Penalty Unit) Instrument 2026, made under subsection 4AA(1A) of the Crimes Act 1914, fixes a penalty unit at $364 from 1 July 2026. On our own multiplication, 2,000 × $364 puts that daily cap at $728,000 — our arithmetic, not a figure either instrument states, and worth a lawyer’s eye, because subsection 4AA(8) of the Crimes Act expresses the indexed amount as applying to offences committed on or after the indexation day while the Do Not Call Register Act caps are civil penalties. Volume does not change the rules; it changes how many times a day you can break them. Our write-up of ACMA’s approach is at Do Not Call Register enforcement and AI voice agents.

Human capacity is the real ceiling

At volume the agent is rarely the bottleneck — the handoff queue is. The rates below are the same illustrative placeholders as before, not published figures. 10,000 attempts at the assumed 25% connect rate is 2,500 conversations; if an assumed 10% ask for a person, that is 250 transfers a day, about 23 an hour across an 11-hour window. At an assumed 12 minutes of handling each, 23 × 12 ÷ 60 = 4.6, so roughly five people available concurrently, before breaks, peak clustering or sick days. If they do not exist, transfers queue, callers hang up and you have paid for an abandonment. Model the human side with the same formula you used for lines, and make the transfer carry context so the handover is short — see context transfer on AI-to-human handoff.

What to alarm on

At 100 calls a day you notice a problem by reading transcripts; at 10,000 you notice because something fired. Alarm on rates and ratios, not counts.

Constraint What breaks How you detect it What to do
Concurrency ceiling Calls rejected or held at the busiest minute Peak concurrent calls as a percentage of your cap; rejects with the provider’s limit error Alarm at 80% of cap; queue rather than fail; buy additional concurrency or move the window
Carrier CPS Batch releases fail in bursts while total volume looks fine Rejects per minute clustered at batch start; SIP 503 CPS responses Rate-limit your dialler below the carrier cap; stagger batch releases
Number reputation Connect rate falls without any change to the script Answer rate per number and per number pool, tracked daily Cap calls per number per day, rotate, retire labelled numbers, fix attestation
Provider rate limits Turn latency spikes, then errors, at peak only p95 turn latency and 429 rate, both segmented by concurrency band Backoff with jitter, provider failover, load test at target concurrency
Retry logic Queue grows faster than it drains; the same records dialled repeatedly Attempts per record per day; ratio of retries to first attempts Cap attempts per record; circuit-break dialling above a reject threshold
Calling windows Calls placed outside permitted hours in another time zone Count of attempts outside the permitted window, computed in the recipient’s local time Schedule per record on recipient local time; hard block, not a warning
DNC scrub freshness Stale washes stop supporting the subsection 11(3) exception Age of the most recent scrub per record, alarmed before 30 days Continuous scrubbing pipeline; block dialling on any record past the threshold
Handoff capacity Transfers queue and callers abandon Transfer queue wait time and abandonment rate by hour Staff to the peak hour, not the daily average; degrade to callback booking

The traces and per-conversation records worth keeping are set out in AI agent observability.

Do you actually need 10,000 a day?

Often, no. If your connect rate is 12% and your conversation-to-meeting rate is under 5%, scaling volume mostly scales the number of people annoyed by your programme while raising every risk on this page at once. Fix the list, the offer and the script at 500 a day, where an error costs two bad calls instead of fifty. If the volume is genuinely there, the discipline is unglamorous: measure line-seconds, size concurrency to peak rather than average, get published limits in writing from every provider in the chain, scrub continuously, schedule in the recipient’s local time, and staff the queue to the peak hour.

Zian has been running outbound acquisition since 2017, across more than 10,000 leads a day, at one point operating campaigns for roughly a hundred businesses in the same vertical simultaneously. Those are lead volumes, not call volumes, and not a concurrency figure — Zian does not publish one. The point of this post is that you should ask any vendor, us included, for the arithmetic rather than the adjective.

Frequently asked questions

How many concurrent lines does 10,000 calls a day need?

Multiply attempts per day by average line-seconds per attempt, divide by the seconds in your calling window, then apply a peak factor. At 10,000 attempts, a 25% connect rate, 180 seconds of talk time, 30 seconds of ring time and an 11-hour window, that is 675,000 line-seconds over 39,600 seconds, or about 17 lines average and about 34 at a 2× peak. Every input in that example is an assumption you should replace with your own measured figures; only the arithmetic is ours.

What are the legal calling hours for outbound calls in Australia?

Section 8 of the Telecommunications (Telemarketing and Research Calls) Industry Standard 2017, made by ACMA under subsection 125A(1) of the Telecommunications Act 1997, prohibits non-research telemarketing calls on a weekday before 9 am or after 8 pm, on a Saturday before 9 am or after 5 pm, and at any time on a Sunday, plus New Year’s Day, Australia Day, Good Friday, Easter Monday, Anzac Day, Christmas Day and Boxing Day, plus any weekday holiday given in lieu of one of those. Subsection 8(4) measures the time at the place that is the relevant account-holder’s usual residential address, so a national list needs per-record local-time scheduling. Checked 2 September 2026 on the Federal Register of Legislation: the Standard is in force and scheduled for repeal on 1 April 2027 under section 50 of the Legislation Act 2003.

What happens when you hit a platform’s concurrency limit?

It depends on the platform, which is why you should read the page before you sign. Vapi documents that new dials wait until a slot frees. Retell AI holds an inbound call briefly and, if no slot opens after about 40 seconds, transfers it to a fallback number or ends it as concurrency_limit_reached. Telnyx returns 403 User channel limit exceeded at the concurrent channel limit and 503 CPS Limit Reached at the calls-per-second limit. All checked 2 September 2026.

How fresh does Do Not Call Register scrubbing have to be?

Subsection 11(3) of the Do Not Call Register Act 2006 ties the exception to information received during the 30-day period ending at the end of the day the call was made. In practice that means a per-record scrub timestamp you can query and block on, not a monthly batch — particularly once retries recycle records days or weeks after the original wash.

Is calls per second or concurrency the tighter limit?

For outbound, usually calls per second. Concurrency is set by how long calls last, which you can estimate; CPS is set by how fast you release batches, which spikes. Twilio publishes a default of 1 CPS per account, raisable to 5 in the Console for accounts with an approved Business Profile, while its SIP Trunking limits page states unlimited concurrent calls, qualified as dependent on carrier support and limited for new accounts without an approved Business Profile. Telnyx publishes 20 CPS per source IP address or SIP username. Both checked 2 September 2026.

Should we scale before the pilot metrics are good?

No. Volume multiplies whatever rate you already have, including the error rate. An illustrative 0.5% compliance error rate at 100 calls a day is one incident every two days; the same rate at 10,000 calls a day is 50 a day. Fix the rate first, then scale.

Running the numbers on your own programme

Sizing an outbound programme and want the concurrency, compliance and handoff maths run against your own numbers rather than a vendor’s brochure? Apply For Partnership.

Related Blogs

Related from Zian AI