Acceptable Hallucination Rate for AI Voice Agents - Zian AI

Acceptable Hallucination Rate for AI Voice Agents

Quick answer: No figure is acceptable everywhere: an acceptable hallucination rate is one whose 95% upper bound, on your own graded calls, sits below the threshold you set for that action. Proving under 1% takes at least 299 graded calls with zero hallucinations. Vectara’s leaderboard (1.8% to 24.2% across 108 models, 22 September 2026) scores written summaries, not calls.

What is an acceptable hallucination rate for a customer-facing AI agent?

An acceptable hallucination rate is not a number you look up. It takes two decisions: a threshold for each kind of thing the agent says, and a measurement on your own graded calls whose 95% upper bound lands below it.

The first decision is how much risk each action can carry. Our guide to confidence thresholds and risk-tiered autonomy for AI sales agents covers that tiering by reversibility and blast radius. This page picks up where it stops: once you have written down “no more than 1% of booking confirmations may misstate the slot”, how do you find out whether your agent meets it?

The second decision is the one that gets skipped. A point estimate such as “we checked 100 calls and found one error, so 1%” does not show that you are under 1%. With 1 error in 100 calls, the Wilson 95% interval (computed by the script on this page) runs from 0.18% to 5.45%. The acceptable rate is the one you can prove, and proof means the upper end of the interval, not the middle.

What counts as a hallucination on a phone call?

Vectara’s model card for its open hallucination detector gives a usable definition: a hallucination is text not supported by a given piece of evidence, and “You always need two pieces of text to determine whether a text is hallucinated or not.” For a voice agent, the first text is what the agent said. The second is what it was entitled to rely on: retrieved passages, tool responses, the system prompt and the caller’s earlier words.

That gives three grading codes, counted separately:

  1. Unsupported factual claim (H1). The agent states something about the world, the product or the account that none of its evidence supports. “Your plan includes free installation” when no retrieved passage says so.
  2. Tool-result misstatement (H2). A tool returned one thing and the agent said another. The calendar API returned 2:30 pm and the agent confirmed 3:30 pm. Or, worse, the booking call returned an error and the agent said “You’re all booked in.” This code is often the expensive one, because the caller acts on it.
  3. Policy violation (P). The agent says something it is not permitted to say, whether or not it is true: offering a discount, making a delivery promise, giving advice reserved for a licensed human. Not strictly a hallucination, since it may be well grounded, but it belongs on the same scorecard because the caller cannot tell the difference.

What is not a hallucination

Each of these has a different fix, so counting them as hallucinations sends the fix to the wrong team.

What happened on the call Code it as Evidence that separates it Where the fix lives
Agent stated a fact no evidence supports H1 hallucination Retrieved passages and prompt contain nothing that supports the claim Prompt, retrieval, answer constraints
Agent restated a tool result wrongly H2 hallucination Logged tool payload differs from what the agent said Read-back rules, structured confirmation
Agent made a forbidden promise P policy violation Claim may be grounded; policy list forbids it Guardrails, human approval tier
Speech recognition misheard the caller, agent answered the wrong question faithfully ASR error, not a hallucination Audio differs from transcript; answer matches the transcript ASR model, vocabulary, confirmation turns
Knowledge base holds the wrong fact, agent relayed it accurately Content defect, not a hallucination Agent’s claim matches the retrieved passage word for word Knowledge-base owner
Retrieval found nothing, agent said it did not know and offered a callback Correct behaviour Empty retrieval, no claim made Nothing to fix; count it in the denominator
Transcript shows words the agent never spoke Transcript defect Recording differs from transcript Transcription pipeline

The ASR row is easy to get wrong: a grader reading only the transcript sees a mismatched answer and codes a hallucination. Our page on whether an AI call transcript is a record of what was actually said documents the ways the two drift apart. Grade disputed turns against the audio.

Why published hallucination rates do not tell you your agent’s rate

A widely cited public number is Vectara’s Hallucination Leaderboard on GitHub. As last updated on 22 September 2026, it lists 108 models, with hallucination rates running from 1.8% (antgroup/finix_s1_32b) to 24.2% (mistralai/ministral-3-3b-2512). The median of the 108 rates is 9.55%, computed by us from the table. For one widely deployed model, openai/gpt-4o-2024-08-06, it shows 9.6%.

Those numbers measure one specific task. Vectara’s methodology section says each model was asked to summarise every document in a private set of more than 7,700 articles, at a temperature of 0 where possible, and that each summary was judged for factual consistency with its source by HHEM-2.3, Vectara’s commercial evaluation model. The unit is a whole summary of a written article.

A phone call differs on nearly every axis: speech passed through recognition, questions rather than a document, tools called mid-turn, and evidence that mixes retrieved passages, API payloads and the caller’s own words.

In fairness to Vectara, it argues the opposite case, and its own words should be on this page. Its README calls summarisation “a good analogue to determine how truthful the models are overall”. It also says that because models in RAG and agentic pipelines act as summarisers of search results, “this leaderboard is also a good indicator for the accuracy of the models when used in RAG or agentic systems.” That is a reasonable basis for picking between models, but it is not a measurement of your deployment.

A second public benchmark shows how wide that gap can be. The τ-bench paper, submitted to arXiv on 17 June 2024, simulates conversations between a user, played by a language model, and an agent that must use domain API tools and follow policy guidelines, in domains including retail. Its abstract reports that “even state-of-the-art function calling agents (like gpt-4o) succeed on <50% of the tasks, and are quite inconsistent (pass^8 <25% in retail).” τ-bench scores whole-task success, not hallucination, so its numbers cannot be compared with Vectara’s either. It measures something closer to a phone agent’s job, and on it even the gpt-4o family fails more than half the time. The paper predates the 2024-08-06 snapshot Vectara lists, so these are not the same model version.

Public figure What it measures Unit counted Why it is not your call rate
Vectara leaderboard: 1.8% to 24.2% across 108 models (22 Sep 2026) Factual consistency of summaries of written articles One summary per document, judged by HHEM-2.3 No speech, no tools, no dialogue; temperature 0 where possible
Vectara leaderboard: 9.6% for gpt-4o-2024-08-06 Same task, one model Same Same
τ-bench: gpt-4o succeeds on under 50% of tasks; pass^8 under 25% in retail (17 Jun 2024) Task completion under domain policy, with simulated users and tools Whole task, checked against expected database state Text, not audio; task failure includes far more than hallucination

A summarisation leaderboard ranks models; only graded calls from your own deployment measure your agent.

How do I measure my voice agent’s hallucination rate?

A team with call recordings, tool logs and a spreadsheet can run this without any vendor.

  1. Log the evidence, not just the transcript. For every turn, keep the retrieved passages, each tool call’s arguments and response, and the prompt version that was live. Without the second text, graders fall back on their own knowledge, which is the error the definition rules out.
  2. Freeze the unit of change. A rate belongs to one configuration: one prompt version, one model, one knowledge-base snapshot. When any of them changes, start a new sample.
  3. Sample at random, then stratify. Draw calls at random within each call type you care about, such as inbound enquiry, booking and follow-up. Do not grade only the calls that were flagged, escalated or complained about. That measures your complaint pipeline. Review flagged calls too, in a separate column.
  4. Grade to codes, blind. Graders mark each checkable claim H1, H2, P or clean, plus a note, against the logged evidence and the audio. They should not see which prompt version or model produced the call. Have a second grader independently code a subset and settle every disagreement before you count.
  5. Pick the denominator before you count (next section), and report the numerator, the denominator and the interval, never the percentage alone.
  6. Compare the upper bound with your threshold. Below it: the claim holds for this configuration. Straddling it: grade more calls. Above it: fix before you scale.

For scripted scenarios before launch, our guide to testing an AI voice agent before go-live works out how many repeat runs a single scenario needs, which it calls the Rule of 59. That is the same zero-failure arithmetic as the next sections, applied to synthetic runs instead of live calls.

Evaluating an AI phone agent for your team? Zian AI runs autonomous sales agents across phone, SMS, email and WhatsApp, with knowledge-base lookups, CRM integrations and private model deployment on your own infrastructure. We are in partnership-application beta.

Apply For Partnership

Per call, per claim or per action: the denominator can change the rate five-fold

“Hallucination rate” has no fixed denominator. Take one set of hypothetical grading results, invented for illustration and not measured on any deployment: 400 graded calls containing 2,600 checkable claims, of which 520 were claims that drove an action (a booking, a price, a commitment). Graders found 14 unsupported claims spread across 11 calls, and 4 of those claims drove an action.

Denominator Count Rate Wilson 95% interval Answers the question
Per call: calls with at least one hallucination 11 / 400 2.75% 1.54% to 4.86% How many callers heard something untrue?
Per claim: all checkable claims 14 / 2,600 0.54% 0.32% to 0.90% How often is any single statement wrong?
Per action-bearing claim 4 / 520 0.77% 0.30% to 1.96% How often is a statement the caller acts on wrong?

Divide 2.75% by 0.54% and the same data gives a per-call rate about five times the per-claim rate. Both are correct. The per-call rate is what callers experience; the per-claim rate is what looks best on a slide. When anyone quotes you a hallucination rate, including a vendor, the first question is “per what?”

Per-claim counting has a second problem: errors cluster. A call where a tool timed out tends to produce several wrong claims. We simulated 4,000 samples of 400 calls, with 6 or 7 claims per call. In 95% of calls each claim had a 0.2% chance of being wrong, and in the other 5% (a “bad path”, such as a failed tool) each claim had an 8% chance. The naive per-claim Wilson 95% interval contained the true per-claim rate of 0.59% in only 92.1% of samples. The per-call interval, which counts each call once, contained the true per-call rate of 3.32% in 95.5%. Those inputs are simulated, not measured, but this is the expected direction whenever errors cluster within calls: the per-claim interval comes out too narrow. If you report per claim, compute the interval at the call level or resample whole calls.

The calculator behind the worked figures uses only the Python standard library. The output shown is from running exactly this code:

import math

def wilson(x, n, z=1.96):
    """95% Wilson score interval for x hallucinations in n graded units."""
    p = x / n
    centre = p + z * z / (2 * n)
    half = z * math.sqrt(p * (1 - p) / n + z * z / (4 * n * n))
    d = 1 + z * z / n
    return (centre - half) / d, (centre + half) / d

def zero_error_upper(n, confidence=0.95):
    """One-sided upper bound when you graded n units and found none."""
    return 1 - (1 - confidence) ** (1 / n)

for label, x, n in [("per call", 11, 400), ("per claim", 14, 2600)]:
    lo, hi = wilson(x, n)
    print(f"{label}: {x}/{n} = {x/n:.2%}  (95% CI {lo:.2%} to {hi:.2%})")
print(f"0 of 300 calls: rate is below {zero_error_upper(300):.2%} (one-sided 95%)")
per call: 11/400 = 2.75%  (95% CI 1.54% to 4.86%)
per claim: 14/2600 = 0.54%  (95% CI 0.32% to 0.90%)
0 of 300 calls: rate is below 0.99% (one-sided 95%)

The Wilson formula is the one in section 7.2.4.1 of the NIST/SEMATECH e-Handbook of Statistical Methods. We checked our implementation against that page’s worked example (a proportion of 0.13 from 200 units) and reproduced its one-sided lower limit of 0.09577.

How many calls do I need to grade?

It depends on what you expect to find.

Case 1: you found zero hallucinations

With zero errors in n graded calls, the exact one-sided 95% upper bound is 1 − 0.051/n. A handy approximation is 3/n, usually called the rule of three. The classic reference on zero counts is Hanley and Lippman-Hand’s 1983 JAMA paper, titled “If nothing goes wrong, is everything all right? Interpreting zero numerators.” Turned around, it tells you how many clean calls you need to show a rate below T.

Threshold you want to prove (T) Rule of three (3/T) Exact clean calls needed (one-sided 95%)
5% 60 59
2% 150 149
1% 300 299
0.5% 600 598
0.1% 3,000 2,995

Read the other way: 100 clean calls only bound the rate below 2.95% (exact) or 3% (rule of three). Thirty clean calls bound it below 9.50%.

Case 2: you expect a few, and need the bound to clear the threshold

Zero is optimistic. A more realistic plan assumes the true rate is half your threshold and asks how many calls give a good chance that the exact upper bound lands below T. The script scanned every sample size for the smallest n from which that chance stays at 80% or higher:

Threshold (T) Assumed true rate Calls to grade Pass if hallucinations ≤ Chance of passing at that n
5% 2.5% 434 14 86.8%
2% 1% 1,091 14 86.2%
1% 0.5% 2,185 14 86.0%

The chance moves up and down in a saw-tooth as n grows, because counts are whole numbers. The first sample sizes to touch 80% were 361, 969 and 1,941, but slightly larger samples dip below 80% again, so plan with the stable figures above. Compare 2,185 with 299: a good but imperfect agent needs about seven times more grading to prove the same 1% claim.

Case 3: “estimate a 1% rate to within ±0.5 points”

The textbook sample-size formula, n = z²p(1 − p)/E², gives 1,521.2, so 1,522 calls. The expected count is 15.2 hallucinations, well past the common rule of thumb that n × p should exceed 5. If you then observe 15 in 1,522, the interval it assumes (the Wald interval) is 0.49% to 1.48%. The Wilson interval is 0.60% to 1.62%, and the exact Clopper-Pearson interval is 0.55% to 1.62%. The upper side overshoots ±0.5 points, because a rate near zero cannot have a symmetrical interval.

The Wald interval’s actual coverage shows why this matters. With a true rate of 1%, we computed the exact probability that each nominal-95% interval contains it:

Calls graded Wald interval actually covers Wilson interval actually covers
100 63.3% 92.1%
300 79.9% 96.7%
1,522 92.5% 94.8%

At 3 hallucinations in 300 calls the Wald interval runs from −0.13% to 2.13%, a negative lower limit for a rate. NIST calls a method that produces an impossible limit “an inferior approach”. The Wilson interval for the same count is 0.34% to 2.90%. For hallucination rates, which live near zero, use Wilson or exact intervals and never the ±1.96 standard-error shortcut.

The 3/T feasibility rule: when sampling cannot prove it

Put the tables together and a decision rule falls out, which we call the 3/T feasibility rule: if you cannot grade at least 3/T calls from one unchanged configuration, sampling cannot demonstrate threshold T, so do not rely on sampling for that action. Gate it instead with a deterministic check (read the tool payload back verbatim, and confirm the slot before committing) or a human approval step.

Threshold you need for the action Minimum clean calls (3/T) Feasibility, and what to do meanwhile
5%: informational answers 60 Almost always feasible; sample and report
1%: bookings and confirmations 300 Feasible for most teams; add structured read-back while the sample builds
0.1%: prices, terms, anything with legal colour 3,000 Rarely feasible between prompt releases; use human approval or a deterministic template

The left-column thresholds are examples to replace with your own; the middle column is arithmetic. It lines up with the risk tiers in our confidence-threshold guide, which put commitments behind a human: at strict thresholds, few grading programmes can prove the agent safe on its own.

What running this yourself actually costs

The cost is grading time, and it scales with the threshold, not with call volume. As an assumption to replace with your own timing: if a grader needs 4 minutes per call to check each claim against logged evidence and audio, the 299 clean calls needed to prove 1% take about 20 hours. The 2,185-call plan for a merely good agent takes about 146 hours. Add a second grader on a subset, and repeat after every change that touches the measured behaviour.

Two things tend to break at volume: evidence logging, when a stack keeps the transcript but not the tool payloads, and grader consistency, as the H1/H2/P line drifts with fatigue. Vectara’s open HHEM-2.1-Open model can score (premise, hypothesis) pairs automatically, returning 0 to 1, where 0 means the hypothesis is not evidenced at all by the premise. That can pre-screen claims for H1, but it is tagged English on Hugging Face and knows nothing about your policies, so validate it against a human-graded subset first. If a grade ever lands on a claim that already reached a customer, our first-hour guide for when an AI agent told a customer something untrue covers what to preserve and whom to call.

Whichever agent platform you evaluate, Zian AI included, the useful question is not “what is your hallucination rate?” It is “can I get per-turn evidence logs, meaning the retrieved passages and tool payloads, so I can measure it myself?” Zian’s agents do research, web and knowledge-base lookups and connect to CRMs such as HubSpot, Salesforce and HighLevel. Each of those lookups is a second text the method above needs, so put the evidence-log question to every vendor on your shortlist, us included. This page quotes no hallucination rate for Zian or any other platform; the point is to measure your own.

Frequently asked questions

What does hallucination rate mean for an AI agent?

It is the share of graded units in which the agent said something its evidence does not support. The evidence is the retrieved passages, tool responses, prompt and caller’s words available at that turn. The rate is meaningless without its unit: per call, per claim or per action-bearing claim can differ roughly five-fold on the same calls.

What is the hallucination rate formula?

Hallucinations found divided by units graded, for example 11 calls with at least one hallucination out of 400 graded calls is 2.75%. Always report the 95% interval with it. For small rates use the Wilson interval given in the NIST/SEMATECH e-Handbook, section 7.2.4.1, not the plus-or-minus 1.96 standard-error shortcut, which can produce a negative lower limit.

What are the hallucination rates by model?

Vectara’s Hallucination Leaderboard, last updated 22 September 2026, lists 108 models with rates from 1.8% to 24.2%, and 9.6% for gpt-4o-2024-08-06. Those rates measure the factual consistency of summaries of more than 7,700 written articles. They are useful for choosing a model, but they are not a phone agent’s rate.

Is the Vectara hallucination leaderboard a good benchmark for a voice agent?

For ranking models, yes. Vectara itself argues it is a good indicator for RAG and agentic systems. For your deployment, no: it scores text summaries at temperature 0 where possible, with no speech recognition, no tool calls and no dialogue. Only graded calls from your own configuration measure your agent.

How many calls do I need to check to know my AI agent’s hallucination rate?

If you find none, 299 graded calls bound the rate below 1% at one-sided 95% confidence, and 59 bound it below 5%. If the agent’s true rate is half your 1% threshold, plan on about 2,185 calls for a stable 80% or better chance of proving it. Start a new sample after any prompt, model or knowledge-base change.

Does a speech recognition mistake count as a hallucination?

No. If the recogniser misheard the caller and the agent answered the misheard question faithfully, that is an ASR error with a different fix. Check disputed turns against the audio, not the transcript. Our page on whether an AI call transcript is a record of what was said covers how the two differ.

Can I use an LLM or HHEM to grade my calls automatically?

As a pre-screen, yes. Vectara’s open HHEM-2.1-Open scores a claim against its evidence from 0 to 1, and it is tagged English on Hugging Face. It does not know your policies and has not been validated on phone calls, so check it against a human-graded subset before trusting its counts.

Where every figure on this page comes from

Figure Who published it Link Date read
Leaderboard last updated 22 September 2026; 108 models; 1.8% (antgroup/finix_s1_32b) to 24.2% (mistralai/ministral-3-3b-2512); 9.6% for openai/gpt-4o-2024-08-06 Vectara github.com/vectara/hallucination-leaderboard 30 September 2026
Median of the 108 listed rates, 9.55% Computed by us from Vectara’s table github.com/vectara/hallucination-leaderboard 30 September 2026
More than 7,700 articles; temperature 0; HHEM-2.3 judge; “good indicator … RAG or agentic systems” Vectara (README, Dataset and Methodology sections) github.com/vectara/hallucination-leaderboard 30 September 2026
“You always need two pieces of text”; HHEM-2.1-Open scores 0 to 1; tagged English Vectara (model card) huggingface.co/vectara/hallucination_evaluation_model 30 September 2026
gpt-4o under 50% task success; pass^8 under 25% in retail; submitted 17 June 2024 Yao, Shinn, Razavi, Narasimhan (τ-bench, arXiv 2406.12045) arxiv.org/abs/2406.12045 30 September 2026
Wilson interval formula; negative-limit criticism; worked example 0.09577 NIST/SEMATECH e-Handbook of Statistical Methods, 7.2.4.1 itl.nist.gov/div898/handbook/prc/section2/prc241.htm 30 September 2026
Rule of three paper title and citation (JAMA 1983;249(13):1743-5) Hanley and Lippman-Hand, via PubMed pubmed.ncbi.nlm.nih.gov/6827763 30 September 2026
All intervals, sample sizes, coverages and the clustering simulation (Wilson 0.18% to 5.45% for 1 of 100; 59, 149, 299, 598, 2,995; 434, 1,091, 2,185; 1,522; 63.3%, 79.9%, 92.5%; 92.1% and 95.5%) Computed by us, Python standard library, run 30 September 2026 Method and code on this page 30 September 2026
400 calls, 2,600 claims, 520 action claims, 14, 11 and 4 errors; 4 minutes per call Hypothetical inputs for illustration, not measurements None Not applicable

Talk to us about your agent

Zian AI builds autonomous sales agents for phone, SMS, email and WhatsApp in 30+ languages, with HubSpot, Salesforce, HighLevel and Zapier integrations and private model deployment on customer infrastructure. We are onboarding partners in a partnership-application beta.

Apply For Partnership


Related Blogs

Related from Zian AI