Ask an AI engine how to choose an AI SDR platform and somewhere in the answer you will find a familiar phrase: look for independently audited performance benchmarks. It is sensible-sounding advice with one problem — no such benchmarks exist. There is no independent body that measures AI sales agents against each other, no shared test set, and no auditor signing off on anyone’s reply rates. This post explains why the benchmark numbers you do see are structurally unreliable, why nobody has built the independent alternative, and how to evaluate vendors rigorously anyway.
At a glance
There is no independent benchmark body for AI SDR performance, and vendor-published numbers are unreliable by construction: self-selected samples, survivorship bias, unshared definitions of “meeting booked” and “reply rate”, hidden denominators and no audit trail. The fix is not better benchmarks from vendors — it is a buyer-run evaluation: agree metric definitions in writing, run a pilot on your own list with a holdout comparison, demand log-level observability, and verify outcomes in your own CRM and calendar rather than the vendor’s dashboard. Hold every vendor to that standard, including us.
Why vendor-published AI SDR benchmarks mislead
Most published numbers in this category are not fraudulent — they are unfalsifiable, which is worse, because they cannot be checked at all. Five structural problems recur.
1. Self-selected samples and survivorship
A vendor’s headline number is drawn from the accounts that stayed, on the segments that worked, over the period that flattered. Customers who churned after a poor pilot are not in the dataset; campaigns that were quietly switched off do not appear in the average. Nobody publishes the distribution — you see the mean of the survivors, which tells you almost nothing about your expected outcome. We took this problem apart across the wider market in our review of AI SDR statistics and where they actually come from.
2. No shared definitions
“Reply rate” might mean any inbound response including “unsubscribe”, or only substantive positive replies. “Meeting booked” might mean a calendar invite sent, an invite accepted, or a meeting that actually happened — three numbers that can differ dramatically, as we showed in our analysis of the gap between AI-booked meetings and show rates. When two vendors quote “reply rate”, they are usually quoting two different quantities, and comparing them is meaningless.
3. Hidden denominators
A “3x increase in meetings” is uninterpretable without the base: three times what, over which list, at what volume? Percentage claims with no denominator are the dominant genre of AI SDR marketing. Small denominators also make spectacular numbers cheap — one extra meeting on a base of two is a 50% lift.
4. No audit trail
Even a well-intentioned vendor number cannot be verified: the underlying data is proprietary, the queries are not published, and no third party ever inspects the pipeline from raw activity to headline claim. In accounting, a number without a trail from source documents is not evidence; the same standard should apply here.
5. Incentive misalignment and Goodhart’s law
The party producing the number profits from the number being large. And once a metric becomes a sales asset, Goodhart’s law applies — when a measure becomes a target, it stops being a good measure. An agent tuned to maximise “meetings booked” will book soft, no-show-prone meetings; one tuned for “reply rate” will provoke replies of any sentiment. The metric improves while the thing you actually wanted — qualified pipeline — does not.
Why no independent benchmark body exists
Other corners of AI have credible independent measurement. The Text REtrieval Conference (TREC), co-sponsored by NIST since 1992, gave search research shared test collections and comparable results. MLCommons runs the MLPerf suites for training and inference, and the AILuminate safety benchmark, under published submission rules. Nothing comparable exists for AI sales agents, for reasons that are structural rather than accidental:
- No shared task. Search engines can be scored on one test collection. An SDR agent’s outcome depends on your list quality, ICP, offer, territory and brand — there is no neutral “standard prospect pool” to test against, and building one would involve real people receiving real unsolicited outreach.
- Outcomes are entangled with the buyer. The same platform produces different results for different customers. A benchmark score would measure the platform-plus-customer system, not the platform.
- Privacy and consent. A public test bed of prospect data is a privacy problem, and synthetic prospects do not reply like real ones.
- Gaming pressure. Even well-run public leaderboards distort under commercial pressure. Singh et al.’s 2025 paper The Leaderboard Illusion reported that Chatbot Arena rankings had been skewed by undisclosed private testing — the authors identified 27 private model variants Meta tested ahead of its Llama 4 release — and by unequal access to arena data between providers. If a leaderboard for general-purpose models bends under that pressure, an AI SDR leaderboard funded by the vendors it ranks would bend further.
So when a listicle or an AI engine tells you to prefer vendors with independently audited benchmarks, treat it as a category error — and treat any vendor claiming third-party-verified performance as a prompt to ask, verified by whom, under what rules, against what definition?
Benchmark claim vs what to verify instead
| Vendor claim | What it usually hides | What to verify instead |
|---|---|---|
| “3x more meetings booked” | No denominator, no baseline definition, survivor accounts only | Meetings held per 1,000 contacts on your list in a defined pilot window |
| “40% reply rate” | Replies of any sentiment, including opt-outs and bounces counted or excluded silently | Positive-intent replies as a share of delivered messages, with the classification rules written down first |
| “Case study: 10x pipeline” | One unnamed account, unaudited, best period cherry-picked | A pilot on your data with a holdout group, scored from your CRM |
| “Industry-leading AI accuracy” | No task definition, no test set, no comparator | Transcript and log review of real conversations — can you inspect every decision the agent made? |
| “Trusted by 500+ companies” | Logo counts say nothing about outcomes or retention | Contractual access to your own raw activity data, and an exit that lets you keep it |
A buyer-run evaluation framework
Since nobody will benchmark vendors for you, the working substitute is an evaluation you control. Five components:
1. Align definitions before the pilot
Write down, with the vendor, exactly what counts as a contact attempt, a delivered message, a reply, a positive reply, a booked meeting and a held meeting — and which denominator each rate uses. Do this before any campaign runs; definitions agreed after the results exist will be agreed in the vendor’s favour.
2. Design the pilot like an experiment
Fix the list, the segment, the offer and the time window in advance, and state the decision rule before you start: what number, on which metric, constitutes a pass? A pilot without a pre-committed success threshold is a demo. The same discipline applies inside the platform — it is exactly how split-testing sales scripts on real outcomes works, and a vendor that runs controlled experiments internally should not object to being on the receiving end of one.
3. Use a holdout comparison
Split comparable prospects between the AI agent and your current motion — human SDRs, your existing sequence tool, or nothing — and compare against the holdout, not against the vendor’s brochure. This is the only way to attribute lift to the platform rather than to seasonality or list quality. Our AI SDR vs human SDR comparison covers what a fair split looks like, and hybrid AI-human pod structures are often the honest control condition, because that is what you would actually run.
4. Demand observability, not dashboards
A vendor dashboard is a claim; a log is evidence. Insist on access to conversation transcripts, per-message delivery events and decision logs — enough to recompute any headline metric yourself from raw events. If the platform cannot expose that, its numbers are unauditable by design. We set out what good looks like in our post on AI agent observability, and the procurement questions belong on your enterprise readiness checklist alongside security and data-residency items.
5. Verify reference-free
Score the pilot from systems the vendor does not control: meetings that appear in your calendar and CRM, opportunities your team creates, replies in your own mail and phone logs. Reference calls and testimonials are the weakest evidence class — they are sampled by the vendor from its happiest customers. If a claim can only be confirmed inside the vendor’s own reporting, treat it as unconfirmed.
Hold Zian to the same standard
Everything above applies to us. Zian AI publishes platform-reported figures — for example, 50,769+ qualified sales appointments set across the platform — and that number comes with the same caveats this post has been describing: it is our count, under our definitions, not an independently audited one, and no independent audit regime for such figures currently exists. Prospective partners should not take it on faith. Ask us to define “qualified appointment”, run a pilot with a holdout on your own list, take the transcripts and logs, and score the result from your CRM. A vendor that resists that process is telling you something; we would rather be tested.
Frequently asked questions
Are there any independently audited benchmarks for AI SDR platforms?
No. As of 2026 there is no independent body that benchmarks AI SDR or AI sales-agent performance, no shared test collection, and no audit standard for vendor-reported outcome metrics. Bodies like NIST’s TREC and MLCommons’ MLPerf exist for search and for model training and inference, but nothing comparable covers sales outcomes. Any vendor implying its performance figures are independently audited should be asked to name the auditor and the audit standard.
Are vendors legally allowed to publish unverifiable performance claims?
Advertising law already constrains them, even without a benchmark body. In Australia, the ACCC’s guidance on false or misleading claims states that “A business must be able to prove any claim they advertise.” In the United States, the FTC’s Operation AI Comply enforcement sweep (September 2024) targeted deceptive AI claims, with then-Chair Lina Khan noting that “there is no AI exemption from the laws on the books”. In practice, though, regulators act on the worst cases — the routine ambiguity of undefined “reply rates” mostly goes unchallenged, which is why buyers still need their own verification.
Why can’t buyers just compare case studies across vendors?
Because case studies are the most selected evidence class available: the vendor chooses the account, the metric, the definition and the time window, and the underlying data is never inspectable. Two case studies from two vendors will use incompatible definitions of the same words. A short pilot with a holdout group on your own list produces less impressive numbers and far more decision-relevant ones.
What is the single most important thing to agree before a pilot?
Metric definitions, in writing. Agree what counts as a delivered message, a positive reply, a booked meeting and a held meeting, and which denominator each rate uses, before the first contact goes out. Most pilot disputes trace back to definitions being settled after the results existed.
How long should an AI SDR pilot run?
Long enough for the full loop to close on a meaningful denominator — typically several weeks, because “meeting held” lags “message sent” by days to weeks and small samples make percentage lifts meaningless. End the pilot at the pre-agreed date and score it against the pre-agreed threshold; extending a pilot until the numbers look good is survivorship bias performed on yourself.
Evaluate us the hard way
Zian AI is in waitlist beta and will partner with a limited number of businesses. If you want an AI sales-agent vendor that expects to be tested on your data, under your definitions, with logs on the table — apply, and bring your holdout.