In 2026, “AI visibility benchmark” became its own content genre. Conductor, Semrush, Profound, Ahrefs and a long tail of agencies have published industry-average citation rates, and marketing teams are quietly turning those averages into targets. Someone reads that AI referral traffic sits near one per cent of sessions, and by the next board deck it is a number their team is accountable for.
That is the same mistake we wrote about in why there are no trustworthy AI SDR benchmarks — only worse, because the thing being measured is non-deterministic. This post applies that evidence standard to AEO and GEO benchmarks themselves. Some are genuinely well documented. Almost none disclose enough for you to compare your number to theirs.
Quick answer: A 2026 AEO benchmark is only comparable to your own data if it publishes its prompt set, samples per prompt, exact date window, named engines and modes, locale, and its definition of a “citation”. Conductor, Semrush, Profound and Ahrefs each disclose some of this; none disclose all of it. Treat published averages as direction, never as targets.
Zian AI runs autonomous sales agents, and we track our own answer-engine visibility the way we would tell any operator to: a fixed prompt set, a fixed schedule, the same engines every time. We publish no benchmark of our own, because one vendor’s prompt list does not deserve to be called an industry average. If you want agents that book meetings rather than dashboards that report on them, Apply For Partnership.
What the 2026 benchmark wave actually published
Four reports are worth naming, because all four are open to read and state at least part of their method. We fetched each at the publisher’s own URL.
Conductor, The 2026 AEO / GEO Benchmarks Report. Conductor analysed 13,770 domains across 10 industries mapped to GICS. Its AI citation index covers 3.5 million unique prompts between May and September 2025, producing 17 million AI-generated responses and over 100 million citations. Its traffic side is separate: 1,215 enterprise customer domains and more than 3.3 billion sessions, of which 35.7 million came from AI sources. A further AI Overviews analysis covers 15 September to 12 October 2025. The headline: “AI referral traffic accounts for 1.08% of all website traffic for these 10 key industries.” Crucially, Conductor states that its traffic and AI Overviews benchmarks come from US data, and it does not publish the prompt list.
Semrush, 2026 AI Visibility Index. 126 million US AI search prompts from January through April 2026 across ChatGPT, Gemini, Google AI Mode and Google AI Overviews, benchmarked across 22 industries. Semrush is one of the few publishers to separate the two metrics that everyone else blends: it reports that “ChatGPT cites an average of 15 sources per response… while Gemini cites an average of 3 sources per response”, and that “On Gemini, the overlap between mentioned brands and cited domains can be as low as 30%”. It also found only 36 global brands held top-100 visibility across all four platforms in every month of the study.
Profound, the Profound Index. Profound is unusual in how it builds its prompt set: it retrieves real user prompts via semantic search across answer engines over a rolling six-month window, then filters, deduplicates, clusters by embedding and selects the top prompt per cluster. That is a defensible construction rule, published openly, and closer to real demand than a keyword-tool export. It updates weekly across 50-plus industries and does not publish the resulting prompt set; it names no target locale, though it does say prompts too narrow to a single region are dropped and that candidates are translated during filtering.
Ahrefs, AI Overview citations. Louise Linehan’s March 2026 analysis of 863,000 keyword SERPs and 4 million AI Overview URLs found that 37.9% of URLs cited in AI Overviews also appeared within the first ten blocks of the same result page, with 31.2% ranking 11-100 and 31.0% not ranking in the top 100 at all. This is the tightest of the four, because the unit of analysis is a SERP rather than a chat session, and SERPs are far more reproducible than chat answers.
The scorecard: apply this to any benchmark report
| Criterion | What a trustworthy report discloses | Red flag | Why it matters |
|---|---|---|---|
| Prompt set | The full list, or a reproducible construction rule, plus the head/long-tail split | “Thousands of prompts” with no list and no rule | Prompt mix moves the citation rate more than your content does. A secret list cannot be reproduced or matched. |
| Samples per prompt | Polls per prompt, and reported variance or confidence intervals | One poll per prompt; a single point estimate with no spread | A single poll of a non-deterministic system is one dice roll reported as a measurement. |
| Window and cadence | Exact start and end dates plus polling frequency | “Recent data”, a season, or a year with no dates | Engines ship model and retrieval changes mid-window. A 2025 window may describe a model that no longer serves traffic. |
| Engine and mode | Each engine named, plus mode: browse vs no-browse, thinking vs standard, AI Overviews vs AI Mode | “Cited by AI” with no engine named | Semrush measured ChatGPT at ~15 sources per response and Gemini at ~3. Averaging them produces a number describing nothing. |
| Geography and language | Locale and query language stated up front | Silence, or a global claim built on one market | Conductor states its traffic and AI Overviews data are US. An Australian query and a US query are different measurements, not the same one with noise. |
| Definition of a citation | Linked source, unlinked brand mention and grounding-payload domain reported separately | One blended “visibility” percentage | Semrush found brand-mention and cited-domain overlap as low as 30% on Gemini. Those are not the same metric. |
| Denominator | Polls attempted, polls that failed or returned unparseable output, and how failures were treated | Only successful polls reported | Counting failed polls as misses drags the rate down; silently dropping them pushes it up. Same raw data, two different headlines. |
| Funding and data source | Publisher’s commercial interest, and whether the data comes from its own customer base | “Industry average” derived entirely from a vendor’s paying customers | Conductor’s traffic benchmark is drawn from 1,215 enterprise customer domains — a real dataset, but an enterprise-skewed one. |
The volatility problem the scorecard exists to catch
Rows two and three are the ones most reports fail. AI answers vary between sessions for reasons unrelated to your content, which we covered in why your brand appears in one ChatGPT answer and vanishes from the next. Mode matters as much: deliberative answers pull in more and different sources than fast ones, the finding behind ChatGPT’s thinking mode citing more sources. A benchmark that ran one poll per prompt, in one mode, on one day, sampled a distribution once and published the mean.
On who paid for it
Conductor, Semrush, Profound and Ahrefs all sell software that improves the number they are benchmarking. Stating that is not an accusation — vendors have the data and academics mostly do not. The right response is to ask whether the sample is the vendor’s own customer base, and whether the metric definition happens to match what the product measures. Zian sells software in an adjacent category and carries the same incentive, which is why we publish no citation-rate benchmark.
A worked trace: the “51% of B2B buyers” stat
Here is the discipline in practice. One of the most-repeated 2026 stats is that 51% of B2B software buyers now begin vendor research in an AI chatbot, up from 29%. We found it in an SEO statistics roundup, which is where most people find it, then traced it.
The trace succeeds. The owner is G2, in a report titled The Answer Economy: How AI Search Is Rewiring B2B Software Buying, announced 15 April 2026. G2’s own wording is: “Half (51%) of B2B software buyers now begin their software research with an AI chatbot more often than with Google, up from 29% in April 2025.” Method: an online survey of 1,076 B2B decision makers, fielded March 2026, global pool spanning North America, EMEA and APAC. No margin of error is published.
Three things the roundups lose in transmission, and they are the interesting part:
- It is self-reported, not observed. Buyers were asked where they start. Nobody watched their browsers.
- The baseline is a different survey wave. The 29% comes from G2’s 2025 Buyer Behavior Report, a separate fielding. “Up from 29%” is a comparison across waves, not a tracked panel.
- The paraphrase drifts, and it drifts at the source. The roundup renders it as buyers researching “in an AI chatbot rather than Google” — but so does G2’s own release, whose lead says buyers “begin their purchasing process in an AI chatbot rather than a traditional search engine”. The finding as reported in the same release is “more often than with Google” — a frequency comparison, not an exclusive one. G2’s own data undercuts the stronger reading: 61% report using AI search and Google in tandem.
The number survives the trace and can be cited. The paraphrase does not. That distinction is the whole job.
Build your own instead
An internal measurement you control beats an external average you cannot reproduce. The minimum viable version is small:
- Write down a fixed prompt set. Twenty to forty prompts a real buyer would type, in a file under version control. Do not edit it to chase good news — version a new set instead.
- Fix the schedule. Same weekday, same time, every week. Cadence consistency matters more than frequency.
- Fix the engines and the modes. Same engines, same mode, every run. If you add an engine, that is a new series, not a continuation of the old one.
- Record the answer text, not a yes/no. Store the full response and the source list. A boolean tells you nothing about why you dropped out, and you cannot re-score history you never kept.
- Log failures separately. Timeouts and unparseable responses go in their own column, never into the numerator or silently out of the denominator.
- Track the trend, not the level. Your absolute rate is an artefact of your prompt set. Its direction over ten weeks is a signal.
This is how we run it at Zian: a fixed prompt list, polled on a schedule across the same engines, with raw answers retained so we can re-score them when our definition of a citation changes. We do not publish the rate, because outside our prompt set it would mean nothing to you. Pair the poll with a share-of-answer view — see share of model — and with the traffic side, which is harder than it looks because AI referral traffic is nearly invisible in GA4. Two independent measurements that disagree beat one that flatters you.
If you would rather spend that discipline on pipeline than on dashboards, our agents are designed to run the outreach while you watch the trend line. Apply For Partnership.
Sources and who owns each figure
| Figure / claim | Owner (organisation) | Where it’s published | Date checked |
|---|---|---|---|
| 13,770 domains; 3.5M prompts, 17M responses, 100M+ citations (May-Sep 2025); 1,215 enterprise domains, 3.3B sessions; AIO window 15 Sep-12 Oct 2025; “AI referral traffic accounts for 1.08% of all website traffic for these 10 key industries”; traffic and AIO benchmarks stated as US data | Conductor | The 2026 AEO / GEO Benchmarks Report | 18 Aug 2026 |
| 126M US AI search prompts, Jan-Apr 2026, 4 platforms, 22 industries; ChatGPT ~15 sources/response vs Gemini ~3; Gemini mention-to-citation overlap “as low as 30%”; only 36 brands held top-100 visibility on all four platforms every month | Semrush | 2026 AI Visibility Index release, 26 Jun 2026 | 18 Aug 2026 |
| Prompt set built from 1.5B+ real user prompts via semantic search over six months, then filtered, deduplicated, clustered and reduced to the top prompt per cluster; updated weekly; 50+ industries | Profound | Introducing the Profound Index, 16 Jun 2026 | 18 Aug 2026 |
| 863K keyword SERPs, 4M AI Overview URLs; 37.9% of cited URLs also appeared in the first 10 blocks; 31.2% at positions 11-100; 31.0% outside the top 100 | Ahrefs (Louise Linehan) | Ahrefs blog, 2 Mar 2026 | 18 Aug 2026 |
| 51% begin research with an AI chatbot more often than Google, up from 29% in April 2025; 61% use AI search and Google in tandem; survey of 1,076 B2B decision makers, March 2026, NA/EMEA/APAC | G2 | G2 newsroom and the 15 Apr 2026 release | 18 Aug 2026 |
| 1,000 completions sampled at temperature 0 from Qwen3-235B-A22B-Instruct-2507 produced 80 unique completions, first diverging at token 103; cause is varying batch size under load | Thinking Machines Lab (Horace He et al.) | Defeating Nondeterminism in LLM Inference, 10 Sep 2025 | 18 Aug 2026 |
Frequently asked questions
Are the 2026 AEO benchmark reports useless?
No. They are useful for direction and for structural findings, and useless as targets. Ahrefs’ analysis of 863,000 keyword SERPs finding that 37.9% of AI Overview citations also appear within the first ten blocks of the same result page is a durable structural insight. “Your industry’s citation rate is X%” is not, because X depends entirely on a prompt list you cannot see.
Doesn’t setting temperature to zero make AI answers reproducible?
No, and this is the technical root of the sampling problem. Thinking Machines Lab sampled 1,000 completions at temperature 0 from Qwen3-235B and got 80 unique completions, first diverging at the 103rd token. The cause is that inference kernels are not batch-invariant, so results shift with server load — a property of the system, not of your query. Any benchmark with one sample per prompt inherits this directly.
How many samples per prompt do I actually need?
Enough to see the spread, which you discover by measuring it rather than by picking a number from a blog. Poll one prompt repeatedly for a week and look at how often the answer set changes. If it is stable, few samples are needed; if it churns, your weekly single-poll number is noise and you should report a range instead of a point.
Can I compare my citation rate to a published industry average?
Only if you rebuild their prompt set, their engines, their modes, their locale and their citation definition — at which point you have reproduced their study rather than compared to it. Since none of the four reports above publish the prompt or keyword list behind their numbers, the honest answer for 2026 is no. Compare your number to your own number from last month.
Why does geography change the result so much?
Because engines ground answers in locale-specific sources and often in locale-specific indexes. Conductor states plainly that its benchmarks come from US data over a five-month period. For an Australian buyer asking about an Australian vendor, that is a different population of candidate sources, so a US-derived average is not a weaker version of your number — it is a different number.
Where should our AI visibility numbers actually live?
In two places that can disagree: a prompt-poll log that measures whether you appear in answers, and a traffic view that measures whether anyone arrives. Start with the fundamentals in our practical guide to answer engine optimisation for SaaS, then reconcile the two series monthly. Divergence between them is usually the most useful thing you will learn all quarter.