AEO Measurement Harness: What to Poll and What to Ignore - Zian AI

AEO Measurement Harness: What to Poll and What to Ignore

An AEO measurement harness is a fixed prompt set polled on a schedule, with every raw answer stored. Google says it rolled its Search Console generative AI performance report out to all websites worldwide as of 31 August 2026, though the same help page still warns not every property has access. It groups impressions by page, country, date and device — no query dimension, and nothing from ChatGPT. Polling closes that gap.

This is the build-and-operate guide: choosing prompts, what to write to disk, which numbers deserve a dashboard, and which movements to ignore. It does not re-explain why the same prompt returns different answers twice in a row — sampling, retrieval churn, index freshness and engine-side reranking are covered in our companion post on AI answer volatility. This one assumes you have accepted that and now have to instrument it.

What Google and Microsoft already report for you

Build nothing until you have checked what the engines hand over free. The gaps in what they publish are the shape of the harness you need.

Google’s generative AI performance report in Search Console carries the note “As of August 31, 2026, we’ve rolled out these insights to all websites worldwide.” It covers AI Overviews and AI Mode and reports impressions grouped by pages, countries, dates and devices — but no query dimension, so you cannot see which prompt produced the impression. Treat the rollout as in progress rather than finished: the same page’s troubleshooting note still says “Not all properties have access to the report, as we’re rolling out over time,” checked 2 September 2026.

Microsoft went the other way. Its 16 June 2026 announcement added Intents, Topics, Citation Share and Compare to the AI Performance report in Bing Webmaster Tools, in preview globally, and it does surface grounding queries — with a caveat worth stealing for your own dashboard: “Importantly, Citation Share is designed as an observational metric – not a ranking system or a competitive scoreboard.”

First-party AI visibility reporting, checked 2 September 2026. “Not published” means we could not find a first-party report page from that vendor on that date, not that one cannot exist.
Engine or surface First-party report What it reports Prompt or query text?
Google Search (AI Overviews, AI Mode) Generative AI performance report, Search Console Impressions by page, country, device, date Not published — no query dimension
Microsoft Copilot, Bing and partner AI experiences AI Performance report, Bing Webmaster Tools (preview) Citation counts, Intents, Topics, Citation Share, Compare Yes — grounding queries
ChatGPT (OpenAI) Not published Not published Not published
Gemini app (Google) Not published Not published Not published

OpenAI’s own crawler documentation publishes inclusion controls — OAI-SearchBot, GPTBot and their IP ranges — but that is a switch, not a report. If ChatGPT matters to your category, nobody is measuring it for you.

Designing the prompt set

A panel is a fixed list of the questions a buyer types before they know your name. Write them from your own sales calls and support tickets, not a keyword tool. Between a dozen and forty works for one product line: fewer and you are measuring the quirks of a few prompts, more and the poll cost starts dictating your cadence.

Include exactly one brand-name prompt, and never count it in your category rate. A prompt like “[your brand] review” tells you one useful thing on day one — that the engine can retrieve your site at all — then repeats it forever. It is a smoke test, not a metric, and blending it into a headline percentage is the easiest way to publish a visibility number that rises while your category position has not moved. Prioritise category and comparison prompts instead: “best X software”, “X versus Y”, “X for [industry]”.

Give every prompt a stable integer ID at creation, never reuse an ID, and never edit a prompt’s text in place — retire the old ID and issue a new one. That discipline is what makes your history readable two quarters from now.

What to record on every poll

The most common harness failure is storing the score instead of the evidence. A row reading 2026-08-20, 1/14, 7% answers no question you will actually have in November. Store the raw material:

  • UTC timestamp, engine identifier and mode (browsing, grounding, reasoning) as separate fields
  • Prompt ID, panel version and the exact prompt text as sent
  • The full answer text, verbatim and never truncated
  • Every cited URL, in the order the engine returned them
  • Two separate booleans: was your domain in the citation list, and was your brand named in the prose — different events with different value, as we set out in cited versus recommended
  • Which competitors were named, so your denominator is not just yourself
  • Run metadata: request country, whether the engine returned zero citations, and whether the poll errored

That last field matters more than it sounds: a poll that failed and a poll that succeeded with no citation look identical in a percentage, and are opposite findings.

The metrics worth tracking, and how each misleads you

Metrics for an AEO poll harness: what each one tells you and where each one lies.
Metric What it tells you How it misleads you
Hit rate over N polls Roughly how present you are across the whole panel One brand-name prompt can carry the whole number, and the rate moves whenever the prompt count moves
Consecutive-hold rate Whether a citation survived into the next poll or was a one-off With one run per poll you cannot separate a genuine drop from sampling variance
Share of prompts ever cited Your ceiling: how much of the category you have reached at least once It only ratchets upwards, so one lucky hit flatters it forever and it never falls to warn you
Per-prompt hit history Which prompts move, when and on which engine — the only view that survives a panel change It is not one number, so it resists dashboards and invites you to report your best prompt
Cited-versus-named split Whether you were a source URL or a brand recommended in the answer text Most tools merge the two into one score

Build the review around per-prompt history. The others are summaries of it, and each can be true and useless in the same week.

A small-n look at our own panel

Label this correctly: a log from one vendor’s panel, not a study and not a benchmark. One poll per engine per cycle, no repeated runs, no controls. An anecdote with timestamps.

We poll a fixed set of 16 prompts against ChatGPT with browsing and Gemini with grounding, on a schedule. From 10 July to 2 September 2026 the brand-name prompt was cited in all 33 ChatGPT polls and all 29 Gemini polls — the smoke test passing, and worth nothing as a category signal. The category prompts are the honest part, and they behave nothing like one another. On Gemini, an Australian voice-agent prompt was cited in 8 of 29 polls, including a run of four consecutive polls in mid-August, after which it dropped out and later returned; a second category prompt held across two consecutive polls; a third appeared once in late July and not since. On ChatGPT only two non-brand prompts were cited at all in that window, each exactly once — the second of them on the day this post was published, which by the rule in the next section is not yet news.

That contrast is the argument for per-prompt history. A four-poll run that ends and a lone appearance are different events, and a single headline rate renders both as the same faint noise. We are not claiming anything we did caused any of it, and we are not going to dress up a weak category position. Our own harness also breaks a rule we are about to give you: one run per engine per cycle sits well below the sample sizes the St Gallen paper recommends, which is why we read the hit list rather than the rate.

What to ignore, and the denominator trap

Single-poll movement. One prompt gaining or losing a citation between two polls is expected behaviour, not news. Keep it out of the weekly report, and do not let anyone attribute it to last week’s content change.

Headline percentage changes. If the number moved but no per-prompt row changed state, nothing happened. Reconcile every rate back to the hit list first.

Any comparison across a changed denominator. You track 14 prompts and one is cited: 1/14 is 7%. You add two comparison prompts, so the same single citation is now 1/16, or 6%. The dashboard shows a one-point decline and somebody asks what broke. Nothing broke — same prompt, same engine, same day. Only the denominator moved.

To keep history comparable across that boundary: stamp a panel version on every poll row; report rates over the intersection of prompts present in both periods, listing new prompts separately as “no prior history”; never restate historical rates against the new denominator, because recomputing 2026 numbers against a 2027 panel invents a trend; and prefer the count to the rate in anything a non-specialist reads. “One category prompt was cited, the same one as last month” is unambiguous. “6%” is not. The same failure mode makes published benchmarks hard to use, which we work through in how to read a 2026 AEO benchmark report.

Cadence, cost and the limits of a small panel

There is a citable answer to “how many runs”. Julius Schulte, Malte Bleeker and Philipp Kaufmann of the University of St Gallen published Don’t Measure Once: Measuring Visibility in AI Search (GEO) on arXiv on 8 April 2026, design fully disclosed: eight prompts in each of four Swiss-German verticals, put daily to ChatGPT, Gemini, Google AI Mode and Perplexity from 24 January to 20 March 2026, plus a second collection from 21 to 25 March 2026 repeating the same prompts up to ten times per engine inside a 24-hour window.

In that repeated-run collection, “Source overlap averages 32–43%; brand overlap ranges from 33% (Sporting Goods) to 48% (Consumer Electronics)” — identical prompts, same engine, same day. The recommendation is explicit: “Practitioners should therefore use at least 7 runs per prompt per day for brand visibility monitoring, and at least 8 runs when source-level coverage matters”, aggregated over a rolling two to four weeks. Take the limitations with it: Swiss servers and locale, Swiss-German consumer verticals rather than B2B SaaS, and brand detection by substring matching. The run counts travel better than the overlap percentages do.

Then the cost: sixteen prompts, two engines and seven runs each is 224 calls a day before anyone analyses anything. Halving the panel or the cadence makes the arithmetic friendlier and the resolution worse. With sixteen prompts polled once, the smallest change the harness can physically observe is one prompt flipping state, about six percentage points; any real effect smaller than that is invisible by construction. Small-n polling cannot detect small effects, and no amount of charting fixes that.

A defensible starting configuration: 20 to 30 category prompts, two or three engines, three runs per prompt per poll, weekly, with no conclusions drawn for eight weeks. That is below the St Gallen floor, and you should say so in your own reporting.

You probably should not build this

If you have not published enough on your category to be retrievable at all, a harness will measure your absence with great precision and change nothing; fix the corpus first. If nobody has committed to reading the log monthly, an unread panel is a running cost that buys a false sense of instrumentation.

If two or three prompts genuinely cover your category, open ChatGPT and Gemini on the first Monday of each month, paste the prompts, save the answer text and citation list to a spreadsheet, and note the date. That is a harness. It costs nothing and will not break when an API changes.

If the real question is whether AI answers send you revenue rather than whether you appear in them, visibility is the wrong instrument — start with measuring AI referral traffic in GA4. Build the harness when you have a stable category prompt set, an intention to change something and the patience to wait two months. To talk through how we run ours, Apply For Partnership.

Frequently asked questions

How many prompts should an AEO panel contain?

More than a handful. The St Gallen team found prompt-level overlap ranging from below 0.2 to above 0.8 inside a single vertical, concluding that “Monitoring based on one or two prompts will reflect the idiosyncrasies of those prompts rather than campaign-level visibility” (arXiv:2604.07585, 8 April 2026). Twenty to thirty category prompts per product line is a reasonable start.

Should I include my own brand name in the prompt set?

One, as a control, and never counted in your category rate. A brand-name prompt confirms the engine can retrieve your site and then tells you nothing further. Mixing it into a headline percentage inflates the rate while measuring nothing about your position in the category.

How often should I poll?

More often per prompt than most teams expect: the St Gallen recommendation is at least seven runs per prompt per day for brand visibility and eight when source coverage matters, aggregated over a rolling two to four weeks. That study ran on Swiss servers against Swiss-German consumer verticals, so treat the run counts as guidance and the overlap percentages as not directly transferable.

Is the “only 30% of brands stay visible” figure a benchmark I can hold my panel to?

No. That line — 30% from one answer to the next, 20% across five consecutive runs — traces to AirOps’ report The 2026 State of AI Search, published 2 December 2025. The report page states the figures but does not publish the prompt set, the number of runs or the collection dates behind them, checked 2 September 2026. Without a disclosed method there is nothing to hold your panel against, so treat it as a directional warning about single polls, not a target.

Do Google or Microsoft report any of this for me?

Partly. Search Console’s generative AI performance report covers AI Overviews and AI Mode and reached all sites worldwide on 31 August 2026, but reports impressions only, with no query dimension. Bing Webmaster Tools’ AI Performance report covers Microsoft Copilot, Bing and select partner AI experiences and does show grounding queries, in preview since 16 June 2026. Neither reports ChatGPT or the Gemini app — the gap a harness exists to fill. For the difference between being cited and being recommended, see cited versus recommended and share of model.

Related Blogs

Related from Zian AI