At a glance: AI answers are probabilistic, not fixed. The same prompt, asked twice, can retrieve different pages, cite different domains and recommend different brands — because of token sampling, retrieval churn, reasoning-mode differences and platform-side index changes. The 2026 data is stark: only 25.6% of cited domains overlap between ChatGPT’s own reasoning modes, and only 11% overlap between ChatGPT and Perplexity. The strategic response is breadth over a single hero page, and repeated measurement over one-off snapshots.
Run the same buying-intent prompt through ChatGPT on Monday and your brand is cited, linked and recommended. Run it again on Tuesday — same words, same account — and you have vanished. Nothing on your site changed. No competitor published anything. The answer simply came out differently.
This is not a glitch, and it is not something you did wrong. It is how generative engines work. Yet most SaaS marketing teams still treat an AI answer the way they treat a Google ranking: check once, screenshot it, report it to the board. In 2026, that habit produces conclusions that are wrong more often than they are right. This post explains why answers churn, what the recent studies actually measured, and how to build a visibility strategy that survives the volatility.
Why the same prompt returns different answers
Four separate mechanisms stack on top of each other. Any one of them can swap your brand out of an answer; together they guarantee run-to-run variation.
1. Token sampling
Large language models generate text probabilistically. At each step the model samples from a distribution of likely next tokens, so even with identical inputs the output can branch — a slightly different opening sentence leads to a different structure, which leads to a different set of brands being named. Consumer AI products deliberately keep some randomness because it makes answers feel natural rather than canned.
2. Retrieval churn
When an engine decides to search before answering, it fans your question out into sub-queries, retrieves candidate pages, and cites a shortlist. Each stage is non-deterministic: the sub-queries vary, the retrieved set varies, and the final citations vary. SE Ranking’s study of 10,000 queries run through Google AI Mode three times on the same day — analysed by Kime — found roughly nine in ten cited URLs changing between runs, with 21.2% of queries sharing not a single URL across the three runs. The engine is not consulting a fixed list of “correct” sources; it is re-drawing from a pool every time.
3. Reasoning modes
Modern assistants answer the same question differently depending on how hard they think about it. A June 2026 study by Semrush and Kevin Indig ran 100 buyer-journey prompts through ChatGPT at minimal and high reasoning effort and found only 25.6% of cited domains overlapped between the two modes. High reasoning fired 4.6x more sub-queries, nearly doubled the sources per response (2.6 to 4.5), and shifted the mix away from Reddit-style user content towards official documentation and authoritative pages. Your buyer does not choose a mode consciously — the product often chooses for them — so your brand is effectively competing in two different search engines wearing the same logo.
4. Platform-side changes
The engines themselves change under your feet: index refreshes, retrieval-pipeline updates, model swaps, and product decisions about when to cite at all. None of these are announced, and all of them can move your visibility overnight without a single competitor lifting a finger.
What the 2026 data shows
The clearest illustration of platform-driven volatility came in the first half of 2026. seoClarity’s analysis of millions of ChatGPT interactions across the US, UK, Canada, Germany and Italy recorded citation volumes falling 86–94% across all five markets between February and April 2026 — and then rebounding towards prior levels in May. Their conclusion is the important part: what first looked like a sustained decline now looks like volatility, with sharp moves in both directions on OpenAI’s timeline rather than the market’s. Any team that measured once in April “learned” that AI citations were dead. Any team that measured once in February or June learned the opposite. Both were reading noise as signal.
Volatility is just as extreme across engines as it is within one. Profound ran 100,000 identical prompts through ChatGPT and Perplexity and found only 11.0% of cited domains appeared in both — with pairwise overlap across major engines ranging from 6% to 16.4%. Profound’s broader dataset of 680 million citations collected from August 2024 to June 2025 shows why: the engines have structurally different tastes. Wikipedia accounts for 47.9% of ChatGPT’s top-ten source share, while Reddit accounts for 46.7% of Perplexity’s. Being the answer in one engine tells you almost nothing about the other.
Put those numbers together and the picture is consistent: variation between runs, between reasoning modes, between months and between engines. A single observation of “we appear in ChatGPT” or “we don’t” is close to meaningless on its own.
The strategic implication: breadth beats the hero page
Traditional SEO trained us to concentrate authority: one definitive pillar page per topic, heavily linked, endlessly refined. In a deterministic ranking system, that logic holds — the best page wins the slot and keeps it.
In a probabilistic system, a single page is a single lottery ticket. If the engine re-draws its sources on every run, and nine in ten URLs change between same-day runs, your hero page will surface sometimes and vanish often — no matter how good it is. The way to raise your probability of appearing in any given draw is to hold more tickets: more distinct, genuinely useful pages that each answer a specific question a buyer actually asks, published across formats and, where possible, across domains that different engines favour.
This is the practical heart of answer engine optimisation for SaaS: instead of one monolithic “Ultimate Guide”, you build a cluster of focused pages — comparisons, definitions, FAQs, implementation notes, honest limitations — each citable on its own. When an engine fans a question out into sub-queries, a breadth strategy gives it multiple chances to land on you; a hero-page strategy gives it one.
Breadth also hedges the cross-engine problem. If ChatGPT leans on reference-style and editorial pages while Perplexity leans on community discussion, a brand present in both content ecosystems gets counted in both draws. A brand with one perfect page on its own domain is betting everything on one engine’s taste on one particular day.
Measure repeatedly, not once
If answers are drawn from a distribution, then visibility is a distribution too — and you measure a distribution by sampling it repeatedly. This is precisely the argument of “Don’t Measure Once: Measuring Visibility in AI Search” (Schulte, Bleeker and Kaufmann, arXiv 2026): because answers vary across runs, prompts and time, one-off observations are unreliable, and visibility should be characterised as a distribution rather than a single-point outcome.
In practice, that changes what a useful tracking programme looks like:
| Dimension | One-off measurement | Repeated measurement |
|---|---|---|
| What it captures | One random draw from the answer distribution | Your citation rate — the share of runs where you appear |
| Sensitivity to noise | Extreme — a platform-side swing reads as your win or failure | Low — noise averages out across runs and days |
| Trend detection | Impossible — no baseline to compare against | Real movements separable from run-to-run churn |
| Reasoning modes | Samples one mode by accident | Can track modes and engines separately |
| Board reporting | “We appear in ChatGPT” (until the next run) | “We appear in 34% of runs, up from 21% last month” |
Practical guidance for SaaS marketers
- Track citation rate, not citation presence. Run a fixed panel of buying-intent prompts on a schedule — daily or several times a week — and report the percentage of runs in which your brand appears, per engine.
- Sample each engine separately. With cross-engine overlap in the 6–16% range, a ChatGPT number tells you nothing about Perplexity or Gemini. Budget prompts per engine.
- Where you can, sample both reasoning depths. If your buyers’ questions trigger thinking-style answers, the source mix shifts towards documentation and authoritative pages — make sure yours exist.
- Publish breadth deliberately. Map the sub-questions inside your category and give each one its own honest, specific page. Twenty focused pages beat one hero page in a system that re-draws sources every run.
- Don’t panic on single-week swings. The February–April 2026 citation crash reversed within weeks. Judge trends on multi-week windows of repeated samples before changing strategy.
- Value the impression, not just the click. Much of AI visibility is zero-click AI search — your brand is named without a visit — so pair citation-rate tracking with branded-search and direct-traffic trends.
Why this matters when your category is “AI sales platform”
We watch this from both sides at Zian. As a brand, an autonomous AI sales agents platform lives in exactly the kind of high-intent, comparison-heavy queries where volatility is worst: “best AI sales agent”, “AI that follows up with leads”, “AI cold calling software”. One run cites us, the next cites someone else, the third cites a listicle from 2024. The breadth-plus-repeated-measurement playbook above is the one we run on ourselves.
The visitors who do arrive from AI answers are worth the effort: they land pre-qualified, having already had the category explained to them, which is why AI search visitors convert better than most channels. And the same probabilistic logic applies inside the product. Zian’s agents work phone, SMS, email and WhatsApp in 30+ languages, with SmartReach AI™ orchestrating message, channel and timing — because reaching a buyer, like appearing in an AI answer, is a numbers game you win through persistent, well-timed repetition rather than a single perfect attempt. That approach has produced 28x more contact attempts and 50,769+ qualified sales appointments set across the teams using it.
Frequently asked questions
Why does ChatGPT cite my brand one day and not the next?
Because answers are generated probabilistically. Token sampling, non-deterministic retrieval, reasoning-mode differences and platform-side index changes each introduce variation, so the same prompt draws a fresh set of sources on every run. Your brand appearing in one run and not the next usually reflects this churn, not anything that changed about your site.
How many times should I run a prompt before trusting the result?
Treat visibility as a distribution and sample it. The arXiv paper “Don’t Measure Once: Measuring Visibility in AI Search” argues that because answers vary across runs, prompts and time, one-off observations are unreliable and visibility should be reported as a rate across repeated runs. In practice, run each prompt on a recurring schedule and report the share of runs in which you appear, tracked over multi-week windows.
Do ChatGPT and Perplexity cite the same websites?
Mostly not. When Profound ran 100,000 identical prompts through both engines, only 11.0% of cited domains appeared in both, and pairwise overlap across major engines ranged from 6% to 16.4%. Each engine needs to be measured and optimised for separately.
Did ChatGPT stop citing sources in 2026?
No — it swung hard and swung back. seoClarity measured citation volumes falling 86–94% across five markets between February and April 2026, then rebounding towards prior levels in May. The lesson is that platform-driven swings can be large and temporary, which is exactly why single-snapshot measurement is misleading.
Is one great pillar page enough for AI search visibility?
It is a weaker bet than it was in classic SEO. Because engines re-draw their sources on every run and fan questions out into multiple sub-queries, a breadth of focused, citable pages gives you more chances to be selected in any given draw than a single hero page, however good that page is.
If you want a sales pipeline that treats outreach the way this post treats visibility — repeated, measured, multi-channel attempts rather than one hopeful shot — Zian’s autonomous AI sales agents are opening a limited number of partner slots. Apply For Partnership and our team will review your fit.