If your prospects prompt ChatGPT in German, does your English blog post stand a chance of being cited? And if you rank well in English, do Spanish-language AI answers ever send anyone your way? Reasonable questions for any company selling internationally — and the engines themselves say almost nothing about them. This post separates what is documented at each engine, what published research has measured, and what a global SaaS should practically do about multilingual answer engine optimisation (AEO).
What each engine actually documents about language
Google: real documentation, but for Search — not for AI answers
Google is the only engine with substantial public documentation touching language, and nearly all of it predates AI answers. Its localised versions guidance explains hreflang: each language version must list itself as well as all other versions, annotations must be reciprocal (“If two pages don’t both point to each other, the tags will be ignored”), and — importantly for AEO speculation — “Google doesn’t use hreflang or the HTML lang attribute to detect the language of a page; instead, we use algorithms to determine the language.” hreflang routes users to the right variant; it does not tell Google what language your page is in. The companion multi-regional site guidance adds the hygiene rules: a single language for content and navigation on each page, no side-by-side translations, no automatic language redirects.
On the AI side, Google’s AI features documentation for AI Overviews and AI Mode says “There are no additional requirements to appear in AI Overviews or AI Mode, nor other special optimizations necessary”, and describes a “query fan-out” technique that surfaces “a wider and more diverse set of helpful links”. What that page does not address, anywhere, is language. The one concrete language fact Google does publish is availability: its AI Overviews help page maintains an explicit list of the countries, territories and languages where AI Overviews appear. The list is long but finite — if a language is not on it, users prompting in that language are not seeing AI Overviews at all.
OpenAI: crawler docs, nothing on language
OpenAI’s bot documentation covers which crawler does what — OAI-SearchBot surfaces sites in ChatGPT’s search features, and opting out has a stated consequence: “Sites that are opted out of OAI-SearchBot will not be shown in ChatGPT search answers, though can still appear as navigational links” — plus user-agent strings, IP ranges and robots.txt handling. Neither this page nor any other public OpenAI developer page we could verify states how ChatGPT search treats a page’s language relative to the prompt’s language.
Perplexity: same story
Perplexity’s crawler documentation describes PerplexityBot as “designed to surface and link websites in search results on Perplexity” and Perplexity-User as visiting pages to answer a specific user’s question. Again: user agents, IP ranges, firewall configuration — and no documentation of language behaviour in retrieval or citation.
Anthropic: location-based localisation, not language-based
Anthropic’s web search tool documentation for the Claude API is the closest any engine comes to documenting localisation: a user_location parameter “allows you to localize search results based on a user’s location”, accepting city, region, country and timezone. Note what is absent: no language parameter, and no statement about how a page’s language affects whether Claude retrieves or cites it. Localisation, as documented, is about place, not language.
Documented vs not documented: the comparison
| Engine / surface | What the owner docs state about language | What remains undocumented |
|---|---|---|
| Google Search (classic) | hreflang routes users to the right language/region variant; language is detected algorithmically, one language per page recommended | — |
| Google AI Overviews / AI Mode | Explicit country-and-language availability list; “no additional requirements” beyond normal Search | Whether a page in one language can be cited for a prompt in another |
| ChatGPT search (OpenAI) | Crawler roles only (OAI-SearchBot, ChatGPT-User, GPTBot) | Any language behaviour in retrieval or citation |
| Perplexity | Crawler roles only (PerplexityBot, Perplexity-User) | Any language behaviour in retrieval or citation |
| Claude web search (Anthropic) | user_location localises results by city/region/country/timezone; citations always attached |
Any language parameter or cross-language retrieval behaviour |
That is the honest state of owner documentation in August 2026: one availability list, one location parameter, and a lot of silence.
What published research has actually measured
Where vendor docs are silent, peer-reviewed research is not. Two studies are worth knowing.
Same-language bias is real and systemic. “Faux Polyglot: A Study on Information Disparity in Multilingual Large Language Models” (Sharma, Murray & Xiao, accepted to NAACL 2025) tested how LLMs behave in cross-language retrieval-augmented generation. The finding: LLMs “displayed systemic bias towards information in the same language as the query language in both document retrieval and answer generation”. And when no documents exist in the prompt’s language, the models prefer documents in high-resource languages — with English the highest-resource of all. For a marketer that reads as two rules: content in the prompt’s language has the inside track, and English content is the fallback used when local-language content doesn’t exist.
Multilingual generation is fragile. “Retrieval-augmented generation in multilingual settings” (Chirkova et al., NAVER Labs Europe) built RAG pipelines across 13 languages and found that “task-specific prompt engineering is needed to enable generation in user languages”, with persistent failure modes including “frequent code-switching in non-Latin alphabet languages, occasional fluency errors, wrong reading of the provided documents, or irrelevant retrieval”. Systems answering non-English prompts work harder and fail more — one more reason clean source pages in the user’s language are valuable to them.
Two caveats, honestly stated. These are lab studies of RAG pipelines and open models, not audits of production ChatGPT, Gemini or Perplexity — the commercial engines may behave differently, and none of them says. And this post’s title question answers itself from the same evidence: non-English content is unlikely to be cited for English prompts, because English is both the query language and the best-resourced fallback. Cross-language traffic mostly flows the other way — English pages cited for non-English prompts where local content is thin.
Practical multilingual AEO for a global SaaS
1. Translate your best pages, not your whole blog. If same-language content wins when it exists, the highest-leverage move is translating the handful of pages that already earn citations in English — properly localised, not machine-dumped — into languages your buyers actually prompt in. Check Google’s AI Overviews availability list first: a buyer language on that list is a strong candidate; one not on it still has classic Search, ChatGPT and Perplexity upside, just not the AI Overviews surface.
2. Keep entity facts identical in every language. Company name, product names, founding facts and numbers must not drift in translation — an engine reconciling your German and English pages should find the same claims, differently worded. This is the multilingual extension of entity consistency for AI search: cross-language inconsistency is indistinguishable, to a machine, from unreliability. Keep proper nouns untranslated everywhere.
3. Do the boring hreflang hygiene. Self-referencing, reciprocal hreflang; one language per page; no side-by-side translations; no auto-redirects by browser language — all straight from Google’s documentation. None of it is documented to influence AI citation, and we won’t pretend it is. But AI Overviews sits on Search infrastructure, and clean language signals cost little.
4. Match transcript language to page language. If you publish call, webinar or demo transcripts, keep each transcript page in one language — the language of the audio — rather than interleaving a translation. Mixed-language pages fight Google’s one-language-per-page preference, and code-switching is precisely the failure mode multilingual RAG research flags. Put translations on separate, hreflang-linked pages, structured for machines like any other — see our guide to writing for agentic parsing.
5. Keep English strong regardless. Until you have local-language content in a market, your English pages are what an engine reaches for when a Portuguese or Japanese prompt finds nothing local. Comprehensive English coverage is the floor under every market you haven’t localised yet.
Multilingual content is not multilingual outreach
One distinction matters here. Everything above concerns multilingual content — what your website says, in which languages, for engines to cite. That is a different problem from multilingual outreach — what happens after a lead exists, in whatever language the lead speaks. Zian’s digital team agents handle the second problem: AI sales agents that work in 30+ languages across phone, SMS, email and WhatsApp, with voice cloning, so a lead generated by your English content can be called back in their own language. Content earns the enquiry, outreach converts it, and the language requirements of the two are independent decisions: translating your site does not make your follow-up multilingual, and multilingual follow-up does not require translating your site.
Sources & ownership
| Figure/claim | Owner (organisation) | Where it’s published | Date checked |
|---|---|---|---|
| hreflang must be reciprocal and self-referencing; Google doesn’t use hreflang or the lang attribute to detect page language | Google (Search Central) | Localised versions guidance | 2026-08-22 |
| One language per page; avoid side-by-side translations and automatic language redirects | Google (Search Central) | Multi-regional site guidance | 2026-08-22 |
| “No additional requirements” for AI Overviews / AI Mode; query fan-out; no language guidance on the page | Google (Search Central) | AI features documentation | 2026-08-22 |
| AI Overviews availability list of countries, territories and languages | Google (Search Help) | AI Overviews help page | 2026-08-22 |
| OAI-SearchBot surfaces sites in ChatGPT search; opted-out sites not shown in search answers; no language documentation | OpenAI | OpenAI bot documentation | 2026-08-22 |
| PerplexityBot surfaces and links websites in Perplexity results; no language documentation | Perplexity | Perplexity crawler documentation | 2026-08-22 |
Claude web search user_location localises results by city/region/country/timezone; citations always enabled; no language parameter |
Anthropic | Web search tool documentation | 2026-08-22 |
| LLMs show systemic same-language bias in retrieval and generation; fall back to high-resource languages when no query-language sources exist | Sharma, Murray & Xiao (NAACL 2025) | arXiv:2407.05502 | 2026-08-22 |
| Multilingual RAG needs task-specific prompting to answer in the user’s language; code-switching and irrelevant retrieval are persistent failure modes | Chirkova et al. (NAVER Labs Europe) | arXiv:2407.01463 | 2026-08-22 |
FAQs
Do AI engines cite non-English pages for English prompts?
Rarely, on the available evidence — and no engine documents this either way. The peer-reviewed Faux Polyglot study (NAACL 2025) found LLMs show systemic bias towards information in the same language as the query in both retrieval and generation, and English is also the high-resource fallback when no same-language sources exist. Both effects favour English sources for English prompts. The likelier cross-language flow is the reverse: English pages cited for non-English prompts where local content is thin.
Should a global SaaS translate its content for AEO?
Translate selectively. Research suggests same-language content has the advantage when it exists, so properly localising your few best-performing pages into languages your buyers actually prompt in beats machine-translating an entire blog. Prioritise languages on Google’s AI Overviews availability list and markets where you have real pipeline; strong English content remains the fallback everywhere else.
Does hreflang affect AI citations?
Nothing published says so. Google documents hreflang as a way to point Search users to the right language or regional variant, and states in its localised versions documentation that it doesn’t use hreflang or the HTML lang attribute to detect a page’s language. No engine documents any hreflang role in AI answer citation. Do it anyway: it is cheap, it serves classic Search, and AI Overviews runs on Search infrastructure.
Which languages do AI Overviews support?
Google maintains the authoritative list on its AI Overviews help page, naming the countries, territories and languages where the feature is available. Check it before investing in a language purely for AI Overviews visibility — if the language isn’t listed, users prompting in it see classic results instead, though ChatGPT and Perplexity may still serve them.
Is translating our website necessary to run multilingual sales outreach?
No — they’re independent. Website translation determines which languages engines can cite you in; outreach language is determined by whoever contacts the lead. Zian’s agents conduct outreach in 30+ languages across phone, SMS, email and WhatsApp regardless of what languages your site publishes in, so an English-only site can still follow up with every lead in the lead’s own language. The reverse also holds: a translated site doesn’t make your sales team multilingual.
Multilingual AEO in 2026 is a discipline practised mostly in the dark: the engines publish availability lists and crawler manuals, not language policies. Work with what is verifiable — same-language content wins where it exists, English is the fallback everywhere else, and entity facts must never drift in translation — and be suspicious of anyone selling you certainty the vendors themselves don’t claim.