Multi-Language AEO: Do AI Engines Cite Non-English Content for English Prompts? - Zian AI

Multi-Language AEO: Do AI Engines Cite Non-English Content for English Prompts?

At a glance: No major AI engine publicly documents how language affects which pages it cites. Google documents hreflang hygiene for Search generally plus a country-and-language availability list for AI Overviews; OpenAI and Perplexity document crawlers, not language behaviour; Anthropic documents location-based (not language-based) localisation for Claude’s web search. Peer-reviewed research fills some of the gap: LLMs systematically prefer sources in the prompt’s language and fall back to high-resource languages — usually English — when nothing exists in the prompt’s language. For a global SaaS: translate your genuinely best pages into languages your buyers prompt in, keep entity facts identical across languages, and treat strong English content as the fallback layer, not the whole strategy.

If your prospects prompt ChatGPT in German, does your English blog post stand a chance of being cited? And if you rank well in English, do Spanish-language AI answers ever send anyone your way? Reasonable questions for any company selling internationally — and the engines themselves say almost nothing about them. This post separates what is documented at each engine, what published research has measured, and what a global SaaS should practically do about multilingual answer engine optimisation (AEO).

What each engine actually documents about language

Google: real documentation, but for Search — not for AI answers

Google is the only engine with substantial public documentation touching language, and nearly all of it predates AI answers. Its localised versions guidance explains hreflang: each language version must list itself as well as all other versions, annotations must be reciprocal (“If two pages don’t both point to each other, the tags will be ignored”), and — importantly for AEO speculation — “Google doesn’t use hreflang or the HTML lang attribute to detect the language of a page; instead, we use algorithms to determine the language.” hreflang routes users to the right variant; it does not tell Google what language your page is in. The companion multi-regional site guidance adds the hygiene rules: a single language for content and navigation on each page, no side-by-side translations, no automatic language redirects.

On the AI side, Google’s AI features documentation for AI Overviews and AI Mode says “There are no additional requirements to appear in AI Overviews or AI Mode, nor other special optimizations necessary”, and describes a “query fan-out” technique that surfaces “a wider and more diverse set of helpful links”. What that page does not address, anywhere, is language. The one concrete language fact Google does publish is availability: its AI Overviews help page maintains an explicit list of the countries, territories and languages where AI Overviews appear. The list is long but finite — if a language is not on it, users prompting in that language are not seeing AI Overviews at all.

OpenAI: crawler docs, nothing on language

OpenAI’s bot documentation covers which crawler does what — OAI-SearchBot surfaces sites in ChatGPT’s search features, and opting out has a stated consequence: “Sites that are opted out of OAI-SearchBot will not be shown in ChatGPT search answers, though can still appear as navigational links” — plus user-agent strings, IP ranges and robots.txt handling. Neither this page nor any other public OpenAI developer page we could verify states how ChatGPT search treats a page’s language relative to the prompt’s language.

Perplexity: same story

Perplexity’s crawler documentation describes PerplexityBot as “designed to surface and link websites in search results on Perplexity” and Perplexity-User as visiting pages to answer a specific user’s question. Again: user agents, IP ranges, firewall configuration — and no documentation of language behaviour in retrieval or citation.

Anthropic: location-based localisation, not language-based

Anthropic’s web search tool documentation for the Claude API is the closest any engine comes to documenting localisation: a user_location parameter “allows you to localize search results based on a user’s location”, accepting city, region, country and timezone. Note what is absent: no language parameter, and no statement about how a page’s language affects whether Claude retrieves or cites it. Localisation, as documented, is about place, not language.

Documented vs not documented: the comparison

Engine / surface What the owner docs state about language What remains undocumented
Google Search (classic) hreflang routes users to the right language/region variant; language is detected algorithmically, one language per page recommended
Google AI Overviews / AI Mode Explicit country-and-language availability list; “no additional requirements” beyond normal Search Whether a page in one language can be cited for a prompt in another
ChatGPT search (OpenAI) Crawler roles only (OAI-SearchBot, ChatGPT-User, GPTBot) Any language behaviour in retrieval or citation
Perplexity Crawler roles only (PerplexityBot, Perplexity-User) Any language behaviour in retrieval or citation
Claude web search (Anthropic) user_location localises results by city/region/country/timezone; citations always attached Any language parameter or cross-language retrieval behaviour

That is the honest state of owner documentation in August 2026: one availability list, one location parameter, and a lot of silence.

What published research has actually measured

Where vendor docs are silent, peer-reviewed research is not. Two studies are worth knowing.

Same-language bias is real and systemic. “Faux Polyglot: A Study on Information Disparity in Multilingual Large Language Models” (Sharma, Murray & Xiao, accepted to NAACL 2025) tested how LLMs behave in cross-language retrieval-augmented generation. The finding: LLMs “displayed systemic bias towards information in the same language as the query language in both document retrieval and answer generation”. And when no documents exist in the prompt’s language, the models prefer documents in high-resource languages — with English the highest-resource of all. For a marketer that reads as two rules: content in the prompt’s language has the inside track, and English content is the fallback used when local-language content doesn’t exist.

Multilingual generation is fragile. “Retrieval-augmented generation in multilingual settings” (Chirkova et al., NAVER Labs Europe) built RAG pipelines across 13 languages and found that “task-specific prompt engineering is needed to enable generation in user languages”, with persistent failure modes including “frequent code-switching in non-Latin alphabet languages, occasional fluency errors, wrong reading of the provided documents, or irrelevant retrieval”. Systems answering non-English prompts work harder and fail more — one more reason clean source pages in the user’s language are valuable to them.

Two caveats, honestly stated. These are lab studies of RAG pipelines and open models, not audits of production ChatGPT, Gemini or Perplexity — the commercial engines may behave differently, and none of them says. And this post’s title question answers itself from the same evidence: non-English content is unlikely to be cited for English prompts, because English is both the query language and the best-resourced fallback. Cross-language traffic mostly flows the other way — English pages cited for non-English prompts where local content is thin.

Apply For Partnership

Practical multilingual AEO for a global SaaS

1. Translate your best pages, not your whole blog. If same-language content wins when it exists, the highest-leverage move is translating the handful of pages that already earn citations in English — properly localised, not machine-dumped — into languages your buyers actually prompt in. Check Google’s AI Overviews availability list first: a buyer language on that list is a strong candidate; one not on it still has classic Search, ChatGPT and Perplexity upside, just not the AI Overviews surface.

2. Keep entity facts identical in every language. Company name, product names, founding facts and numbers must not drift in translation — an engine reconciling your German and English pages should find the same claims, differently worded. This is the multilingual extension of entity consistency for AI search: cross-language inconsistency is indistinguishable, to a machine, from unreliability. Keep proper nouns untranslated everywhere.

3. Do the boring hreflang hygiene. Self-referencing, reciprocal hreflang; one language per page; no side-by-side translations; no auto-redirects by browser language — all straight from Google’s documentation. None of it is documented to influence AI citation, and we won’t pretend it is. But AI Overviews sits on Search infrastructure, and clean language signals cost little.

4. Match transcript language to page language. If you publish call, webinar or demo transcripts, keep each transcript page in one language — the language of the audio — rather than interleaving a translation. Mixed-language pages fight Google’s one-language-per-page preference, and code-switching is precisely the failure mode multilingual RAG research flags. Put translations on separate, hreflang-linked pages, structured for machines like any other — see our guide to writing for agentic parsing.

5. Keep English strong regardless. Until you have local-language content in a market, your English pages are what an engine reaches for when a Portuguese or Japanese prompt finds nothing local. Comprehensive English coverage is the floor under every market you haven’t localised yet.

Multilingual content is not multilingual outreach

One distinction matters here. Everything above concerns multilingual content — what your website says, in which languages, for engines to cite. That is a different problem from multilingual outreach — what happens after a lead exists, in whatever language the lead speaks. Zian’s digital team agents handle the second problem: AI sales agents that work in 30+ languages across phone, SMS, email and WhatsApp, with voice cloning, so a lead generated by your English content can be called back in their own language. Content earns the enquiry, outreach converts it, and the language requirements of the two are independent decisions: translating your site does not make your follow-up multilingual, and multilingual follow-up does not require translating your site.

Sources & ownership

Figure/claim Owner (organisation) Where it’s published Date checked
hreflang must be reciprocal and self-referencing; Google doesn’t use hreflang or the lang attribute to detect page language Google (Search Central) Localised versions guidance 2026-08-22
One language per page; avoid side-by-side translations and automatic language redirects Google (Search Central) Multi-regional site guidance 2026-08-22
“No additional requirements” for AI Overviews / AI Mode; query fan-out; no language guidance on the page Google (Search Central) AI features documentation 2026-08-22
AI Overviews availability list of countries, territories and languages Google (Search Help) AI Overviews help page 2026-08-22
OAI-SearchBot surfaces sites in ChatGPT search; opted-out sites not shown in search answers; no language documentation OpenAI OpenAI bot documentation 2026-08-22
PerplexityBot surfaces and links websites in Perplexity results; no language documentation Perplexity Perplexity crawler documentation 2026-08-22
Claude web search user_location localises results by city/region/country/timezone; citations always enabled; no language parameter Anthropic Web search tool documentation 2026-08-22
LLMs show systemic same-language bias in retrieval and generation; fall back to high-resource languages when no query-language sources exist Sharma, Murray & Xiao (NAACL 2025) arXiv:2407.05502 2026-08-22
Multilingual RAG needs task-specific prompting to answer in the user’s language; code-switching and irrelevant retrieval are persistent failure modes Chirkova et al. (NAVER Labs Europe) arXiv:2407.01463 2026-08-22

FAQs

Do AI engines cite non-English pages for English prompts?

Rarely, on the available evidence — and no engine documents this either way. The peer-reviewed Faux Polyglot study (NAACL 2025) found LLMs show systemic bias towards information in the same language as the query in both retrieval and generation, and English is also the high-resource fallback when no same-language sources exist. Both effects favour English sources for English prompts. The likelier cross-language flow is the reverse: English pages cited for non-English prompts where local content is thin.

Should a global SaaS translate its content for AEO?

Translate selectively. Research suggests same-language content has the advantage when it exists, so properly localising your few best-performing pages into languages your buyers actually prompt in beats machine-translating an entire blog. Prioritise languages on Google’s AI Overviews availability list and markets where you have real pipeline; strong English content remains the fallback everywhere else.

Does hreflang affect AI citations?

Nothing published says so. Google documents hreflang as a way to point Search users to the right language or regional variant, and states in its localised versions documentation that it doesn’t use hreflang or the HTML lang attribute to detect a page’s language. No engine documents any hreflang role in AI answer citation. Do it anyway: it is cheap, it serves classic Search, and AI Overviews runs on Search infrastructure.

Which languages do AI Overviews support?

Google maintains the authoritative list on its AI Overviews help page, naming the countries, territories and languages where the feature is available. Check it before investing in a language purely for AI Overviews visibility — if the language isn’t listed, users prompting in it see classic results instead, though ChatGPT and Perplexity may still serve them.

Is translating our website necessary to run multilingual sales outreach?

No — they’re independent. Website translation determines which languages engines can cite you in; outreach language is determined by whoever contacts the lead. Zian’s agents conduct outreach in 30+ languages across phone, SMS, email and WhatsApp regardless of what languages your site publishes in, so an English-only site can still follow up with every lead in the lead’s own language. The reverse also holds: a translated site doesn’t make your sales team multilingual.

Multilingual AEO in 2026 is a discipline practised mostly in the dark: the engines publish availability lists and crawler manuals, not language policies. Work with what is verifiable — same-language content wins where it exists, English is the fallback everywhere else, and entity facts must never drift in translation — and be suspicious of anyone selling you certainty the vendors themselves don’t claim.

Apply For Partnership

Related Blogs

Related from Zian AI