llms.txt in 2026: Does Any AI Engine Actually Read It? - Zian AI

llms.txt in 2026: Does Any AI Engine Actually Read It?

At a glance: No major AI engine — OpenAI, Anthropic, Google or Perplexity — has ever documented consuming llms.txt from third-party websites; their crawler docs name exactly one control file, robots.txt. Confusingly, several of those companies publish llms.txt for their own docs, and Chrome’s Lighthouse now audits your site for one. Serving the file is cheap, but treat claims that it drives AI citations as unproven — and if a WordPress plugin auto-generated yours, audit what it’s advertising.

llms.txt has become the cheapest box to tick in answer engine optimisation: WordPress plugins generate it — All in One SEO by default — agencies list it in every AEO audit, and Lighthouse now checks whether you serve one. A good moment to ask the question the sales decks skip — has any AI engine ever said, in its own documentation, that it reads this file? We went looking at the owners’ pages.

What llms.txt actually is — a proposal, not a standard

llmstxt.org describes itself, in its own subtitle, as “A proposal to standardise on using an /llms.txt file to provide information to help agents use a website.” Jeremy Howard published it on 3 September 2024; a v2 revision landed on 10 August 2026. The idea: a Markdown file at your site root — site name, blockquote summary, curated link lists — so an agent with a limited context window can orient itself without crawling everything. As the proposal puts it, “Agents are best served by concise, expert-level information gathered in a single, accessible location.”

Be fair about what it claims: the proposal doesn’t promise citations or rankings, only to help agents use a site efficiently. It is also, two years in, still a proposal. No standards body has adopted it, and none of the four companies whose engines dominate AI answers has committed to reading it.

Producing is not consuming

The most common evidence offered for llms.txt adoption is that the AI labs themselves publish the file. True — and it proves the opposite of what it’s cited for. OpenAI’s developer docs serve their own llms.txt; so do Anthropic’s Claude platform docs and Perplexity’s developer docs. That’s producing — a courtesy to the coding agents that read API docs all day. Consuming is a different act entirely, and the one that would matter for your visibility: an engine’s crawler fetching your llms.txt and using it to decide what to read, index or cite. Check each company’s crawler documentation and that claim is nowhere:

  • OpenAI’s crawler documentation for OAI-SearchBot, ChatGPT-User and GPTBot discusses one control file: robots.txt. llms.txt appears on that page only as a link to OpenAI’s own documentation index — never as an input to any of the three bots.
  • Anthropic’s crawler policy says its bots respect “do not crawl” signals via industry-standard robots.txt directives. No mention of ClaudeBot reading llms.txt — even though Anthropic publishes one for its own docs.
  • Google Search Central’s guidance on AI features is the bluntest: “You don’t need to create new machine readable files, AI text files, or markup to appear in these features.” That sentence sits in Google’s own documentation for AI Overviews and AI Mode.
  • Perplexity’s bot documentation for PerplexityBot and Perplexity-User documents robots.txt as the only control — from a docs site serving its own llms.txt at the root.

So the honest one-line summary of the ecosystem in August 2026: the AI labs write llms.txt files; none of them has documented reading yours.

What John Mueller actually said (and where)

Google’s John Mueller is the most-quoted voice on this topic — usually accurately, rarely with a source. The venue was an April 2025 r/TechSEO thread asking what llms.txt adoption looked like. Mueller’s reply, in full:

“AFAIK none of the AI services have said they’re using LLMs.TXT (and you can tell when you look at your server logs that they don’t even check for it). To me, it’s comparable to the keywords meta tag – this is what a site-owner claims their site is about … (Is the site really like that? well, you can check it. At that point, why not just check the site directly?)”

Both halves of that comment are checkable: the documentation review above answers the first, and your own server logs answer the second — a test worth running before paying anyone for llms.txt work.

Mueller returned to the topic on Bluesky in July 2025, saying the file wouldn’t normally be duplicate content but that “using noindex for it could make sense, as sites might link to it and it could otherwise become indexed, which would be weird for users.” Personal comments, not formal Google policy — but they point the same direction as the official documentation.

The Lighthouse wrinkle: Google’s products don’t agree with each other

Here’s the genuinely confusing 2026 development. While Google Search’s documentation says you don’t need AI text files, Chrome’s Lighthouse added an “agentic browsing” audit that checks for llms.txt — calling it “an emerging convention,” flagging your page if the file returns a server error, and marking the audit Not Applicable if you don’t serve one, “as providing the file is optional at the moment.”

Read that carefully before updating your roadmap. Lighthouse checks whether browser-based agents could orient themselves on your site; it says nothing about rankings or citations, and a missing file doesn’t fail the audit. But it is the first Google-owned product to treat llms.txt as worth measuring — exactly the kind of weak, ambiguous signal that gets inflated into “Google now requires llms.txt” by the time it reaches a sales deck. It doesn’t. The same company’s search documentation still says the opposite.

What each owner actually documents

Organisation Documents consuming llms.txt from your site? What their docs do say Publishes its own llms.txt?
OpenAI No Crawler docs cover robots.txt for OAI-SearchBot, ChatGPT-User and GPTBot Yes — developer docs index
Anthropic No Crawler policy: bots honour robots.txt directives Yes — Claude platform docs
Google (Search) No “You don’t need to create new machine readable files, AI text files, or markup” for AI Overviews / AI Mode
Google (Chrome Lighthouse) n/a — audits presence only Calls it “an emerging convention”; absent file = Not Applicable, “optional at the moment”
Perplexity No Bot docs cover robots.txt for PerplexityBot and Perplexity-User Yes — developer docs index

Meanwhile, WordPress is generating these files at scale

The supply side hasn’t waited for the demand side. All in One SEO ships an llms.txt generator that is “enabled by default”, producing both llms.txt and a heavier llms-full.txt from your selected post types and taxonomies. Yoast SEO’s functional specification describes a file “updated weekly by a scheduled action” containing roughly the five most recently updated items per content type, cornerstone content first.

Put those two facts together — plugins generating files at scale (in AIOSEO’s case by default), engines documenting no consumption — and you get the current oddity: a very large number of websites now serve a manifest that, per the owners’ own documentation, no major engine has committed to reading. That’s not a reason to panic. It is a reason to stop treating “we added llms.txt” as an AEO deliverable worth paying for, and to start scrutinising the file you already have.

Zian sits on the other side of this equation: autonomous sales agents built to work leads over phone, SMS, email and WhatsApp, whatever discovery files the engines eventually standardise on. If you’d rather spend energy on pipeline than speculative crawl files, Apply For Partnership.

The real risk: what your auto-generated file is advertising

Here’s the part almost nobody checks. An auto-generated llms.txt is built from whatever public content your CMS holds — not from what you’d choose to show a prospect. The entire argument for serving one is that some agent, someday, reads it as your site’s self-description. If you believe that enough to serve the file, believe it enough to read the file. Failure modes we’d audit for on any site:

  • Theme and page-builder leftovers. Demo templates, sample pages and builder library items registered as public post types can be swept into a generated file. If your llms.txt lists “Home Copy 2” or a template gallery, an agent sees that as your curated best content.
  • Placeholder text. Pages still carrying lorem ipsum or “your headline here” from the original build — invisible to you because nothing links to them, but a generator that enumerates post types doesn’t care what humans link to.
  • Stale offers and CTAs. Files built from “most recently updated” content can pin last year’s webinar, an expired promotion or a retired pricing page as a top-level link for months.
  • Contradictions with your index directives. Content you’ve deliberately noindexed can still be advertised in the manifest — telling one system “don’t surface this” and another “start here.”
  • Entity drift. A plugin-default description (“Just another WordPress site”) or an old tagline undercuts the brand-fact consistency that actually matters for AI answers.

The audit takes five minutes: fetch yourdomain.com/llms.txt (and llms-full.txt if it exists), read every line as if you were a buyer’s research agent, and exclude anything you wouldn’t put on a sales call. Then grep your access logs for llms.txt and see which user agents, if any, ever request it — your personal answer to this article’s title question, updated live. If you keep the file, follow Mueller’s suggestion and serve it with a noindex header.

Verdict: serve it if you like, audit it if you do, believe nobody who sells it as a ranking lever

llms.txt costs almost nothing to serve — the strongest argument for it, and the weakest argument anyone should need. An audited file is a harmless hedge on a future where agentic browsers standardise on the convention, a future Lighthouse’s audit gestures at without promising. What the evidence does not support, anywhere in the owners’ documentation, is that llms.txt drives citations, rankings or AI traffic today. Anyone selling it as a measurable AEO win is selling ahead of the evidence — the same pattern we’ve documented in AEO benchmark reports and MCP-for-marketers hype.

The work that visibly feeds AI answers remains unglamorous: answer-shaped pages, consistent entity facts, clean HTML that agents can parse, and measurement via share of model rather than vibes — the programme in our guide to answer engine optimisation for SaaS. Zian’s agents are built on the same philosophy applied to pipeline: verifiable outcomes over speculative optimisation. See what that looks like on your funnel — Apply For Partnership.

Frequently asked questions

What is llms.txt and who created it?

llms.txt is a proposed convention — a Markdown file at a website’s root summarising the site and its key pages for AI agents. Jeremy Howard proposed it in September 2024 at llmstxt.org, which describes it as “a proposal to standardise” the practice. It is not a ratified standard, and no major AI engine documents reading it from third-party sites.

Does Google use llms.txt?

Google’s own documentation says no: Search Central’s AI features guidance states “You don’t need to create new machine readable files, AI text files, or markup to appear in these features.” Google’s John Mueller also wrote in an April 2025 Reddit thread that no AI service had said it uses llms.txt, comparing it to the keywords meta tag. Chrome’s Lighthouse audits for the file’s presence, but calls providing it “optional at the moment.”

Should I delete the llms.txt file my SEO plugin generated?

Not necessarily — the file is harmless if accurate, and removing it buys you nothing. The better move is to audit it: remove theme-template leftovers, placeholder pages and stale offers, check it doesn’t advertise noindexed content, and consider serving it with a noindex header. Then check your server logs to see whether any AI crawler ever requests it.

Is llms.txt the same kind of file as robots.txt?

No, and the difference matters. robots.txt is a decades-old access-control convention that OpenAI, Anthropic, Google and Perplexity all document honouring in their crawler docs. llms.txt is the reverse — an invitation rather than a restriction, with no documented consumer among the major engines and no enforcement role at all.

Could llms.txt matter in the future?

Possibly. Chrome Lighthouse’s agentic-browsing audit is the first Google-owned product to measure the file, and the AI labs publish it for their own developer docs. If a major engine ever documents fetching llms.txt from third-party sites, the calculus changes overnight — and it will be announced in crawler documentation, not a vendor webinar. Until then, treat it as cheap optionality, not a channel.

Sources and ownership

Claim / quote Owner Verified at
llms.txt proposal wording, dates (3 Sep 2024; v2 10 Aug 2026) and quotes llmstxt.org (Jeremy Howard) llmstxt.org
Mueller April 2025 quote (“AFAIK none of the AI services…”) John Mueller via Reddit r/TechSEO reddit.com/r/TechSEO
Mueller noindex comment, 21 Jul 2025 John Mueller via Bluesky bsky.app/profile/johnmu.com
“You don’t need to create new machine readable files, AI text files, or markup…” Google Search Central developers.google.com
Lighthouse audit behaviour and “optional at the moment” wording Chrome for Developers (Google) developer.chrome.com
OpenAI crawler docs (robots.txt only; own llms.txt linked) OpenAI developers.openai.com
Anthropic crawler policy (robots.txt; no llms.txt consumption) Anthropic support.claude.com
Perplexity bot docs (robots.txt only; own llms.txt served) Perplexity docs.perplexity.ai
AIOSEO generator “enabled by default”; llms.txt + llms-full.txt All in One SEO aioseo.com
Yoast file “updated weekly by a scheduled action”; ~5 latest items per type Yoast (developer portal) developer.yoast.com

Related Blogs

Related from Zian AI