Blocking AI Crawlers: Did It Drop Me From Google? - Zian AI

Blocking AI Crawlers: Did It Drop Me From Google?

Quick answer: probably nothing. Cloudflare reports that “less than 1% of Cloudflare sites choose to block Search bots”, and on 15 September 2026 existing settings carried over untouched. What did change is the meaning of one word: Block now stops Googlebot, Bingbot and Applebot as well. The risk starts the next time somebody clicks it.

First, the odds: this is rarely why a site drops out of Google

The honest starting position is that you are unlikely to have done this to yourself. Cloudflare published its own adoption figures alongside the September 2026 change: “less than 1% of Cloudflare sites choose to block Search bots”, against “17% of sites choose to enable some mechanism to block training”. Search-bot blocking is a rounding error. Training-bot blocking is one site in six.

Google’s own guide to debugging drops in Google Search traffic lists the causes it expects you to investigate: algorithmic update, technical issues, security issues, spam issues, seasonality and changing interests, and site moves and migrations. A crawler block is not a named entry on that list; it falls inside “technical issues”, alongside a misplaced noindex and a server outage.

So work the cheap tests first, before you touch a robots.txt line, a Cloudflare bot rule or a Google-Extended token. Each of these takes under ten minutes, three of them need nothing but Search Console, and all of them are worth doing before you conclude that Googlebot, Bingbot, Applebot or your AI Overviews visibility went anywhere.

Candidate cause What the Search Console chart looks like The test that rules it in or out
A Google ranking update Clicks and impressions fall together over days, not hours. Pages still get crawled. Compare your break date against Google’s published ranking updates page. Google’s framing is positional: “Small drop in position? For example, dropping from position 2 to 4” versus “dropping from the top 10 results to position 29”. If average position moved and crawling did not, it is not a block.
A technical change that was not a bot rule A cliff. Pages move into Page indexing exclusions with a named reason. Run URL Inspection on three URLs that used to rank. Read the exclusion reason verbatim. A stray noindex in a template, a redirect loop or a 5xx run all name themselves here.
Seasonality or falling demand The same dip as last year, at the same time of year. Switch the Performance report to Last 16 months, then check the two or three queries that lost the most clicks in Google Trends. Google’s guidance is explicit that this is to “understand if the drop was only for your website or throughout the web”.
A reporting artefact An impossible shape: one property flat, another collapsed. Check the Search Console Data Anomalies page for your date, and confirm you are comparing the same property type. A domain property and a URL-prefix property do not report the same numbers.
A manual action or security issue Abrupt, site-wide, often with a message in the account. Open the Manual Actions report and the Security Issues report. Both are one click and both are usually empty.
A crawler rule that caught Googlebot Crawl stats fall to near zero first, then indexing degrades over days. The Four-Timestamp Test below. Grep your access log for reverse-DNS verified Googlebot before you believe anything else.

The quotable version: if verified Googlebot is still fetching your pages and getting a 200, your AI-crawler settings did not cause the drop, whatever the dashboard says.

What Cloudflare changed on 15 September 2026

On 15 September 2026 Cloudflare published “Have it both ways: stay discoverable in search while disallowing AI training”, by Bryan Becker. One paragraph in it matters more than the rest:

“Block and Block on pages with ads now apply to mixed-use crawlers, including Applebot, Bingbot, and Googlebot, so either setting impacts search as well as training. To stop training and keep search, use Disallow AI Training.”

A mixed-use crawler is one crawler doing two jobs. Before this change, Cloudflare carved those crawlers out of Block precisely because of the collateral damage. In Cloudflare’s own words, those two settings “previously did not apply to mixed-use crawlers because blocking them could also affect search discoverability”. That carve-out is gone. Block now means block.

Three things this change is not, each of them easy to get wrong on a skim:

  • Cloudflare did not start blocking AI crawlers by default. On what existing customers had to do, the post says: “Nothing, in almost every case. Your current settings carry over on their own.”
  • The migration did not switch anyone’s search off. Sites that had used the legacy Block AI Bots control were mapped to Search: Allow. Sites that had configured the granular controls kept the practical effect of their choice, because “Previous Training selections of Block or Block on pages with ads will migrate to Disallow AI Training.”
  • The new presets apply to domains onboarding from 15 September, not to yours. A new ad-monetised domain is offered Search: Allow, Training: Disallow AI Training, Agent: Block on pages with ads. A new non-ad domain is offered Allow, Allow, Allow.

Which means the date most people need is not 15 September 2026. It is the next time anyone in your organisation opened the bot controls and selected Block, believing it meant what it meant in August.

The new setting that resolves the trade-off is Disallow AI Training. Cloudflare limits it deliberately: “Disallow AI Training is only available as a setting for Training, not Search or Agent.” Two older controls are on the way out. Cloudflare says Block AI Bots “will be deprecated in favor of the more granular Search, Training, and Agent controls”, and that “Managed Robots.txt will be deprecated in favor of Bot Preference Sync.”

Which Cloudflare setting stops what: the blast-radius table

Every cell below is Cloudflare’s own description of its own product, read on 22 September 2026. Applebot, Bingbot and Googlebot are the three crawlers Cloudflare designates Accountable and mixed-use, so they are the ones a Block decision now reaches.

Setting What it stops Google, Bing and Apple web search Stops AI training? Effect on AI answers When you would actually want it
Allow Nothing, “unless blocked by another setting or a WAF rule” Unaffected No Full eligibility Any site whose buyers research in search or in an assistant
Disallow AI Training (Training control only) “Every other training crawler”, including the training-only crawlers of Amazon, Anthropic, Meta and OpenAI Unaffected. “Accountable mixed-use crawlers remain allowed for search” Yes, via a robots.txt preference published by Bot Preference Sync Retained, because search crawling continues The default answer for a marketing or docs site that does not want to feed training corpora
Block on pages with ads “Crawlers, including mixed-use crawlers, are blocked only on pages detected to be serving an ad” Affected on ad-serving pages On those pages Degraded on those pages Ad-funded publishing, where the page only earns when a human arrives
Block “All crawlers, including mixed-use crawlers, are blocked” Stopped. “It will stop Applebot, Bingbot, and Googlebot from reaching your site — search included” Yes Removed, because AI answers depend on the same crawl Staging hosts, gated research, anything you did not want indexed in the first place
Block AI Bots (legacy) Cloudflare says it “will be deprecated in favor of the more granular Search, Training, and Agent controls” Migrated to Search: Allow Migrated to Disallow AI Training Retained after migration Nothing. Move to the granular controls and read what you selected
Managed Robots.txt (legacy) Cloudflare says it “will be deprecated in favor of Bot Preference Sync”. “Customers who enabled Managed Robots.txt will migrate to the new system” Unaffected by the migration itself Expressed through Bot Preference Sync instead Unaffected by the migration itself Nothing. Confirm what Bot Preference Sync is now publishing in your robots.txt

One asymmetry is worth stating on its own, because it is the reason this page exists. Google’s documentation for AI features says a page is eligible to be shown as a supporting link in AI Overviews or AI Mode when it is “indexed and eligible to be shown in Google Search with a snippet”, and adds: “There are no additional technical requirements.” Search eligibility and AI-answer eligibility are the same eligibility. Block the crawler and you lose both in one move.

Apply For Partnership

The Four-Timestamp Test: was it the block, or something else?

Causes precede effects, and a crawler block leaves four timestamps in three different systems. Collect all four before you touch a setting. We call it the Four-Timestamp Test, and the rule is one sentence: if T1 does not sit at or before T2, and T2 before T3, and T3 at or before T4, your crawler settings did not cause the drop.

  1. T1, the change date. When was a bot or WAF rule last edited? Cloudflare’s audit log carries the timestamp and the user. If your robots.txt lives in the repo, git log the file. A change date after the traffic break is an alibi, not a cause.
  2. T2, the last verified Googlebot 200. Not the last request with a Googlebot user-agent string, which is worth nothing. Reverse-resolve the address and confirm it lands inside googlebot.com, then forward-resolve back. Our method for this is in verifying AI bot traffic with reverse DNS and published IP ranges.
  3. T3, the first day Page indexing shows a new exclusion reason. Read the reason verbatim: Google’s own Page indexing documentation names “URL blocked by robots.txt”, “Server error (5xx)” and “Blocked due to access forbidden (403)” as separate reasons, and they are three different diagnoses. If you want the AI-surface view alongside it, Search Console’s generative AI performance report is the companion.
  4. T4, the day clicks broke. Take it from the Performance report at the 16-month range so you can see whether the same shape appeared a year ago.

Two failure modes the test is designed to catch. If T1 is after T4, somebody changed a setting in response to the drop and is now reading their own fix as the cause. If T2 shows verified Googlebot still collecting 200s through T4, the crawler is getting in and the problem is ranking, indexing or reporting rather than access.

Two blocks, two different symptoms: robots.txt versus the edge

“I blocked AI crawlers” describes two mechanisms with different signatures, and telling them apart narrows the fix considerably.

Symptom robots.txt Disallow that caught Googlebot Edge or WAF block returning 403 or 5xx
Server log Googlebot still fetches /robots.txt, then stops fetching pages Googlebot keeps requesting pages and keeps receiving the error status
Page indexing reason “Indexed, though blocked by robots.txt” or “URL blocked by robots.txt” “Server error (5xx)” or “Blocked due to access forbidden (403)”
Do URLs stay in the index? Often yes, without a useful snippet. Google: “A page that’s disallowed in robots.txt can still be indexed if linked to from other sites”, and “Because of the robots.txt rule, any snippet shown in Google Search results for the page will probably be very limited.” Not for long. Google: “if Googlebot observes these status codes on the same URL for multiple days, the URL may be dropped from Google’s index.”
Speed of the fall Gradual, as snippets and freshness decay Fast, and Google recommends against serving those codes “longer than 1-2 days”
Where the fix lives The file, plus whatever generates it The CDN or WAF rule, plus Bot Preference Sync if it is writing your robots.txt

A fast, total disappearance from Google is an edge-layer story. A slow leak of snippets and rankings is a robots.txt story. Google’s own Page indexing guidance gives you the signature to look for: “If you see a drop in total indexed pages without a corresponding increase in errors, you might be blocking access to your existing pages via robots.txt, ‘noindex’ or a required login. Look for a spike in non-indexed URLs that corresponds to your drop in indexed pages.” If neither shape matches your chart, go back to the first table.

What a user-agent string is worth: 117 requests from a crawler that has no user agent

Every step above depends on identifying who actually knocked. Here is why we insist on verification rather than a grep for a name.

Method and window, stated so you can reproduce it: zian.ai nginx access logs, one domain, all requests of any status, 8 September 2026 00:00 UTC to 21 September 2026 23:59 UTC inclusive — fourteen complete days, 166,653 requests. User agents were matched on token, then checked against each operator’s published IP ranges or reverse DNS. Single site, one fortnight. Do not read it as a measurement of the web.

Agent claimed in the user-agent string Requests Verified against the operator’s own published identity
Applebot 9,983 Not checked
Bytespider 7,096 No published IP list to check against
meta-externalagent 5,442 Not checked
bingbot 2,433 Not checked
Amazonbot 2,182 Not checked
OAI-SearchBot 1,883 1,634 inside openai.com/searchbot.json. 13.2% were not.
Googlebot 1,352 1,213 reverse-resolved into googlebot.com (89.7%). Of the remainder, 54 requests came from 18 googleusercontent.com addresses, which are Google Cloud customers rather than Googlebot.
ClaudeBot 1,245 Not checked
ChatGPT-User 1,077 816 inside openai.com/chatgpt-user.json. 24.2% were not.
PerplexityBot 1,048 Not checked
GPTBot 643 443 inside openai.com/gptbot.json. 31.1% were not.
Google-Extended 117 Zero can be genuine. Google documents that “Google-Extended doesn’t have a separate HTTP request user agent string”, so nothing named Google-Extended ever arrives.

That last row is the whole argument in one line. Google-Extended is a robots.txt token, not a crawler. It cannot appear in a log. It appeared 117 times in ours in a fortnight.

A single address, 207.175.172.107, sent 1,070 requests across those fourteen days while claiming thirteen different AI crawler identities: Amazonbot, Applebot, Bytespider, ChatGPT-User, ClaudeBot, Claude-SearchBot, Claude-User, GPTBot, Google-Extended, meta-externalagent, OAI-SearchBot, Perplexity-User and PerplexityBot. Eighteen other addresses behaved the same way, fifteen of them also claiming Googlebot — and all nineteen reverse-resolve into bc.googleusercontent.com, which is Google Compute Engine, not Googlebot. A dashboard counting user-agent strings would have reported thirteen AI crawlers discovering us. It was one scanner.

The practical consequence for your diagnosis: a bot count that trusts the user-agent string can be wrong by a third in either direction. Before you conclude that Googlebot stopped visiting, verify that it was ever visiting. The taxonomy behind the names is in our breakdown of training crawlers versus on-demand fetchers, which explains why blocking one class costs nothing and blocking the other costs citations.

How do I block AI training but stay in Google?

You can have both, and each operator publishes its own lever. Nothing below is a Cloudflare feature; these work whether or not you use a CDN, and Cloudflare’s Disallow AI Training setting is a way of publishing several of them for you.

Operator The lever What the operator says it costs you in search What it does not do
Google Disallow: / for the Google-Extended token in robots.txt Nothing. “Google-Extended does not impact a site’s inclusion in Google Search nor is it used as a ranking signal in Google Search.” Does not remove you from AI Overviews or AI Mode. Those run on ordinary Googlebot crawling.
Apple Disallow for Applebot-Extended in robots.txt Nothing. “Site rules for Applebot-Extended are not considered in ranking for Search.” “Applebot-Extended does not crawl webpages”, so it will never appear in your logs and cannot be verified there.
Microsoft The NOARCHIVE robots meta tag Nothing in web results: “content with the NOCACHE tag or NOARCHIVE tag will still appear in our search results.” It does cost you the chat surface. Microsoft: “Content tagged NOARCHIVE will not be included in Bing Chat answers, not be linked to in the answers.” Cloudflare reports a robots.txt no-training preference for Bing is “targeted for early 2027”.
OpenAI Disallow: / for GPTBot, while allowing OAI-SearchBot Nothing, because they are separate agents with separate tokens. Allow a lag: “it can take ~24 hours from a site’s robots.txt update for our systems to adjust”.
Perplexity Disallow for the relevant token Blocking PerplexityBot is what removes you from Perplexity’s results, so leave it allowed. It will not stop the user-triggered fetcher. Perplexity’s own docs on Perplexity-User: “Since a user requested the fetch, this fetcher generally ignores robots.txt rules.”

If you want the full per-agent file rather than the principles, we publish one: the AI crawler allowlist, written as a robots.txt for B2B SaaS. And if your goal is to limit how much of your page a Google AI answer shows rather than to block training at all, Google names the controls: nosnippet, data-nosnippet, max-snippet and noindex. Those trade search visibility for snippet control, which is a different bargain and should be made deliberately.

The method above is complete and free. What it costs to run is attention: operators add agents, tokens get renamed, Bot Preference Sync writes into a file you thought you owned, and the scanner traffic in the table above means somebody has to re-verify identities rather than trust a name. That is a quarterly job for a person who already has one. Zian AI works the other side of the same funnel, engaging and qualifying the buyers your visibility earns across phone, SMS, email and WhatsApp, in partnership-application beta.

Apply For Partnership

Frequently asked questions

Does blocking AI bots hurt SEO?

It depends entirely on which bots. Blocking a training-only crawler such as GPTBot costs nothing in search, because it is a separate agent from the one that surfaces you. Blocking a mixed-use crawler does hurt, and since 15 September 2026 Cloudflare’s Block setting includes them: “Block and Block on pages with ads now apply to mixed-use crawlers, including Applebot, Bingbot, and Googlebot, so either setting impacts search as well as training.”

My traffic fell the same week I clicked Block in Cloudflare. Is that proof?

No, it is a coincidence worth testing. Run the Four-Timestamp Test: the change must land on or before the last verified Googlebot 200, which must precede the new exclusion reason in Page indexing, which must precede or match the day clicks broke. Crawl stats falling to near zero while the settings change sits after the break means you are looking at your own reaction, not the cause.

Did Cloudflare start blocking AI crawlers on my site automatically?

No. Cloudflare’s answer to what existing customers needed to do was “Nothing, in almost every case. Your current settings carry over on their own.” Sites on the legacy Block AI Bots control were migrated to Search: Allow, and previous Training selections of Block migrated to the new Disallow AI Training setting rather than to Block. The new presets apply to domains onboarding from 15 September 2026.

Will disallowing Google-Extended lower my rankings?

No. Google’s crawler documentation states that Google-Extended “does not impact a site’s inclusion in Google Search nor is it used as a ranking signal in Google Search”. It controls whether your content trains and grounds Gemini models, and it is a robots.txt token rather than a crawler, so disallowing it will never show up in your access log.

Can I stop Bing using my content for training without leaving Bing?

Partly. Microsoft’s published control is the NOARCHIVE meta tag, and Microsoft states that “content with the NOCACHE tag or NOARCHIVE tag will still appear in our search results”. The catch is on the same page: “Content tagged NOARCHIVE will not be included in Bing Chat answers, not be linked to in the answers.” You keep web results and give up the chat surface, which is a real cost and not one Cloudflare’s setting can avoid for you today.

I unblocked Googlebot. How long until I come back?

Nobody publishes a recovery time, so treat anyone who quotes you one with suspicion. What is documented is the mechanism: Google warns against serving 500, 503 or 429 to crawlers for longer than “1-2 days” because “if Googlebot observes these status codes on the same URL for multiple days, the URL may be dropped from Google’s index”. Recovery therefore requires a successful recrawl of every affected URL. Confirm with URL Inspection on a sample, then watch the Crawl stats report climb before you watch clicks.

Where every figure on this page comes from

Figure Who published it Link Date read
“less than 1% of Cloudflare sites choose to block Search bots”; “17% of sites choose to enable some mechanism to block training” Cloudflare blog.cloudflare.com 22 Sep 2026
Block and Block on pages with ads now apply to mixed-use crawlers, including Applebot, Bingbot and Googlebot (15 September 2026) Cloudflare blog.cloudflare.com 22 Sep 2026
“Nothing, in almost every case. Your current settings carry over on their own.”; the two migration tables; the new-domain presets Cloudflare blog.cloudflare.com 22 Sep 2026
Definitions of Allow, Disallow AI Training, Block on pages with ads and Block; deprecation of Block AI Bots and Managed Robots.txt; the Accountable designation Cloudflare blog.cloudflare.com 22 Sep 2026
Bing robots.txt no-training preference “targeted for early 2027” Cloudflare, reporting Microsoft’s plans blog.cloudflare.com 22 Sep 2026
AI Overviews and AI Mode eligibility: “indexed and eligible to be shown in Google Search with a snippet”, “There are no additional technical requirements”; the nosnippet, data-nosnippet, max-snippet and noindex controls Google developers.google.com 22 Sep 2026
“Google-Extended does not impact a site’s inclusion in Google Search nor is it used as a ranking signal in Google Search”; “Google-Extended doesn’t have a separate HTTP request user agent string” Google developers.google.com 22 Sep 2026
The named causes of a search traffic drop; the position examples “position 2 to 4” and “top 10 results to position 29”; and “understand if the drop was only for your website or throughout the web” Google developers.google.com 22 Sep 2026
“if Googlebot observes these status codes on the same URL for multiple days, the URL may be dropped from Google’s index”; the “1-2 days” guidance for 500, 503 and 429 Google developers.google.com 22 Sep 2026
“A page that’s disallowed in robots.txt can still be indexed if linked to from other sites.” Google developers.google.com 22 Sep 2026
The Page indexing reason strings “URL blocked by robots.txt”, “Indexed, though blocked by robots.txt”, “Server error (5xx)” and “Blocked due to access forbidden (403)”; the limited-snippet note; the drop signature “If you see a drop in total indexed pages without a corresponding increase in errors…” Google support.google.com 22 Sep 2026
Applebot-Extended: “does not crawl webpages”, “Webpages that disallow Applebot-Extended can still be included in search results”, “Site rules for Applebot-Extended are not considered in ranking for Search” Apple support.apple.com 22 Sep 2026
“Content tagged NOARCHIVE will not be included in Bing Chat answers, not be linked to in the answers”; “content with the NOCACHE tag or NOARCHIVE tag will still appear in our search results” Microsoft (Bing Webmaster Blog) blogs.bing.com 22 Sep 2026
GPTBot, OAI-SearchBot, OAI-AdsBot and ChatGPT-User as separate agents; “it can take ~24 hours from a site’s robots.txt update for our systems to adjust”; the published IP files used to verify our own logs OpenAI developers.openai.com 22 Sep 2026
“Since a user requested the fetch, this fetcher generally ignores robots.txt rules.” Perplexity docs.perplexity.ai 22 Sep 2026
All crawler request counts, the IP-verification percentages, the 117 Google-Extended claims and the thirteen identities from 207.175.172.107 Zian AI, first-party nginx logs for zian.ai, 8 to 21 September 2026 inclusive, 166,653 requests, all statuses Not externally published 22 Sep 2026

Related Blogs

Related from Zian AI