Cloudflare retired Managed robots.txt on September 15. Most sites that used it now block no AI crawlers in robots.txt

If you turned on Cloudflare's Managed robots.txt to keep GPTBot, ClaudeBot and Google-Extended out, check your live robots.txt today. In my census of the Tranco top 5,000, most sites that served Cloudflare's managed block on September 7 were serving a robots.txt with no AI-crawler rules at all by September 22. Most of them still refuse GPTBot and ClaudeBot at Cloudflare's edge.

Want to know if YOUR site's AI visibility changes?

Enter your domain and email. One email if a crawler you rely on gets newly blocked, or an existing block goes away. No spam.

What Cloudflare changed

Cloudflare's press release of September 15, 2026 says it is replacing its Managed robots.txt feature with a new one called Bot Preference Sync. The Bot Preference Sync blog post (first published August 21, updated September 15) gives the new feature's marker as # BEGIN Cloudflare Bot Preference Sync and says existing Managed robots.txt users would be prompted to review and move over. At the time of writing, the Managed robots.txt docs page still describes the old feature and does not mention the replacement.

The old feature prepended a block that starts with # BEGIN Cloudflare Managed Content and disallows eight crawlers: Amazonbot, Applebot-Extended, Bytespider, CCBot, ClaudeBot, Google-Extended, GPTBot and meta-externalagent. How that block combines with rules you already had is covered in the managed block guide.

What I measured

On September 7, 79 sites in my crawler census served a robots.txt whose blocked crawlers were exactly those eight. On September 22 I fetched all 79 again with the same fetcher and parser:

45
of the 54 still readable now block none of the eight crawlers in robots.txt
4
still serve the Managed Content block
5
block some or all of the eight with their own rules, without Cloudflare's marker

The other 25 did not return a readable robots.txt on September 22, so they are left out of the counts. Across all 2,771 census sites that had a readable robots.txt on September 7, 6 still carry the old Managed Content marker and none carry the Bot Preference Sync marker.

I checked a sample by hand against archived copies. patreon.com served the managed block in the Wayback Machine's capture of September 15 at 12:46 UTC and not in its capture of September 16. blender.org had it on September 10 and not on September 17. marinetraffic.com (captured September 2) and hugedomains.com (captured September 4) both had it and neither blocks any of the eight today. ko-fi.com, findagrave.com and kick.com blocked all eight on September 7 and block none today.

Beyond the top 5,000: Common Crawl's own robots.txt fetches

My census covers the top 5,000 sites only, so on September 23 I checked a much wider set using the robots.txt files that Common Crawl saves during each crawl. Its September crawl fetched robots.txt files from September 4 to September 17, a window that spans the change. I took random samples of those files by fetch date and counted how many carried the Managed Content marker.

Fetchedrobots.txt filesWith the Managed Content blockBlocking all eight crawlers
September 5 to 1326,639658 (2.47%)701 (2.63%)
September 1421,570521 (2.42%)575 (2.67%)
September 1521,425424 (1.98%)477 (2.23%)
September 16 to 1726,0848 (0.03%)100 (0.38%)

Until September 14, about one robots.txt file in 40 carried the block. By September 16 it was about one in 3,000. No file from any of these days carried the new Bot Preference Sync marker.

I also followed individual sites. In 60 random files from Common Crawl's August crawl, fetched between August 7 and 19, I found 754 sites serving the managed block with all eight crawlers disallowed. I fetched each of them again on September 23. Of the 496 that returned a robots.txt, 468 now block none of the eight and 13 still block all eight. Another 162 now return no robots.txt at all (108 a 404 and 54 an HTML page), which suggests Cloudflare had been writing their whole file. The other 96 did not respond. I spot-checked four of the 404 sites and all four are still served by Cloudflare. Six sites I checked by hand returned the same file to a browser, to CCBot and to GPTBot, so the result does not depend on which crawler asks.

Common Crawl's crawl leans toward smaller sites, so these percentages describe the sites it visits and not the whole web.

A separate measurement found the same drop. The team behind Torumata fetched robots.txt from the 22,846 most-visited domains in ten European countries and reported on the Cloudflare Community forum that 769 of them served the managed block on September 14 and only 39 still did on September 16. SeenSure's before-and-after study of 1,046 sites looked at a different layer, how often sites refuse a request from each AI crawler, and found that refusals of search and agent crawlers such as OAI-SearchBot fell sharply on Cloudflare-served sites after the change.

On September 23 IT-Administrator, a German magazine for system and network administrators, added these Common Crawl, panel and edge figures to its article on Cloudflare's new crawler settings (in German) and asked Cloudflare for a statement.

At the edge, most of these sites still turn GPTBot away

A robots.txt file is only what a site asks crawlers to do. Cloudflare can also refuse a crawler at its edge, so on September 23 I checked whether these sites still turn AI crawlers away there. I requested the homepage of each of the 468 Common Crawl panel sites that now block none of the eight in robots.txt, once with a Chrome user agent and once each with the GPTBot, ClaudeBot and PerplexityBot user agents. 422 answered the Chrome request normally from Cloudflare. On 336 of those 422, GPTBot and ClaudeBot got a 403 while PerplexityBot got the page. I then tried OAI-SearchBot, ChatGPT-User and Claude-SearchBot on the 337 sites with this pattern (the 336 plus one not served by Cloudflare), and each of them got the page on at least 322.

As a control I took 273 Cloudflare sites from the same August crawl files that never served the managed block and blocked none of the eight. 17 of them showed the same pattern.

This matches the migration in Cloudflare's September 15 post, which says sites with the old Block AI bots setting were moved to Disallow AI Training. That setting blocks training-only crawlers such as GPTBot and ClaudeBot and leaves search crawlers allowed. On most of these sites the robots.txt lines went away and the edge refusal of training crawlers stayed, so the robots.txt file now shows less than the site actually blocks.

There are two limits. The crawler user agents came from my own IP address, so this only shows rules that match on the user agent. A rule that checks a crawler's real IP range would not show up. I also have no edge measurement from before September 15, so I cannot say whether any of these refusals are new. How the settings work now is in the guide to Cloudflare's AI crawler settings.

Across the top 5,000: robots.txt allows the AI search crawler, the server refuses it

On September 23 I ran a similar test across all 4,999 sites in my census, this time for the three AI search crawlers: OAI-SearchBot (ChatGPT search), PerplexityBot and Claude-SearchBot. 2,429 sites served their homepage to a normal browser request and had a live robots.txt that allowed at least one of the three. On a first pass, 119 of them refused a request that named one of those crawlers in its user agent.

A refusal on its own proves less than it seems to. 53 of the 119 also refused the same request when it named Googlebot, so they most likely turn away any request that claims to be a crawler without coming from that crawler's own network, and the real crawlers may well get through. 10 refused my request even with no crawler named in it. That leaves 56 sites, 2.3% of the 2,429, that refused a request naming an AI search crawler while serving the identical request naming Googlebot or naming no crawler at all. In every one of those cases the crawler's own published user agent string was refused too. 31 of the 56 are served by Cloudflare.

48
of 2,311 sites whose robots.txt allows PerplexityBot refuse it
20
of 2,400 whose robots.txt allows OAI-SearchBot refuse it
9
of 2,392 whose robots.txt allows Claude-SearchBot refuse it

I checked seven of the 56 by hand. envato.com, redfin.com and ticketmaster.com refuse PerplexityBot, morningstar.com refuses OAI-SearchBot, thehindu.com refuses Claude-SearchBot, and ipcc.ch refuses all three. toyota.com's robots.txt has its own group for OAI-SearchBot, under a comment that describes it as a verified AI search crawler, and its server refused a request naming OAI-SearchBot while it served one naming Googlebot.

All 56 look fully open to a check that reads robots.txt alone, and that includes my own census. The limits above apply here too. The test came from one IP address, so a site that lets the real crawler in from its own network and refuses the name from anywhere else looks the same as one that refuses the crawler outright. The word "verified" in toyota.com's comment may mean exactly that. The free checker now runs both controls, a request naming Googlebot and one naming no crawler, before it tells anyone that their server refuses an AI crawler.

Later that day I sent the 56 sites the same requests again, adding GPTBot and ClaudeBot, the two crawlers that collect training data. 53 still refused at least one of the AI search crawlers, and 48 of those 53 refused GPTBot and ClaudeBot as well. On most of these sites, then, the search crawler is turned away by a list of AI bots that covers the training crawlers too. An owner who set out to keep AI training off the site may have kept it out of AI search answers at the same time. Only three of the 53 served both training crawlers. One of them, uchicago.edu, refuses any request whose user agent contains the word SearchBot, so OAI-SearchBot and Claude-SearchBot get a 403 while GPTBot, ClaudeBot and Googlebot get the homepage, and its robots.txt allows every crawler. You can see it with curl -I -A SearchBot https://www.uchicago.edu/. The same one-IP limit applies to this test.

What this does not tell you

Apart from the user-agent tests above, this is a measurement of robots.txt files. It does not show the edge rules that check a crawler's IP range, which are part of AI Crawl Control, or what a site owner sees in the dashboard. Cloudflare's blog describes a prompt to review the change, and I cannot see whether owners were prompted or what they chose. What I can see is that the block stopped being served on almost every site I track.

It also means that anyone comparing robots.txt files from before and after September 15 will see dozens of well-known sites apparently deciding to let GPTBot, ClaudeBot and Google-Extended back in. Nearly all of those changes came from Cloudflare no longer serving the block. My change log and next census run keep this cluster out of the counts of site decisions for that reason.

What to do if this was your site

Open https://yourdomain.com/robots.txt in a browser, or run it through the checker. If the Managed Content block is gone and you still want those crawlers kept out, you have two routes. You can turn on Bot Preference Sync in the Cloudflare dashboard and confirm its marker shows up in the live file, or you can write your own User-agent and Disallow lines for each crawler, which keeps working whatever Cloudflare changes next. The user-agent token list has the exact name to write for each crawler, and the crawler directory says what each one does.

Before copying the old block back in, think about what each crawler does. GPTBot and ClaudeBot collect training data. OAI-SearchBot, Claude-SearchBot and PerplexityBot fetch pages for AI search answers and citations, and blocking those is what takes a site out of ChatGPT and Perplexity answers. Google-Extended controls Gemini training and does not affect AI Overviews. The old managed block never touched the search crawlers, and a hand-written replacement should leave them alone too unless you mean to be invisible in AI answers.

The data

The top-5,000 robots.txt figures above come from one re-fetch on September 22, 2026, compared with the September 7 census baseline (Tranco top 5,000, 2,771 sites with a readable robots.txt). Per-site results are below.

The Common Crawl sample, the 754-site panel and the edge test are published separately so anyone can check or reproduce them: per-site results, the exact Common Crawl files sampled, the control sites and the code are in this Zenodo record (CC BY 4.0).

All 54 sites readable on September 22
SiteSeptember 22, 2026
255md.comBlocks none of the eight
acronis.comBlocks none of the eight
animanch.comBlocks none of the eight
aps.orgBlocks none of the eight
ay267.comBlocks none of the eight
bblaa.comBlocks none of the eight
bef77.comBlocks none of the eight
bgtee.comBlocks none of the eight
bisp.gov.pkBlocks none of the eight
blender.orgBlocks none of the eight
crn77.comBlocks none of the eight
dogdrip.netBlocks none of the eight
findagrave.comBlocks none of the eight
global-e.comBlocks none of the eight
hai8g.comBlocks none of the eight
hugedomains.comBlocks none of the eight
intelbras.com.brBlocks none of the eight
kick.comBlocks none of the eight
kir2kos.netBlocks none of the eight
ko-fi.comBlocks none of the eight
loteriadehoy.comBlocks none of the eight
manhwaweb.comBlocks none of the eight
marinetraffic.comBlocks none of the eight
momon-ga.comBlocks none of the eight
moviesda33.comBlocks none of the eight
moviesda34.comBlocks none of the eight
moviezwap.landBlocks none of the eight
myinstants.comBlocks none of the eight
nya2.comBlocks none of the eight
omg10.comBlocks none of the eight
opensea.ioBlocks none of the eight
oregonstate.eduBlocks none of the eight
rexify.com.ngBlocks none of the eight
rm358.comBlocks none of the eight
shahvani.comBlocks none of the eight
si.eduBlocks none of the eight
toyhou.seBlocks none of the eight
vaticannews.vaBlocks none of the eight
vdy.toBlocks none of the eight
wiki.ggBlocks none of the eight
wnacg.comBlocks none of the eight
x-arxx.comBlocks none of the eight
xn--12c1ezaww.comBlocks none of the eight
xnxx.healthBlocks none of the eight
xxxhindi.toBlocks none of the eight
bpexch.liveOwn rules, blocks 8 of the eight
lectormangass.netOwn rules, blocks 5 of the eight
pussyspace.comOwn rules, blocks 4 of the eight
scan-manga.comOwn rules, blocks 8 of the eight
sportfilm800.comOwn rules, blocks 8 of the eight
twkan.comStill serves the Managed Content block
xnhau.bioStill serves the Managed Content block
xnhau.xxxStill serves the Managed Content block
xnhau.youStill serves the Managed Content block
Want to know if this changes?