Cloudflare AI Crawl Control: what it costs and what it blocks on real sites

AI Crawl Control is the part of Cloudflare's dashboard that shows which AI crawlers visit a site and lets the owner block them one at a time. It comes with every plan, including Free. I tested the top 5,000 sites from outside on September 23, 2026. Of the 760 that run on Cloudflare and allow the AI search crawlers in robots.txt, 31 still refuse at least one of those crawlers at the server, and 27 of the 31 refuse PerplexityBot.

Want to know if YOUR site's AI visibility changes?

Enter your domain and email. One email if a crawler you rely on gets newly blocked, or an existing block goes away. No spam.

What it is

AI Crawl Control used to be called AI Audit. Cloudflare launched AI Audit in September 2024 and renamed it AI Crawl Control when it left beta in August 2025. It lists the AI crawlers that fetch a site's pages, and the Crawlers tab has an Action column where each one can be set to allow or block. A Directives tab shows crawlers that asked for paths the site's robots.txt disallows. On some plans it can also charge crawlers through pay per crawl.

A block here is a rule at Cloudflare's network edge, so it applies whether or not the crawler reads robots.txt. According to Cloudflare's docs, the blocked crawler gets 403 Forbidden, or 402 Payment Required if the site picked that response to say the crawler has to pay.

What it costs

Cloudflare doesn't sell it separately. The getting-started page lists what every plan gets: AI crawler detection via user agent strings, an analytics window of at most 24 hours, and allow/block controls. Enterprise plans with Bot Management add detection through Bot Management's detection IDs, longer analytics windows, and pay per crawl. A custom response code and message is limited to paid plans, and so are referral counts. Pay per crawl is still invite-only. Cloudflare's own pages call it a closed beta in one place and a private beta in another, and interested sites apply through a signup form.

The detection method matters for anyone checking a site from outside. Below Enterprise, AI Crawl Control recognizes a crawler by the name in its user agent. A test request that names a crawler should therefore meet the same rule the crawler itself meets, and that is what makes the measurement further down possible.

GoDaddy and AI Crawl Control

In April 2026 GoDaddy and Cloudflare announced that GoDaddy would build AI Crawl Control into its hosting, so site owners could allow crawlers, block them, or signal that payment is required. In GoDaddy's Managed WordPress the feature is called AI Bot Crawler Activity. A June GoDaddy post describes it as a beta with one-click allow and block controls for each crawler. If you host with GoDaddy and are looking for an AI crawl control bar, that feature is the likely place. I couldn't find a GoDaddy help page that uses the word "bar", so I can't confirm what that screen looks like.

How it fits with the Search, Agent and Training settings

Since July 1, 2026 every Cloudflare plan also has three broader settings, Search, Agent and Training, each of which can allow a category, block it everywhere or block it on pages with ads (changelog). On September 15 Training gained a fourth option, Disallow AI Training. It blocks training-only crawlers such as GPTBot and ClaudeBot at the edge, keeps Googlebot, Bingbot and Applebot crawling for search, and writes a no-training preference for those three into robots.txt (Cloudflare's post). Those settings act on whole categories. AI Crawl Control is where a site sets one crawler differently. The settings guide covers the categories and the defaults in detail.

Cloudflare's crawler reference files GPTBot and ClaudeBot under AI Crawler, which means training. OAI-SearchBot, Claude-SearchBot and PerplexityBot are under AI Search. ChatGPT-User, Claude-User and Perplexity-User are under AI Assistant, the crawlers that fetch a page because a person asked for it just then.

What it blocks on real sites

On September 23, 2026 I requested the homepage of every site in the top 5,000 census, once as a normal browser and once each with a user agent naming OAI-SearchBot, PerplexityBot or Claude-SearchBot (the user agent also names this checker, so nobody mistakes it for the real crawler). 2,429 sites served the browser request and allow at least one of the three crawlers in robots.txt. A 401, a 403 or a Cloudflare challenge counted as a refusal only if the same site also served two control requests, one naming no crawler and one naming Googlebot. That leaves out sites that turn away every unverified crawler claim. The full method, the limits and all 56 sites are in the edge section of the Cloudflare data page.

I split the sites by the server header on the browser response. A site behind Cloudflare that rewrites that header lands in the other column.

Top-5,000 sites that allow the crawler in robots.txtOn CloudflareOther servers
Sites tested7601,669
Server refuses at least one AI search crawler31 (4.1%)25 (1.5%)
PerplexityBot refused27 of 728 (3.7%)21 of 1,583 (1.3%)
OAI-SearchBot refused7 of 755 (0.9%)13 of 1,645 (0.8%)
Claude-SearchBot refused8 of 749 (1.1%)1 of 1,643 (0.1%)

A Cloudflare site in this set was about 2.7 times as likely as the rest to refuse an AI search crawler that its own robots.txt allows, and PerplexityBot accounts for most of the gap. OpenAI's search crawler was refused at about the same low rate on and off Cloudflare. The server header shows a site is on Cloudflare. It can't show which setting did the refusing, whether that was AI Crawl Control, one of the category settings or a custom firewall rule.

I re-checked five of the Cloudflare sites on September 23: britannica.com, merriam-webster.com, teacherspayteachers.com, envato.com and laracasts.com. The robots.txt of each allows PerplexityBot, OAI-SearchBot and Claude-SearchBot. Each server returned 403 to requests naming PerplexityBot, GPTBot or ClaudeBot, and served the page to requests naming OAI-SearchBot, Claude-SearchBot, Googlebot or no crawler at all. A training-crawler block plus a single-crawler block on PerplexityBot produces exactly that result.

There is a likely reason PerplexityBot stands out. In August 2025 Cloudflare published evidence that Perplexity was also crawling under undeclared user agents to get past blocks, and said it had "de-listed them as a verified bot and added heuristics to our managed rules that block this stealth crawling". The data can't show why any one site blocked PerplexityBot. A site owner who read that post and then blocked PerplexityBot in AI Crawl Control while leaving the other search crawlers alone would produce the pattern above. It also means the site's robots.txt still invites a crawler its server turns away, which is worth fixing in one place or the other.

A second measurement shows the training-only block working as Cloudflare describes it. Of 422 Cloudflare sites that served Cloudflare's Managed robots.txt block in August, 336 now refuse GPTBot and ClaudeBot at the server while serving PerplexityBot, OAI-SearchBot, ChatGPT-User and Claude-SearchBot. In a comparison group of 273 Cloudflare sites from the same August sample that never used the managed block, 17 do (details).

Does Cloudflare block ChatGPT?

On the sites I tested, Cloudflare almost never blocked ChatGPT's search crawler. OAI-SearchBot was refused at the server by 7 of 755 eligible Cloudflare sites in the top 5,000, about the same rate as everywhere else. Of the 337 sites in the second measurement that refuse GPTBot and ClaudeBot (336 on Cloudflare and one elsewhere), 322 served OAI-SearchBot and 324 served ChatGPT-User.

What Cloudflare blocks by default is narrower. Since September 15, a new domain that says it runs ads starts with Training on Disallow AI Training, which blocks GPTBot, OpenAI's training crawler, and Agent on block on pages with ads. Cloudflare files ChatGPT-User under AI Assistant, the category closest to Agent, so on those sites it may be turned away from pages that carry ads. A new domain without ads starts with everything allowed, and existing domains kept the settings they already had. ChatGPT search uses OAI-SearchBot, so a GPTBot block on its own keeps a site out of OpenAI's training data and leaves it available to ChatGPT's answers.

Checking your own site

Related: Cloudflare's AI crawler settings since September 15, the Managed robots.txt retirement data, and how to show up in Perplexity.

Want to know if this changes?