What is CCBot and should you block it?
CCBot is Common Crawl's training crawler. Blocking it opts your pages out of model training and costs you nothing in search, as long as the file is written correctly.
What CCBot does
Common Crawl is a non-profit that publishes an open archive of the web. CCBot is its crawler. The archive is free to anyone, and it is well known that many published training datasets were built from it, which is why it sits on most AI block lists even though Common Crawl itself trains nothing.
On robots.txt: CCBot respects robots.txt. Common Crawl's own instruction for opting out is the two-line block below.
What blocking it costs
It costs nothing in search, if the file is written correctly. Blocking CCBot tells Common Crawl not to use your pages for training. It does not remove you from any search product.
The one way to lose search visibility while blocking CCBot is a catch-all group that disallows everything. A crawler that is not named in the file falls into the catch-all, and that is how sites block OAI-SearchBot and PerplexityBot without meaning to.
How many of the top 5,000 sites block it
In my September 2026 census of the Tranco top 5,000, 2,778 sites returned a robots.txt. 596 of them block CCBot (21.5%), which makes it the 1st most-blocked of the 20 crawlers I checked. 543 files mention it by name; the rest of the mentions either allow it explicitly or restrict only part of the site.
For comparison, 726 sites (26.1%) block at least one training crawler, against 393 (14.1%) that block at least one AI search crawler. Most sites that opt out of training are doing it deliberately; most sites that block AI search are not. The full census is here.
The robots.txt lines
To block it:
User-agent: CCBot Disallow: /
Keep the catch-all group permissive, or give the AI search crawlers their own groups, so this block does not take them down with it.
How to verify a request really came from CCBot
Common Crawl publishes no address list and warns that some crawlers falsely identify as CCBot. Its advice is a reverse DNS lookup on the requesting address.
Source: Common Crawl's crawler documentation. Census: the September 2026 robots.txt census, raw data at /data/census.json.
All 20 crawlers. Check your own robots.txt against all of them with the checker.