Every site in the top 5,000 that blocks an AI search crawler
368 of the 2,771 sites with a robots.txt block OAI-SearchBot, PerplexityBot or Claude-SearchBot, so they cannot be cited by that engine. The full list with each site's result, from a 2026-09-07 census.
The shorter list covers 32 names most people recognise. This is all of them. Of the 2,771 sites in the census that served a robots.txt, 368 block at least one of the three crawlers that feed AI answers, and 200 block all three. A red mark means that engine is not allowed to read the site root, so it cannot cite the site. Every domain links to its own live result.
Counted per crawler: OAI-SearchBot, which feeds ChatGPT citations, is blocked by 240 of them. PerplexityBot is blocked by 358. Claude-SearchBot is blocked by 249.
Two things worth knowing before you read anything into a row. Blocking a training crawler such as GPTBot is a separate decision and does not appear here, because it does not affect citations. And the Tranco list ranks raw hostnames, so it includes CDN and API endpoints nobody visits, such as gstatic.com and fbcdn.net. 2 of the 368 rows look like infrastructure rather than sites you would browse. They are real robots.txt results so they stay in the table, but they are not evidence about a publisher.
Measured on 2026-09-07 against each site's live robots.txt, using the same matcher Google publishes, so wildcards and dollar anchors are handled the way a real crawler handles them. The raw per-domain data is on the census page and in the open dataset.
Not on this list and want to be sure? Check your own site. If you are on it by accident, this guide covers how the mistake usually happens and how to undo it.