The 296 sites that block AI search but not Google: deliberate or accidental?

A reader's critique of my own "296 of 369" number, tested against the census data by rank.

Want to know if YOUR site's AI visibility changes?

Enter your domain and email. One email if a crawler you rely on gets newly blocked, or an existing block goes away. No spam.

The critique

Mike, who runs viewfy.ai, left a comment on the Dev.to post that started this project: my "296 of 369" number, sites that block at least one AI search crawler while still allowing Googlebot, treats those 296 as one population. His point, using Facebook and Amazon as examples: a large platform blocking ChatGPT and Perplexity is very plausibly a deliberate policy decision, not an accident of the same catch-all rule that sweeps up a small business's WordPress site. Calling the whole group "usually by accident," which is what several pages on this site said until now, glosses over that difference. He is right that the number needs the split. Here it is.

The split, by rank

On 2026-09-07 I requested robots.txt from the Tranco top 5,000 sites and recorded, per site, whether it blocks any of OAI-SearchBot, PerplexityBot or Claude-SearchBot, and separately whether it blocks Googlebot. 369 sites block at least one of the three AI search crawlers. Split by Tranco rank:

Rank bandBlocks an AI search crawler...and still allows GooglebotShare
1-100151386.7%
101-1,000826882.9%
1,001-5,00027221579.0%
All 36936929680.2%

The gradient runs the direction Mike's critique would predict, but it is modest, not sharp: 86.7% of the top-100 blockers still allow Googlebot, against 79.0% below rank 1,000, an 8-point spread. The other side of that same split, sites that block Googlebot along with an AI search crawler, the pattern that looks least like an accident, runs from 13.3% at the top to 21.0% in the long tail. Rank correlates with something here, but not enough to sort the 296 cleanly into two bins.

The household names in the top tiers

What rank alone does not show is which specific sites make up that top-100 count. All 13 of the top-100 sites in the 296 are names anyone would recognize: facebook.com, instagram.com, twitter.com, amazon.com, whatsapp.net, x.com, tiktok.com, whatsapp.com, pinterest.com, yahoo.com, msn.com, chatgpt.com and baidu.com. Ranks 101-1,000 add nytimes.com, cnn.com, theguardian.com, forbes.com, bbc.com, bbc.co.uk, ebay.com, imdb.com, reuters.com and dozens more, mostly media publishers and large platforms. Every one of these has the legal and engineering resources to write a robots.txt group on purpose, and several, the news publishers especially, are exactly the kind of site with a live commercial interest in whether AI answers reuse their reporting without a visit. Calling this group's block an accident asks more of coincidence than it should.

Why rank alone is not a clean answer

The one time this whole project received a direct answer from a site owner complicates a simple "big sites decide, small sites don't" story. In September, timeanddate.com, ranked 2,109, well outside the top 1,000, replied to an outreach email: "We have intentionally blocked Perplexity." That is a confirmed deliberate block from the part of the population a rank cutoff would have called the accidental long tail. Robots.txt intent does not sort neatly by traffic rank; it sorts, imperfectly, by whether a site has looked at the question at all, and plenty of sites far outside the top 1,000 have.

What this changes

The honest fix is not a cleaner cutoff, since one does not exist in this data. It is dropping the blanket framing that treated all 296 the same way. Recognizability, whether a domain is a platform or publisher with an obvious reason to have a content policy, is a better signal than raw rank, though it still is not proof for any single domain: my census reads the contents of a robots.txt file, and no field in that file records who wrote a rule or why. The pages on this site that cited the 296 figure now link here instead of asserting an intent this data cannot establish.

Raised by Mike of viewfy.ai in a comment on the original Dev.to article. The raw per-site data behind this split is the same dataset as everywhere else on this site: the census page, raw JSON at /data/census.json, CC BY 4.0.

Related: which well-known sites block AI search, how to show up in Perplexity, should you block GPTBot.

Want to know if this changes?