What changed: AI crawler policy shifts

1132 sites (every current blocker, the search/training asymmetry set, and the Tranco top 500) re-crawled and diffed against the 2026-09-07 baseline. This page tracks the real drift: sites that newly blocked or newly opened to an AI-search crawler, and larger training-crawler policy shifts. Updated when the census is re-run, not on a fixed schedule yet.

5
sites newly block an AI-search crawler (PerplexityBot, Claude-SearchBot or OAI-SearchBot) since 2026-09-07
7
sites newly opened to an AI-search crawler in the same window
33
sites added Applebot-Extended alone, including netflix.com, pinterest.com, imdb.com, lnkd.in, redditmedia.com, toyota.com, wsj.com and marketwatch.com

Newly blocked from AI search

These sites allowed at least one AI-search crawler as of 2026-09-07 and now block it. If you rely on being cited by ChatGPT search, Perplexity or Claude, a change like this on your own site is exactly what $19/mo monitoring would have caught the day it happened.

RankSiteNewly blocks
1441pussyspace.comPerplexityBot Monitor this site →
1459hdtube.pornClaude-SearchBot Monitor this site →
1647sportfilm800.comClaude-SearchBot, OAI-SearchBot, PerplexityBot Monitor this site →
1725lectormangass.netPerplexityBot Monitor this site →
1739bpexch.liveClaude-SearchBot, OAI-SearchBot, PerplexityBot Monitor this site →

Newly opened to AI search

The opposite direction: these sites blocked an AI-search crawler as of 2026-09-07 and removed the block, becoming newly citable.

RankSiteNewly allows
305weather.comPerplexityBot Monitor this site →
456sohu.comClaude-SearchBot, OAI-SearchBot, PerplexityBot Monitor this site →
1230gazzetta.itPerplexityBot Monitor this site →
1564ladies.deClaude-SearchBot, OAI-SearchBot, PerplexityBot Monitor this site →
2529glassdoor.comPerplexityBot Monitor this site →
3222flipboard.comClaude-SearchBot, PerplexityBot Monitor this site →
4057kompas.comClaude-SearchBot Monitor this site →

Training-crawler policy shift: Applebot-Extended

33 sites added Applebot-Extended, Apple's AI-training crawler, to their robots.txt in this same window and changed nothing else -- a deliberate policy decision, not template churn (see method below). Applebot-Extended controls AI-training opt-out only; it does not affect whether these sites can still be cited, the way the search-crawler tables above do.

RankSite
41netflix.com Monitor this site →
57pinterest.com Monitor this site →
181flickr.com Monitor this site →
254meraki.com Monitor this site →
256imdb.com Monitor this site →
361wsj.com Monitor this site →
363netflix.net Monitor this site →
429nflxvideo.net Monitor this site →
700pinimg.com Monitor this site →
814academia.edu Monitor this site →
833theverge.com Monitor this site →
1270amap.com Monitor this site →
1381ilmeteo.it Monitor this site →
1879marketwatch.com Monitor this site →
2052lnkd.in Monitor this site →
2138thesun.co.uk Monitor this site →
2289vox.com Monitor this site →
2452redditmedia.com Monitor this site →
2539news.com.au Monitor this site →
2774instacart.com Monitor this site →
2925thetimes.co.uk Monitor this site →
2955theregister.com Monitor this site →
3212sas.com Monitor this site →
3285thetimes.com Monitor this site →
3349nymag.com Monitor this site →
3729realestate.com.au Monitor this site →
3819toyota.com Monitor this site →
3847wordreference.com Monitor this site →
4568barrons.com Monitor this site →
4611theregister.co.uk Monitor this site →
4700boxofficemojo.com Monitor this site →
4713turbify.com Monitor this site →
3587ebsco.com Monitor this site →

Everything else

17 smaller or mixed single-crawler changes in the same window (a training crawler added or removed here, a live-fetch agent there) did not meet either bar above. The full list, plus every field behind the tables on this page, is in the raw data.

Get the data

The full event list, including "everything else," is at /changes.json (CC BY 4.0, CORS-open). Subscribe to /changes.xml for the same events as an RSS feed. Method: re-crawl of the high-signal slice described above using the same crawler this site's AI crawler census uses, diffed field-by-field against the named baseline; cdn_template_churn (shared hosting-template noise) is classified and excluded before anything reaches this page.