What each crawler is for, what blocking it costs, how many of the top 5,000 sites block it, the exact robots.txt lines, and how to verify it. Numbers from the September 2026 census of 2,778 sites with a robots.txt.
The short version: the three AI search crawlers decide whether you get cited, the on-demand fetchers act for one person at a time, and the training crawlers can be blocked at no cost to search. Most sites that are missing from AI answers blocked a search crawler by accident while trying to block a training one.
AI search crawlers (these decide whether you get cited)
Each page quotes the same census: 4999 of the Tranco top 5,000 fetched, 2,778 with a robots.txt, a crawler counted as blocked when the most specific group that applies to it disallows the site root. Method, tables and the raw JSON are on the data page. Vendor claims on each page link to the vendor's own documentation.