What is meta-externalagent and should you block it?

meta-externalagent is Meta's training crawler. Blocking it opts your pages out of model training and costs you nothing in search, as long as the file is written correctly.

What meta-externalagent does

Meta's crawler documentation lists Meta-ExternalAgent as the crawler for use cases such as training AI models or improving products by indexing content directly. Meta-ExternalFetcher, a separate agent, handles fetches a person triggers.

On robots.txt: meta-externalagent respects robots.txt, per Meta's crawler documentation.

What blocking it costs

It costs nothing in search, if the file is written correctly. Blocking meta-externalagent tells Meta not to use your pages for training. It does not remove you from any search product.

The one way to lose search visibility while blocking meta-externalagent is a catch-all group that disallows everything. A crawler that is not named in the file falls into the catch-all, and that is how sites block OAI-SearchBot and PerplexityBot without meaning to.

How many of the top 5,000 sites block it

In my September 2026 census of the Tranco top 5,000, 2,778 sites returned a robots.txt. 502 of them block meta-externalagent (18.1%), which makes it the 5th most-blocked of the 20 crawlers I checked. 415 files mention it by name; the rest of the mentions either allow it explicitly or restrict only part of the site.

For comparison, 726 sites (26.1%) block at least one training crawler, against 393 (14.1%) that block at least one AI search crawler. Most sites that opt out of training are doing it deliberately; most sites that block AI search are not. The full census is here.

The robots.txt lines

To block it:

User-agent: meta-externalagent
Disallow: /

Keep the catch-all group permissive, or give the AI search crawlers their own groups, so this block does not take them down with it.

How to verify a request really came from meta-externalagent

Meta's documentation describes how to verify its crawlers; it publishes address ranges through its developer site rather than a single JSON file.

Source: Meta's crawler documentation. Census: the September 2026 robots.txt census, raw data at /data/census.json.

All 20 crawlers. Check your own robots.txt against all of them with the checker.