What is Diffbot and should you block it?

Diffbot is Diffbot's training crawler. Blocking it opts your pages out of model training and costs you nothing in search, as long as the file is written correctly.

What Diffbot does

Diffbot is a commercial web data company. Its crawler feeds its knowledge graph and runs crawls on behalf of customers, and the extracted data is sold, including to companies building AI products.

On robots.txt: Diffbot honours robots.txt according to Diffbot's documentation.

What blocking it costs

It costs nothing in search, if the file is written correctly. Blocking Diffbot tells Diffbot not to use your pages for training. It does not remove you from any search product.

The one way to lose search visibility while blocking Diffbot is a catch-all group that disallows everything. A crawler that is not named in the file falls into the catch-all, and that is how sites block OAI-SearchBot and PerplexityBot without meaning to.

How many of the top 5,000 sites block it

In my September 2026 census of the Tranco top 5,000, 2,778 sites returned a robots.txt. 395 of them block Diffbot (14.2%), which makes it the 11th most-blocked of the 20 crawlers I checked. 266 files mention it by name; the rest of the mentions either allow it explicitly or restrict only part of the site.

For comparison, 726 sites (26.1%) block at least one training crawler, against 393 (14.1%) that block at least one AI search crawler. Most sites that opt out of training are doing it deliberately; most sites that block AI search are not. The full census is here.

The robots.txt lines

To block it:

User-agent: Diffbot
Disallow: /

Keep the catch-all group permissive, or give the AI search crawlers their own groups, so this block does not take them down with it.

How to verify a request really came from Diffbot

Diffbot does not publish a single address list. Verification is by the user-agent string and Diffbot's documentation.

Source: Diffbot's crawler documentation. Census: the September 2026 robots.txt census, raw data at /data/census.json.

All 20 crawlers. Check your own robots.txt against all of them with the checker.