What is cohere-ai and should you block it?
cohere-ai is Cohere's training crawler. Blocking it opts your pages out of model training and costs you nothing in search, as long as the file is written correctly.
What cohere-ai does
cohere-ai is the token that appears in block lists for Cohere, a model company. Cohere has not published crawler documentation that describes a user agent by that name, so what it fetches and how often is not on record.
On robots.txt: cohere-ai behaviour is undocumented.
What blocking it costs
It costs nothing in search, if the file is written correctly. Blocking cohere-ai tells Cohere not to use your pages for training. It does not remove you from any search product.
The one way to lose search visibility while blocking cohere-ai is a catch-all group that disallows everything. A crawler that is not named in the file falls into the catch-all, and that is how sites block OAI-SearchBot and PerplexityBot without meaning to.
How many of the top 5,000 sites block it
In my September 2026 census of the Tranco top 5,000, 2,778 sites returned a robots.txt. 396 of them block cohere-ai (14.3%), which makes it the 10th most-blocked of the 20 crawlers I checked. 292 files mention it by name; the rest of the mentions either allow it explicitly or restrict only part of the site.
For comparison, 726 sites (26.1%) block at least one training crawler, against 393 (14.1%) that block at least one AI search crawler. Most sites that opt out of training are doing it deliberately; most sites that block AI search are not. The full census is here.
The robots.txt lines
To block it:
User-agent: cohere-ai Disallow: /
Keep the catch-all group permissive, or give the AI search crawlers their own groups, so this block does not take them down with it.
How to verify a request really came from cohere-ai
There is no published address list or documentation, so a request claiming to be cohere-ai cannot be verified. The token is in the census because it is in the block lists sites copy from.
No vendor documentation to link. Census: the September 2026 robots.txt census, raw data at /data/census.json.
All 20 crawlers. Check your own robots.txt against all of them with the checker.