robots.txt is the public, machine-readable file a website uses to tell automated crawlers which it wants to stay out. CipherCue reads it for every tracked organisation and records, per named AI crawler, whether the site declares a rule disallowing it. Browse by crawler to see which organisations opt each one out.
AI crawlers, by organisations observed blocking them
| GPTBot · OpenAI |
|
17,401 |
| Amazonbot · Amazon |
|
16,355 |
| ClaudeBot · Anthropic |
|
16,305 |
| Meta-ExternalAgent · Meta |
|
15,812 |
| Google-Extended · Google |
|
15,535 |
| anthropic-ai · Anthropic |
|
15,395 |
| Applebot-Extended · Apple |
|
15,222 |
| ChatGPT-User · OpenAI |
|
14,998 |
| cohere-ai · Cohere |
|
14,896 |
| FacebookBot · Meta |
|
14,796 |
| Claude-Web · Anthropic |
|
14,708 |
| PerplexityBot · Perplexity |
|
14,700 |
| OAI-SearchBot · OpenAI |
|
14,139 |
| DuckAssistBot · DuckDuckGo |
|
13,920 |
| Perplexity-User · Perplexity |
|
13,790 |
| Claude-SearchBot · Anthropic |
|
13,766 |
| MistralAI-User · Mistral |
|
13,742 |
| CCBot · Common Crawl |
|
5,503 |
| Bytespider · ByteDance |
|
5,069 |
Counts are organisations in the CipherCue directory observed disallowing each crawler. robots.txt is a stated preference, not a guarantee the crawler honours it.
We use privacy-friendly analytics (Matomo, self-hosted in the EU) to understand how visitors use the site. Nothing is tracked unless you agree. Privacy policy.