CIPHERCUE

← Directory

Which companies block AI crawlers in robots.txt?

robots.txt is the public, machine-readable file a website uses to tell automated crawlers which it wants to stay out. CipherCue reads it for every tracked organisation and records, per named AI crawler, whether the site declares a rule disallowing it. Browse by crawler to see which organisations opt each one out.

GPTBot · OpenAI
17,401
Amazonbot · Amazon
16,355
ClaudeBot · Anthropic
16,305
Meta-ExternalAgent · Meta
15,812
Google-Extended · Google
15,535
anthropic-ai · Anthropic
15,395
Applebot-Extended · Apple
15,222
ChatGPT-User · OpenAI
14,998
cohere-ai · Cohere
14,896
FacebookBot · Meta
14,796
Claude-Web · Anthropic
14,708
PerplexityBot · Perplexity
14,700
OAI-SearchBot · OpenAI
14,139
DuckAssistBot · DuckDuckGo
13,920
Perplexity-User · Perplexity
13,790
Claude-SearchBot · Anthropic
13,766
MistralAI-User · Mistral
13,742
CCBot · Common Crawl
5,503
Bytespider · ByteDance
5,069

Counts are organisations in the CipherCue directory observed disallowing each crawler. robots.txt is a stated preference, not a guarantee the crawler honours it.

How this is observed: CipherCue reads each organisation's public robots.txt and records the disallow rules it declares per named user-agent. See the AI-crawler methodology.