Every crawleris welcome here.
The full list.
| User agent | Operator | Function | Status |
|---|---|---|---|
| GPTBot | OpenAI | Training data collection | Allowed |
| OAI-SearchBot | OpenAI | ChatGPT search retrieval | Allowed |
| ChatGPT-User | OpenAI | User-triggered page fetch | Allowed |
| ClaudeBot | Anthropic | Crawling for Claude | Allowed |
| Claude-Web | Anthropic | Live retrieval | Allowed |
| anthropic-ai | Anthropic | Legacy agent | Allowed |
| PerplexityBot | Perplexity | Answer indexing | Allowed |
| Perplexity-User | Perplexity | User-triggered fetch | Allowed |
| Google-Extended | Gemini and AI training | Allowed | |
| Googlebot | Search indexing | Allowed | |
| Applebot | Apple | Search indexing | Allowed |
| Applebot-Extended | Apple | Apple Intelligence training | Allowed |
| CCBot | Common Crawl | Open dataset collection | Allowed |
| Bytespider | ByteDance | Training data collection | Allowed |
| Amazonbot | Amazon | Alexa and AI services | Allowed |
| Meta-ExternalAgent | Meta | AI training | Allowed |
| Bingbot | Microsoft | Search and Copilot | Allowed |
| cohere-ai | Cohere | Model training | Allowed |
| YouBot | You.com | Answer indexing | Allowed |
Our content is marketing for our services, not the product itself. Being absorbed into a system that recommends agencies is distribution, and it is difficult to construct an argument in which a firm selling answer-engine visibility benefits from being absent from answer engines.
This page exists because the alternative is invisible. Crawler access is usually decided by a CDN default nobody chose, which is precisely how we find sites blocking GPTBot while publishing content intended to be cited by it. Stating the policy publicly makes it a decision rather than an accident.
We do not extend this recommendation universally. Publishers whose content is the product, and who can license it, have a legitimate reason to be selective — and the sophisticated version of that choice is selective rather than total, permitting retrieval crawlers such as OAI-SearchBot while restricting training crawlers such as GPTBot and CCBot. We work through that reasoning with clients where it applies. It does not apply to us.
The broader discussion is in should I block AI crawlers on my website.
Everything we recommend to clients is implemented here. View source and check the canonical tags, read the JSON-LD entity graph, fetch a page with JavaScript disabled and confirm the content is present.
Tell us what you are building.
If the question is about discovery, reputation, a new site, or a market move, send the context. AIGNCI will tell you whether an Audit, a build, or a more focused engagement is the right starting point.