Should I block AI crawlers on my website?
You may already be blocking them without knowing
This is the most common version of the problem, and it rarely results from a decision anyone made.
CDN and security products increasingly ship AI-crawler blocking enabled by default, and managed rule sets add user agents over time without notifying the site owner. The practical consequence is that a business can spend a year producing content while being invisible to the systems it hoped would cite it.
Checking robots.txt is not sufficient, because a CDN can block a crawler that robots.txt explicitly welcomes. Both layers must be checked separately, and the CDN layer is where the surprises live. We have audited agencies selling AI visibility that were blocking GPTBot on their own domains.
Understanding what each crawler does
The choice is not binary, because different user agents serve different functions. Blocking training crawlers while allowing retrieval crawlers is a coherent position; blocking everything rarely is.
| User agent | Operator | Function |
|---|---|---|
| GPTBot | OpenAI | Gathers content for model training |
| OAI-SearchBot | OpenAI | Powers ChatGPT search retrieval |
| ChatGPT-User | OpenAI | Fetches pages when a user prompt triggers browsing |
| ClaudeBot | Anthropic | Crawling for Claude |
| PerplexityBot | Perplexity | Indexing for Perplexity answers |
| Google-Extended | Controls Gemini and AI training use, not Search indexing | |
| Applebot-Extended | Apple | Controls Apple Intelligence training use |
| CCBot | Common Crawl | Open dataset used by many model builders |
Who should allow everything
Service businesses, professional firms, local businesses, SaaS companies, e-commerce brands, and anyone whose revenue depends on being found and recommended.
The logic is straightforward. Your content is not the product — it is marketing for the product. Being absorbed into a system that recommends vendors is distribution, not theft. A model that has never encountered your brand cannot name it, and there is no mechanism by which being absent from training data helps a business that wants to be discovered.
The exception worth noting is that allowing crawlers is necessary but not sufficient. Access gets you considered; corroboration gets you recommended.
Who has a genuine reason to be selective
Publishers whose content is the product, and who can license it. If subscriptions or licensing fees are your revenue, allowing free training-data absorption undercuts the asset directly. Several major publishers now negotiate paid licensing instead, which is a rational commercial position.
Sites with genuinely proprietary research or paywalled work products, where the content itself has standalone commercial value.
In both cases the sophisticated approach is selective rather than total: block training crawlers such as GPTBot and CCBot while permitting retrieval crawlers such as OAI-SearchBot and PerplexityBot, so you remain citable in live answers with attribution while not donating your archive to the next training run.
Declaring your policy publicly
Whichever choice you make, stating it explicitly is better than leaving it to inference. An llms.txt file describing your organization in machine-readable prose, plus a human-readable access policy page, removes ambiguity about intent.
It is also a small credibility signal. Very few sites do it, and it demonstrates that access was decided rather than defaulted. Ours is published at /ai-access.
Does blocking GPTBot remove me from ChatGPT entirely?
No. GPTBot gathers training data, while OAI-SearchBot and ChatGPT-User handle retrieval for live answers. Blocking GPTBot alone keeps you out of future training data while remaining citable in ChatGPT search results. Blocking all three removes you from both paths.
Does Google-Extended affect my Google Search rankings?
No. Google-Extended controls whether your content is used for Gemini and AI model training. It has no effect on Googlebot's crawling or on Search indexing and ranking. They are separate controls and can be set independently.
How do I check whether I am blocking AI crawlers?
Check two layers. First read your robots.txt for Disallow rules against the AI user agents. Second, check your CDN or security provider's bot management settings, since managed rules can block crawlers regardless of what robots.txt permits. Cloudflare's AI Crawler Control under Security and Bots is a common source of unintended blocks.
Tell us what you are building.
If the question is about discovery, reputation, a new site, or a market move, send the context. AIGNCI will tell you whether an Audit, a build, or a more focused engagement is the right starting point.