Most businesses should not treat “block AI crawlers” as one decision. OpenAI, Anthropic, and Google document separate controls for search visibility, user-triggered fetching, and model training. If your goal is to appear in AI answers while limiting training use, allow the search crawlers and decide on training bots separately. This article shows how to make that call without accidentally removing your site from ChatGPT search or confusing Google Search AI features with Gemini training controls.
Key takeaways
Search access and training access are different controls. OpenAI states that OAI-SearchBot and GPTBot settings are independent. Anthropic documents the same split across Claude-SearchBot, Claude-User, and ClaudeBot.
Google-Extended is not the Google Search AI Overviews switch. Google documents Google-Extended as a control for Gemini training and grounding in Gemini Apps and Vertex AI. It does not affect inclusion or ranking in Google Search.
Blanket AI blocks often create findability gaps. A CDN rule or robots.txt Disallow meant for scrapers can also block the crawlers that surface pages in assistant search.
Use a separation ladder before editing robots.txt. Decide what outcome you want—search citations, training opt-out, private-path protection, or full lockdown—then map each bot to that outcome.
Why “block AI” is the wrong starting question
Security teams hear that AI systems scrape the web. Marketing teams hear that AI search can send referral traffic and citations. Legal teams hear training-risk concerns. Those are three different problems. Merging them into one robots.txt rule usually solves none of them cleanly.
Google’s guidance for AI features in Search says AI Overviews and AI Mode rely on indexed, snippet-eligible pages and do not add separate technical requirements beyond Search fundamentals. That experience is governed through Google Search crawling and preview controls—not through a generic “AI bot” label.
Meanwhile, OpenAI’s crawler documentation says a site can allow OAI-SearchBot so it can appear in ChatGPT search results while disallowing GPTBot to signal that content should not be used for training foundation models. If you block every OpenAI agent together, you may opt out of search discovery while believing you only blocked training.
What each major control actually does
Before changing robots.txt, map the agents you are likely to see to the outcome they affect. Keep classic Googlebot available if you want Google Search—including AI Overviews and AI Mode eligibility—to remain intact.
Control | Primary documented job | What blocking typically trades away |
Googlebot | Crawling for Google Search and related Search surfaces | Search visibility overall, including eligibility for AI features that require indexed, snippet-eligible pages |
Google-Extended | Standalone robots.txt token for Gemini model training and grounding in Gemini Apps / Vertex AI grounding with Google Search | Use of crawled content for those Gemini training and grounding products—not Google Search inclusion or ranking |
OAI-SearchBot | Surfacing websites in ChatGPT search features | Appearance in ChatGPT search answers; OpenAI notes navigational links may still appear in some cases |
GPTBot | Crawling that may be used for training OpenAI generative AI foundation models | Training-related crawl access; does not, by itself, manage ChatGPT Search opt-out |
ChatGPT-User | User-initiated fetches in ChatGPT / Custom GPTs | OpenAI states robots.txt rules may not apply because the action is user-initiated; Search opt-out should use OAI-SearchBot |
Claude-SearchBot | Indexing content to improve Claude search result quality | Anthropic says disabling it may reduce visibility and accuracy in user search results |
ClaudeBot | Collecting web content that could contribute to model training | A signal that future materials should be excluded from Anthropic training datasets |
Claude-User | Fetching pages when Claude users ask questions | Anthropic says disabling it may reduce visibility for user-directed web search |
Sources for the table: Google’s common crawlers, OpenAI crawlers, and Anthropic’s crawler help article.
The AI Crawler Separation Ladder
Use this ladder to choose a policy posture before you edit robots.txt or CDN bot rules. Climb only as far as your risk and visibility goals require.
Level | Policy posture | When it fits |
L0 — Accidental blanket block | Wildcard bot challenges, aggressive WAF rules, or one Disallow that catches search and training agents together | Almost never intentional. Audit and unwind if AI search visibility matters. |
L1 — Separate training from search | Allow search crawlers; disallow selected training tokens such as GPTBot, ClaudeBot, and optionally Google-Extended | Default for many SMEs that want citations without treating all AI use as identical |
L2 — Explicit search allowlist | Named Allow groups for OAI-SearchBot and Claude-SearchBot, plus intact Googlebot, with training decisions documented separately | Teams that previously blocked broadly and need a recoverable, auditable policy |
L3 — Path-level protection | Public marketing pages crawlable; private paths (/account, /api, draft tools, customer data areas) disallowed across agents | SaaS and ecommerce sites with public content and sensitive authenticated areas |
L4 — Visibility lockdown | Search crawlers disallowed because legal, contractual, or brand policy requires minimizing AI retrieval | Use only when leadership accepts reduced AI-answer visibility as a deliberate cost |
Most growth-focused sites land at L1 or L2. L4 is a business decision, not a default security setting. If you already decided to stay in Google’s Search generative AI features, do not confuse that Search Console choice with Google-Extended; they control different products. See also our earlier decision guide on keeping or opting out of Google generative AI Search features.
Visibility Tradeoff Scorecard
Score each factor from 0 to 2. Total the points to choose a starting posture. Recheck after any CDN, hosting, or robots.txt change.
Factor | 0 points | 1 point | 2 points |
Buyer discovery in AI search | AI answers are irrelevant to demand | Useful for research, not purchase paths | Category and comparison questions influence shortlists |
Training sensitivity | Public commodity content; low concern | Mixed public/IP content | Proprietary methods, gated research, or strict content licensing stance |
Policy clarity | No owner for robots.txt / bot rules | Marketing and IT disagree | Named owner and written crawl policy |
Current access evidence | Unknown; logs and robots never reviewed | robots.txt reviewed; CDN rules unclear | robots.txt, CDN/WAF, and key paths verified |
Surface priority | No AI surfaces matter yet | One surface matters (for example ChatGPT or Google AI features) | Multiple surfaces matter and conflict if treated as one switch |
Compliance constraint | No contractual AI restriction | Industry caution without hard rule | Legal/compliance requires limited retrieval or training exposure |
How to interpret the score: 0–4 usually means fix accidental blocks and keep Googlebot healthy before specializing. 5–8 typically supports L1 or L2: allow search crawlers, decide training tokens deliberately. 9–12 supports a documented split with path-level protection (L3). If compliance alone forces a lockdown, treat L4 as an executive decision and measure the visibility cost.
Worked example: the accidental ChatGPT opt-out
Consider a hypothetical regional professional-services firm. Organic rankings are stable. Leadership asks ChatGPT for “best [category] firms in [city]” and sees competitors. An IT ticket from months earlier added a bot-management rule labeled “block AI scrapers.” The rule challenged several AI user agents, including OAI-SearchBot.
A separation review would show:
Googlebot remained allowed, so classic SEO looked healthy.
OAI-SearchBot was effectively blocked, so ChatGPT search retrieval was impaired.
GPTBot was also blocked—which may have been the intended training preference—but was not separated from search access.
No owner had documented which AI surfaces the business wanted to participate in.
The fix is not a content rewrite first. Restore intentional search access, keep or refine the training opt-out, verify the live robots.txt group matching, and only then retest category prompts. That sequence matches the broader findability logic in why rankings can coexist with AI invisibility.
Implementation guidance: change robots.txt without creating new gaps
Inventory outcomes first. Write the desired state for Google Search AI features, ChatGPT search, Claude search, and training opt-outs as separate lines.
Read the live robots.txt for every hostname. Subdomains can differ. Anthropic notes that opt-outs must be set per subdomain you want excluded.
Create named groups, do not rely on folklore. A GPTBot Disallow does not control OAI-SearchBot. A Google-Extended Disallow does not remove Google Search AI feature eligibility.
Keep private paths private. Disallow authenticated, account, API, and draft paths across agents even when public pages are allowed.
Check CDN and WAF rules separately. robots.txt permission is useless if edge security challenges the same agents.
Prefer robots.txt over IP blocks for opt-outs. Anthropic warns that blocking crawler IPs can impede reading robots.txt and may not reliably express the intended opt-out.
Allow time and re-verify. OpenAI notes search systems can take about 24 hours to adjust after a robots.txt update. Google recrawl and processing can take longer depending on refresh priority.
Track referrals where platforms support it. After restoring access, review analytics for assistant referral patterns and branded prompt probes so marketing and IT share one evidence set.
A practical L1 starting pattern
The following is a hypothetical public-site pattern for teams that want search discovery while signaling training opt-outs. Adapt paths to your site. Do not copy it into production without reviewing private routes and legal requirements.
Allow Googlebot for Search. Leave Google-Extended as an explicit business choice. Allow OAI-SearchBot and Claude-SearchBot if those surfaces matter. Disallow GPTBot and ClaudeBot if training opt-out is the goal. Review Claude-User separately because Anthropic treats it as a visibility-affecting user-fetch control, while OpenAI notes ChatGPT-User may not follow robots.txt the same way.
Common failure patterns
One “AI bots” firewall rule. Collapses search and training into a single deny.
Assuming Google-Extended opts you out of AI Overviews. Google documents that Google-Extended does not impact Google Search inclusion or ranking.
Editing robots.txt but ignoring the CDN. Permission on disk, denial at the edge.
Disallowing Googlebot to “stop AI.” That damages Search broadly, including the AI features that depend on Search eligibility.
Confusing user-initiated fetchers with index crawlers. OpenAI separates ChatGPT-User from OAI-SearchBot; Anthropic documents Claude-User as its own control.
No retest after the change. Policy edits without prompt probes, log checks, or analytics review leave the business guessing.
Limitations and counterarguments
Allowing search crawlers does not guarantee citations. Google is explicit that eligibility is not a promise of inclusion, and AI Overviews do not trigger on every query. OpenAI and Anthropic document access controls, not ranking guarantees. Prompt tests are directional evidence, not a complete measurement system.
Training opt-outs are also imperfect signals. They express preference through published robots.txt controls; they are not a substitute for legal review of licensing, confidential content handling, or contractual restrictions. If your strongest pages contain proprietary methods you do not want broadly retrieved, path-level disallow or keeping that material offline may matter more than a homepage training token.
Finally, surface priority still matters. If ChatGPT is unimportant to your buyers and Google AI features dominate, spend less time on OAI-SearchBot nuance and more on Search eligibility, extractable source pages, and Search Console evidence. For prioritization across surfaces, see which AI answer surface to prioritize first.
Recommended next steps
Write a one-page crawl policy that separates search visibility from training preferences.
Audit robots.txt and CDN/WAF rules for Googlebot, Google-Extended, OAI-SearchBot, GPTBot, Claude-SearchBot, ClaudeBot, and Claude-User.
Restore intentional search access if you find an accidental blanket block.
Protect private paths, then retest category prompts and referral analytics after the change window.
If you want help turning crawler policy, source-page quality, and AI-answer measurement into one operating plan, review Oasbit’s generative engine optimization services. When the blockers are clear and you need a practical roadmap, book a growth strategy session.




