Short answer
To control AI crawlers with robots.txt, add a User-agent group for each bot, followed by Allow: / or Disallow: /. To block AI training but stay in AI answers, disallow GPTBot, ClaudeBot, Google-Extended, Applebot-Extended, meta-externalagent and CCBot, and allow OAI-SearchBot, Claude-SearchBot and PerplexityBot. Never disallow Googlebot or Bingbot.
Key takeaways
- robots.txt controls AI crawlers bot by bot. Training, search and user-fetch bots have separate names, so you can treat them differently.
- A crawler that has its own group ignores the
User-agent: *group entirely. This is the most common robots.txt mistake. - Google-Extended and Applebot-Extended only opt you out of AI training. They don't affect Google Search, AI Overviews or Siri.
- If robots.txt returns a server error, crawlers may treat your whole site as off-limits. Check that it returns a 200 status.
- robots.txt is only one layer. CDN and firewall rules can still block crawlers it allows.
In this article
- How robots.txt works for AI crawlers
- Template 1: allow all AI crawlers (recommended for most B2B sites)
- Template 2: allow AI search, block AI training
- Template 3: block all AI crawlers but keep search engines
- Template 4: block AI training on specific folders only
- The rules that trip people up
- How to test your robots.txt
- Frequently asked questions
How robots.txt works for AI crawlers
robots.txt is a plain-text file at the root of your host (https://www.example.com/robots.txt) that tells crawlers which paths they may fetch. Its rules are standardized in RFC 9309, and the major AI companies say their main crawlers follow it.
The file is made of groups. Each group starts with one or more User-agent lines naming crawlers, followed by Allow and Disallow rules:
User-agent: GPTBot
User-agent: ClaudeBot
Disallow: /
User-agent: *
Allow: /
Before you copy a template, decide what you want from each type of crawler. Our AI crawler field guide explains all 15. In short:
- Search crawlers (OAI-SearchBot, Claude-SearchBot, PerplexityBot, Googlebot, Bingbot, Applebot) decide whether you can appear in AI answers. Blocking them makes you invisible.
- User fetchers (ChatGPT-User, Claude-User, Perplexity-User) open pages when a user asks about them. Several ignore robots.txt anyway.
- Training crawlers and tokens (GPTBot, ClaudeBot, Google-Extended, Applebot-Extended, meta-externalagent, CCBot) decide whether your content is used to train models. Blocking them doesn't remove you from AI search.
Template 1: allow all AI crawlers (recommended for most B2B sites)
Use this if you want to be found and recommended everywhere. Technically, User-agent: * with Allow: / is enough. Naming the AI crawlers makes your intent explicit and protects them if someone later adds a broad Disallow to the * group.
robots.txt: allow all AI crawlers
# Search engines and AI assistants: all welcome
User-agent: Googlebot
User-agent: Bingbot
User-agent: Applebot
User-agent: OAI-SearchBot
User-agent: ChatGPT-User
User-agent: GPTBot
User-agent: Claude-SearchBot
User-agent: Claude-User
User-agent: ClaudeBot
User-agent: PerplexityBot
User-agent: Perplexity-User
User-agent: Google-Extended
User-agent: Applebot-Extended
User-agent: meta-externalagent
User-agent: CCBot
Disallow: /admin/
Allow: /
User-agent: *
Disallow: /admin/
Allow: /
Sitemap: https://www.example.com/sitemap.xml
Notice that Disallow: /admin/ appears twice. That's deliberate, and the common mistakes section explains why.
Template 2: allow AI search, block AI training
Use this if you want to appear in ChatGPT, Claude, Perplexity, Google and Copilot answers but don't want your content used to train future models. It's the most common choice for companies with proprietary content.
robots.txt: allow AI search, block AI training
# AI search and user-requested fetches: allowed
User-agent: OAI-SearchBot
User-agent: ChatGPT-User
User-agent: Claude-SearchBot
User-agent: Claude-User
User-agent: PerplexityBot
User-agent: Perplexity-User
Allow: /
# AI model training: blocked
# (Google-Extended and Applebot-Extended are opt-out tokens;
# Googlebot and Applebot still crawl for Search and Siri)
User-agent: GPTBot
User-agent: ClaudeBot
User-agent: Google-Extended
User-agent: Applebot-Extended
User-agent: meta-externalagent
User-agent: CCBot
Disallow: /
# Everyone else, including Googlebot, Bingbot and Applebot
User-agent: *
Allow: /
Sitemap: https://www.example.com/sitemap.xml
Trade-off: models learn about companies partly from training data. Blocking training can make assistants less aware of your brand when they answer without searching the web. For most B2B companies whose website is marketing material, that's a reason to allow training (Template 1).
Template 3: block all AI crawlers but keep search engines
Use this if your content is the product (paid research, data, courses) and AI visibility isn't a goal. Googlebot and Bingbot stay allowed, so you remain in Google and Bing. Note that this includes AI Overviews and Copilot, which use the same crawlers.
robots.txt: block all AI crawlers, keep Google and Bing
# All AI crawlers and training tokens: blocked
User-agent: OAI-SearchBot
User-agent: ChatGPT-User
User-agent: GPTBot
User-agent: Claude-SearchBot
User-agent: Claude-User
User-agent: ClaudeBot
User-agent: PerplexityBot
User-agent: Perplexity-User
User-agent: Google-Extended
User-agent: Applebot-Extended
User-agent: meta-externalagent
User-agent: CCBot
Disallow: /
# Traditional search engines: allowed
User-agent: *
Allow: /
Sitemap: https://www.example.com/sitemap.xml
Remember that ChatGPT-User and Perplexity-User may ignore these rules, because their providers treat them as acting for a user. To enforce a block, you'll need your CDN or firewall. See how Cloudflare's AI crawler controls work.
Template 4: block AI training on specific folders only
Use this if most of your site is marketing you want AI to learn from, but some sections hold material you don't want in training data, such as a research library, gated whitepapers or customer documentation.
robots.txt: block training on selected folders
# Training crawlers: everything except protected folders
User-agent: GPTBot
User-agent: ClaudeBot
User-agent: Google-Extended
User-agent: Applebot-Extended
User-agent: meta-externalagent
User-agent: CCBot
Disallow: /research/
Disallow: /customer-docs/
Allow: /
# Everyone else, including all AI search crawlers
User-agent: *
Allow: /
Sitemap: https://www.example.com/sitemap.xml
Paths are case-sensitive and match by prefix, so Disallow: /research/ covers /research/2026-report but not /Research/.
The rules that trip people up
1. A named group replaces the * group
A crawler follows only the group that names it most specifically. If GPTBot has its own group, it ignores everything in User-agent: *. So this setup lets GPTBot into /admin/:
User-agent: *
Disallow: /admin/
User-agent: GPTBot
Allow: /
Repeat any rules that should apply to everyone inside every named group. That's why Template 1 lists Disallow: /admin/ twice.
2. The longest matching rule wins
When both an Allow and a Disallow match a URL, the rule with the longer path wins, whatever the order. With Disallow: / and Allow: /blog/, crawlers can read the blog and nothing else. If two matching rules are the same length, RFC 9309 says crawlers should use the least restrictive one, which is Allow.
3. A server error can block everything
If robots.txt returns a 4xx error such as 404, crawlers treat the site as fully allowed. If it returns a 5xx server error or times out, RFC 9309 tells crawlers to assume they're disallowed from the whole site. A misconfigured CDN or an overloaded server can therefore take you out of AI search. Check that /robots.txt returns a 200 status.
4. One file per host
Rules at www.example.com/robots.txt don't apply to example.com, blog.example.com or shop.example.com. Each host serves its own file.
5. Applebot copies your Googlebot rules
If your robots.txt doesn't mention Applebot but does mention Googlebot, Applebot follows the Googlebot rules. A Googlebot-specific block therefore also affects Siri and Spotlight.
6. Disallow doesn't mean "don't show"
A disallowed URL can still appear in search results (without a description) if other sites link to it. To keep a page out of results, allow crawling and add a noindex robots meta tag, which crawlers can only read if they're allowed to fetch the page. Our own robots.txt explains this choice in its comments.
7. Your CDN may rewrite the file
Cloudflare's managed robots.txt, when it's switched on, adds its own rules to the top of your file, including Disallow: / for GPTBot, ClaudeBot, Google-Extended, CCBot and others. Always check the live URL, not the copy in your codebase.
How to test your robots.txt
- Open the live file in a private window:
https://www.yourdomain.com/robots.txt. Confirm a200status, a plain-text content type, and the rules you expect. - Run the free LLM Readiness Check. It parses your live robots.txt using RFC 9309 rules and shows, for each of the 15 AI crawlers, whether it's allowed and which group matched.
- Check Google Search Console's robots.txt report to see when Google last fetched the file and whether it hit errors.
- Watch your logs for a few days after a change. OpenAI says robots.txt changes take about 24 hours to reach its systems, and Google generally caches the file for up to a day.
Frequently asked questions
How do I block GPTBot but stay in ChatGPT search?
Add a group for GPTBot with Disallow: / and make sure OAI-SearchBot is allowed, either with its own Allow: / group or through User-agent: * with Allow: /. GPTBot collects training data. OAI-SearchBot decides whether you appear in ChatGPT search answers.
Does robots.txt stop AI companies from using my content?
It stops compliant crawlers from fetching it in future. OpenAI, Anthropic, Google, Microsoft, Apple, Perplexity, Meta and Common Crawl say their main crawlers obey robots.txt. It doesn't remove content already collected, and several user-triggered fetchers (ChatGPT-User, Perplexity-User, Google-Agent) may ignore it because a person asked for the page.
Do I need to list AI crawlers if I allow everything?
No. User-agent: * with Allow: / (or an empty Disallow:) already allows every crawler. Listing the AI crawlers explicitly is optional. It documents your intent and makes it harder for someone to block them by accident later.
Why is my Disallow rule for everyone not applying to GPTBot?
Because GPTBot has its own group. A crawler follows only the most specific group that names it and ignores the User-agent: * group completely. Repeat any sitewide rules, such as Disallow: /admin/, inside each named group.
How long until AI crawlers see my robots.txt changes?
Usually within a day. OpenAI says changes take about 24 hours to reach its systems, and Google generally caches robots.txt for up to 24 hours. Other crawlers re-read the file on their own schedules.
Does robots.txt on www cover my other subdomains?
No. Each host needs its own file. Rules at www.example.com/robots.txt don't apply to blog.example.com or shop.example.com.
Can I block AI training on only part of my site?
Yes. Use path rules in the training crawlers' group, for example Disallow: /research/, and they'll skip only that folder. See Template 4 in this guide.
Free tool
Test your robots.txt against 15 AI crawlers
The free LLM Readiness Check reads your live robots.txt, shows which AI crawlers are allowed or blocked and which rule matched, and gives you a snippet to paste.
Sources
- IETF RFC 9309: Robots Exclusion Protocol
- OpenAI: Overview of OpenAI crawlers
- Anthropic: Does Anthropic crawl data from the web, and how can site owners block the crawler?
- Perplexity: Perplexity crawlers
- Google Search Central: Google's common crawlers
- Google Search Central: How Google interprets the robots.txt specification
- Apple Support: About Applebot
- Meta for Developers: Meta web crawlers
- Common Crawl: CCBot
- Cloudflare Docs: Managed robots.txt