Short answer

AI crawlers are bots that AI companies use to read websites. They do three different jobs: training (GPTBot, ClaudeBot, CCBot), search indexing for AI answers (OAI-SearchBot, Claude-SearchBot, PerplexityBot, Googlebot, Bingbot) and live fetches when a user asks a question (ChatGPT-User, Claude-User, Perplexity-User). If you block the search and live-fetch bots, you drop out of AI answers. Blocking training bots doesn't do that.

Key takeaways

  • Each major AI company runs separate bots for training, search and live user requests, and robots.txt controls each one separately.
  • Blocking a training bot (GPTBot, ClaudeBot, Google-Extended, Applebot-Extended, CCBot) does not remove you from AI search answers.
  • Blocking a search bot (OAI-SearchBot, Claude-SearchBot, PerplexityBot, Googlebot, Bingbot) does remove you from that engine's answers.
  • Google-Extended and Applebot-Extended aren't separate crawlers. They're opt-out tokens, and Google-Extended has no effect on AI Overviews.
  • robots.txt isn't the only gatekeeper. Since July 2025, Cloudflare blocks AI crawlers by default on new domains.
In this article
  1. The three jobs AI crawlers do
  2. The full list: 15 AI crawlers and what blocking each one does
  3. Engine by engine: what you need to know
  4. Recommended robots.txt setup for a B2B company
  5. robots.txt isn't the only gatekeeper
  6. How to check which AI crawlers can reach your site
  7. Frequently asked questions

The three jobs AI crawlers do

People talk about "AI bots" as if they were one thing, and that leads to expensive mistakes. A company blocks "AI" to protect its content and then disappears from ChatGPT's recommendations without realizing why. Every major AI company now splits its crawling by purpose:

  • Training crawlers collect public web content that may be used to train future models. Blocking them keeps your future content out of training data, but it doesn't affect whether you're cited in answers today.
  • Search crawlers build the index an AI assistant searches when it answers a question. This is the index that decides whether you can be cited, linked and recommended. Blocking them makes you invisible to that assistant's search.
  • User-triggered fetchers visit a specific page because a user asked the assistant to, for example "summarize this vendor's pricing page". Several companies say these fetchers don't follow robots.txt, because a person initiated the request.

For a B2B company that wants to be recommended by AI, the search crawlers and user fetchers are the ones that matter. Whether to allow training crawlers is a business decision, which we cover below and in more depth in training vs search vs user bots.

The full list: 15 AI crawlers and what blocking each one does

These are the 15 crawlers our LLM Readiness Check tests, across 8 AI engines. The "robots.txt" column reflects each company's own documentation as of September 2026.

Enginerobots.txt tokenJobObeys robots.txtIf you block it
ChatGPTOAI-SearchBotSearchYesYour pages won't appear in ChatGPT search answers
ChatGPTChatGPT-UserUser fetchMay not applyUnreliable: OpenAI says robots.txt rules may not apply to user-initiated visits
ChatGPTGPTBotTrainingYesContent excluded from OpenAI model training. No effect on ChatGPT search
ClaudeClaude-SearchBotSearchYesYour pages aren't indexed for Claude's search results
ClaudeClaude-UserUser fetchYesClaude can't open your pages when a user asks about them
ClaudeClaudeBotTrainingYesFuture content excluded from Anthropic model training
PerplexityPerplexityBotSearchYesYour pages won't be surfaced or linked in Perplexity answers
PerplexityPerplexity-UserUser fetchGenerally noLittle effect: Perplexity says this fetcher generally ignores robots.txt
Google AI Overviews & GeminiGooglebotSearchYesYou drop out of Google Search entirely, including AI Overviews and AI Mode. Never do this
Google AI Overviews & GeminiGoogle-ExtendedTraining and Gemini grounding (token only)YesExcluded from Gemini training and grounding in Gemini Apps and Vertex AI. No effect on Search or AI Overviews
Microsoft CopilotBingbotSearchYesYou drop out of Bing, and Copilot can't ground answers in your pages
Meta AImeta-externalagentTraining and indexingYesContent excluded from Meta's AI training and product indexing
Apple IntelligenceApplebotSearchYesOut of Siri, Spotlight and Safari suggestions
Apple IntelligenceApplebot-ExtendedTraining (token only)YesContent excluded from Apple foundation-model training. Still eligible for Apple search features
Common CrawlCCBotOpen training datasetYesExcluded from future Common Crawl archives, which many AI models are trained on

Engine by engine: what you need to know

ChatGPT (OpenAI): OAI-SearchBot, ChatGPT-User, GPTBot

OpenAI runs the clearest three-way split. OAI-SearchBot decides whether you can appear in ChatGPT search answers. OpenAI's documentation says sites that opt out won't be shown in search answers, and that robots.txt changes take about 24 hours to reach its systems. GPTBot is for training only. ChatGPT-User handles visits users ask for inside ChatGPT and custom GPTs, and OpenAI says robots.txt rules "may not apply" to it.

OpenAI also runs OAI-AdsBot, which checks landing pages submitted as ChatGPT ads. It visits only pages submitted as ads and doesn't follow robots.txt.

User agent (OAI-SearchBot)

Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/131.0.0.0 Safari/537.36; compatible; OAI-SearchBot/1.4; +https://openai.com/searchbot

Claude (Anthropic): Claude-SearchBot, Claude-User, ClaudeBot

Anthropic follows the same pattern. ClaudeBot collects training data, Claude-SearchBot indexes content for search, and Claude-User fetches pages when a user asks Claude a question. Unlike OpenAI and Perplexity, Anthropic says all three respect robots.txt, including Crawl-delay. Anthropic also advises against blocking its IP addresses, because that stops its bots from reading your robots.txt in the first place.

Perplexity: PerplexityBot, Perplexity-User

PerplexityBot is a search crawler. Perplexity says it's used to surface and link websites in results, not to train foundation models, and recommends allowing it. Perplexity-User fetches pages when a user asks a question and, in Perplexity's words, "generally ignores robots.txt rules".

Google AI Overviews, AI Mode and Gemini: Googlebot, Google-Extended

This is the most misunderstood pair. AI Overviews and AI Mode are part of Google Search, so they use the regular Googlebot crawl and index. Google-Extended is not a separate crawler but a token that controls whether content Google has already crawled can be used to train Gemini models and to ground answers in Gemini Apps and Vertex AI. Google states that Google-Extended "does not impact a site's inclusion in Google Search nor is it used as a ranking signal."

Common mistake: blocking Google-Extended to "opt out of AI Overviews". It doesn't work. The only controls for AI Overviews are Search controls such as nosnippet or noindex, and those also affect your regular search listings.

Google also runs user-triggered fetchers that generally ignore robots.txt, including Google-Agent (used by agents running on Google infrastructure to browse and take actions for users) and Google-GeminiNotebook (for URLs users add as sources).

Microsoft Copilot: Bingbot

Copilot doesn't have its own crawler. It grounds answers in Bing's index, so Bingbot access is what counts. That makes Bing more important for B2B than its search share suggests, because Copilot is built into the Microsoft 365 tools your buyers use at work. In February 2026, Bing's webmaster guidelines added two directives worth knowing: NOARCHIVE keeps a page out of Copilot answers and grounding, and NOCACHE limits Copilot to using only the URL, title and snippet.

Meta AI: meta-externalagent

meta-externalagent crawls for purposes including training AI models and indexing content for Meta's products, and it follows robots.txt. Its sibling meta-externalfetcher fetches individual links at a user's request and, according to Meta, may bypass robots.txt.

Apple Intelligence: Applebot, Applebot-Extended

Applebot powers search in Siri, Spotlight and Safari. One detail catches people out: if your robots.txt doesn't mention Applebot but does mention Googlebot, Applebot follows the Googlebot rules. Applebot-Extended doesn't crawl at all. It only tells Apple not to use content Applebot has already collected to train its foundation models, and pages that disallow it can still appear in Apple's search features.

Common Crawl: CCBot

CCBot builds Common Crawl's free, open archive of the web, which researchers and many AI developers use as training data. Blocking it keeps your future content out of new archives, but not out of archives that already exist. Common Crawl warns that other bots sometimes pretend to be CCBot, so verify its IPs before trusting the user agent.

For most mid-market B2B companies, being recommended by AI is worth far more than keeping marketing pages out of training data. Our default recommendation is:

  • Always allow the search crawlers and user fetchers: OAI-SearchBot, ChatGPT-User, Claude-SearchBot, Claude-User, PerplexityBot, Perplexity-User, Googlebot, Bingbot and Applebot.
  • Usually allow the training crawlers as well. Models that have seen your content are more likely to know your brand, products and category when a buyer asks without triggering a web search.
  • Block training only if you publish content you sell (research, data or paid courses), or if legal has a specific concern. Use Disallow on the paths that matter rather than across the whole site.

robots.txt: allow AI search, block AI training

# AI search and user-requested fetches: allowed
User-agent: OAI-SearchBot
User-agent: ChatGPT-User
User-agent: Claude-SearchBot
User-agent: Claude-User
User-agent: PerplexityBot
User-agent: Perplexity-User
Allow: /

# AI model training: blocked
User-agent: GPTBot
User-agent: ClaudeBot
User-agent: Google-Extended
User-agent: Applebot-Extended
User-agent: meta-externalagent
User-agent: CCBot
Disallow: /

# Everyone else, including Googlebot, Bingbot and Applebot
User-agent: *
Allow: /

Sitemap: https://www.example.com/sitemap.xml

Listing several User-agent lines above one set of rules is valid under the robots.txt standard (RFC 9309). If you'd rather allow everything, you don't need AI-specific rules at all: User-agent: * with Allow: / covers every crawler. Naming the AI bots explicitly, as our own robots.txt does, makes your intent clear and protects you if someone later adds a broad Disallow. For more setups, see our copy-paste robots.txt templates.

robots.txt isn't the only gatekeeper

Often, the reason a site is invisible to AI isn't robots.txt at all. It's a layer in front of the site:

  • CDN and firewall defaults. On July 1, 2025, Cloudflare started asking every new domain whether to allow AI crawlers, with blocking as the default. A site can have a perfect robots.txt and still return errors to OAI-SearchBot. See how Cloudflare's AI crawler controls work.
  • Security plugins and bot protection that challenge or block non-browser traffic, or block whole cloud IP ranges.
  • JavaScript-only content. Most AI crawlers don't run JavaScript, so a page whose text loads client-side can look empty to them even when access is allowed. See how to test and fix it.
  • IP blocking. As Anthropic points out, blocking a crawler's IPs also stops it from reading your robots.txt, so your rules can't be applied.

How to check which AI crawlers can reach your site

  1. Run the free LLM Readiness Check. It reads your robots.txt, reports which of the 15 crawlers are allowed or blocked and which rule matched, and gives you a snippet to paste. The free Chrome extension runs the same check on any page you're viewing.
  2. Look at your server or CDN logs for the user-agent tokens above. If you see Googlebot but never OAI-SearchBot or PerplexityBot, something upstream may be blocking them.
  3. Verify the IP addresses. OpenAI (openai.com/gptbot.json, openai.com/searchbot.json, openai.com/chatgpt-user.json), Perplexity (perplexity.com/perplexitybot.json), Anthropic (claude.com/crawling/bots.json) and Common Crawl (index.commoncrawl.org/ccbot.json) publish their IP ranges. Google, Bing and Apple support reverse DNS verification.
  4. Test from the outside. Request a page with a crawler's user agent (for example with curl -A "OAI-SearchBot") and check you get a 200 status with real content, not a challenge page. Some firewalls check IPs as well as user agents, so a clean result here isn't a guarantee.

Frequently asked questions

Should I block GPTBot?

Only if you don't want your content used to train OpenAI's models. Blocking GPTBot does not remove you from ChatGPT search answers, because those rely on OAI-SearchBot. Most B2B companies that want to be recommended by AI allow both.

Does blocking Google-Extended remove my site from AI Overviews?

No. Google says Google-Extended does not affect inclusion in Google Search, and AI Overviews and AI Mode are part of Search. Google-Extended controls Gemini model training and grounding in Gemini Apps and Vertex AI. The only way to keep a page out of AI Overviews is to use Search controls such as nosnippet or noindex, which also affect regular results.

Do AI crawlers obey robots.txt?

The major training and search crawlers from OpenAI, Anthropic, Perplexity, Google, Microsoft, Apple, Meta and Common Crawl say they do. User-triggered fetchers are different: OpenAI says robots.txt rules may not apply to ChatGPT-User, Perplexity says Perplexity-User generally ignores them, and Meta and Google say the same of their user-triggered fetchers. Anthropic says Claude-User respects robots.txt.

What's the difference between GPTBot and OAI-SearchBot?

GPTBot collects content that may be used to train OpenAI's models. OAI-SearchBot finds and indexes pages so they can appear, with links, in ChatGPT search answers. They're controlled separately in robots.txt, so you can allow one and block the other.

Which AI crawler matters most for Microsoft Copilot?

Bingbot. Copilot grounds its answers in Bing's index, so if Bingbot can't crawl a page, Copilot can't cite it. Bing's NOARCHIVE and NOCACHE meta tags also limit how Copilot can use your content.

How long does it take for a robots.txt change to take effect?

It depends on the crawler. OpenAI says changes can take about 24 hours to reach its systems. Other crawlers re-read robots.txt on their own schedules, usually within a day or two.

How can I tell if a visit really came from an AI crawler?

Check the request's IP address against the ranges the company publishes, for example openai.com/searchbot.json, perplexity.com/perplexitybot.json, claude.com/crawling/bots.json or index.commoncrawl.org/ccbot.json. User-agent strings are easy to fake.

Free tool

Which of these 15 crawlers can reach your site?

The free LLM Readiness Check tests your robots.txt against all 15 crawlers in this guide and gives you a ready-to-paste fix. No sign-up.

Run the free check
Iulian Grecu

Iulian Grecu

Founder of AGI Search Labs. More than a decade in search, analytics and performance marketing (GA4, server-side tagging, Google Ads). Google Partners Digital Champion 2023. LinkedIn