Short answer
Training bots (GPTBot, ClaudeBot, CCBot) collect content to train AI models. Search bots (OAI-SearchBot, Claude-SearchBot, PerplexityBot) index pages so assistants can cite them. User bots (ChatGPT-User, Claude-User, Perplexity-User) open a page when someone asks about it. To be recommended by AI, always allow search and user bots. Allowing training bots is a business decision, and most B2B companies should allow them.
Key takeaways
- OpenAI, Anthropic and Perplexity each run separate bots for training, search and user requests, and robots.txt treats them separately.
- Search bots decide whether you can be cited. Blocking one makes you invisible in that assistant's search answers.
- User bots fetch pages on demand. Several of them ignore robots.txt because a person asked for the page.
- Training bots only affect future model training. Blocking them doesn't remove you from AI search.
- A fourth type is growing: browser agents that act for users. They look like normal browsers, and the better-behaved ones sign their requests.
In this article
- Why the bot type matters more than the company
- The bot types side by side
- Training bots: what they do and when to block them
- Search bots: the ones that decide whether you're cited
- User bots: the live lookups
- The fourth type: AI browser agents
- Decision guide: what to allow, by business type
- Four traps to avoid
- Frequently asked questions
Why the bot type matters more than the company
A common robots.txt request goes like this: "block OpenAI, we don't want them using our content." The result is often a site that disappears from ChatGPT's answers while competitors keep getting recommended. OpenAI alone runs three different bots, and only one of them, GPTBot, is about training.
Every major AI company now splits its crawling by purpose. The question isn't "do we allow OpenAI?" but "which jobs do we want AI to do with our site?" There are three classic jobs, plus a newer fourth.
The bot types side by side
| Training bots | Search bots | User bots | Browser agents | |
|---|---|---|---|---|
| Job | Collect content to train future models | Build the index an assistant searches when answering | Fetch a specific page because a user asked | Navigate sites and complete tasks for a user |
| Examples | GPTBot, ClaudeBot, meta-externalagent, CCBot; tokens Google-Extended, Applebot-Extended | OAI-SearchBot, Claude-SearchBot, PerplexityBot, Googlebot, Bingbot, Applebot | ChatGPT-User, Claude-User, Perplexity-User, Google-Agent | ChatGPT agent (cloud browser) and other agentic browsers |
| When it visits | On its own schedule | On its own schedule | Live, during a conversation | Live, during a task |
| Obeys robots.txt | Yes | Yes | Varies: often no | Not documented |
| If you block it | Content left out of future training. No direct effect on AI answers | You can't be cited in that assistant's search answers | The assistant can't read your page when a buyer asks about it | Agents can't compare, quote or book with you |
Training bots: what they do and when to block them
Training bots collect public web content that may become training data for future models. GPTBot (OpenAI), ClaudeBot (Anthropic), meta-externalagent (Meta) and CCBot (Common Crawl, whose open archive many models are trained on) are actual crawlers. Google-Extended and Applebot-Extended are tokens instead. They don't crawl. They tell Google and Apple not to use content their search crawlers have already collected for model training. Google-Extended also covers grounding in Gemini Apps and Vertex AI.
Blocking training bots has no direct effect on whether assistants cite you in search answers today. The trade-off is less obvious: models answer many questions from what they learned in training, without searching the web. If your company, products and category language aren't in that training data, the model knows less about you when it answers that way.
Block training when your content is what you sell (paid research, proprietary data, courses), when legal or contractual obligations require it, or for specific folders such as customer documentation. Allow it when your site is mainly marketing: product pages, service pages, case studies and guides that exist to make you better known.
Search bots: the ones that decide whether you're cited
When a buyer asks ChatGPT, Claude or Perplexity "who are the best contract manufacturers for medical devices in Ohio?", the assistant searches an index and builds its answer from the pages it finds. Search bots build those indexes:
- OAI-SearchBot: OpenAI says sites that opt out won't be shown in ChatGPT search answers.
- Claude-SearchBot: Anthropic says blocking it prevents your content from being indexed for search.
- PerplexityBot: Perplexity says it surfaces and links sites in results and isn't used to train foundation models.
- Googlebot and Bingbot: they feed Google Search (including AI Overviews and AI Mode) and Bing (which grounds Microsoft Copilot).
- Applebot: it powers search in Siri, Spotlight and Safari.
For a company that wants AI to recommend it, there's no good reason to block any of these. If you take one thing from this guide, check that all six are allowed, in robots.txt and in your firewall.
User bots: the live lookups
User bots visit a page because a person in a conversation asked for it: "check this vendor's pricing page" or "what does this company say about lead times?" They fetch one page, not your whole site.
Their robots.txt behavior differs by company, which is a common source of confusion:
- ChatGPT-User: OpenAI says robots.txt rules "may not apply" because the action is user-initiated.
- Perplexity-User: Perplexity says it "generally ignores robots.txt rules".
- Google-Agent and other Google user-triggered fetchers: Google says they generally ignore robots.txt.
- meta-externalfetcher: Meta says it may bypass robots.txt.
- Claude-User: Anthropic says it respects robots.txt, so blocking it does stop Claude from fetching your pages for users.
In practice, user bots work for you. They show up when a buyer is evaluating you specifically. Blocking them means the assistant can't check your pricing, specs or certifications at that moment, and it may answer from older or third-party information instead.
The fourth type: AI browser agents
Agents such as ChatGPT's agent mode don't just fetch a page. They run a real browser, click through your site, fill in forms and complete tasks. Because they use a full browser, they often look like an ordinary Chrome visitor, and unlike most AI crawlers, they run JavaScript.
The better-behaved agents identify themselves cryptographically. OpenAI's help center says ChatGPT's cloud browser signs its requests using HTTP Message Signatures (RFC 9421), with the header Signature-Agent: "https://chatgpt.com". Cloudflare, Akamai, HUMAN and Vercel recognize it automatically. That matters because bot protection that challenges "anything automated" can block the agent a buyer sent to request a quote from you.
Cloudflare's taxonomy reflects this shift. Since July 2026 it groups bots by behavior, including Search, Agent and Training, and it no longer separates "AI search" from traditional search.
Decision guide: what to allow, by business type
| Your situation | Search bots | User bots | Training bots |
|---|---|---|---|
| B2B company whose site is marketing (manufacturing, logistics, professional services) | Allow | Allow | Allow |
| B2B company with a gated research library or customer docs | Allow | Allow | Block those folders only |
| Publisher or data business whose content is the product | Allow if AI referrals are valuable to you | Case by case | Block |
| Regulated firm (healthcare, finance) with legal concerns | Allow on public marketing pages | Allow | Follow legal's guidance; block by folder if needed |
Our robots.txt templates cover each of these setups.
Four traps to avoid
- Blocking Google-Extended to leave AI Overviews. It doesn't work. Google says Google-Extended doesn't affect inclusion in Google Search, and AI Overviews are part of Search.
- Blocking "Training" in Cloudflare. Cloudflare says multi-purpose crawlers such as Googlebot, Applebot and BingBot "will be blocked by customers who have selected to block Training." Use robots.txt tokens instead. See our Cloudflare guide.
- Forgetting Applebot's fallback. If robots.txt doesn't mention Applebot but does mention Googlebot, Applebot follows the Googlebot rules.
- Assuming robots.txt blocks user bots. For most providers it doesn't. If you truly need to keep them out, you'll need firewall rules.
Frequently asked questions
What's the difference between GPTBot and OAI-SearchBot?
GPTBot collects content that may be used to train OpenAI's models. OAI-SearchBot indexes pages so they can be shown and linked in ChatGPT search answers. You can block GPTBot and still appear in ChatGPT search, as long as OAI-SearchBot is allowed.
Does blocking training bots hurt my AI visibility?
Not directly. Blocking training bots doesn't remove you from AI search answers, which use search crawlers. It may make assistants less familiar with your brand when they answer from what the model already knows, without searching the web. That's the main reason most B2B companies allow training.
Can I block user bots like ChatGPT-User?
Not reliably with robots.txt. OpenAI says robots.txt rules may not apply to ChatGPT-User, Perplexity says Perplexity-User generally ignores them, and Google says its user-triggered fetchers generally ignore them. Anthropic says Claude-User does respect robots.txt. To enforce a block you'd need firewall rules, but blocking user bots stops assistants from reading your pages when a buyer asks about you.
Is Googlebot a training bot?
Google treats Googlebot as its search crawler and offers the separate Google-Extended token to opt out of Gemini training and grounding. Cloudflare, however, classes Googlebot as multi-purpose, so blocking the Training category in Cloudflare's settings can block Googlebot too.
What are AI browser agents?
Tools such as ChatGPT's agent mode that use a real browser to navigate sites and complete tasks for a user. They often look like ordinary Chrome visitors. ChatGPT's cloud browser signs its requests (Signature-Agent: "https://chatgpt.com") so CDNs like Cloudflare, Akamai and Vercel can recognize it.
What should a B2B company allow?
Allow all search bots and user bots. Allow training bots unless your content is something you sell, such as paid research or data. In that case block training only on those folders.
Free tool
See which bot types your site lets in
The free LLM Readiness Check groups 15 AI crawlers by job (search, user and training) and shows which ones your robots.txt allows. No sign-up.
Sources
- OpenAI: Overview of OpenAI crawlers
- OpenAI Help: ChatGPT agent allowlisting (HTTP message signatures)
- Anthropic: Does Anthropic crawl data from the web, and how can site owners block the crawler?
- Perplexity: Perplexity crawlers
- Google Search Central: Google's common crawlers
- Google Search Central: Google's user-triggered fetchers
- Apple Support: About Applebot
- Meta for Developers: Meta web crawlers
- Cloudflare Docs: Verified bot categories and behavior classifications
- Cloudflare Blog: Your site, your rules: new AI traffic options for all customers (July 1, 2026)