Short answer

Cloudflare can block ChatGPT and other AI crawlers even when your robots.txt allows them. AI Crawl Control lets you allow or block each AI crawler, and since July 2026 Cloudflare's security settings also sort AI bots into Search, Agent and Training. To stay visible in AI answers, keep Search and Agent bots allowed. Don't block "Training" without checking the effect on Googlebot and Bingbot.

Key takeaways

  • Cloudflare enforces AI crawler blocks in front of your server, so a clean robots.txt doesn't guarantee OAI-SearchBot or PerplexityBot can get in.
  • AI Crawl Control (formerly AI Audit) is on all plans. It shows which AI crawlers visit, how often, and whether they respect robots.txt, with an Allow or Block action for each one.
  • From September 15, 2026, new Cloudflare domains block Training and Agent bots on pages with ads by default. Search bots stay allowed.
  • Cloudflare treats Googlebot, Applebot and BingBot as multi-purpose (search plus training). Blocking the Training category can block them too.
  • For most B2B sites: allow Search and Agent, decide on Training case by case, leave Pay Per Crawl off, and check the Metrics tab monthly.
In this article
  1. Why Cloudflare settings matter for AI search
  2. AI Crawl Control: see and control every AI crawler
  3. The Search, Agent and Training settings (and the September 15, 2026 defaults)
  4. Managed robots.txt
  5. Pay Per Crawl: should a B2B company charge AI crawlers?
  6. Recommended Cloudflare setup for a B2B site
  7. How to check whether Cloudflare is blocking AI crawlers
  8. Frequently asked questions

Why Cloudflare settings matter for AI search

Cloudflare sits in front of a large share of the web as a CDN and firewall. Every request, including requests from AI crawlers, passes through it before it reaches your server. That means Cloudflare can answer an AI crawler with a 403 Forbidden long before your robots.txt or your content comes into play.

This is how companies end up invisible in ChatGPT without anyone deciding it. Someone ticks a box that says "block AI bots" to protect content from scraping, or accepts a default during onboarding, and months later sales asks why competitors show up in AI answers and they don't. In our audits we check the CDN layer as well as robots.txt, because the free LLM Readiness Check and most SEO tools only read robots.txt.

Cloudflare now has four separate controls that affect AI crawlers. They overlap, so it helps to know what each one does.

AI Crawl Control: see and control every AI crawler

AI Crawl Control (formerly AI Audit) is Cloudflare's dashboard for AI crawler traffic. Cloudflare says it's available on all plans and works automatically. To open it, select your account and domain in the Cloudflare dashboard and choose AI Crawl Control. The main tabs:

  • Overview: a snapshot of AI crawler activity on your domain.
  • Crawlers: every AI crawler Cloudflare has seen, with its operator, category, request trend and robots.txt violations. The Action column lets you set each one to Allow or Block, or Charge if you're in the Pay Per Crawl beta.
  • Metrics: requests broken down by date, crawler, operator, status code, hostname and path. Since February 2026 it also includes referral analytics, which shows the visitors AI platforms send you.
  • Directives (renamed from "Robots.txt" in April 2026): the health of your robots.txt and a list of the crawlers that ignore it.

When you block a crawler here, Cloudflare creates or updates a WAF custom rule on your domain, and blocked requests get a 403 (or 402 Payment Required with Pay Per Crawl). Free plans detect crawlers by their user agent. Paid plans can customize the block response, and plans with Bot Management get more accurate detection.

Cloudflare groups crawlers into categories. A few examples from its bot reference:

Cloudflare categoryExamplesWhat blocking costs a B2B site
AI SearchOAI-SearchBot, Claude-SearchBot, PerplexityBot, ApplebotYou disappear from those assistants' search answers
AI AssistantChatGPT-User, Claude-User, Perplexity-UserAssistants can't open your pages when a buyer asks about you
AI Crawler (training)GPTBot, ClaudeBot, Meta-ExternalAgent, CCBotFuture content stays out of model training. No direct effect on AI search answers
Search EngineGooglebot, BingBotNever block: you lose Google (including AI Overviews), Bing and Copilot

For the full list of crawlers and what each one does, see our AI crawler field guide.

The Search, Agent and Training settings (and the September 15, 2026 defaults)

On July 1, 2026, Cloudflare replaced its single "Block AI bots" toggle with three separate categories, available on every plan including Free. You'll find them under Security → Settings:

  • Search: crawlers that index your content so an AI can answer questions about it later.
  • Agent: automated activity acting in real time for a person, such as the fetchers ChatGPT or Claude use when a user asks about a page.
  • Training: crawlers that take content to train or fine-tune a model.

For each category you can allow, block everywhere, or block only on pages that display ads.

From September 15, 2026, domains newly added to Cloudflare get these defaults: Training and Agent bots blocked on pages with ads, and Search bots allowed. Cloudflare's reasoning is that an ad signals a page is meant for a person to visit. Most B2B websites don't run display ads, so the ad-page default rarely affects them, but it's worth confirming.

The Googlebot catch. Cloudflare classifies crawlers by everything they do. In its announcement it says that "multi-purpose crawlers such as Googlebot, Applebot, and BingBot will be blocked by customers who have selected to block Training."

So a well-meant "block Training" setting can also cut off Google Search, AI Overviews, Bing, Copilot and Siri. If you want to keep your content out of training, use robots.txt tokens such as Google-Extended and Applebot-Extended instead. They opt you out of training without blocking the search crawler.

The older Block AI bots setting (Security Settings → Block AI bots) still appears on some accounts and is being deprecated in favor of the new categories. It blocks verified training bots and some unverified bots that behave like them, and it excludes mixed-purpose bots used for both training and search.

Managed robots.txt

Managed robots.txt is a separate, optional setting (Security Settings → Bot traffic → "Set your preference to block training in robots.txt"). It's available on all plans and is off unless someone turns it on. When it's on, Cloudflare puts its own rules in front of your existing robots.txt and serves both as one file. The rules include:

  • Disallow: / for a list of training crawlers: Amazonbot, Applebot-Extended, Bytespider, CCBot, ClaudeBot, Google-Extended, GPTBot and meta-externalagent.
  • A Content Signals line for all crawlers: Content-signal: search=yes, ai-train=no, use=reference.

It doesn't block the AI search crawlers (OAI-SearchBot, Claude-SearchBot, PerplexityBot), so it's a reasonable choice if you've decided against training. But because it changes robots.txt without any edit to the file in your codebase, it can confuse whoever maintains the site. If you turn it on, write it down, and check the live file at yourdomain.com/robots.txt rather than the copy in your repository.

Pay Per Crawl: should a B2B company charge AI crawlers?

Pay Per Crawl (in beta) lets a site charge AI crawlers per request. Crawlers that don't pay get a 402 Payment Required response. It runs after Cloudflare's bot protections, so Cloudflare notes that you have to turn off Block AI Bots for it to work.

It makes sense for publishers, whose content is the product. For a manufacturer, distributor or professional services firm it usually doesn't: the website's job is to get the company recommended, and making crawlers pay to read it works against that. We'll cover the exceptions in a separate guide.

  1. Security → Settings, AI traffic: set Search to Allow and Agent to Allow. Leave Training on Allow unless you've made a deliberate decision to block it (see the Googlebot catch above).
  2. AI Crawl Control → Crawlers: confirm that OAI-SearchBot, ChatGPT-User, Claude-SearchBot, Claude-User, PerplexityBot, Perplexity-User and Applebot show Allow.
  3. Legacy Block AI bots: if it's still shown, set it to "Allow (do not block)" unless you mean to block training.
  4. Managed robots.txt: turn it on only if you've decided to block training, and record that decision.
  5. Pay Per Crawl: leave it off.
  6. Security Events: filter by the user agents above and check that AI crawlers aren't being challenged by Bot Fight Mode, rate limits or custom WAF rules. A challenge page counts as a block to a crawler. Our step-by-step WAF rules guide has the exact rules.

How to check whether Cloudflare is blocking AI crawlers

  • AI Crawl Control → Metrics: filter by status code. A spike of 403 responses to OAI-SearchBot or PerplexityBot is the clearest sign of a problem.
  • Directives tab: compare what your robots.txt allows with what Cloudflare is actually doing.
  • Test from outside: curl -I -A "OAI-SearchBot" https://www.yourdomain.com/ should return 200. It isn't conclusive, because Cloudflare also verifies crawlers by IP, but a 403 or a challenge page tells you something is wrong.
  • Run the free LLM Readiness Check to confirm your robots.txt allows the 15 major AI crawlers, then check Cloudflare with the steps above.

Frequently asked questions

Does Cloudflare block AI crawlers by default?

It depends on when the domain joined Cloudflare and what it shows. Since July 2025, new domains are asked during setup whether to allow AI crawlers. Since September 15, 2026, new domains get default policies that block Training and Agent bots on pages that display ads, while Search bots stay allowed. Existing domains keep their current settings unless someone changes them.

Does Cloudflare block Googlebot?

Not on its own. But Cloudflare says that multi-purpose crawlers such as Googlebot, Applebot and BingBot, which combine search with AI training, "will be blocked by customers who have selected to block Training." If you choose to block the Training category, check that you haven't also blocked the crawlers behind Google Search, Bing and Copilot.

Where is AI Crawl Control in the Cloudflare dashboard?

Select your account and domain, then open AI Crawl Control. The Crawlers tab lists every AI crawler seen on your site, with an Allow or Block action for each. The Metrics tab shows requests over time, and the Directives tab shows which crawlers ignore your robots.txt.

Is AI Crawl Control free?

Yes. Cloudflare says it is available on all plans. Free plans detect AI crawlers by user agent. Paid plans add custom block responses, and plans with Bot Management get more accurate detection.

What is Cloudflare's managed robots.txt?

An optional setting that adds Cloudflare's own rules to the top of your robots.txt. It disallows training crawlers such as GPTBot, ClaudeBot, Google-Extended, Applebot-Extended, meta-externalagent and CCBot, and adds a Content-Signal line (search=yes, ai-train=no). It doesn't block AI search crawlers.

Should a B2B company use Pay Per Crawl?

Usually not. Pay Per Crawl charges AI crawlers per request, which makes sense for publishers whose content is the product. A B2B company's website exists to get the company recommended, so making AI crawlers pay to read it works against that goal.

Free tool

Is something blocking AI crawlers from your site?

The free LLM Readiness Check tests your robots.txt against 15 AI crawlers and checks that your homepage is readable without JavaScript. No sign-up.

Run the free check
Iulian Grecu

Iulian Grecu

Founder of AGI Search Labs. More than a decade in search, analytics and performance marketing (GA4, server-side tagging, Google Ads). Google Partners Digital Champion 2023. LinkedIn