Short answer
To set up Cloudflare for AI crawlers, first check Security Events for blocked AI bots. Then use WAF custom rules to (1) skip bot protection and rate limits for verified search and AI search crawlers, (2) block requests that claim to be GPTBot, ClaudeBot or PerplexityBot but aren't verified, and (3) block training crawlers only on folders you want to protect. Bot Fight Mode can't be skipped, so check it separately.
Key takeaways
- AI Crawl Control's per-crawler blocks run as WAF custom rules, before Cloudflare's bot products, so rule order matters.
cf.client.botidentifies verified bots, andcf.verified_bot_categorytells you what kind of bot it is ("AI Search", "AI Crawler" and so on).- Bot Fight Mode (Free plan) can't be bypassed with Skip rules. Super Bot Fight Mode (Pro and up) can.
- A rule blocking unverified requests that use AI crawler user agents stops scrapers posing as GPTBot without touching the real crawlers.
- The whole setup fits in Cloudflare's 5 free custom rules.
In this article
- How Cloudflare's layers stack up
- Step 1: Find out what's being blocked today
- Step 2: Set crawler actions in AI Crawl Control
- Step 3: Deal with Bot Fight Mode (Free plan)
- Step 4: Let verified search crawlers skip bot protection and rate limits
- Step 5: Block bots that pretend to be AI crawlers
- Step 6: Block training crawlers only where you need to
- Step 7: Rate-limit the unverified scrapers
- Step 8: Verify and keep an eye on it
- The full rule set at a glance
- Frequently asked questions
How Cloudflare's layers stack up
Before writing rules, it helps to know the order in which Cloudflare evaluates a request from an AI crawler. Cloudflare's documentation describes this sequence:
- WAF custom rules, including the rule called AI Crawl Control that the dashboard creates when you block a crawler (under Security → Security rules).
- Bot solutions: Bot Fight Mode, Super Bot Fight Mode or Bot Management, depending on your plan.
- Pay Per Crawl, if you use it.
Separately, the AI traffic settings under Security → Settings (Search, Agent and Training, added in July 2026) control whole categories. We cover those, and the Googlebot catch that comes with blocking "Training", in our guide to Cloudflare AI Crawl Control. This guide is about the rules layer underneath.
Two fields do most of the work in the rules below:
cf.client.bot: true when the request comes from a known good bot that Cloudflare has verified by IP range, reverse DNS or cryptographic signature. It gives the same information ascf.bot_management.verified_bot.cf.verified_bot_category: the category of that verified bot. Values relevant here are"Search Engine Crawler"(Googlebot, Bingbot),"AI Search"(OAI-SearchBot, Claude-SearchBot, PerplexityBot, Applebot),"AI Assistant"(ChatGPT-User, Claude-User, Perplexity-User) and"AI Crawler"(GPTBot, ClaudeBot, Meta-ExternalAgent, CCBot).
About the new taxonomy: in July 2026 Cloudflare introduced behavior classifications (Search, Agent, Training and others) and stopped treating "AI search" separately from traditional search. Cloudflare says these classifications aren't fields in WAF custom rules, and that the original categories above keep working for backward compatibility. Use them in rules.
Step 1: Find out what's being blocked today
Go to Security → Analytics (or Security → Events) and filter by user agent for OAI-SearchBot, PerplexityBot, Claude-SearchBot and ChatGPT-User. For each blocked or challenged request, note which service acted: a custom rule, managed rules, rate limiting, or Bot Fight Mode / Super Bot Fight Mode. Then open AI Crawl Control → Metrics and filter by status code. A run of 403 responses to AI search crawlers tells you which layer to fix first.
Step 2: Set crawler actions in AI Crawl Control
In AI Crawl Control → Crawlers, set the AI search and assistant crawlers to Allow. Block only the crawlers you've decided against. Cloudflare turns your blocks into a single WAF rule named "AI Crawl Control". You can edit that rule directly, for example to exempt certain paths, but edits made in the WAF aren't shown back in the AI Crawl Control dashboard, so keep a note of any changes.
Step 3: Deal with Bot Fight Mode (Free plan)
Bot Fight Mode is Cloudflare's free bot protection. It challenges traffic that matches known bot patterns. Cloudflare's documentation is explicit: "You cannot bypass or skip Bot Fight Mode using WAF custom rules or Page Rules." Skip, Bypass and Allow actions have no effect on it.
So if Step 1 showed Bot Fight Mode challenging AI search crawlers, you have two options: turn Bot Fight Mode off (Security → Settings), or upgrade to a plan with Super Bot Fight Mode, where you can allow verified bots and use Skip rules. For most B2B marketing sites, turning it off and relying on the targeted rules below works well.
On Pro and higher plans, open Super Bot Fight Mode and set Verified bots to Allow.
Step 4: Let verified search crawlers skip bot protection and rate limits
Create a custom rule under Security → Security rules → Create rule → Custom rule and put it first in the list. In the expression editor, paste:
Rule 1: verified search and AI search crawlers, action Skip
(cf.client.bot and cf.verified_bot_category in {"Search Engine Crawler" "AI Search" "AI Assistant"})
Set the action to Skip and choose All Super Bot Fight Mode rules and All rate limiting rules. Leave managed rules (the WAF's attack protection) on: verified crawlers don't need to skip them, and skipping them only widens your attack surface. This rule makes sure Googlebot, Bingbot, Applebot, OAI-SearchBot, Claude-SearchBot, PerplexityBot and the user fetchers aren't challenged or throttled by mistake.
Skip doesn't work for Bot Fight Mode (see Step 3), so this rule matters on Pro plans and above.
Step 5: Block bots that pretend to be AI crawlers
Anyone can put "GPTBot" in a user agent, and Common Crawl warns that fake CCBots are common. Real AI crawlers are verified by Cloudflare, so you can safely block requests that use their names but aren't verified:
Rule 2: impostor AI crawlers, action Block
(
lower(http.user_agent) contains "gptbot" or
lower(http.user_agent) contains "oai-searchbot" or
lower(http.user_agent) contains "chatgpt-user" or
lower(http.user_agent) contains "claudebot" or
lower(http.user_agent) contains "claude-searchbot" or
lower(http.user_agent) contains "claude-user" or
lower(http.user_agent) contains "perplexitybot" or
lower(http.user_agent) contains "perplexity-user" or
lower(http.user_agent) contains "ccbot"
) and not cf.client.bot
Set the action to Block and place it after Rule 1. Before turning it on, check Security Events for a few days to confirm that real crawler traffic shows as verified on your plan. If you see legitimate requests matching, use Managed Challenge instead while you investigate.
Step 6: Block training crawlers only where you need to
If you want AI training kept out of specific sections, such as a research library or customer docs, without affecting the rest of the site:
Rule 3: training crawlers on protected folders, action Block
(cf.verified_bot_category eq "AI Crawler" and (
starts_with(http.request.uri.path, "/research/") or
starts_with(http.request.uri.path, "/customer-docs/")
))
Pair this with matching Disallow lines in robots.txt (see Template 4) so well-behaved crawlers don't request those URLs in the first place. Blocking training this way is safer than the account-wide "Training" setting, which Cloudflare says can also block multi-purpose crawlers such as Googlebot, BingBot and Applebot.
Step 7: Rate-limit the unverified scrapers
Aggressive scrapers that don't identify themselves are the real load problem. Under Security rules → Rate limiting rules, create a rule that targets unverified automated traffic and leaves people and verified bots alone. For example, match not cf.client.bot on paths that scrapers hammer (search pages, product listings, faceted filters), with a threshold far above what any human would reach. Start with Managed Challenge rather than Block.
If a verified crawler is overloading your server, slow it down instead of blocking it. Anthropic supports Crawl-delay in robots.txt, and you can also contact the crawler's operator. Their documentation pages list contact details.
Step 8: Verify and keep an eye on it
- Security Events: filter by each of your rules for a week and check that nothing legitimate is being caught.
- AI Crawl Control → Metrics: AI search crawlers should show
200responses, and the referral analytics should start showing visits from AI platforms. - Directives tab: check which crawlers ignore your robots.txt. Those are candidates for a WAF block.
- Outside view: run the free LLM Readiness Check to confirm robots.txt is clean, then track AI referrals in GA4.
The full rule set at a glance
| Order | Rule | Matches | Action |
|---|---|---|---|
| 1 | Verified search crawlers | cf.client.bot and category Search Engine Crawler, AI Search or AI Assistant | Skip Super Bot Fight Mode and rate limiting |
| 2 | Impostor AI crawlers | AI crawler name in user agent and not verified | Block |
| 3 | Training on protected folders | Category AI Crawler on chosen paths | Block |
| Auto | AI Crawl Control | Crawlers you set to Block in the dashboard | Block (403) |
| Rate limit | Unverified scrapers | Not verified, high request rate on heavy paths | Managed Challenge |
Frequently asked questions
Does Cloudflare Bot Fight Mode block AI crawlers?
It can challenge automated traffic, and its related Block AI bots option blocks verified training bots. Bot Fight Mode can't be skipped with WAF custom rules or Page Rules, so if Security Events show it challenging AI search crawlers, your options are to turn it off or move to a plan with Super Bot Fight Mode, where verified bots can be allowed and Skip rules work.
What is cf.verified_bot_category?
A field in Cloudflare's rules language that holds the category of a verified bot, such as "Search Engine Crawler", "AI Search", "AI Assistant" or "AI Crawler". You can use it in WAF custom rules, for example cf.verified_bot_category eq "AI Crawler" to target training crawlers.
How do I block fake GPTBot traffic in Cloudflare?
Create a custom rule that matches requests whose user agent contains an AI crawler name but which Cloudflare doesn't recognize as a verified bot (not cf.client.bot), and set the action to Block. Real crawlers verified by IP or signature pass. Impostors using the same user agent are blocked.
How many WAF custom rules do I get on the Free plan?
Cloudflare lists 5 custom rules on Free, 20 on Pro, 100 on Business and 1,000 on Enterprise. The setup in this guide uses three or four rules.
Can I use Cloudflare's new Search, Agent and Training classes in WAF rules?
Not directly. Cloudflare says the behavior classifications introduced in July 2026 aren't individual fields in WAF custom rules. Use the original verified bot categories in rules, which Cloudflare keeps for backward compatibility, and use Security → Settings for the Search, Agent and Training controls.
Should I rate-limit AI crawlers?
Rate-limit unverified automated traffic, not verified AI search crawlers. If a verified crawler is genuinely overloading your server, Anthropic supports Crawl-delay in robots.txt, and a generous rate limit is better than a block.
Free tool
Check your robots.txt before you touch the firewall
The free LLM Readiness Check confirms that robots.txt allows the 15 major AI crawlers, so you know any remaining blocks are coming from your firewall.
Sources
- Cloudflare Docs: WAF custom rules (plans, actions)
- Cloudflare Docs: Skip options in custom rules
- Cloudflare Docs: Bot Fight Mode
- Cloudflare Docs: Super Bot Fight Mode
- Cloudflare Docs: Verified bots
- Cloudflare Docs: Verified bot categories
- Cloudflare Docs: cf.client.bot field
- Cloudflare Docs: AI Crawl Control with Cloudflare WAF
- Cloudflare Docs: AI Crawl Control with Cloudflare Bots (order of operations)
- Cloudflare Docs: AI Crawl Control bot reference