Skip to content
AI & SEARCH / FIELD NOTE 178

Managing 40+ AI Crawlers in 2026: My robots.txt After Two Pivots

Reading map: The 40+ Count: Where That Number Comes From; A Working Taxonomy of AI Crawlers; The Full robots.txt: Every Named Bot, May 2026; Cloudflare AI Bot Rules: The Config I Run
A reading map of this field note. Download SVG ↓

Published 19 May 2026. I manage SEO and server-side access controls for several content sites. This is the actual robots.txt configuration I use, the reasoning behind it, and the two times I changed course after being wrong.

The 40+ Count: Where That Number Comes From

Everyone cites "40+ AI crawlers" now and almost nobody explains how they count them. Here's my methodology, because the number matters less than the methodology.

I maintain a running list drawn from three sources: the Dark Visitors database (which tracks AI User-Agents with excellent documentation), my own server log analysis across eight domains I control directly, and the aggregated community data from the ai-robots-txt GitHub repository. Cross-referencing these gives me a working list of User-Agent strings associated with AI systems — training crawlers, retrieval systems, summarization bots, research scrapers, and the in-between category of "we call it research but it's probably training."

My current list has 43 distinct strings that are active — meaning I've seen traffic from them in the last 90 days. Another 11 appear historically but have gone quiet. That 43 number will be outdated by the time you read this. New products launch. Companies rebrand their bots. The count in April 2025 was around 31. The growth is real.

The reason this matters: when I hear someone say "I blocked all AI crawlers," my first question is always "how many do you have on your deny list?" I've seen robots.txt files that block 6 User-Agents and call it comprehensive. That might have been adequate in 2023. It's meaningfully incomplete now.

A Working Taxonomy of AI Crawlers

Not all AI crawlers want the same thing, and treating them identically is a mistake I see constantly. My working taxonomy has four categories.

Training crawlers harvest content to train or fine-tune language models. GPTBot, Google-Extended, CCBot, and DataForSeoBot fall here. These bots build the model's underlying knowledge. Once content is in a training set, it's very hard to remove — there are documented cases of content owners requesting removal after discovering training use, and the process is slow and inconsistently honored.

Retrieval crawlers index content for real-time citation and summarization — they're the engine behind AI answer products that quote sources. PerplexityBot is the canonical example. Blocking these hurts your discoverability in AI-powered search products. These bots can drive actual referral traffic back to your site.

Product feature crawlers serve specific AI-powered features within existing platforms. ChatGPT's browsing functionality, Bing Copilot's live web access, Apple's AI summary features — these use crawlers that differ from the training crawlers of the same company. OpenAI runs both GPTBot (training) and OAI-SearchBot (browsing). They behave differently and you might want to treat them differently.

Shadow crawlers are the genuinely uncomfortable category. These are scrapers that present no clear identity, rotate User-Agents, or present as generic bots (Python-requests, curl, unnamed CDN traffic). Some are AI-related. Some are competitive intelligence tools. A few are almost certainly training scrapers operating without disclosure. You can't block them cleanly via robots.txt because they don't identify themselves. This is where Cloudflare-level controls become necessary.

The Full robots.txt: Every Named Bot, May 2026

Below is the robots.txt I run on content sites where I want to block training crawlers while allowing retrieval and search engine crawlers. Comments explain each choice. Copy this, but read the comments before deploying.

# robots.txt — AI Crawler Configuration
# Updated: May 2026
# Strategy: Block training crawlers. Allow retrieval/search crawlers.
# For enforcement beyond named bots, use Cloudflare WAF rules.

# ── SEARCH ENGINE CRAWLERS (Always Allow) ──────────────────────────
User-agent: Googlebot
Allow: /

User-agent: Googlebot-Image
Allow: /

User-agent: Googlebot-Video
Allow: /

User-agent: Bingbot
Allow: /

User-agent: Slurp
Allow: /

User-agent: DuckDuckBot
Allow: /

User-agent: Baiduspider
Allow: /

User-agent: YandexBot
Allow: /

# ── AI TRAINING CRAWLERS (Block) ───────────────────────────────────
# These crawlers are used primarily for model training.
# Blocking them does not affect search rankings.

# OpenAI training crawler
User-agent: GPTBot
Disallow: /

# Anthropic training crawler
User-agent: ClaudeBot
Disallow: /

# Google AI training (distinct from Googlebot — does NOT affect search)
User-agent: Google-Extended
Disallow: /

# Common Crawl — feeds many training datasets
User-agent: CCBot
Disallow: /

# Meta AI training
User-agent: FacebookBot
Disallow: /

# Amazon AI/Alexa training
User-agent: Amazonbot
Disallow: /

# Apple AI training
User-agent: Applebot-Extended
Disallow: /
# Note: Applebot (without -Extended) powers Spotlight/Safari — allow that
User-agent: Applebot
Allow: /

# Cohere AI training
User-agent: cohere-ai
Disallow: /

# AI21 Labs
User-agent: AI2Bot
Disallow: /

# Diffbot (scraping/AI data pipeline)
User-agent: Diffbot
Disallow: /

# DataForSEO (data resale, feeds AI tools)
User-agent: DataForSeoBot
Disallow: /

# ImagesiftBot (image training)
User-agent: ImagesiftBot
Disallow: /

# Omgili/Webz.io (news data for AI)
User-agent: omgili
Disallow: /

User-agent: omgilibot
Disallow: /

# Bytespider (ByteDance/TikTok AI)
User-agent: Bytespider
Disallow: /

# Petalbot (Huawei AI — training signals)
User-agent: PetalBot
Disallow: /

# ISSCyberRiskCrawler
User-agent: ISSCyberRiskCrawler
Disallow: /

# ── AI RETRIEVAL CRAWLERS (Conditionally Allow) ─────────────────────
# These crawlers power AI answer products that may cite and link back.
# Consider allowing if referral traffic from AI products matters to you.
# Comment out the Allow and uncomment Disallow to block.

# Perplexity — active referral traffic, cites sources
User-agent: PerplexityBot
Allow: /
# Disallow: /

# You.com AI search
User-agent: YouBot
Allow: /
# Disallow: /

# OpenAI browsing product (separate from GPTBot training)
User-agent: OAI-SearchBot
Allow: /
# Disallow: /

# Bing AI (Copilot) — distinct from Bingbot, but shares index
User-agent: Bingbot
Allow: /

# ── AGGRESSIVE SCRAPERS (Block) ────────────────────────────────────
# These have shown disregard for crawl limits or disavow disclosures.

User-agent: SemrushBot
Disallow: /

User-agent: AhrefsBot
Disallow: /

User-agent: MJ12bot
Disallow: /

User-agent: DotBot
Disallow: /

# ── MISCELLANEOUS AI/DATA CRAWLERS (Block) ─────────────────────────
User-agent: anthropic-ai
Disallow: /

User-agent: Claude-Web
Disallow: /

User-agent: Kangaroo Bot
Disallow: /

User-agent: VelenPublicWebCrawler
Disallow: /

User-agent: Scrapy
Disallow: /

User-agent: newspaper
Disallow: /

# ── CRAWL DELAY FOR LEGITIMATE BOTS ───────────────────────────────
# Only applies to bots that respect Crawl-delay (many don't).
# Prevents log floods from high-frequency crawlers.

User-agent: *
Crawl-delay: 10

# ── SITEMAP ────────────────────────────────────────────────────────
Sitemap: https://yourdomain.com/sitemap.xml

Cloudflare AI Bot Rules: The Config I Run

robots.txt handles named, compliant bots. For everything else — shadow crawlers, bots that ignore robots.txt, high-volume scrapers — I use Cloudflare's WAF and Bot Management.

# Cloudflare WAF Custom Rules — AI Bot Management
# Dashboard path: Security → WAF → Custom Rules
# These are expression-based rules; order matters.

# ── RULE 1: BLOCK KNOWN BAD AI BOT USER-AGENTS ────────────────────
# Expression:
(http.user_agent contains "GPTBot" and not cf.client.bot) or
(http.user_agent contains "CCBot") or
(http.user_agent contains "Bytespider") or
(http.user_agent contains "cohere-ai") or
(http.user_agent contains "anthropic-ai") or
(http.user_agent contains "Google-Extended") or
(http.user_agent contains "Amazonbot") or
(http.user_agent contains "Diffbot") or
(http.user_agent contains "DataForSeoBot") or
(http.user_agent contains "ImagesiftBot") or
(http.user_agent contains "omgili")

# Action: Block (returns 403)
# Note: The "not cf.client.bot" exception for GPTBot prevents false
# positives where Cloudflare's own verified bot checks might conflict.

# ── RULE 2: CHALLENGE SUSPICIOUS HEADLESS BROWSERS ────────────────
# Targets scrapers that rotate UAs but show headless signals.
# Expression:
(cf.bot_score lt 10) and
(not cf.verified_bot) and
(http.request.uri.path contains "/blog" or
 http.request.uri.path contains "/articles" or
 http.request.uri.path contains "/docs")

# Action: JS Challenge (managed challenge for borderline cases)

# ── RULE 3: RATE LIMIT AI-ADJACENT PATHS ──────────────────────────
# Apply to /api/, /feed/, /rss/, /sitemap.xml, /llms.txt
# Cloudflare Rate Limiting (separate from WAF custom rules):
# Threshold: 60 requests / 10 minutes per IP
# Action: Block for 1 hour

# ── RULE 4: VERIFIED BOT ALLOW LIST ───────────────────────────────
# Cloudflare maintains a verified bot list. Always allow these.
# Toggle in Bot Management: "Allow Verified Bots" = ON
# This ensures Googlebot, Bingbot, etc. are never blocked by rules above.

# ── CLOUDFLARE AI BOTS TOGGLE (Bot Management Feature) ────────────
# If you have Bot Management (Enterprise or add-on):
# Security → Bots → AI Scrapers and Crawlers → Block
# This is a single toggle that blocks Cloudflare's internal list of
# known AI scrapers. Use it as a complement to custom rules, not a
# replacement — the built-in list lags community lists by 4–8 weeks.

# ── RULE 5: TARPIT FOR PERSISTENT BAD BOTS ────────────────────────
# Cloudflare Workers — deploy at edge for persistent scrapers
# See article: /179-anti-ai-scraping-2026.html for full tarpit patterns

Pivot One: Why I Stopped Blocking Everything

My first instinct, in mid-2024, was to block every bot I could identify as AI-related. Total lockdown. I had a client who was genuinely worried about training use and I shared that concern. We deployed a robots.txt that denied about 22 User-Agents. We also turned on Cloudflare's AI scrapers toggle and added some aggressive WAF rules.

Then Perplexity traffic showed up in the logs. Or rather — it stopped showing up. We'd blocked PerplexityBot in our sweep. At the time, Perplexity was sending the client's site about 340 sessions per month. Not huge, but not nothing, and it was growing. Those sessions were hitting the site's SaaS pricing page at a 3.1x higher rate than average organic sessions. Perplexity users apparently arrive already contextualized about what they want.

We unblocked PerplexityBot within 10 days. I also unblocked OAI-SearchBot (OpenAI's browsing crawler, distinct from GPTBot). The traffic came back. The lesson I drew: retrieval crawlers that power cited-answer products are categorically different from training crawlers. Blunt blocking ignores that distinction.

Pivot Two: Why I Started Blocking Training Crawlers Specifically

After Pivot One, I went too permissive. I removed several training crawlers from the deny list because a client argued that being in OpenAI's training data would increase "brand awareness" — their phrase, not mine. I thought the argument was thin but I honored it.

Then the AP's public statements about their content licensing deal with OpenAI came out. The AP confirmed they'd been in negotiations partly because OpenAI had used their archive without prior agreement. The AP's deal, signed in 2024, included back-licensing for past training use — meaning OpenAI compensated them retroactively as part of the forward licensing agreement. The implication for my client: content already used in training has no leverage in future licensing negotiations. It's been consumed.

The strategic calculus changed. If my client ever wanted to be in a licensing conversation — which, as content sites mature and AI licensing becomes a real revenue line, becomes increasingly plausible — having given away training access for free destroyed that negotiating position. We put the training crawlers back on the deny list. The client understood the argument once I framed it that way.

Related reading: Content Licensing for AI in 2026 goes deep on why that negotiating position matters.

The BCA Framework: Block, Channel, Allow

After both pivots, I arrived at a decision framework I now call BCA — Block, Channel, Allow — which I apply when auditing any client's AI crawler posture.

Block applies to training crawlers. These bots take content for model training. If you have not agreed to this use and might someday want licensing revenue, block them. The cost of allowing is permanent; the cost of blocking is recoverable if you later choose to allow or license.

Channel applies to crawlers where you want selective access. You might allow the main site but disallow subscriber-only content. You might allow your blog but block your API documentation. robots.txt path-level rules handle this, as does serving different llms.txt content. See the llms.txt piece for the disclosure layer.

Allow applies to retrieval crawlers that power products with genuine referral potential, and obviously to all legitimate search engine crawlers. This isn't charity — it's distribution. Being cited in Perplexity answers is a form of traffic acquisition.

The BCA framework also has a fourth implicit category: Tarpit — for shadow crawlers and persistent bad actors that bypass robots.txt entirely. That's handled at the WAF layer. Full patterns in the anti-AI scraping piece.

Two Positions I Hold That Most Practitioners Reject

First: blocking AI crawlers in robots.txt is not primarily a technical measure. It's a legal measure. Compliant bots honor robots.txt, so the technical protection only works for compliant bots — which are precisely the bots that might honor a cease-and-desist or licensing request anyway. The robots.txt block creates a clear, public, timestamped record that you did not consent to crawling. That record is what matters if you ever litigate or negotiate.

Second: the obsession with blocking AI crawlers will look naive in three years. The better strategic position for most content sites is not to wall off AI systems but to structure access so it generates revenue. The publishers who fought access entirely are mostly losing. The publishers who negotiated structured access — AP, Reddit, several news organizations — are getting paid. Blocking is a holding action. Licensing is an endgame.

The Mistake Worth Naming

In January 2025, I deployed a robots.txt for a client that used wildcard matching incorrectly. The intent was to block all AI bots not explicitly allowed. The execution was a User-agent: * Disallow block placed before the specific Allow rules for Googlebot. Because of how some crawlers process robots.txt sequentially versus hierarchically, the result was a period where the site's crawler access was ambiguous for about nine days until I caught it in GSC coverage reports.

The actual ranking impact was minimal — Google continued crawling based on cached robots.txt. But it was a real mistake made under time pressure. The lesson: always test robots.txt changes against Google's robots.txt Tester in Search Console before deploying. Always. And deploy changes to a staging environment first when the site has more than 50,000 indexed pages.

How I Monitor Bot Traffic Now

I run a monthly log analysis across all sites I manage. The key metrics I track for AI crawlers specifically:

Total requests by User-Agent, filtered to my deny list. If a bot I've blocked is still generating server hits, it means it's ignoring robots.txt — which itself is useful information about that bot's compliance posture.

Crawl rate (requests per hour) for unblocked retrieval bots. PerplexityBot in particular can hit crawl rates that affect server performance on smaller hosts. I rate-limit at the Cloudflare level as a secondary control.

New User-Agents not on my known list. Once a quarter I review log entries for User-Agent strings I haven't seen before. Roughly 2–3 new AI-adjacent strings appear per quarter. Some are new products; some are existing products that updated their bot name; some are hard to identify at all.

The Dark Visitors website is the best single resource I've found for keeping the list current. Their database is actively maintained and includes documentation on each bot's declared purpose, compliance statements, and known behavior. I check it monthly.

The Honest State of Control in 2026

Here's the uncomfortable truth about bot management in May 2026: you have good control over compliant bots and almost no control over non-compliant ones. The 40+ named crawlers on my deny list are mostly the polite ones. The scrapers that genuinely don't care what your robots.txt says are running through rotating proxies, consumer ISP IP ranges, and headless browsers that are nearly indistinguishable from real users at the HTTP layer.

The WAF helps. Tarpitting helps. Rate limiting helps. None of it is complete. The realistic goal is not to achieve perfect exclusion but to create enough friction that casual scrapers move on to easier targets, to create a legal record of non-consent for the compliant bots, and to structure your access so that the bots you want (retrieval, search) get what they need while the ones you don't want (training scrapers) face enough barriers that the signal is clear.

That's achievable. Complete control over what AI systems do with your content is not.


Filed under: Technical SEO, AI Crawlers, Server Configuration | Last verified: May 2026 | Reading time: ~16 minutes

YOUR READING CHECKLIST

Make the ideas stick.

Mark the sections you’ve worked through. Saved in this browser.

0 of 4 reviewed
Andrii Stanetskyi
ABOUT THE AUTHOR

Andrii Stanetskyi

Head of SEO / Technical SEO Lead based in Tallinn, Estonia. Technical architecture, enterprise eCommerce, Python automation, and AI-assisted workflows.

More about Andrii ↗
LET’S FIND THE REAL BOTTLENECK

A clearer picture.
A practical next step.

Get a focused SEO audit or a consultation on your next technical decision. We’ll agree on the scope and fee before any work begins.

01 / Diagnose02 / Prioritize03 / Plan
How can I help?

Scope and fee agreed before any work begins.

Choose your language

Explore SEO services in 26 languages. Journal articles retain their original language.

ENEnglish↗DEDeutsch↗FRFrançais↗ESEspañol↗ITItaliano↗PTPortuguês↗NLNederlands↗PLPolski↗SVSvenska↗DADansk↗FISuomi↗NONorsk↗ETEesti↗LVLatviešu↗LTLietuvių↗CSČeština↗RORomână↗HUMagyar↗ELΕλληνικά↗BGБългарски↗HRHrvatski↗SKSlovenčina↗SLSlovenščina↗RUРусский↗UKУкраїнська↗TRTürkçe↗
LET’S WORK ON YOUR WEBSITE
A CLEAR NEXT STEP

Let’s talk
about your site.

A focused SEO audit or a conversation about a specific challenge. Tell me where you are and what you want to change.

Andrii Stanetskyi
Andrii StanetskyiHead of SEO / Technical SEO Lead
[email protected] ↗
How can I help?

Scope and fee agreed before any work begins.