Published 19 May 2026. I've been testing scraping defenses on production sites since late 2023. This piece is about what still works. Several things I recommended in 2024 have since been bypassed; those are in here too.
The Defense Problem in 2026
AI scraping defense in 2026 is an arms race where the scraping side has structural advantages. Scrapers have access to the same AI tools that make fingerprint detection difficult — they can generate varied browser profiles, rotate residential proxies, mimic real human timing patterns, and dynamically adjust to detection signals. The defense side is always playing catch-up.
That's the honest context. Anyone selling you a complete solution is selling you something that will last six months before the scrapers adapt. The goal is not total protection — it's cost imposition. Make scraping your site expensive enough in time, infrastructure, and engineering resources that marginal scrapers move on. Serious, well-funded scrapers (the ones doing large-scale AI training data acquisition) are a harder problem that requires legal and contractual approaches alongside technical ones.
What follows is what I've found to actually work, as of May 2026, after testing on production environments. I'll be specific about where each technique fails.
What Failed: Techniques That No Longer Hold
Let me get the obituary section out of the way.
Pure User-Agent blocking was never adequate, but it's even more inadequate now. Scrapers have been spoofing Googlebot since 2019. In 2025, I documented scrapers cycling through 40+ User-Agent strings within a single scraping session on one of my sites. UA-based blocking catches compliant bots that don't need catching.
IP-based blocking erodes constantly. Residential proxy networks have expanded dramatically. In 2024, I was seeing scraping traffic from IPs in the same /24 ranges as major ISPs in Eastern Europe and Southeast Asia — real residential addresses, impossible to block without collateral damage to actual readers. The largest AI training data operations appear to use commercial proxy networks that rotate millions of IPs.
Simple CAPTCHAs are effectively dead for serious scrapers. Automated CAPTCHA solving services exist at scale, and multimodal AI has made visual CAPTCHAs trivial to solve programmatically. hCaptcha Enterprise holds up better than most, but adds enough friction that conversion rates suffer on content sites.
Rate limiting by IP alone is bypassable through the same proxy rotation that defeats IP blocking. Worth doing as a basic hygiene measure, but not a primary defense.
JavaScript rendering requirements (forcing JS execution to reveal content) now fail against headless Chrome-based scrapers. Playwright and Puppeteer are widely used in scraping pipelines. The render cost adds a modest penalty but sophisticated scrapers absorb it.
Tarpit Theory: Why Slowing Is Better Than Blocking
A tarpit, in network security, is a mechanism that absorbs attacker resources by responding slowly and seemingly legitimately. Applied to web scraping, a tarpit serves a scraper a response that looks valid — correct HTTP status codes, plausible content-type headers, structural HTML — but either delays delivery, feeds garbage data, or loops the scraper into infinite pagination. The scraper continues "successfully" but wastes its allocated resources.
Why tarpit instead of block? Three reasons.
First, blocking tells the scraper it's been detected. A 403 or a connection reset is a signal to change tactics. A slow, apparently successful response tells the scraper nothing is wrong. It keeps pulling. You're wasting its time without alerting its operators to change the approach.
Second, tarpitting is harder to detect from the scraper's monitoring dashboards. If you're running a scraping operation and seeing 403s, you investigate. If you're seeing 200 OKs with data coming in, you might not notice for days that the data is degraded or fabricated.
Third, for scrapers running on per-request cost models (cloud compute, proxy bandwidth), a tarpit imposes real financial cost. A block costs them one request. A 30-second delayed response costs them 30 seconds of proxy time and a held connection slot.
The tarpit approach does carry one significant risk: you need accurate detection before serving tarpit content. If you tarpit real users, you've created a site that's broken for your actual audience. Detection accuracy is everything.
Cloudflare Workers Tarpit: The Full Pattern
This is the Cloudflare Workers-based tarpit I currently run for clients on the Business and Enterprise plans. It requires Workers (available at $5/month on the paid plan) and ideally Bot Management for the fingerprint score.
// Cloudflare Workers — AI Scraper Tarpit
// Deploy at: Workers & Pages → Create Worker
// Add Route: yourdomain.com/* → this worker (except static assets)
// Last updated: May 2026
addEventListener('fetch', event => {
event.respondWith(handleRequest(event.request))
})
async function handleRequest(request) {
const url = new URL(request.url)
const ua = request.headers.get('User-Agent') || ''
const cfBot = request.cf?.botScore ?? 100
const isVerifiedBot = request.cf?.isBot ?? false
// ── STEP 1: ALWAYS PASS VERIFIED BOTS AND HIGH HUMAN SCORES ──────
// cf.botScore: 1 = definitely bot, 99 = definitely human
// Verified bots (Googlebot, Bingbot, etc.) are never tarpitted.
if (isVerifiedBot || cfBot > 30) {
return fetch(request)
}
// ── STEP 2: HONEYPOT URL CHECK ────────────────────────────────────
// Honeypot paths are in sitemap and llms.txt as plausible but
// infrequently visited by real users. Real users almost never hit them.
const honeypotPaths = [
'/archive/legacy/',
'/data/export/',
'/bulk-content/',
'/full-text-index/',
'/.well-known/ai-content.json',
]
const isHoneypot = honeypotPaths.some(p => url.pathname.startsWith(p))
// ── STEP 3: KNOWN BAD USER-AGENTS ────────────────────────────────
const badBotPatterns = [
'GPTBot', 'CCBot', 'Bytespider', 'Google-Extended',
'anthropic-ai', 'Claude-Web', 'cohere-ai', 'Amazonbot',
'Diffbot', 'DataForSeoBot', 'ImagesiftBot', 'omgili',
'python-requests', 'Go-http-client', 'libwww-perl',
'Scrapy', 'axios/', 'node-fetch', 'Wget/', 'curl/'
]
const isBadUA = badBotPatterns.some(p =>
ua.toLowerCase().includes(p.toLowerCase())
)
// ── STEP 4: DECIDE RESPONSE TYPE ─────────────────────────────────
const shouldTarpit = isHoneypot || isBadUA || cfBot < 10
if (!shouldTarpit) {
return fetch(request)
}
// ── STEP 5: TARPIT RESPONSE GENERATION ───────────────────────────
// Option A: Slow drip response (most effective for connection cost)
// Option B: Fake infinite pagination (most effective for data poisoning)
// Option C: Immediate 200 with degraded content (stealth option)
// We randomize to prevent pattern detection by scraper operators
const roll = Math.random()
if (roll < 0.33) {
return slowDripResponse(url)
} else if (roll < 0.66) {
return fakeInfinitePagination(url)
} else {
return degradedContentResponse(request, url)
}
}
// ── TARPIT VARIANT A: SLOW DRIP ───────────────────────────────────
// Returns a chunked response with artificial delays between chunks.
// Holds the scraper's connection open for 30-60 seconds.
// Real browsers handle this fine; scrapers with short timeouts give up.
async function slowDripResponse(url) {
const { readable, writable } = new TransformStream()
const writer = writable.getWriter()
const encoder = new TextEncoder()
const fakeHTML = generateFakePage(url.pathname)
const chunks = chunkString(fakeHTML, 512)
// Non-blocking: stream chunks with delays
const streamChunks = async () => {
await writer.write(encoder.encode('
Fastly Edge Config: A Different Architecture
Fastly's VCL-based configuration takes a different approach from Workers. It's less flexible than JavaScript-based Workers, but it's faster (VCL runs closer to bare metal) and integrates more naturally with Fastly's edge dictionary and rate-limiting systems.
// Fastly VCL — AI Bot Detection and Tarpit Routing
// Placed in vcl_recv subroutine
sub vcl_recv {
// ── DECLARE BOT INDICATOR ─────────────────────────────────────────
set req.http.X-Bot-Detected = "false";
// ── KNOWN AI BOT USER-AGENTS ──────────────────────────────────────
if (req.http.User-Agent ~ "(?i)(GPTBot|CCBot|Bytespider|Google-Extended|anthropic-ai|Claude-Web|cohere-ai|Amazonbot|Diffbot|DataForSeoBot|ImagesiftBot|omgili|python-requests|libwww-perl|Scrapy)") {
set req.http.X-Bot-Detected = "true";
}
// ── HONEYPOT PATH DETECTION ───────────────────────────────────────
if (req.url ~ "^/(archive/legacy|data/export|bulk-content|full-text-index)/") {
set req.http.X-Bot-Detected = "true";
}
// ── FASTLY EDGE RATE LIMITING ─────────────────────────────────────
// Requires Fastly Rate Limiting feature
// Rate limit: 30 req/min for detected bots, 200 req/min for humans
if (req.http.X-Bot-Detected == "true") {
if (ratelimit.check_rate(req.http.Fastly-Client-IP, "ai_bot_limit", 30, 60s)) {
// Rate limit exceeded — serve 429 immediately
error 429 "Too Many Requests";
}
// Route to tarpit backend (a slow-responding origin or synthetic VCL response)
set req.backend = F_tarpit_backend;
return(pass);
}
#FASTLY recv
return(pass);
}
// ── TARPIT SYNTHETIC RESPONSE ─────────────────────────────────────
// Fastly can generate synthetic responses without hitting origin.
// This burns scraper time without origin cost.
sub vcl_error {
if (obj.status == 429) {
set obj.http.Content-Type = "text/html; charset=UTF-8";
set obj.http.Retry-After = "3600";
synthetic {"
Request rate exceeded. Please retry after one hour.
"};
return(deliver);
}
}
// ── EDGE DICTIONARY FOR DYNAMIC BOT LIST ─────────────────────────
// Fastly Edge Dictionary allows real-time updates without deploys.
// Maintain a dictionary "ai_bots" with User-Agent substrings as keys.
// Update via Fastly API when new bots are identified.
// In vcl_recv, check: if (table.lookup(ai_bots, regsuball(req.http.User-Agent, "^(\w+).*$", "\1")) == "block")
// ── RESPONSE TIME MANIPULATION ────────────────────────────────────
// Fastly's sleep() VCL function can introduce response delays.
// Available in some Fastly configurations:
// if (req.http.X-Bot-Detected == "true") { sleep(5000); }
// (5000ms = 5 second delay before response — significant for bot operations)
Honeypot Links and Semantic Poisoning
Two techniques that work well in combination with tarpits.
Honeypot links are URLs embedded in your HTML that are invisible to real users (CSS display: none or positioned off-screen) but visible to scrapers parsing raw HTML. When a scraper follows a honeypot link, it triggers a flag. The IP, session, or fingerprint gets added to a deny list. I typically embed two to three honeypot links per page, pointing to paths like /archive/legacy/?ref=hpot. Real users don't see them; scrapers that crawl the HTML tree follow them.
The failure mode for honeypots: sophisticated scrapers now run visual rendering before extraction. If the scraper renders the page like a browser and only extracts visible content, the hidden links are never followed. This is why you also want to put honeypot paths in your llms.txt as "important pages" — scrapers reading your llms.txt and following its links will hit the honeypot.
Semantic poisoning is the darker version of degraded content. Instead of serving garbage, you serve content that is plausible but factually incorrect. The idea is that AI training datasets contaminated with convincingly wrong information produce worse models. I have mixed feelings about this approach. It works as a deterrent — if scrapers know your site contaminates data, they may avoid it. But if the poisoned content escapes into models trained on it, you've contributed misinformation to the ecosystem. I deploy it only for deliberately meaningless content: procedurally generated text with no factual claims.
TLS and HTTP/2 Fingerprinting
This is the most technically interesting defensive layer, and the one most practitioners haven't deployed yet.
TLS fingerprinting (specifically JA3 and JA4 fingerprinting) identifies the TLS client library being used to make a request, based on the specific combination of cipher suites, extensions, and elliptic curves offered in the ClientHello message. Real Chrome browsers produce a specific, recognizable JA3 hash. Python's requests library produces a different hash. Playwright running headless Chrome produces yet another — slightly different from real Chrome because of how it initializes the TLS stack.
Cloudflare exposes JA3 fingerprints via request.cf.tlsClientAuth and related fields in Workers. Fastly exposes them through its bot detection signals. The practical application: a request claiming to be Chrome 124 on macOS but carrying a JA3 hash consistent with Python-requests is almost certainly a scraper. You don't need to block it immediately — but you can route it to the tarpit.
HTTP/2 fingerprinting is similar in concept. The order and values of HTTP/2 SETTINGS frames, the WINDOW_UPDATE values, and the HEADERS frame priority signals differ between real browsers and automation tools. Akamai pioneered this; it's now implementable at the Cloudflare Workers layer with some engineering work, though not out-of-the-box.
The reason this matters: TLS/HTTP2 fingerprinting survives User-Agent spoofing. A scraper that correctly sets its User-Agent to "Mozilla/5.0 Chrome/124..." but is running on a Python HTTP library will still have the wrong TLS fingerprint. That mismatch is detectable without requiring the scraper to make any behavioral mistake.
The FDR Framework: Fingerprint, Degrade, Record
My current framework for deploying these techniques has three stages, which I abbreviate FDR.
Fingerprint: Identify suspicious requests using multiple signals — UA string, bot score, TLS fingerprint, honeypot triggers, behavioral timing. No single signal is enough. The combination produces a confidence score. I treat anything above a 70% confidence of being a non-compliant bot as a tarpit candidate.
Degrade: Don't block — degrade. Serve a slow response, degraded content, or fake pagination. The scraper operator doesn't know they're being tarpitted. They continue operating, believing they're collecting data. Their effective data quality drops while their costs stay constant.
Record: Log every tarpitted request. IP, UA, timestamp, path, TLS fingerprint. This log has two uses. First, it feeds your dynamic deny list — IPs that trigger tarpits repeatedly get added to a harder block. Second, the log is a legal record of who attempted to scrape you and when. In content licensing negotiations and any future litigation, having a timestamped record of unauthorized access attempts has real value.
The FDR framework connects to the broader access control strategy in Managing 40+ AI Crawlers and the legal/contractual layer in Content Licensing for AI in 2026.
Two Things the Industry Gets Wrong
First: the obsession with blocking is misplaced for most sites. Blocking named AI crawlers is table-stakes hygiene, but it addresses the compliant minority of the scraping problem. The volume-weighted majority of suspicious bot traffic I see in logs is from unnamed scrapers, rotated proxies, and headless browsers — none of which care about your robots.txt. Practitioners spend hours perfecting their robots.txt and five minutes on WAF configuration. The ratio should be inverted.
Second: semantic poisoning as a primary defense is ethically and strategically problematic. I see it recommended increasingly, and I'm skeptical. If your poisoned content ends up in a major LLM's training data, you've added misinformation to a system that millions of people rely on. You've also taken on legal and reputational risk. The technical argument for it is sound; the strategic argument ignores second-order effects. Use tarpit delays and fake pagination instead — they waste scraper resources without contaminating the information environment.
A Mistake I Made in Production
In October 2024, I deployed a Cloudflare Workers tarpit for a content client that was experiencing heavy scraping. The Worker was running the slow-drip variant I described above. It worked well — until I misconfigured the verified-bot exemption.
The isVerifiedBot check I'd implemented relied on request.cf.isBot, which I misread as "is a verified legitimate bot" — it actually means "Cloudflare believes this is a bot of any kind." The logic was inverted. I was passing through unverified bots while tarpitting Cloudflare's verified bot list, which includes Googlebot.
GSC started showing coverage drops after nine days. It took me three days to identify the Workers config as the cause because the error wasn't obvious — Googlebot was receiving slow-drip HTML responses that technically rendered, just very slowly. Google's crawler has a timeout of roughly 5 seconds for response initiation; our drip started within that window, so it wasn't outright rejected. But the crawl efficiency collapsed.
The fix was simple once identified. The lesson: always add explicit Googlebot log monitoring when deploying any edge Workers that might intercept crawler traffic. Check GSC coverage reports within 48 hours of any Workers deployment. And read the Cloudflare documentation on request.cf.isBot versus cf.verified_bot — they are not the same.
The Arms Race Position in May 2026
Anti-scraping defense in 2026 is a cost-imposition game, not a containment game. You are not going to stop a well-resourced AI company from scraping your site if they decide to do it. They have more resources, more engineering, and more motivation than any individual publisher has to stop them. What you can do is make it expensive, noisy, and legally risky.
Expensive: tarpits, delay responses, fake pagination. Costs the scraper time and money.
Noisy: detailed logging, honeypots, fingerprint tracking. Creates evidence you can use later.
Legally risky: clear robots.txt directives, llms.txt non-consent declarations, terms of service that explicitly prohibit automated data extraction. The Computer Fraud and Abuse Act, various EU data protection laws, and an increasing number of AI-specific regulations create exposure for scrapers who ignore explicit disallowances.
The combination of technical friction and legal exposure is more effective than either alone. Technical controls without legal records can be circumvented silently. Legal records without technical controls look like paper threats. Together, they create the conditions under which a licensing conversation becomes more attractive to an AI company than continued unauthorized access.
That's the actual goal. Not to block AI permanently — to make the calculus favor paying for access over stealing it.
Filed under: Technical SEO, Security, AI Crawlers, Cloudflare | Last verified: May 2026 | Reading time: ~17 minutes
