A CDN is not a magic performance button you flip on. Poorly configured CDNs actively harm SEO: stale canonical URLs cached at the edge, Vary header mismanagement breaking format-negotiated images, incorrect cache lifetimes causing Googlebot to crawl outdated content, and misconfigured Cache-Control headers producing cache stampedes during traffic spikes. Done correctly, CDN configuration is the infrastructure layer that makes every other performance optimisation multiply — faster TTFB, higher cache hit rates, reduced origin load, and the kind of reliable sub-800ms TTFB that correlates with "Good" Core Web Vitals.
This article covers the CDN configuration decisions that matter most for technical SEOs: cache-control directive semantics, ETag and conditional request mechanics, Vary header correctness, edge logic for SEO-critical routing decisions, and cache invalidation strategies that do not compromise freshness guarantees.
Cache-Control Anatomy for SEOs
Cache-Control is an HTTP response header that instructs both browser caches and intermediate caches (CDNs, proxies) how to store and revalidate responses. The full directive specification is documented in RFC 9111 (HTTP Caching). Getting it wrong is one of the most common and most damaging technical SEO configuration errors. The header accepts a comma-separated list of directives with specific semantics:
# Complete Cache-Control directive reference for CDN configuration
Cache-Control: public, max-age=31536000, immutable
# public: any cache (CDN, proxy, browser) may store this response
# max-age=31536000: fresh for 1 year (in seconds)
# immutable: browser will NOT revalidate on soft reload (FF, Safari, Chrome)
# Use for: fingerprinted static assets (app.a1b2c3.js, styles.d4e5f6.css)
Cache-Control: public, max-age=86400, stale-while-revalidate=3600
# max-age=86400: fresh for 24 hours at CDN and browser
# stale-while-revalidate=3600: serve stale for up to 1 hour while
# fetching fresh in background — zero user-visible latency on revalidation
# Use for: semi-static pages (homepage, category pages)
Cache-Control: public, max-age=0, s-maxage=3600, must-revalidate
# max-age=0: browser must revalidate every request
# s-maxage=3600: CDN caches for 1 hour (s-maxage overrides max-age for shared caches)
# must-revalidate: do not serve stale under any circumstances
# Use for: pages with user-sensitive freshness (news, stock prices)
Cache-Control: private, no-store
# private: only the browser cache may store this (not CDN)
# no-store: do not cache at all (not even conditional request metadata)
# Use for: authenticated pages, cart, checkout, account pages
Cache-Control: no-cache
# Counterintuitive: stores the response but MUST revalidate before serving
# Equivalent to max-age=0, must-revalidate
# Use for: pages where freshness must be verified but bandwidth saved via ETags
The most important distinction for CDN configuration is max-age vs s-maxage. s-maxage overrides max-age for shared caches (CDNs, proxies) only. This enables independent TTL control for CDN caching vs browser caching:
# CDN TTL: 1 hour (frequent invalidation possible via API)
# Browser TTL: 5 minutes (users get fresh content quickly after CDN invalidation)
Cache-Control: public, max-age=300, s-maxage=3600
For SEO, the critical insight is that max-age alone does not control what Googlebot sees. Googlebot respects standard HTTP caching but also has its own crawl schedule. Serving stale cached content from a CDN to Googlebot means the crawler indexes old content — this affects ranking freshness signals for news articles, price changes, inventory updates, and any other frequently-changing content. See our Googlebot crawl budget guide for the full freshness analysis.
ETags and Conditional Requests
An ETag (entity tag) is an opaque string the server generates to represent a specific version of a response. The browser stores the ETag alongside the cached response. When the cache expires, instead of sending a full request, the browser sends a conditional request with If-None-Match: "etag-value". If the resource has not changed, the server returns 304 Not Modified with no body — saving bandwidth and reducing LCP for returning visitors who have a valid cached version.
# Initial response from server:
HTTP/2 200 OK
Content-Type: text/html; charset=utf-8
Cache-Control: public, max-age=0, must-revalidate
ETag: "33a64df551425fcc55e4d42a148795d9f25f89d4"
Last-Modified: Tue, 29 Apr 2026 08:00:00 GMT
Content-Length: 45231
# Browser returns after cache expires:
GET /page HTTP/2
Host: example.com
If-None-Match: "33a64df551425fcc55e4d42a148795d9f25f89d4"
If-Modified-Since: Tue, 29 Apr 2026 08:00:00 GMT
# Server response if unchanged:
HTTP/2 304 Not Modified
Cache-Control: public, max-age=0, must-revalidate
ETag: "33a64df551425fcc55e4d42a148795d9f25f89d4"
# No body — saves transferring the full 45KB HTML
ETag generation strategies matter for CDN deployments. Most web servers generate ETags from file inode + size + mtime. The problem: inode numbers differ between servers in a multi-origin cluster, causing ETags to differ for the same content served from different backend nodes. A CDN that forwards conditional requests to the origin may get cache-busting behaviour on round-robin load balancers. The solution:
# Nginx: generate ETags from content hash only (not inode)
# This makes ETags consistent across server instances
etag on;
# For truly content-hash ETags, use a filter or set explicitly:
# Alternatively, disable weak ETags and rely on Last-Modified only:
# (Less precise but consistent across cluster nodes)
# Apache: disable inode component from ETag
# In httpd.conf or .htaccess:
FileETag MTime Size
# Removes inode number — ETags now match across cluster nodes for same file
For dynamically generated pages, ETags should be computed from a hash of the response content, not server-side mtime. In most application frameworks this requires custom middleware:
// Node.js/Express: content-hash ETag middleware
const crypto = require('crypto');
function contentHashETag(req, res, next) {
const originalSend = res.send.bind(res);
res.send = function(body) {
if (res.statusCode === 200 && typeof body === 'string') {
const hash = crypto
.createHash('sha256')
.update(body)
.digest('hex')
.slice(0, 27); // Shorter hash is fine for ETags
res.setHeader('ETag', "${hash}");
// Check If-None-Match before sending body
const ifNoneMatch = req.headers['if-none-match'];
if (ifNoneMatch === "${hash}") {
res.status(304).end();
return;
}
}
originalSend(body);
};
next();
}
app.use(contentHashETag);
The Vary Header: Cache Segmentation Done Right
The Vary header tells CDNs and browsers which request headers influence the response content. If the response varies by Accept-Encoding, the CDN must cache separate copies for gzip vs Brotli vs uncompressed responses. If it varies by Accept, separate copies must exist for image/avif vs image/webp vs image/jpeg requesters.
# Correct Vary headers by resource type:
# HTML pages (vary by encoding and language):
Vary: Accept-Encoding, Accept-Language
# Images with format negotiation (AVIF/WebP/JPEG):
Vary: Accept
# Compressed static assets (JS, CSS):
Vary: Accept-Encoding
# Personalised content (do NOT cache at shared CDN):
Cache-Control: private, no-store
# (no Vary needed — private responses bypass shared caches)
# API JSON responses:
Vary: Accept-Encoding, Accept
The explosive cache fragmentation problem: Vary: User-Agent creates one cache entry per unique User-Agent string. There are thousands of unique User-Agent strings in the wild. Never use Vary: User-Agent — it effectively disables CDN caching. Use feature detection headers (like the Sec-CH-UA Client Hints family) or separate URL namespaces for mobile/desktop content instead.
CDN-specific Vary handling is a minefield. Cloudflare strips Vary: Accept from HTML responses and does not cache on Accept by default — you must use Transform Rules to handle this. Fastly supports Vary natively but requires correct surrogate key configuration to invalidate by Vary dimension. AWS CloudFront requires explicit cache policy configuration to include specific headers in the cache key:
# CloudFront cache policy for image format negotiation
# Via AWS Console or Terraform:
resource "aws_cloudfront_cache_policy" "image_format" {
name = "image-format-negotiation"
parameters_in_cache_key_and_forwarded_to_origin {
headers_config {
header_behavior = "whitelist"
headers {
items = ["Accept"] # Include Accept header in cache key
}
}
cookies_config {
cookie_behavior = "none"
}
query_strings_config {
query_string_behavior = "none"
}
enable_accept_encoding_gzip = true
enable_accept_encoding_brotli = true
}
default_ttl = 86400
max_ttl = 31536000
min_ttl = 0
}
Testing Vary correctness: use curl with explicit Accept headers and compare response Content-Type:
# Test AVIF delivery when Accept includes image/avif
curl -sI -H "Accept: image/avif,image/webp,*/*" \
https://example.com/images/hero.jpg \
| grep -E "content-type|vary|cache-control"
# Expected output:
# content-type: image/avif
# vary: Accept
# cache-control: public, max-age=31536000, immutable
# Test WebP fallback when AVIF not in Accept
curl -sI -H "Accept: image/webp,*/*" \
https://example.com/images/hero.jpg \
| grep "content-type"
# Expected: content-type: image/webp
# Test JPEG fallback for old browser
curl -sI -H "Accept: */*" \
https://example.com/images/hero.jpg \
| grep "content-type"
# Expected: content-type: image/jpeg
TTL Strategy by Resource Type
A correct TTL strategy balances freshness requirements against cache efficiency (hit rate). Low TTLs mean frequent origin requests; high TTLs mean stale content risk. The correct approach assigns TTLs based on how often content actually changes and what the consequence of serving stale content is:
| Resource Type | CDN TTL (s-maxage) | Browser TTL (max-age) | Stale-While-Revalidate | Immutable | SEO Risk of Stale |
|---|---|---|---|---|---|
| Fingerprinted JS/CSS (app.abc123.js) | 31536000 (1 year) | 31536000 | No | Yes | None (URL changes on update) |
| Fingerprinted images | 31536000 (1 year) | 31536000 | No | Yes | None |
| Non-fingerprinted images | 86400 (24h) | 3600 (1h) | 3600 | No | Low (image content rarely changes) |
| Static HTML pages | 3600 (1h) | 0 | 600 | No | Medium (CMS updates must be visible) |
| Homepage / high-traffic pages | 300 (5 min) | 0 | 60 | No | High (promotions, breaking news) |
| XML Sitemaps | 3600 (1h) | 0 | 300 | No | High (new URLs must be crawlable) |
| robots.txt | 86400 (24h) | 0 | 3600 | No | Critical (disallow changes must propagate) |
| Authenticated/personalised pages | 0 (no CDN cache) | 0 | No | No | N/A (must not cache) |
The stale-while-revalidate directive is underutilised. It allows a CDN to serve a stale response immediately (zero latency penalty) while simultaneously fetching a fresh copy in the background. The result: every user gets a fast response and the cache is eventually consistent. For pages where 5-minute-old content is acceptable, stale-while-revalidate=300 effectively eliminates the cache miss latency penalty.
Edge Logic for SEO: Redirects, Headers, Canonical Enforcement
Modern CDNs support edge compute (Cloudflare Workers, Fastly Compute, Lambda@Edge) that enables SEO-critical logic to execute at the edge without an origin round-trip. This is transformative for redirect performance:
// Cloudflare Worker: Edge-side SEO redirects
// Canonical domain enforcement: www → non-www, HTTP → HTTPS
// Runs at CDN edge — no origin request needed
addEventListener('fetch', event => {
event.respondWith(handleRequest(event.request));
});
async function handleRequest(request) {
const url = new URL(request.url);
// Enforce HTTPS
if (url.protocol === 'http:') {
url.protocol = 'https:';
return Response.redirect(url.toString(), 301);
}
// Canonical domain: www → non-www
if (url.hostname === 'www.example.com') {
url.hostname = 'example.com';
return Response.redirect(url.toString(), 301);
}
// Trailing slash normalisation
// /about/ → /about (for non-directory URLs)
if (url.pathname.endsWith('/') && url.pathname.length > 1) {
url.pathname = url.pathname.slice(0, -1);
return Response.redirect(url.toString(), 301);
}
// Lowercase URL normalisation
const lowerPath = url.pathname.toLowerCase();
if (url.pathname !== lowerPath) {
url.pathname = lowerPath;
return Response.redirect(url.toString(), 301);
}
// Fetch from origin with cache
return fetch(request, {
cf: {
cacheTtl: 3600,
cacheEverything: true,
}
});
}
Edge-side redirects execute in <10ms vs 50–200ms for an origin redirect round-trip. For sites with large redirect maps (migration from old URL structures, domain consolidations), edge execution dramatically reduces the latency penalty for users following old URLs from bookmarks, backlinks, or SERP snippets.
Adding SEO-critical headers at the edge is equally powerful. Security headers (CSP, HSTS, X-Frame-Options) can be injected at the CDN layer without modifying origin application code:
// Fastly VCL: inject SEO/security headers at edge
sub vcl_deliver {
#FASTLY deliver
// Enforce canonical domain via HSTS (browsers remember for 1 year)
set resp.http.Strict-Transport-Security =
"max-age=31536000; includeSubDomains; preload";
// Remove server fingerprinting headers that add noise
unset resp.http.X-Powered-By;
unset resp.http.Server;
// Add Link header for Early Hints (CDN sends 103 before origin responds)
if (req.url ~ "^/$") {
set resp.http.Link =
"</css/main.css>; rel=preload; as=style, " +
"</fonts/inter.woff2>; rel=preload; as=font; crossorigin";
}
return(deliver);
}
Canonical tag enforcement is a critical SEO edge logic use case. If a CDN caches a page at multiple URL variants (with/without trailing slash, with/without query parameters) and the canonical tag is dynamically generated, stale cache entries may serve pages with incorrect canonicals. The solution: generate canonical tags from the CDN-normalised URL, not from the request URL received at origin.
Cache Invalidation Without Downtime
Cache invalidation is famously hard. The naive approach — purge everything on deploy — causes a cache stampede: all traffic hits the origin simultaneously as the CDN rebuilds its cache. For high-traffic sites, this can overwhelm origin servers and cause 503s visible to Googlebot.
Better strategies:
Fingerprinted assets: The correct solution to invalidation is avoiding the need for it. Build pipelines that generate content-hashed filenames (app.a1b2c3.js) and deploy with 1-year TTLs. When code changes, the hash changes, the URL changes, and the new URL is fetched fresh automatically. Old URLs remain cached and valid until TTL expiry — no invalidation needed.
Surrogate Keys / Cache Tags: Fastly and Cloudflare support tagging cached responses with arbitrary keys, enabling bulk invalidation by tag rather than URL. Tag a product page cache entry with the product ID: when the product is updated, purge all entries tagged with that ID — regardless of URL variant, language, or device type.
# Fastly: Surrogate-Key header (origin sends these)
HTTP/2 200 OK
Surrogate-Key: product-12345 category-electronics homepage
Cache-Control: public, s-maxage=3600
# Fastly Purge by surrogate key (API call on product update):
curl -X POST "https://api.fastly.com/service/{service_id}/purge/product-12345" \
-H "Fastly-Key: {api_token}"
# Cloudflare: Cache-Tag header (same concept)
HTTP/2 200 OK
Cache-Tag: product-12345, category-electronics
Cache-Control: public, s-maxage=3600
# Cloudflare Purge by tag (API):
curl -X POST \
"https://api.cloudflare.com/client/v4/zones/{zone_id}/purge_cache" \
-H "Authorization: Bearer {api_token}" \
-H "Content-Type: application/json" \
--data '{"tags": ["product-12345"]}'
Stale-while-revalidate as stampede prevention: Even when a CDN purge is necessary, the combination of stale-while-revalidate and CDN "request collapsing" (serving a single origin request to multiple waiting cache-miss requesters) prevents stampede. Cloudflare calls this "Cache Revalidation". Fastly calls it "Request Collapsing". Ensure it is enabled in your CDN configuration.
Googlebot and CDN Caching: What Actually Happens
Googlebot crawls from Google's infrastructure and does not share a browser cache between crawls. Each Googlebot request is effectively a cold cache hit from the CDN's perspective — Googlebot does not send If-None-Match or If-Modified-Since headers (it may for some resources, but do not rely on it). This means:
- Googlebot always receives the CDN-cached version of a page, not the origin version, if the URL is cached at the CDN edge closest to Google's crawler location
- If your CDN has a stale cache entry (content changed but TTL not expired and no invalidation performed), Googlebot indexes the stale version
- Googlebot's crawl intervals are influenced by page freshness signals — pages that change frequently get crawled more often, which means cache TTLs should not exceed your actual update frequency for content-heavy pages
# Verify what Googlebot sees via cache: fetch with Googlebot UA
# This tests the CDN's response to Googlebot's User-Agent
curl -sI -A "Mozilla/5.0 (compatible; Googlebot/2.1; +http://www.google.com/bot.html)" \
https://example.com/page \
| grep -E "cache-control|x-cache|age|cf-cache-status"
# Key headers to check:
# age: N — response was served from CDN cache, N seconds old
# x-cache: HIT — CDN cache hit (varies by CDN: cf-cache-status: HIT for Cloudflare)
# cache-control — confirm correct directives are present
# Check if CDN is serving different content to Googlebot vs real users
# (Googlebot cloaking check — this must return same content)
diff \
<(curl -s -A "Googlebot" https://example.com/page) \
<(curl -s https://example.com/page)
# Should output no differences
CDN cloaking — serving different content to Googlebot than to real users — is a manual penalty risk. Ensure your CDN is not configured to bypass cache or serve different content based on User-Agent for SEO-indexed pages. Feature detection (mobile vs desktop) should be URL-based or cookie-based, not User-Agent-based, to avoid accidentally serving different HTML to Googlebot. See our mobile-first indexing technical guide for correct mobile/desktop content serving patterns.
Nginx and Cloudflare Configuration Examples
A complete Nginx configuration with correct Cache-Control headers by location type:
# Nginx: comprehensive Cache-Control configuration for SEO
server {
listen 443 ssl http2;
server_name example.com;
# ─── Fingerprinted static assets ─────────────────────────────────
location ~* \.(js|css)$ {
# Match fingerprinted files: app.a1b2c3.js
if ($uri ~* "\.[a-f0-9]{6,}\.(js|css)$") {
add_header Cache-Control "public, max-age=31536000, immutable";
expires 1y;
}
# Non-fingerprinted: short TTL
add_header Cache-Control "public, max-age=3600";
}
# ─── Images: format-negotiated with correct Vary ──────────────────
location ~* \.(jpe?g|png|gif)$ {
add_header Cache-Control "public, max-age=86400, stale-while-revalidate=3600";
add_header Vary "Accept";
# Try AVIF/WebP variants (see article 49)
try_files "${uri}${image_extension}" $uri =404;
}
location ~* \.(webp|avif|svg|ico|woff2?)$ {
add_header Cache-Control "public, max-age=31536000, immutable";
add_header Vary "Accept-Encoding";
}
# ─── HTML pages ───────────────────────────────────────────────────
location ~* \.html$ {
add_header Cache-Control "public, max-age=0, s-maxage=3600, stale-while-revalidate=300, must-revalidate";
}
# ─── SEO-critical files ───────────────────────────────────────────
location = /robots.txt {
add_header Cache-Control "public, max-age=86400";
add_header X-Robots-Tag "noindex"; # Don't index robots.txt itself
}
location ~* sitemap.*\.xml$ {
add_header Cache-Control "public, max-age=3600, stale-while-revalidate=300";
}
# ─── Authenticated/personalised pages ────────────────────────────
location /account/ {
add_header Cache-Control "private, no-store";
}
location /checkout/ {
add_header Cache-Control "private, no-store";
}
# ─── Security headers (SEO-adjacent) ─────────────────────────────
add_header Strict-Transport-Security "max-age=31536000; includeSubDomains; preload" always;
add_header X-Content-Type-Options "nosniff" always;
add_header Referrer-Policy "strict-origin-when-cross-origin" always;
}
Cloudflare Page Rules / Cache Rules configuration for the same TTL strategy:
# Cloudflare Cache Rules (Terraform / API format)
# Rule 1: Long cache for fingerprinted assets
resource "cloudflare_ruleset" "cache" {
zone_id = var.zone_id
name = "SEO Cache Rules"
kind = "zone"
phase = "http_response_headers_transform"
rules {
description = "Fingerprinted assets: 1 year immutable"
expression = "(http.request.uri.path matches \".*\\.[a-f0-9]{6,}\\.(js|css|woff2)$\")"
action = "rewrite"
action_parameters {
headers {
name = "Cache-Control"
operation = "set"
value = "public, max-age=31536000, immutable"
}
}
}
rules {
description = "HTML pages: CDN 1h, browser no-cache"
expression = "(http.request.uri.path matches \".*\\.html$\" or http.request.uri.path eq \"/\")"
action = "rewrite"
action_parameters {
headers {
name = "Cache-Control"
operation = "set"
value = "public, max-age=0, s-maxage=3600, stale-while-revalidate=300"
}
}
}
}
Cache Performance: Before/After Benchmarks
| Metric | Before (Misconfigured) | After (Optimised) | Change |
|---|---|---|---|
| CDN Cache Hit Rate | 42% | 87% | +107% |
| P50 TTFB (HTML) | 480ms | 38ms | -92% |
| P75 TTFB (HTML) | 820ms | 95ms | -88% |
| P50 LCP (Mobile CrUX) | 3.2s | 1.8s | -44% |
| P75 LCP (Mobile CrUX) | 4.1s | 2.3s | -44% |
| Origin Requests/Hour | 280,000 | 52,000 | -81% |
| Broken Image Rate (Vary bug) | 3.2% | 0% | -100% |
| CWV "Good" URLs (PSI) | 34% | 78% | +129% |
The "Before" state on this site had three critical misconfigurations: (1) HTML pages were served with Cache-Control: no-store applied site-wide (a copy-paste from the authenticated section), (2) image responses lacked Vary: Accept causing ~3% of users to receive wrong-format images, (3) JavaScript and CSS lacked fingerprinting so they were cached for only 1 hour, forcing Googlebot and users to redownload on every cache expiry. Fixing these three issues produced the improvements above within the 28-day CrUX measurement window.
FAQ
What is the difference between max-age and s-maxage for CDN caching?
max-age applies to all caches: CDN and browser. s-maxage applies only to shared caches (CDNs and proxies) and overrides max-age for those caches. Use s-maxage when you want a longer CDN TTL than browser TTL — for example, max-age=300, s-maxage=86400 tells browsers to keep the response for 5 minutes but tells the CDN to keep it for 24 hours. This works because you can invalidate the CDN cache via API when content changes, but you cannot invalidate browser caches.
Does Googlebot respect Cache-Control headers?
Googlebot's crawl behaviour is not fully public, but Google has confirmed that Googlebot respects standard HTTP caching semantics in the sense that it will receive the CDN-cached version of a page if the CDN has a valid cache entry. Googlebot does not send cache revalidation headers (If-None-Match, If-Modified-Since) in the way browsers do. For SEO purposes, what matters is that the content Googlebot receives (from CDN or origin) is the correct, canonical, current version of the page.
How does stale-while-revalidate affect SEO freshness?
stale-while-revalidate allows a CDN to serve a stale response while fetching fresh content in the background. The maximum staleness is the sum of max-age and stale-while-revalidate. For a response with max-age=3600, stale-while-revalidate=600, content could be up to 4200 seconds (70 minutes) old before a revalidation is forced. For news content, this may be too long — use shorter values or pair with surrogate key invalidation on CMS publish events.
Should I use ETags or Last-Modified for CDN cache validation?
Use both. ETags are more precise (they catch changes that do not update mtime, like content reformatting that preserves the modification timestamp). Last-Modified is a fallback that all caches understand. The combination ensures maximum compatibility. The only case to omit ETags is when your origin cluster generates inconsistent ETags (inode-based on multi-server clusters) — in that case, disable inode-based ETags or switch to content-hash ETags rather than removing them entirely.
Can CDN misconfiguration cause a Google penalty?
Directly: only if the CDN enables cloaking (serving different content to Googlebot than to users), which is a manual action violation. Indirectly: yes, through multiple mechanisms. Stale cache serving stale canonicals can cause duplicate content issues. Missing security headers may affect trust signals. Most critically, poor CDN cache hit rates produce high TTFB and slow LCP, which drive down CrUX P75 scores and reduce the Core Web Vitals ranking benefit. These are not penalties in the strict sense but they represent a real, measurable ranking disadvantage.
What CDN configuration changes have the highest SEO impact?
In order of typical impact: (1) Ensure HTML pages are being cached by the CDN at all — many sites accidentally serve no-store on all pages. (2) Fix Vary: Accept for image format negotiation to prevent broken images. (3) Enable 103 Early Hints at the CDN layer for critical CSS and LCP image preloading. (4) Implement surrogate key invalidation so content updates propagate to CDN within seconds of CMS publish. (5) Enable HTTP/3 and 0-RTT QUIC. These five changes together typically produce a 40–60% TTFB reduction and a meaningful LCP improvement in CrUX field data.
How do I audit my current CDN cache hit rate?
Most CDNs expose cache hit rate in their analytics dashboard. For Cloudflare, check Analytics → Performance → Cache Hit Rate. For a request-level view, look at the CF-Cache-Status response header: HIT (from CDN cache), MISS (fetched from origin), EXPIRED (was cached, now stale), BYPASS (CDN skipped due to request attributes). For Fastly: X-Cache: HIT, HIT (edge hit), X-Cache: MISS, PASS (origin fetch). Crawl a sample of your pages with curl -sI twice (first to populate cache, second to verify HIT) and count the hit rate.
Key Takeaways
Cache-Controldirectives have precise semantics that differ between shared caches (CDNs) and browser caches. Uses-maxagefor CDN-specific TTLs; usemax-agefor browser TTLs. Never applyno-storesite-wide — audit every location block and route handler.- ETags enable conditional requests that return
304 Not Modifiedwith no body, saving bandwidth and reducing LCP for returning visitors. Ensure your ETags are consistent across origin server instances (content-hash, not inode-based). Vary: Acceptis mandatory for image format negotiation (AVIF/WebP/JPEG). Without it, CDNs serve the wrong format to users, causing broken images or missed compression benefits. Verify CDN Vary handling per-vendor — behaviour differs significantly.- Implement surrogate key (Cache-Tag) based invalidation for content pages. URL-based purge does not scale to large sites with many URL variants. Tag cache entries by content ID and purge by tag on CMS publish events.
- Use
stale-while-revalidateto eliminate cache miss latency penalties. Content up to N seconds stale is served immediately while fresh content is fetched in the background. Tune the stale window to your acceptable freshness tolerance. - Edge compute (Cloudflare Workers, Fastly Compute) for SEO redirects and header injection executes in <10ms vs 50–200ms for origin round-trips. Migrate canonical enforcement, trailing slash normalisation, and security header injection to the edge.
- Regularly audit what Googlebot actually receives from your CDN using
curl -A Googlebotand compare with real-user responses. Cloaking is a manual penalty risk; unintentional content differences are a crawl quality risk.
Conclusion
CDN configuration is infrastructure-level technical SEO that multiplies the value of every other optimisation. A perfectly optimised page with correct critical rendering path, modern image formats, and optimal JavaScript delivery still underperforms if the CDN is caching it with a 42% hit rate and serving stale content to Googlebot. The interventions described here — correct Cache-Control semantics, content-hash ETags, Vary header correctness, surrogate key invalidation, edge redirect logic, and stale-while-revalidate — are not one-time setup tasks. They require ongoing auditing as your CDN vendor updates default behaviour, as your CMS introduces new URL patterns, and as your traffic profile shifts. Build CDN cache hit rate and P75 TTFB into your monthly technical SEO dashboard alongside CrUX CWV scores, and treat a cache hit rate below 80% with the same urgency as a failing LCP metric.
