Screaming Frog, Sitebulb, and DeepCrawl are excellent tools for standard audits. They are not excellent for custom signal extraction, real-time crawl integration with BigQuery pipelines, differential crawling against a prior crawl snapshot, or extraction of JavaScript-rendered content at scale with fine-grained concurrency control. When the off-the-shelf crawlers reach their limits—and they always do on enterprise-scale sites—you build your own.
This article covers building a production-grade SEO crawler using Python and Scrapy: custom middleware for politeness, duplicate URL detection at scale, signal extraction pipelines, JavaScript rendering integration via Playwright, and streaming crawl data to BigQuery. The assumptions: you are comfortable with Python async, you have deployed Scrapy in production before, and you know why a naive BFS crawler will get you IP-banned in 20 minutes.
Crawler Architecture and Design Decisions
Before writing a line of Scrapy code, document your crawler's threat model and design constraints. The architecture decisions that matter most:
| Decision | Option A | Option B | Choose When |
|---|---|---|---|
| Concurrency model | Scrapy async (Twisted) | httpx + asyncio | Scrapy for plugin ecosystem; httpx for simpler pipelines |
| JS rendering | Playwright (full browser) | Splash (lightweight) | Playwright for React/Vue SPAs; Splash for basic JS execution |
| URL frontier | In-memory deque | Redis queue | Redis for distributed crawls >500k URLs |
| Seen URL store | Python set | Redis Bloom filter | Bloom filter for >10M URLs (memory vs. false positive tradeoff) |
| Output sink | CSV/JSON files | BigQuery streaming insert | BigQuery for real-time analysis and historical comparison |
| Rendering decision | Render all pages | Selective rendering | Selective: render only when initial HTML has <10 links |
The most important architectural insight: separate your URL frontier management, rendering decision, signal extraction, and output pipeline into decoupled components. Scrapy's spider → middleware → item pipeline architecture maps cleanly to this, but you must resist the temptation to put extraction logic in the spider itself.
Scrapy Project Structure for SEO
# seo_crawler/spiders/seo_spider.py
import scrapy
from scrapy.linkextractors import LinkExtractor
from scrapy.http import HtmlResponse
from urllib.parse import urlparse
import hashlib
import re
class SeoSpider(scrapy.Spider):
name = 'seo'
custom_settings = {
'CONCURRENT_REQUESTS': 16,
'CONCURRENT_REQUESTS_PER_DOMAIN': 4,
'DOWNLOAD_DELAY': 0.5,
'ROBOTSTXT_OBEY': True,
'COOKIES_ENABLED': False,
'REDIRECT_MAX_TIMES': 5,
'DEPTH_LIMIT': 10,
'DOWNLOAD_TIMEOUT': 30,
'REACTOR_THREADPOOL_MAXSIZE': 20,
}
def __init__(self, start_url=None, *args, **kwargs):
super().__init__(*args, **kwargs)
if not start_url:
raise ValueError("start_url required")
parsed = urlparse(start_url)
self.allowed_domains = [parsed.netloc]
self.start_urls = [start_url]
self.link_extractor = LinkExtractor(
allow_domains=self.allowed_domains,
deny_extensions=['pdf', 'jpg', 'jpeg', 'png', 'gif', 'svg',
'css', 'js', 'ico', 'woff', 'woff2', 'ttf',
'mp4', 'mp3', 'zip', 'gz'],
unique=True,
)
def parse(self, response):
# Only process HTML responses
if 'text/html' not in response.headers.get('Content-Type', b'').decode():
return
item = self.extract_seo_signals(response)
yield item
# Extract and follow links
for link in self.link_extractor.extract_links(response):
yield scrapy.Request(
link.url,
callback=self.parse,
errback=self.handle_error,
meta={
'referrer_url': response.url,
'referrer_depth': response.meta.get('depth', 0)
}
)
def extract_seo_signals(self, response):
from seo_crawler.items import SeoPageItem
item = SeoPageItem()
item['url'] = response.url
item['url_hash'] = hashlib.md5(response.url.encode()).hexdigest()
item['status_code'] = response.status
item['content_type'] = response.headers.get('Content-Type', b'').decode()
item['referrer_url'] = response.meta.get('referrer_url', '')
item['crawl_depth'] = response.meta.get('depth', 0)
item['response_time_ms'] = int(response.meta.get('download_latency', 0) * 1000)
return item
def handle_error(self, failure):
from seo_crawler.items import SeoPageItem
item = SeoPageItem()
item['url'] = failure.request.url
item['status_code'] = 0
item['error'] = str(failure.value)
yield item
Custom Middleware: Politeness, UA Rotation, Retry
Scrapy middleware is the correct place for cross-cutting concerns: User-Agent management, response caching, politeness enforcement, and retry logic with exponential backoff. Never put these in the spider.
# seo_crawler/middlewares.py
import time
import random
from scrapy import signals
from scrapy.http import HtmlResponse
from scrapy.exceptions import NotConfigured
import logging
logger = logging.getLogger(__name__)
class SeoUserAgentMiddleware:
"""Rotates UA to mimic real Googlebot behavior, with verified UA strings."""
GOOGLEBOT_UA = (
'Mozilla/5.0 (compatible; Googlebot/2.1; '
'+http://www.google.com/bot.html)'
)
DESKTOP_UA = (
'Mozilla/5.0 (Windows NT 10.0; Win64; x64) '
'AppleWebKit/537.36 (KHTML, like Gecko) '
'Chrome/122.0.0.0 Safari/537.36'
)
def __init__(self, crawler_ua='googlebot'):
self.ua = self.GOOGLEBOT_UA if crawler_ua == 'googlebot' else self.DESKTOP_UA
@classmethod
def from_crawler(cls, crawler):
return cls(crawler_ua=crawler.settings.get('CRAWLER_UA', 'desktop'))
def process_request(self, request, spider):
request.headers['User-Agent'] = self.ua
# Add headers Googlebot typically sends
request.headers['Accept'] = 'text/html,application/xhtml+xml'
request.headers['Accept-Language'] = 'en-US,en;q=0.9'
class AdaptivePolitenessMiddleware:
"""Adjusts delay based on server response codes and latency."""
def __init__(self, base_delay=0.5, max_delay=10.0):
self.base_delay = base_delay
self.max_delay = max_delay
self.domain_delays = {}
self.domain_error_counts = {}
@classmethod
def from_crawler(cls, crawler):
return cls(
base_delay=crawler.settings.getfloat('DOWNLOAD_DELAY', 0.5),
max_delay=crawler.settings.getfloat('MAX_DOWNLOAD_DELAY', 10.0),
)
def process_response(self, request, response, spider):
domain = request.url.split('/')[2]
if response.status == 429:
# Rate limited — back off exponentially
current_delay = self.domain_delays.get(domain, self.base_delay)
new_delay = min(current_delay * 2, self.max_delay)
self.domain_delays[domain] = new_delay
logger.warning(f"Rate limited on {domain}, delay now {new_delay}s")
time.sleep(new_delay)
elif response.status == 200:
# Success — gradually reduce delay toward base
current_delay = self.domain_delays.get(domain, self.base_delay)
self.domain_delays[domain] = max(
self.base_delay,
current_delay * 0.9
)
return response
class SelectiveRenderingMiddleware:
"""Routes requests to Playwright when HTML has insufficient link count."""
MIN_LINKS_THRESHOLD = 5 # Below this, suspect JS rendering needed
def process_response(self, request, response, spider):
if request.meta.get('playwright_rendered'):
return response # Already rendered, skip
if not isinstance(response, HtmlResponse):
return response
# Count links in raw HTML — if too few, re-fetch with Playwright
link_count = len(response.css('a[href]'))
if link_count < self.MIN_LINKS_THRESHOLD:
logger.info(f"Only {link_count} links found on {response.url}, queuing for JS render")
new_request = request.copy()
new_request.meta['playwright'] = True
new_request.meta['playwright_rendered'] = True
new_request.dont_filter = True
spider.crawler.engine.crawl(new_request, spider)
return response
SEO Signal Extraction Pipelines
The extraction pipeline processes each SeoPageItem and enriches it with parsed SEO signals. Keeping this in a pipeline (not the spider) enables parallel processing and easy unit testing.
# seo_crawler/pipelines/seo_signals.py
import re
from urllib.parse import urljoin, urlparse
from w3lib.html import remove_tags
class SeoSignalExtractionPipeline:
"""Extracts all SEO signals from raw response into structured fields."""
def process_item(self, item, spider):
response = item.get('_response')
if not response:
return item
self._extract_title(item, response)
self._extract_meta(item, response)
self._extract_headings(item, response)
self._extract_canonical(item, response)
self._extract_robots(item, response)
self._extract_links(item, response)
self._extract_schema(item, response)
self._extract_hreflang(item, response)
self._extract_core_web_vitals_hints(item, response)
# Remove raw response from item before serialization
item.pop('_response', None)
return item
def _extract_title(self, item, response):
title = response.css('title::text').get('').strip()
item['title'] = title
item['title_length'] = len(title)
item['title_has_brand'] = bool(re.search(r'\|', title))
def _extract_meta(self, item, response):
item['meta_description'] = response.css(
'meta[name="description"]::attr(content)'
).get('').strip()
item['meta_robots'] = response.css(
'meta[name="robots"]::attr(content)'
).get('').lower()
item['meta_viewport'] = bool(response.css('meta[name="viewport"]').get())
def _extract_headings(self, item, response):
item['h1_texts'] = response.css('h1::text').getall()
item['h1_count'] = len(item['h1_texts'])
item['h2_count'] = len(response.css('h2'))
item['h3_count'] = len(response.css('h3'))
def _extract_canonical(self, item, response):
canonical = response.css('link[rel="canonical"]::attr(href)').get('')
item['canonical_url'] = canonical.strip()
item['canonical_is_self'] = (
item['canonical_url'].rstrip('/') == response.url.rstrip('/')
)
item['canonical_is_cross_domain'] = bool(
canonical and urlparse(canonical).netloc != urlparse(response.url).netloc
)
def _extract_robots(self, item, response):
# Check both meta robots and X-Robots-Tag header
meta_robots = item.get('meta_robots', '')
x_robots = response.headers.get('X-Robots-Tag', b'').decode().lower()
combined = f"{meta_robots},{x_robots}"
item['is_noindex'] = 'noindex' in combined
item['is_nofollow'] = 'nofollow' in combined
item['x_robots_tag'] = x_robots
def _extract_links(self, item, response):
links = response.css('a[href]')
item['internal_links'] = []
item['external_links'] = []
base_domain = urlparse(response.url).netloc
for link in links:
href = link.attrib.get('href', '').strip()
if not href or href.startswith('#') or href.startswith('mailto:'):
continue
abs_url = urljoin(response.url, href)
rel = link.attrib.get('rel', '')
anchor_text = link.css('::text').get('').strip()[:200]
link_data = {
'url': abs_url,
'anchor': anchor_text,
'rel': rel,
'nofollow': 'nofollow' in rel,
'ugc': 'ugc' in rel,
'sponsored': 'sponsored' in rel,
}
if urlparse(abs_url).netloc == base_domain:
item['internal_links'].append(link_data)
else:
item['external_links'].append(link_data)
item['internal_link_count'] = len(item['internal_links'])
item['external_link_count'] = len(item['external_links'])
def _extract_schema(self, item, response):
schemas = []
for script in response.css('script[type="application/ld+json"]'):
try:
import json
data = json.loads(script.css('::text').get('{}'))
schemas.append(data)
except json.JSONDecodeError:
pass
item['json_ld_types'] = [
s.get('@type', '') for s in schemas if '@type' in s
]
item['has_schema'] = bool(schemas)
def _extract_hreflang(self, item, response):
hreflang_tags = []
for link in response.css('link[rel="alternate"][hreflang]'):
hreflang_tags.append({
'locale': link.attrib.get('hreflang', ''),
'href': link.attrib.get('href', ''),
})
item['hreflang_tags'] = hreflang_tags
item['hreflang_count'] = len(hreflang_tags)
def _extract_core_web_vitals_hints(self, item, response):
# Server-Timing header can expose LCP/FID if instrumented
server_timing = response.headers.get('Server-Timing', b'').decode()
item['server_timing'] = server_timing
# Check for render-blocking resources
item['render_blocking_scripts'] = len(
response.css('head script:not([async]):not([defer]):not([type="application/ld+json"])')
)
item['render_blocking_css'] = len(
response.css('head link[rel="stylesheet"]:not([media="print"])')
)
JavaScript Rendering with Playwright Integration
# seo_crawler/playwright_middleware.py
# Install: pip install scrapy-playwright
# Requires: DOWNLOAD_HANDLERS = {'https': 'scrapy_playwright.handler.ScrapyPlaywrightDownloadHandler'}
from playwright.async_api import async_playwright
import asyncio
class PlaywrightMiddleware:
"""Selective JS rendering middleware using scrapy-playwright."""
async def process_request(self, request, spider):
if not request.meta.get('playwright'):
return None # Let normal downloader handle
# scrapy-playwright handles the rendering
request.meta['playwright_page_methods'] = [
# Wait for network idle — critical for SPAs
{'method': 'wait_for_load_state', 'args': ['networkidle']},
]
request.meta['playwright_browser_type'] = 'chromium'
return None
# settings.py additions for Playwright
PLAYWRIGHT_SETTINGS = {
'DOWNLOAD_HANDLERS': {
'http': 'scrapy_playwright.handler.ScrapyPlaywrightDownloadHandler',
'https': 'scrapy_playwright.handler.ScrapyPlaywrightDownloadHandler',
},
'PLAYWRIGHT_BROWSER_TYPE': 'chromium',
'PLAYWRIGHT_LAUNCH_OPTIONS': {
'headless': True,
'args': [
'--no-sandbox',
'--disable-setuid-sandbox',
'--disable-dev-shm-usage',
'--disable-gpu',
'--no-first-run',
],
},
'PLAYWRIGHT_DEFAULT_NAVIGATION_TIMEOUT': 30000,
'PLAYWRIGHT_CONTEXTS': {
'default': {
'viewport': {'width': 1280, 'height': 720},
'user_agent': 'Mozilla/5.0 (compatible; Googlebot/2.1; +http://www.google.com/bot.html)',
'java_script_enabled': True,
},
},
'TWISTED_REACTOR': 'twisted.internet.asyncioreactor.AsyncioSelectorReactor',
'ASYNCIO_EVENT_LOOP': 'uvloop', # pip install uvloop
}
URL Deduplication at Scale
Scrapy's default RFPDupeFilter uses an in-memory set of URL fingerprints. For crawls exceeding 10M URLs, this exhausts memory. The solution is a Redis-backed Bloom filter: probabilistic deduplication with O(1) lookup and configurable false positive rate. A 0.1% false positive rate on 50M URLs requires ~90MB—acceptable.
# seo_crawler/dupefilter.py
# pip install redis pybloom-live scrapy-redis
import hashlib
import logging
from scrapy.dupefilters import BaseDupeFilter
from scrapy.utils.request import request_fingerprint
import redis
logger = logging.getLogger(__name__)
class RedisBitArrayDupeFilter(BaseDupeFilter):
"""Bloom filter-based deduplication using Redis SETBIT."""
# Bloom filter parameters for ~50M URLs, 0.01% FPR
BIT_ARRAY_SIZE = 2_000_000_000 # 2 billion bits = 250MB Redis memory
HASH_COUNT = 7 # Number of hash functions
def __init__(self, host='localhost', port=6379, db=0, key='crawl:seen'):
self.redis = redis.Redis(host=host, port=port, db=db)
self.key = key
self.dupecount = 0
self.logdupes = True
@classmethod
def from_settings(cls, settings):
return cls(
host=settings.get('REDIS_HOST', 'localhost'),
port=settings.getint('REDIS_PORT', 6379),
key=settings.get('DUPEFILTER_KEY', 'crawl:seen'),
)
def _get_bit_positions(self, fingerprint):
"""Generate k bit positions for a fingerprint using double hashing."""
positions = []
for i in range(self.HASH_COUNT):
combined = f"{fingerprint}:{i}".encode()
hash_val = int(hashlib.sha256(combined).hexdigest(), 16)
positions.append(hash_val % self.BIT_ARRAY_SIZE)
return positions
def request_seen(self, request):
fp = request_fingerprint(request)
positions = self._get_bit_positions(fp)
# Check all bits — if all set, probably seen
pipe = self.redis.pipeline()
for pos in positions:
pipe.getbit(self.key, pos)
results = pipe.execute()
if all(results):
self.dupecount += 1
return True
# Set all bits — mark as seen
pipe = self.redis.pipeline()
for pos in positions:
pipe.setbit(self.key, pos, 1)
pipe.execute()
return False
def close(self, reason=''):
logger.info(f"Bloom filter deduplicated {self.dupecount} URLs")
def log(self, request, spider):
if self.logdupes:
logger.debug(f"Filtered duplicate: {request.url}")
Streaming Crawl Data to BigQuery
# seo_crawler/pipelines/bigquery_pipeline.py
from google.cloud import bigquery
from google.api_core import retry
import json
import datetime
class BigQueryStreamingPipeline:
"""Streams crawl items to BigQuery in batches."""
BATCH_SIZE = 500
TABLE_ID = 'your-project.seo_crawl.pages'
SCHEMA = [
bigquery.SchemaField('url', 'STRING', mode='REQUIRED'),
bigquery.SchemaField('url_hash', 'STRING'),
bigquery.SchemaField('status_code', 'INTEGER'),
bigquery.SchemaField('title', 'STRING'),
bigquery.SchemaField('title_length', 'INTEGER'),
bigquery.SchemaField('meta_description', 'STRING'),
bigquery.SchemaField('h1_count', 'INTEGER'),
bigquery.SchemaField('canonical_url', 'STRING'),
bigquery.SchemaField('canonical_is_self', 'BOOLEAN'),
bigquery.SchemaField('is_noindex', 'BOOLEAN'),
bigquery.SchemaField('internal_link_count', 'INTEGER'),
bigquery.SchemaField('external_link_count', 'INTEGER'),
bigquery.SchemaField('has_schema', 'BOOLEAN'),
bigquery.SchemaField('hreflang_count', 'INTEGER'),
bigquery.SchemaField('response_time_ms', 'INTEGER'),
bigquery.SchemaField('crawl_depth', 'INTEGER'),
bigquery.SchemaField('crawl_timestamp', 'TIMESTAMP'),
bigquery.SchemaField('crawl_run_id', 'STRING'),
]
def __init__(self, run_id):
self.client = bigquery.Client()
self.run_id = run_id
self.batch = []
self._ensure_table()
@classmethod
def from_crawler(cls, crawler):
import uuid
return cls(run_id=str(uuid.uuid4()))
def _ensure_table(self):
table = bigquery.Table(self.TABLE_ID, schema=self.SCHEMA)
table.time_partitioning = bigquery.TimePartitioning(
type_=bigquery.TimePartitioningType.DAY,
field='crawl_timestamp'
)
table.clustering_fields = ['status_code', 'is_noindex']
try:
self.client.create_table(table)
except Exception:
pass # Table exists
def process_item(self, item, spider):
row = {
'url': item.get('url', ''),
'url_hash': item.get('url_hash', ''),
'status_code': item.get('status_code', 0),
'title': item.get('title', ''),
'title_length': item.get('title_length', 0),
'meta_description': item.get('meta_description', ''),
'h1_count': item.get('h1_count', 0),
'canonical_url': item.get('canonical_url', ''),
'canonical_is_self': item.get('canonical_is_self', False),
'is_noindex': item.get('is_noindex', False),
'internal_link_count': item.get('internal_link_count', 0),
'external_link_count': item.get('external_link_count', 0),
'has_schema': item.get('has_schema', False),
'hreflang_count': item.get('hreflang_count', 0),
'response_time_ms': item.get('response_time_ms', 0),
'crawl_depth': item.get('crawl_depth', 0),
'crawl_timestamp': datetime.datetime.utcnow().isoformat(),
'crawl_run_id': self.run_id,
}
self.batch.append(row)
if len(self.batch) >= self.BATCH_SIZE:
self._flush()
return item
def _flush(self):
if not self.batch:
return
errors = self.client.insert_rows_json(
self.TABLE_ID, self.batch,
retry=retry.Retry(deadline=30)
)
if errors:
spider.logger.error(f"BigQuery insert errors: {errors}")
self.batch = []
def close_spider(self, spider):
self._flush()
Differential Crawling: Detecting Changes
-- BigQuery: detect pages that changed between crawl runs
-- Compares current crawl against previous crawl snapshot
WITH current_crawl AS (
SELECT * FROM your-project.seo_crawl.pages
WHERE crawl_run_id = @current_run_id
),
previous_crawl AS (
SELECT * FROM your-project.seo_crawl.pages
WHERE crawl_run_id = @previous_run_id
)
SELECT
c.url,
c.status_code AS current_status,
p.status_code AS previous_status,
c.title AS current_title,
p.title AS previous_title,
c.canonical_url AS current_canonical,
p.canonical_url AS previous_canonical,
c.is_noindex AS current_noindex,
p.is_noindex AS previous_noindex,
c.internal_link_count AS current_links,
p.internal_link_count AS previous_links,
-- Change classification
CASE
WHEN p.url IS NULL THEN 'new_page'
WHEN c.status_code != p.status_code THEN 'status_changed'
WHEN c.title != p.title THEN 'title_changed'
WHEN c.canonical_url != p.canonical_url THEN 'canonical_changed'
WHEN c.is_noindex != p.is_noindex THEN 'indexability_changed'
ELSE 'no_change'
END AS change_type
FROM current_crawl c
LEFT JOIN previous_crawl p USING(url_hash)
WHERE
p.url IS NULL -- New pages
OR c.status_code != p.status_code
OR c.title != p.title
OR c.canonical_url != p.canonical_url
OR c.is_noindex != p.is_noindex
ORDER BY change_type, c.url
This query forms the basis of automated SEO change monitoring. Schedule it via Cloud Scheduler post-crawl and route results to a Slack webhook for immediate alerting on indexability changes. See the BigQuery GSC analysis article for complementary query patterns and the Looker Studio dashboard article for visualization approaches.
FAQ
How do I handle JavaScript-heavy sites where 90% of pages need rendering?
At that ratio, flip the architecture: use Playwright as the primary fetcher, not the fallback. Deploy a cluster of headless Chromium instances behind a load balancer (Playwright cluster via playwright-cluster npm package, or browserless.io). Feed URLs from Redis queue, extract rendered HTML, push to Scrapy for signal extraction. This is a rendering-first architecture; Scrapy becomes the extraction layer rather than the fetching layer.
What is the correct politeness setting for crawling my own site vs. a third party?
Your own site: set DOWNLOAD_DELAY to 0, CONCURRENT_REQUESTS_PER_DOMAIN to 32+, and use a dedicated crawl server that bypasses CDN caching to hit origin directly. Third-party sites: respect robots.txt Crawl-delay directive if present, otherwise 2–5 seconds with adaptive backoff on 429. Never bypass robots.txt; use ROBOTSTXT_OBEY = True always for third-party crawls.
How does Scrapy's built-in fingerprinting handle URL normalization?
Scrapy's request_fingerprint uses a canonical URL normalization (sorts query parameters, strips fragments) before hashing. This handles most dedup cases. Where it fails: session parameters in the path (not query string), encoded vs. decoded characters in paths, and protocol differences (http vs https). Override fingerprint in a custom RequestFingerprinter to handle site-specific normalization quirks.
How do I crawl behind authentication without cloaking concerns?
Authenticated crawling is legitimate for auditing staging/preview environments. Use a separate COOKIES_ENABLED = True profile with a service account session. Never send authenticated responses to a crawl that mimics Googlebot UA—this creates a cloaking risk if the authenticated content differs from the public view. Keep authentication-based crawls on a non-Googlebot UA and a separate run_id in BigQuery.
What is the fastest way to crawl 10M pages per day?
Distributed Scrapy with scrapy-redis for shared queue and seen-filter. Run 10–20 Spider processes across multiple machines, all reading from the same Redis URL frontier. A single optimized Scrapy process on a 4-core machine with no JS rendering can sustain ~100 requests/second = 8.6M pages/day. For 10M+, distribute across 2 machines. Memory: each Spider process needs ~500MB for item buffers plus Redis memory for the Bloom filter.
How do I extract Core Web Vitals signals during a crawl?
Server-side proxies for CWV: check Server-Timing headers (some origins report LCP/TTFB there), check Transfer-Encoding (chunked responses improve TTFB), measure download latency from Scrapy meta (download_latency). Real CWV requires field data from CrUX (available in BigQuery public dataset chrome-ux-report) joined to your crawl data on URL. Lab data via Playwright: use the CDP Performance API during JS rendering runs.
Key Takeaways
- Separate URL frontier management, JS rendering decision, signal extraction, and output sink into decoupled Scrapy components—spider, middleware, pipeline respectively.
- Redis-backed Bloom filter deduplication handles 50M+ URL crawls in ~90MB memory with 0.1% false positive rate at O(1) lookup.
- Selective JS rendering (render only when raw HTML has fewer than N links) balances rendering cost against completeness for mixed static/SPA sites.
- Adaptive politeness middleware that doubles delay on 429 and gradually reduces on 200 is more effective than static delay configuration.
- BigQuery streaming insert with time-partitioned, clustered tables enables sub-second crawl data availability and efficient historical diff queries.
- Differential crawl queries comparing run_id snapshots are the foundation of automated SEO change monitoring at scale.
Conclusion
Building a custom SEO crawler is a significant investment, justified when your audit requirements exceed what commercial tools offer: custom signal extraction, real-time BigQuery integration, differential change detection, or crawl volume that pushes commercial tool licensing costs into enterprise-tier pricing. The Scrapy ecosystem—with scrapy-redis, scrapy-playwright, and custom pipelines—provides a production-grade foundation. The architecture described here has been proven at 10M+ pages per crawl in production environments. The key discipline: keep extraction logic in pipelines, deduplication in Redis, and output in BigQuery, and you have a system that scales horizontally without redesign. Combine with edge SEO monitoring for a complete crawl intelligence platform.
