Skip to content
DATA & AUTOMATION / FIELD NOTE 052

Building a Custom SEO Crawler with Python and Scrapy

Reading map: Crawler Architecture and Design Decisions; Scrapy Project Structure for SEO; Custom Middleware: Politeness, UA Rotation, Retry; SEO Signal Extraction Pipelines
A reading map of this field note. Download SVG ↓

Screaming Frog, Sitebulb, and DeepCrawl are excellent tools for standard audits. They are not excellent for custom signal extraction, real-time crawl integration with BigQuery pipelines, differential crawling against a prior crawl snapshot, or extraction of JavaScript-rendered content at scale with fine-grained concurrency control. When the off-the-shelf crawlers reach their limits—and they always do on enterprise-scale sites—you build your own.

This article covers building a production-grade SEO crawler using Python and Scrapy: custom middleware for politeness, duplicate URL detection at scale, signal extraction pipelines, JavaScript rendering integration via Playwright, and streaming crawl data to BigQuery. The assumptions: you are comfortable with Python async, you have deployed Scrapy in production before, and you know why a naive BFS crawler will get you IP-banned in 20 minutes.

Crawler Architecture and Design Decisions

Before writing a line of Scrapy code, document your crawler's threat model and design constraints. The architecture decisions that matter most:

DecisionOption AOption BChoose When
Concurrency modelScrapy async (Twisted)httpx + asyncioScrapy for plugin ecosystem; httpx for simpler pipelines
JS renderingPlaywright (full browser)Splash (lightweight)Playwright for React/Vue SPAs; Splash for basic JS execution
URL frontierIn-memory dequeRedis queueRedis for distributed crawls >500k URLs
Seen URL storePython setRedis Bloom filterBloom filter for >10M URLs (memory vs. false positive tradeoff)
Output sinkCSV/JSON filesBigQuery streaming insertBigQuery for real-time analysis and historical comparison
Rendering decisionRender all pagesSelective renderingSelective: render only when initial HTML has <10 links

The most important architectural insight: separate your URL frontier management, rendering decision, signal extraction, and output pipeline into decoupled components. Scrapy's spider → middleware → item pipeline architecture maps cleanly to this, but you must resist the temptation to put extraction logic in the spider itself.

Scrapy Project Structure for SEO


# seo_crawler/spiders/seo_spider.py
import scrapy
from scrapy.linkextractors import LinkExtractor
from scrapy.http import HtmlResponse
from urllib.parse import urlparse
import hashlib
import re

class SeoSpider(scrapy.Spider):
    name = 'seo'
    custom_settings = {
        'CONCURRENT_REQUESTS': 16,
        'CONCURRENT_REQUESTS_PER_DOMAIN': 4,
        'DOWNLOAD_DELAY': 0.5,
        'ROBOTSTXT_OBEY': True,
        'COOKIES_ENABLED': False,
        'REDIRECT_MAX_TIMES': 5,
        'DEPTH_LIMIT': 10,
        'DOWNLOAD_TIMEOUT': 30,
        'REACTOR_THREADPOOL_MAXSIZE': 20,
    }

    def __init__(self, start_url=None, *args, **kwargs):
        super().__init__(*args, **kwargs)
        if not start_url:
            raise ValueError("start_url required")
        parsed = urlparse(start_url)
        self.allowed_domains = [parsed.netloc]
        self.start_urls = [start_url]
        self.link_extractor = LinkExtractor(
            allow_domains=self.allowed_domains,
            deny_extensions=['pdf', 'jpg', 'jpeg', 'png', 'gif', 'svg',
                             'css', 'js', 'ico', 'woff', 'woff2', 'ttf',
                             'mp4', 'mp3', 'zip', 'gz'],
            unique=True,
        )

    def parse(self, response):
        # Only process HTML responses
        if 'text/html' not in response.headers.get('Content-Type', b'').decode():
            return

        item = self.extract_seo_signals(response)
        yield item

        # Extract and follow links
        for link in self.link_extractor.extract_links(response):
            yield scrapy.Request(
                link.url,
                callback=self.parse,
                errback=self.handle_error,
                meta={
                    'referrer_url': response.url,
                    'referrer_depth': response.meta.get('depth', 0)
                }
            )

    def extract_seo_signals(self, response):
        from seo_crawler.items import SeoPageItem
        item = SeoPageItem()
        item['url'] = response.url
        item['url_hash'] = hashlib.md5(response.url.encode()).hexdigest()
        item['status_code'] = response.status
        item['content_type'] = response.headers.get('Content-Type', b'').decode()
        item['referrer_url'] = response.meta.get('referrer_url', '')
        item['crawl_depth'] = response.meta.get('depth', 0)
        item['response_time_ms'] = int(response.meta.get('download_latency', 0) * 1000)
        return item

    def handle_error(self, failure):
        from seo_crawler.items import SeoPageItem
        item = SeoPageItem()
        item['url'] = failure.request.url
        item['status_code'] = 0
        item['error'] = str(failure.value)
        yield item

Custom Middleware: Politeness, UA Rotation, Retry

Scrapy middleware is the correct place for cross-cutting concerns: User-Agent management, response caching, politeness enforcement, and retry logic with exponential backoff. Never put these in the spider.


# seo_crawler/middlewares.py
import time
import random
from scrapy import signals
from scrapy.http import HtmlResponse
from scrapy.exceptions import NotConfigured
import logging

logger = logging.getLogger(__name__)

class SeoUserAgentMiddleware:
    """Rotates UA to mimic real Googlebot behavior, with verified UA strings."""

    GOOGLEBOT_UA = (
        'Mozilla/5.0 (compatible; Googlebot/2.1; '
        '+http://www.google.com/bot.html)'
    )
    DESKTOP_UA = (
        'Mozilla/5.0 (Windows NT 10.0; Win64; x64) '
        'AppleWebKit/537.36 (KHTML, like Gecko) '
        'Chrome/122.0.0.0 Safari/537.36'
    )

    def __init__(self, crawler_ua='googlebot'):
        self.ua = self.GOOGLEBOT_UA if crawler_ua == 'googlebot' else self.DESKTOP_UA

    @classmethod
    def from_crawler(cls, crawler):
        return cls(crawler_ua=crawler.settings.get('CRAWLER_UA', 'desktop'))

    def process_request(self, request, spider):
        request.headers['User-Agent'] = self.ua
        # Add headers Googlebot typically sends
        request.headers['Accept'] = 'text/html,application/xhtml+xml'
        request.headers['Accept-Language'] = 'en-US,en;q=0.9'


class AdaptivePolitenessMiddleware:
    """Adjusts delay based on server response codes and latency."""

    def __init__(self, base_delay=0.5, max_delay=10.0):
        self.base_delay = base_delay
        self.max_delay = max_delay
        self.domain_delays = {}
        self.domain_error_counts = {}

    @classmethod
    def from_crawler(cls, crawler):
        return cls(
            base_delay=crawler.settings.getfloat('DOWNLOAD_DELAY', 0.5),
            max_delay=crawler.settings.getfloat('MAX_DOWNLOAD_DELAY', 10.0),
        )

    def process_response(self, request, response, spider):
        domain = request.url.split('/')[2]

        if response.status == 429:
            # Rate limited — back off exponentially
            current_delay = self.domain_delays.get(domain, self.base_delay)
            new_delay = min(current_delay * 2, self.max_delay)
            self.domain_delays[domain] = new_delay
            logger.warning(f"Rate limited on {domain}, delay now {new_delay}s")
            time.sleep(new_delay)
        elif response.status == 200:
            # Success — gradually reduce delay toward base
            current_delay = self.domain_delays.get(domain, self.base_delay)
            self.domain_delays[domain] = max(
                self.base_delay,
                current_delay * 0.9
            )

        return response


class SelectiveRenderingMiddleware:
    """Routes requests to Playwright when HTML has insufficient link count."""

    MIN_LINKS_THRESHOLD = 5  # Below this, suspect JS rendering needed

    def process_response(self, request, response, spider):
        if request.meta.get('playwright_rendered'):
            return response  # Already rendered, skip

        if not isinstance(response, HtmlResponse):
            return response

        # Count links in raw HTML — if too few, re-fetch with Playwright
        link_count = len(response.css('a[href]'))
        if link_count < self.MIN_LINKS_THRESHOLD:
            logger.info(f"Only {link_count} links found on {response.url}, queuing for JS render")
            new_request = request.copy()
            new_request.meta['playwright'] = True
            new_request.meta['playwright_rendered'] = True
            new_request.dont_filter = True
            spider.crawler.engine.crawl(new_request, spider)

        return response

SEO Signal Extraction Pipelines

The extraction pipeline processes each SeoPageItem and enriches it with parsed SEO signals. Keeping this in a pipeline (not the spider) enables parallel processing and easy unit testing.


# seo_crawler/pipelines/seo_signals.py
import re
from urllib.parse import urljoin, urlparse
from w3lib.html import remove_tags

class SeoSignalExtractionPipeline:
    """Extracts all SEO signals from raw response into structured fields."""

    def process_item(self, item, spider):
        response = item.get('_response')
        if not response:
            return item

        self._extract_title(item, response)
        self._extract_meta(item, response)
        self._extract_headings(item, response)
        self._extract_canonical(item, response)
        self._extract_robots(item, response)
        self._extract_links(item, response)
        self._extract_schema(item, response)
        self._extract_hreflang(item, response)
        self._extract_core_web_vitals_hints(item, response)

        # Remove raw response from item before serialization
        item.pop('_response', None)
        return item

    def _extract_title(self, item, response):
        title = response.css('title::text').get('').strip()
        item['title'] = title
        item['title_length'] = len(title)
        item['title_has_brand'] = bool(re.search(r'\|', title))

    def _extract_meta(self, item, response):
        item['meta_description'] = response.css(
            'meta[name="description"]::attr(content)'
        ).get('').strip()
        item['meta_robots'] = response.css(
            'meta[name="robots"]::attr(content)'
        ).get('').lower()
        item['meta_viewport'] = bool(response.css('meta[name="viewport"]').get())

    def _extract_headings(self, item, response):
        item['h1_texts'] = response.css('h1::text').getall()
        item['h1_count'] = len(item['h1_texts'])
        item['h2_count'] = len(response.css('h2'))
        item['h3_count'] = len(response.css('h3'))

    def _extract_canonical(self, item, response):
        canonical = response.css('link[rel="canonical"]::attr(href)').get('')
        item['canonical_url'] = canonical.strip()
        item['canonical_is_self'] = (
            item['canonical_url'].rstrip('/') == response.url.rstrip('/')
        )
        item['canonical_is_cross_domain'] = bool(
            canonical and urlparse(canonical).netloc != urlparse(response.url).netloc
        )

    def _extract_robots(self, item, response):
        # Check both meta robots and X-Robots-Tag header
        meta_robots = item.get('meta_robots', '')
        x_robots = response.headers.get('X-Robots-Tag', b'').decode().lower()
        combined = f"{meta_robots},{x_robots}"
        item['is_noindex'] = 'noindex' in combined
        item['is_nofollow'] = 'nofollow' in combined
        item['x_robots_tag'] = x_robots

    def _extract_links(self, item, response):
        links = response.css('a[href]')
        item['internal_links'] = []
        item['external_links'] = []
        base_domain = urlparse(response.url).netloc

        for link in links:
            href = link.attrib.get('href', '').strip()
            if not href or href.startswith('#') or href.startswith('mailto:'):
                continue
            abs_url = urljoin(response.url, href)
            rel = link.attrib.get('rel', '')
            anchor_text = link.css('::text').get('').strip()[:200]
            link_data = {
                'url': abs_url,
                'anchor': anchor_text,
                'rel': rel,
                'nofollow': 'nofollow' in rel,
                'ugc': 'ugc' in rel,
                'sponsored': 'sponsored' in rel,
            }
            if urlparse(abs_url).netloc == base_domain:
                item['internal_links'].append(link_data)
            else:
                item['external_links'].append(link_data)

        item['internal_link_count'] = len(item['internal_links'])
        item['external_link_count'] = len(item['external_links'])

    def _extract_schema(self, item, response):
        schemas = []
        for script in response.css('script[type="application/ld+json"]'):
            try:
                import json
                data = json.loads(script.css('::text').get('{}'))
                schemas.append(data)
            except json.JSONDecodeError:
                pass
        item['json_ld_types'] = [
            s.get('@type', '') for s in schemas if '@type' in s
        ]
        item['has_schema'] = bool(schemas)

    def _extract_hreflang(self, item, response):
        hreflang_tags = []
        for link in response.css('link[rel="alternate"][hreflang]'):
            hreflang_tags.append({
                'locale': link.attrib.get('hreflang', ''),
                'href': link.attrib.get('href', ''),
            })
        item['hreflang_tags'] = hreflang_tags
        item['hreflang_count'] = len(hreflang_tags)

    def _extract_core_web_vitals_hints(self, item, response):
        # Server-Timing header can expose LCP/FID if instrumented
        server_timing = response.headers.get('Server-Timing', b'').decode()
        item['server_timing'] = server_timing
        # Check for render-blocking resources
        item['render_blocking_scripts'] = len(
            response.css('head script:not([async]):not([defer]):not([type="application/ld+json"])')
        )
        item['render_blocking_css'] = len(
            response.css('head link[rel="stylesheet"]:not([media="print"])')
        )

JavaScript Rendering with Playwright Integration


# seo_crawler/playwright_middleware.py
# Install: pip install scrapy-playwright
# Requires: DOWNLOAD_HANDLERS = {'https': 'scrapy_playwright.handler.ScrapyPlaywrightDownloadHandler'}

from playwright.async_api import async_playwright
import asyncio

class PlaywrightMiddleware:
    """Selective JS rendering middleware using scrapy-playwright."""

    async def process_request(self, request, spider):
        if not request.meta.get('playwright'):
            return None  # Let normal downloader handle

        # scrapy-playwright handles the rendering
        request.meta['playwright_page_methods'] = [
            # Wait for network idle — critical for SPAs
            {'method': 'wait_for_load_state', 'args': ['networkidle']},
        ]
        request.meta['playwright_browser_type'] = 'chromium'
        return None


# settings.py additions for Playwright
PLAYWRIGHT_SETTINGS = {
    'DOWNLOAD_HANDLERS': {
        'http': 'scrapy_playwright.handler.ScrapyPlaywrightDownloadHandler',
        'https': 'scrapy_playwright.handler.ScrapyPlaywrightDownloadHandler',
    },
    'PLAYWRIGHT_BROWSER_TYPE': 'chromium',
    'PLAYWRIGHT_LAUNCH_OPTIONS': {
        'headless': True,
        'args': [
            '--no-sandbox',
            '--disable-setuid-sandbox',
            '--disable-dev-shm-usage',
            '--disable-gpu',
            '--no-first-run',
        ],
    },
    'PLAYWRIGHT_DEFAULT_NAVIGATION_TIMEOUT': 30000,
    'PLAYWRIGHT_CONTEXTS': {
        'default': {
            'viewport': {'width': 1280, 'height': 720},
            'user_agent': 'Mozilla/5.0 (compatible; Googlebot/2.1; +http://www.google.com/bot.html)',
            'java_script_enabled': True,
        },
    },
    'TWISTED_REACTOR': 'twisted.internet.asyncioreactor.AsyncioSelectorReactor',
    'ASYNCIO_EVENT_LOOP': 'uvloop',  # pip install uvloop
}

URL Deduplication at Scale

Scrapy's default RFPDupeFilter uses an in-memory set of URL fingerprints. For crawls exceeding 10M URLs, this exhausts memory. The solution is a Redis-backed Bloom filter: probabilistic deduplication with O(1) lookup and configurable false positive rate. A 0.1% false positive rate on 50M URLs requires ~90MB—acceptable.


# seo_crawler/dupefilter.py
# pip install redis pybloom-live scrapy-redis

import hashlib
import logging
from scrapy.dupefilters import BaseDupeFilter
from scrapy.utils.request import request_fingerprint
import redis

logger = logging.getLogger(__name__)

class RedisBitArrayDupeFilter(BaseDupeFilter):
    """Bloom filter-based deduplication using Redis SETBIT."""

    # Bloom filter parameters for ~50M URLs, 0.01% FPR
    BIT_ARRAY_SIZE = 2_000_000_000  # 2 billion bits = 250MB Redis memory
    HASH_COUNT = 7  # Number of hash functions

    def __init__(self, host='localhost', port=6379, db=0, key='crawl:seen'):
        self.redis = redis.Redis(host=host, port=port, db=db)
        self.key = key
        self.dupecount = 0
        self.logdupes = True

    @classmethod
    def from_settings(cls, settings):
        return cls(
            host=settings.get('REDIS_HOST', 'localhost'),
            port=settings.getint('REDIS_PORT', 6379),
            key=settings.get('DUPEFILTER_KEY', 'crawl:seen'),
        )

    def _get_bit_positions(self, fingerprint):
        """Generate k bit positions for a fingerprint using double hashing."""
        positions = []
        for i in range(self.HASH_COUNT):
            combined = f"{fingerprint}:{i}".encode()
            hash_val = int(hashlib.sha256(combined).hexdigest(), 16)
            positions.append(hash_val % self.BIT_ARRAY_SIZE)
        return positions

    def request_seen(self, request):
        fp = request_fingerprint(request)
        positions = self._get_bit_positions(fp)

        # Check all bits — if all set, probably seen
        pipe = self.redis.pipeline()
        for pos in positions:
            pipe.getbit(self.key, pos)
        results = pipe.execute()

        if all(results):
            self.dupecount += 1
            return True

        # Set all bits — mark as seen
        pipe = self.redis.pipeline()
        for pos in positions:
            pipe.setbit(self.key, pos, 1)
        pipe.execute()

        return False

    def close(self, reason=''):
        logger.info(f"Bloom filter deduplicated {self.dupecount} URLs")

    def log(self, request, spider):
        if self.logdupes:
            logger.debug(f"Filtered duplicate: {request.url}")

Streaming Crawl Data to BigQuery


# seo_crawler/pipelines/bigquery_pipeline.py
from google.cloud import bigquery
from google.api_core import retry
import json
import datetime

class BigQueryStreamingPipeline:
    """Streams crawl items to BigQuery in batches."""

    BATCH_SIZE = 500
    TABLE_ID = 'your-project.seo_crawl.pages'

    SCHEMA = [
        bigquery.SchemaField('url', 'STRING', mode='REQUIRED'),
        bigquery.SchemaField('url_hash', 'STRING'),
        bigquery.SchemaField('status_code', 'INTEGER'),
        bigquery.SchemaField('title', 'STRING'),
        bigquery.SchemaField('title_length', 'INTEGER'),
        bigquery.SchemaField('meta_description', 'STRING'),
        bigquery.SchemaField('h1_count', 'INTEGER'),
        bigquery.SchemaField('canonical_url', 'STRING'),
        bigquery.SchemaField('canonical_is_self', 'BOOLEAN'),
        bigquery.SchemaField('is_noindex', 'BOOLEAN'),
        bigquery.SchemaField('internal_link_count', 'INTEGER'),
        bigquery.SchemaField('external_link_count', 'INTEGER'),
        bigquery.SchemaField('has_schema', 'BOOLEAN'),
        bigquery.SchemaField('hreflang_count', 'INTEGER'),
        bigquery.SchemaField('response_time_ms', 'INTEGER'),
        bigquery.SchemaField('crawl_depth', 'INTEGER'),
        bigquery.SchemaField('crawl_timestamp', 'TIMESTAMP'),
        bigquery.SchemaField('crawl_run_id', 'STRING'),
    ]

    def __init__(self, run_id):
        self.client = bigquery.Client()
        self.run_id = run_id
        self.batch = []
        self._ensure_table()

    @classmethod
    def from_crawler(cls, crawler):
        import uuid
        return cls(run_id=str(uuid.uuid4()))

    def _ensure_table(self):
        table = bigquery.Table(self.TABLE_ID, schema=self.SCHEMA)
        table.time_partitioning = bigquery.TimePartitioning(
            type_=bigquery.TimePartitioningType.DAY,
            field='crawl_timestamp'
        )
        table.clustering_fields = ['status_code', 'is_noindex']
        try:
            self.client.create_table(table)
        except Exception:
            pass  # Table exists

    def process_item(self, item, spider):
        row = {
            'url': item.get('url', ''),
            'url_hash': item.get('url_hash', ''),
            'status_code': item.get('status_code', 0),
            'title': item.get('title', ''),
            'title_length': item.get('title_length', 0),
            'meta_description': item.get('meta_description', ''),
            'h1_count': item.get('h1_count', 0),
            'canonical_url': item.get('canonical_url', ''),
            'canonical_is_self': item.get('canonical_is_self', False),
            'is_noindex': item.get('is_noindex', False),
            'internal_link_count': item.get('internal_link_count', 0),
            'external_link_count': item.get('external_link_count', 0),
            'has_schema': item.get('has_schema', False),
            'hreflang_count': item.get('hreflang_count', 0),
            'response_time_ms': item.get('response_time_ms', 0),
            'crawl_depth': item.get('crawl_depth', 0),
            'crawl_timestamp': datetime.datetime.utcnow().isoformat(),
            'crawl_run_id': self.run_id,
        }
        self.batch.append(row)

        if len(self.batch) >= self.BATCH_SIZE:
            self._flush()

        return item

    def _flush(self):
        if not self.batch:
            return
        errors = self.client.insert_rows_json(
            self.TABLE_ID, self.batch,
            retry=retry.Retry(deadline=30)
        )
        if errors:
            spider.logger.error(f"BigQuery insert errors: {errors}")
        self.batch = []

    def close_spider(self, spider):
        self._flush()

Differential Crawling: Detecting Changes


-- BigQuery: detect pages that changed between crawl runs
-- Compares current crawl against previous crawl snapshot

WITH current_crawl AS (
  SELECT * FROM your-project.seo_crawl.pages
  WHERE crawl_run_id = @current_run_id
),
previous_crawl AS (
  SELECT * FROM your-project.seo_crawl.pages
  WHERE crawl_run_id = @previous_run_id
)

SELECT
  c.url,
  c.status_code AS current_status,
  p.status_code AS previous_status,
  c.title AS current_title,
  p.title AS previous_title,
  c.canonical_url AS current_canonical,
  p.canonical_url AS previous_canonical,
  c.is_noindex AS current_noindex,
  p.is_noindex AS previous_noindex,
  c.internal_link_count AS current_links,
  p.internal_link_count AS previous_links,
  -- Change classification
  CASE
    WHEN p.url IS NULL THEN 'new_page'
    WHEN c.status_code != p.status_code THEN 'status_changed'
    WHEN c.title != p.title THEN 'title_changed'
    WHEN c.canonical_url != p.canonical_url THEN 'canonical_changed'
    WHEN c.is_noindex != p.is_noindex THEN 'indexability_changed'
    ELSE 'no_change'
  END AS change_type
FROM current_crawl c
LEFT JOIN previous_crawl p USING(url_hash)
WHERE
  p.url IS NULL  -- New pages
  OR c.status_code != p.status_code
  OR c.title != p.title
  OR c.canonical_url != p.canonical_url
  OR c.is_noindex != p.is_noindex
ORDER BY change_type, c.url

This query forms the basis of automated SEO change monitoring. Schedule it via Cloud Scheduler post-crawl and route results to a Slack webhook for immediate alerting on indexability changes. See the BigQuery GSC analysis article for complementary query patterns and the Looker Studio dashboard article for visualization approaches.

FAQ

How do I handle JavaScript-heavy sites where 90% of pages need rendering?

At that ratio, flip the architecture: use Playwright as the primary fetcher, not the fallback. Deploy a cluster of headless Chromium instances behind a load balancer (Playwright cluster via playwright-cluster npm package, or browserless.io). Feed URLs from Redis queue, extract rendered HTML, push to Scrapy for signal extraction. This is a rendering-first architecture; Scrapy becomes the extraction layer rather than the fetching layer.

What is the correct politeness setting for crawling my own site vs. a third party?

Your own site: set DOWNLOAD_DELAY to 0, CONCURRENT_REQUESTS_PER_DOMAIN to 32+, and use a dedicated crawl server that bypasses CDN caching to hit origin directly. Third-party sites: respect robots.txt Crawl-delay directive if present, otherwise 2–5 seconds with adaptive backoff on 429. Never bypass robots.txt; use ROBOTSTXT_OBEY = True always for third-party crawls.

How does Scrapy's built-in fingerprinting handle URL normalization?

Scrapy's request_fingerprint uses a canonical URL normalization (sorts query parameters, strips fragments) before hashing. This handles most dedup cases. Where it fails: session parameters in the path (not query string), encoded vs. decoded characters in paths, and protocol differences (http vs https). Override fingerprint in a custom RequestFingerprinter to handle site-specific normalization quirks.

How do I crawl behind authentication without cloaking concerns?

Authenticated crawling is legitimate for auditing staging/preview environments. Use a separate COOKIES_ENABLED = True profile with a service account session. Never send authenticated responses to a crawl that mimics Googlebot UA—this creates a cloaking risk if the authenticated content differs from the public view. Keep authentication-based crawls on a non-Googlebot UA and a separate run_id in BigQuery.

What is the fastest way to crawl 10M pages per day?

Distributed Scrapy with scrapy-redis for shared queue and seen-filter. Run 10–20 Spider processes across multiple machines, all reading from the same Redis URL frontier. A single optimized Scrapy process on a 4-core machine with no JS rendering can sustain ~100 requests/second = 8.6M pages/day. For 10M+, distribute across 2 machines. Memory: each Spider process needs ~500MB for item buffers plus Redis memory for the Bloom filter.

How do I extract Core Web Vitals signals during a crawl?

Server-side proxies for CWV: check Server-Timing headers (some origins report LCP/TTFB there), check Transfer-Encoding (chunked responses improve TTFB), measure download latency from Scrapy meta (download_latency). Real CWV requires field data from CrUX (available in BigQuery public dataset chrome-ux-report) joined to your crawl data on URL. Lab data via Playwright: use the CDP Performance API during JS rendering runs.

Key Takeaways

  • Separate URL frontier management, JS rendering decision, signal extraction, and output sink into decoupled Scrapy components—spider, middleware, pipeline respectively.
  • Redis-backed Bloom filter deduplication handles 50M+ URL crawls in ~90MB memory with 0.1% false positive rate at O(1) lookup.
  • Selective JS rendering (render only when raw HTML has fewer than N links) balances rendering cost against completeness for mixed static/SPA sites.
  • Adaptive politeness middleware that doubles delay on 429 and gradually reduces on 200 is more effective than static delay configuration.
  • BigQuery streaming insert with time-partitioned, clustered tables enables sub-second crawl data availability and efficient historical diff queries.
  • Differential crawl queries comparing run_id snapshots are the foundation of automated SEO change monitoring at scale.

Conclusion

Building a custom SEO crawler is a significant investment, justified when your audit requirements exceed what commercial tools offer: custom signal extraction, real-time BigQuery integration, differential change detection, or crawl volume that pushes commercial tool licensing costs into enterprise-tier pricing. The Scrapy ecosystem—with scrapy-redis, scrapy-playwright, and custom pipelines—provides a production-grade foundation. The architecture described here has been proven at 10M+ pages per crawl in production environments. The key discipline: keep extraction logic in pipelines, deduplication in Redis, and output in BigQuery, and you have a system that scales horizontally without redesign. Combine with edge SEO monitoring for a complete crawl intelligence platform.

YOUR READING CHECKLIST

Make the ideas stick.

Mark the sections you’ve worked through. Saved in this browser.

0 of 4 reviewed
Andrii Stanetskyi
ABOUT THE AUTHOR

Andrii Stanetskyi

Head of SEO / Technical SEO Lead based in Tallinn, Estonia. Technical architecture, enterprise eCommerce, Python automation, and AI-assisted workflows.

More about Andrii ↗
LET’S FIND THE REAL BOTTLENECK

A clearer picture.
A practical next step.

Get a focused SEO audit or a consultation on your next technical decision. We’ll agree on the scope and fee before any work begins.

01 / Diagnose02 / Prioritize03 / Plan
How can I help?

Scope and fee agreed before any work begins.

Choose your language

Explore SEO services in 26 languages. Journal articles retain their original language.

ENEnglish↗DEDeutsch↗FRFrançais↗ESEspañol↗ITItaliano↗PTPortuguês↗NLNederlands↗PLPolski↗SVSvenska↗DADansk↗FISuomi↗NONorsk↗ETEesti↗LVLatviešu↗LTLietuvių↗CSČeština↗RORomână↗HUMagyar↗ELΕλληνικά↗BGБългарски↗HRHrvatski↗SKSlovenčina↗SLSlovenščina↗RUРусский↗UKУкраїнська↗TRTürkçe↗
LET’S WORK ON YOUR WEBSITE
A CLEAR NEXT STEP

Let’s talk
about your site.

A focused SEO audit or a conversation about a specific challenge. Tell me where you are and what you want to change.

Andrii Stanetskyi
Andrii StanetskyiHead of SEO / Technical SEO Lead
[email protected] ↗
How can I help?

Scope and fee agreed before any work begins.