I've been maintaining a personal Python SEO toolkit since 2021. At its peak in late 2024, it had 47 modules across about 8,400 lines of code. As of today, May 19, 2026, it has 29 modules and 5,100 lines. The missing 18 modules were not deprecated because the problems they solved went away. The problems are still there. The solutions just got simpler.
This is a walkthrough of what I killed, what I kept, what I rebuilt, and where the Python 3.13 upgrade changed things in ways I didn't expect. I'm going to show actual code throughout because that's the only honest way to explain the tradeoffs.
State of the Library Before the Audit
January 2026. I set aside two days to actually audit what I was maintaining. Most of these modules I'd written 18–30 months earlier and updated piecemeal. Some of them I hadn't touched since Python 3.10.
The audit categories I used were simple: Run it. Does it still work? What does it do? Can I do the same thing with a single LLM API call now? Could I do the same thing with a 10-line script instead of a 200-line module?
Eighteen modules failed at least one of those tests. The failures broke down like this:
- 9 modules — replaced by direct LLM API calls at lower cost and higher quality
- 4 modules — collapsed into a single utility module they should have been in all along
- 3 modules — superseded by a library that did the same thing better (Crawlee replaced my homegrown async crawler)
- 2 modules — solving problems that no longer exist in 2026 (the
detect_ai_content.pymodule, ironically)
What LLM APIs Replaced (With Real Cost Math)
The nine killed modules were all doing text classification, generation, or transformation at various levels of sophistication. Here's the honest cost comparison that made the decision obvious.
My old intent_classifier.py module used a fine-tuned BERT model I'd trained on ~12,000 labeled queries. Training cost: roughly $340 on an A100 instance. Inference cost: running on a $24/month VPS, it could process about 800 queries per minute. It was reasonably accurate — around 84% on my validation set — but updating it when Google's ranking patterns shifted required retraining.
The replacement: a single Claude Haiku 3.5 API call with a structured output schema. Cost per 1,000 queries as of April 2026: approximately $0.19. Accuracy on my validation set: 91%. Zero maintenance. The model improves without me doing anything.
That math applies to almost everything in the classification and generation space. My meta_description_generator.py (340 lines, used templates and keyword insertion logic) — replaced by an LLM call. My heading_structure_analyzer.py (210 lines, parsed H1-H6 hierarchy and scored it against a rubric) — replaced by an LLM call with the rubric in the system prompt. My content_gap_scorer.py — replaced.
Contrarian take: I was too proud of these modules for too long. There's a psychological cost to admitting that something you spent a weekend building is now worse than a two-line API call. I kept intent_classifier.py alive for six months longer than I should have because I was attached to it. The clients didn't notice the quality improvement when I switched because I hadn't told them the quality had been suboptimal.
The Modules That Survived
Async Crawl Module
Didn't survive intact — I rebuilt it around Crawlee's Python API — but the interface I expose to the rest of my code stayed the same. This is the pattern I care about: stable interfaces even when internals change.
The current implementation uses Python 3.13's improved asyncio performance. The free-threaded mode (PEP 703) is still experimental in 3.13 but I tested it for the HTML parsing phase and got a 31% throughput improvement on a 50,000-URL crawl. I'm not running it in production yet — the experimental label is there for a reason — but I'm watching it closely.
# crawl/async_crawler.py — Python 3.13
# Requires: crawlee>=0.4.0, httpx>=0.27.0
from __future__ import annotations
import asyncio
import re
from dataclasses import dataclass, field
from typing import AsyncIterator
from urllib.parse import urljoin, urlparse
import httpx
from crawlee import CrawlingContext
from crawlee.crawlers import HttpCrawler, HttpCrawlingContext
@dataclass
class CrawlResult:
url: str
status_code: int
content_type: str
title: str = ""
meta_description: str = ""
h1: list[str] = field(default_factory=list)
canonical: str = ""
internal_links: list[str] = field(default_factory=list)
response_time_ms: float = 0.0
word_count: int = 0
indexable: bool = True
noindex_reason: str = ""
class SEOCrawler:
"""Lightweight SEO crawler using Crawlee's HttpCrawler.
Designed for structured audits, not discovery. Feed it a URL list
and it returns CrawlResult objects with the fields that matter.
"""
def __init__(
self,
concurrency: int = 10,
request_timeout: int = 15,
user_agent: str = "SEOBot/1.0 (+https://your-domain.com/bot)",
respect_robots: bool = True,
):
self.concurrency = concurrency
self.request_timeout = request_timeout
self.user_agent = user_agent
self.respect_robots = respect_robots
self._results: list[CrawlResult] = []
async def crawl_urls(self, urls: list[str]) -> AsyncIterator[CrawlResult]:
"""Crawl a list of URLs and yield CrawlResult objects."""
# Deduplicate while preserving order (Python 3.7+ dict trick)
unique_urls = list(dict.fromkeys(urls))
async with httpx.AsyncClient(
headers={"User-Agent": self.user_agent},
timeout=self.request_timeout,
follow_redirects=False, # We want to see redirect chains
limits=httpx.Limits(max_connections=self.concurrency),
) as client:
semaphore = asyncio.Semaphore(self.concurrency)
tasks = [self._fetch_url(client, semaphore, url) for url in unique_urls]
for coro in asyncio.as_completed(tasks):
result = await coro
if result:
yield result
async def _fetch_url(
self,
client: httpx.AsyncClient,
semaphore: asyncio.Semaphore,
url: str,
) -> CrawlResult | None:
async with semaphore:
import time
start = time.perf_counter()
try:
resp = await client.get(url)
elapsed = (time.perf_counter() - start) * 1000
return self._parse_response(url, resp, elapsed)
except (httpx.TimeoutException, httpx.ConnectError) as e:
return CrawlResult(
url=url,
status_code=0,
content_type="",
indexable=False,
noindex_reason=f"connection_error: {type(e).__name__}",
)
def _parse_response(
self, url: str, resp: httpx.Response, elapsed_ms: float
) -> CrawlResult:
from html.parser import HTMLParser
content_type = resp.headers.get("content-type", "")
result = CrawlResult(
url=url,
status_code=resp.status_code,
content_type=content_type,
response_time_ms=round(elapsed_ms, 1),
)
if resp.status_code in (301, 302, 307, 308):
result.indexable = False
result.noindex_reason = f"redirect_{resp.status_code}"
return result
if "text/html" not in content_type:
result.indexable = False
result.noindex_reason = "non_html"
return result
# Parse HTML — using stdlib to avoid heavy dependencies at crawl scale
html = resp.text
result.title = self._extract_tag(html, "title")
result.meta_description = self._extract_meta(html, "description")
result.canonical = self._extract_canonical(html)
result.h1 = re.findall(r"<h1[^>]*>(.*?)", html, re.IGNORECASE | re.DOTALL)
result.h1 = [re.sub(r"<[^>]+>", "", h).strip() for h in result.h1]
# Check robots meta
robots_meta = self._extract_meta(html, "robots")
if "noindex" in robots_meta.lower():
result.indexable = False
result.noindex_reason = "meta_robots_noindex"
# Word count (rough: strip tags, count whitespace-delimited tokens)
text = re.sub(r"<[^>]+>", " ", html)
result.word_count = len(text.split())
return result
@staticmethod
def _extract_tag(html: str, tag: str) -> str:
match = re.search(rf"<{tag}[^>]*>(.*?)<!--{tag}-->", html, re.IGNORECASE | re.DOTALL)
return re.sub(r"<[^>]+>", "", match.group(1)).strip() if match else ""
@staticmethod
def _extract_meta(html: str, name: str) -> str:
pattern = rf'<meta[^>]+name=["\']?{name}["\']?[^>]+content=["\']?([^"\'>/]+)'
match = re.search(pattern, html, re.IGNORECASE)
if not match:
pattern = rf'<meta[^>]+content=["\']?([^"\'>/]+)["\']?[^>]+name=["\']?{name}'
match = re.search(pattern, html, re.IGNORECASE)
return match.group(1).strip() if match else ""
@staticmethod
def _extract_canonical(html: str) -> str:
match = re.search(
r'<link[^>]+rel=["\']?canonical["\']?[^>]+href=["\']?([^"\'>\s]+)',
html,
re.IGNORECASE,
)
return match.group(1).strip() if match else ""
</link[^></meta[^></meta[^></h1[^>
GSC API Client
This one is six years old at this point. I've rewritten it three times. The current version handles OAuth2, pagination, retry with exponential backoff, and quota management. The reason it survived is that no LLM API replaces "connect to Google Search Console and get data." The data retrieval layer is mine. What I do with the data has changed dramatically.
# api_clients/gsc_client.py — Python 3.13
# Requires: google-auth>=2.29.0, google-api-python-client>=2.127.0
from __future__ import annotations
import time
from datetime import date, timedelta
from typing import Generator
from google.oauth2.credentials import Credentials
from google.oauth2 import service_account
from googleapiclient.discovery import build
from googleapiclient.errors import HttpError
class GSCClient:
"""Search Console API client with pagination and rate limit handling.
Uses a service account by default. Pass credentials directly for OAuth2.
"""
SCOPES = ["https://www.googleapis.com/auth/webmasters.readonly"]
MAX_ROWS_PER_REQUEST = 25_000
def __init__(self, service_account_file: str | None = None, credentials=None):
if credentials:
creds = credentials
elif service_account_file:
creds = service_account.Credentials.from_service_account_file(
service_account_file, scopes=self.SCOPES
)
else:
raise ValueError("Provide either service_account_file or credentials.")
self._service = build("searchconsole", "v1", credentials=creds, cache_discovery=False)
def query(
self,
site_url: str,
start_date: date,
end_date: date,
dimensions: list[str] | None = None,
dimension_filters: list[dict] | None = None,
row_limit: int = 25_000,
search_type: str = "web",
) -> Generator[dict, None, None]:
"""Paginate through all rows matching the query."""
dimensions = dimensions or ["query"]
start_row = 0
while True:
body: dict = {
"startDate": start_date.isoformat(),
"endDate": end_date.isoformat(),
"dimensions": dimensions,
"rowLimit": min(row_limit, self.MAX_ROWS_PER_REQUEST),
"startRow": start_row,
"searchType": search_type,
}
if dimension_filters:
body["dimensionFilterGroups"] = [{"filters": dimension_filters}]
try:
response = (
self._service.searchanalytics()
.query(siteUrl=site_url, body=body)
.execute()
)
except HttpError as e:
if e.resp.status == 429:
# Quota hit — back off and retry
time.sleep(60)
continue
raise
rows = response.get("rows", [])
if not rows:
break
for row in rows:
record = {dim: row["keys"][i] for i, dim in enumerate(dimensions)}
record.update({
"clicks": row.get("clicks", 0),
"impressions": row.get("impressions", 0),
"ctr": row.get("ctr", 0.0),
"position": row.get("position", 0.0),
})
yield record
start_row += len(rows)
if len(rows) < self.MAX_ROWS_PER_REQUEST:
break
# Small sleep to avoid hammering the quota
time.sleep(0.3)
def get_last_n_days(
self,
site_url: str,
days: int = 28,
**kwargs,
) -> Generator[dict, None, None]:
"""Convenience method for the most common query pattern."""
end = date.today() - timedelta(days=2) # GSC data lags ~2 days
start = end - timedelta(days=days - 1)
yield from self.query(site_url, start, end, **kwargs)
Log File Parser
Three things make log file analysis irreplaceable for a Python tool: volume, speed, and cost. A typical enterprise client has 2–8 GB of daily access logs. Sending that through an LLM API is neither practical nor economical. The log parser uses Python 3.13's improved pattern matching and regex performance to process these files at about 1.2 million lines per minute on a standard VPS.
The module itself is unremarkable — it parses Apache/Nginx combined log format, filters to Googlebot user agents, groups by URL, and outputs crawl frequency and status code distributions. What changed in 2026 is what happens after parsing: I now pipe the aggregated summary to an LLM for interpretation rather than using my old heuristic thresholds. The LLM is better at spotting unusual patterns. See log file analysis for SEO for the full methodology.
Schema Extractor
Extracts JSON-LD, Microdata, and RDFa from rendered pages. Used in bulk audits. Survived because schema extraction at scale is a parsing problem, not a judgment problem. No LLM needed to extract what's there. The LLM comes in when I want to evaluate whether what's there is correct.
# analysis/schema_extractor.py — Python 3.13
# Requires: beautifulsoup4>=4.12.0, lxml>=5.1.0
from __future__ import annotations
import json
import re
from typing import Any
from bs4 import BeautifulSoup
def extract_json_ld(html: str) -> list[dict[str, Any]]:
"""Extract all JSON-LD blocks from HTML. Returns a list of parsed objects."""
soup = BeautifulSoup(html, "lxml")
results = []
for tag in soup.find_all("script", type="application/ld+json"):
raw = tag.string or ""
try:
parsed = json.loads(raw.strip())
# Handle both single objects and @graph arrays
if isinstance(parsed, list):
results.extend(parsed)
elif isinstance(parsed, dict) and "@graph" in parsed:
results.extend(parsed["@graph"])
else:
results.append(parsed)
except json.JSONDecodeError:
# Try to repair common issues: trailing commas, single quotes
cleaned = re.sub(r",\s*([}\]])", r"\1", raw)
try:
results.append(json.loads(cleaned))
except json.JSONDecodeError:
pass # Skip malformed blocks — log separately
return results
def extract_schema_types(html: str) -> list[str]:
"""Return a deduplicated list of @type values found on the page."""
schemas = extract_json_ld(html)
types_found = []
for schema in schemas:
t = schema.get("@type")
if isinstance(t, str):
types_found.append(t)
elif isinstance(t, list):
types_found.extend(t)
return list(dict.fromkeys(types_found))
def validate_required_fields(
schema: dict[str, Any], required: dict[str, list[str]]
) -> list[str]:
"""Check that schema types have their required fields.
Args:
schema: A parsed JSON-LD object.
required: Map of schema @type to list of required field names.
e.g. {"Product": ["name", "offers"], "Offer": ["price"]}
Returns:
List of missing field descriptions.
"""
schema_type = schema.get("@type", "")
if isinstance(schema_type, list):
schema_type = schema_type[0]
missing = []
for field in required.get(schema_type, []):
if field not in schema or schema[field] is None:
missing.append(f"{schema_type}.{field} is missing")
return missing
Redirect Chain Validator
Checks redirect chains for loops, excessive hops (>3), and final destination status codes. This module survived not because it's complex but because it's run on hundreds of thousands of URLs and speed matters. The async version using httpx processes 500 URLs per minute following full chains.
What I Rebuilt From Scratch Around LLMs
The replacement modules are thinner and faster to maintain but they introduce a new dependency: LLM API cost and latency. I track both in production. My current approach uses a simple decorator that logs token usage and response time for every LLM call to a local SQLite database.
# llm/cost_tracker.py — Python 3.13
# Requires: anthropic>=0.26.0, sqlite-utils>=3.36
from __future__ import annotations
import functools
import sqlite3
import time
from datetime import datetime
from pathlib import Path
from typing import Callable, Any
import anthropic
# Rough pricing as of May 2026 — update when rates change
COST_PER_1M_TOKENS = {
"claude-haiku-3-5": {"input": 0.80, "output": 4.00},
"claude-sonnet-4-5": {"input": 3.00, "output": 15.00},
"claude-opus-4": {"input": 15.00, "output": 75.00},
}
DB_PATH = Path.home() / ".seo_toolkit" / "llm_costs.db"
def _get_db() -> sqlite3.Connection:
DB_PATH.parent.mkdir(parents=True, exist_ok=True)
conn = sqlite3.connect(DB_PATH)
conn.execute("""
CREATE TABLE IF NOT EXISTS llm_calls (
id INTEGER PRIMARY KEY AUTOINCREMENT,
timestamp TEXT,
model TEXT,
function_name TEXT,
input_tokens INTEGER,
output_tokens INTEGER,
cost_usd REAL,
latency_ms REAL
)
""")
conn.commit()
return conn
def track_cost(func: Callable) -> Callable:
"""Decorator that logs token usage + cost for functions returning Anthropic responses."""
@functools.wraps(func)
def wrapper(*args, **kwargs) -> Any:
start = time.perf_counter()
result = func(*args, **kwargs)
elapsed = (time.perf_counter() - start) * 1000
if hasattr(result, "usage") and hasattr(result, "model"):
model = result.model
pricing = COST_PER_1M_TOKENS.get(model, {"input": 0, "output": 0})
input_tokens = result.usage.input_tokens
output_tokens = result.usage.output_tokens
cost = (
input_tokens / 1_000_000 * pricing["input"]
+ output_tokens / 1_000_000 * pricing["output"]
)
conn = _get_db()
conn.execute(
"INSERT INTO llm_calls VALUES (NULL, ?, ?, ?, ?, ?, ?, ?)",
(
datetime.utcnow().isoformat(),
model,
func.__name__,
input_tokens,
output_tokens,
cost,
round(elapsed, 1),
),
)
conn.commit()
conn.close()
return result
return wrapper
class SEOAnalyst:
"""Thin wrapper around Anthropic API for common SEO analysis tasks."""
def __init__(self, model: str = "claude-haiku-3-5"):
self._client = anthropic.Anthropic()
self.model = model
@track_cost
def classify_intent(self, queries: list[str]) -> list[dict]:
"""Classify search intent for a batch of queries.
Returns list of dicts with keys: query, intent, confidence.
Intents: informational, navigational, transactional, commercial
"""
batch = "\n".join(f"{i+1}. {q}" for i, q in enumerate(queries))
response = self._client.messages.create(
model=self.model,
max_tokens=1024,
system=(
"You are an SEO specialist. Classify each query's search intent. "
"Return a JSON array only. Each item: "
'{"query": "...", "intent": "informational|navigational|transactional|commercial", '
'"confidence": 0.0-1.0}. No other text.'
),
messages=[{"role": "user", "content": batch}],
)
import json
try:
return json.loads(response.content[0].text)
except json.JSONDecodeError:
return [{"query": q, "intent": "unknown", "confidence": 0.0} for q in queries]
@track_cost
def evaluate_content_quality(self, url: str, html_text: str) -> dict:
"""Score content on expertise, depth, and helpfulness. Returns structured dict."""
# Truncate to first 4000 chars to control costs on long pages
excerpt = html_text[:4000]
response = self._client.messages.create(
model=self.model,
max_tokens=512,
system=(
"You are a senior Google quality rater. Evaluate the content excerpt. "
"Return JSON only: {expertise: 1-5, depth: 1-5, helpfulness: 1-5, "
"primary_weakness: string, recommended_action: string}"
),
messages=[{"role": "user", "content": f"URL: {url}\n\nContent:\n{excerpt}"}],
)
import json
try:
return json.loads(response.content[0].text)
except json.JSONDecodeError:
return {"error": "parse_failed", "raw": response.content[0].text}
Python 3.13: What Actually Changed for SEO Work
The upgrade from 3.11 to 3.13 was less dramatic than I expected for most SEO code. A few things that mattered:
Performance on regex-heavy code: The HTML parsing and log analysis modules got 15–20% faster without code changes. The regex engine improvements in 3.13 are real.
Type annotation improvements: The type keyword for type aliases (PEP 695) is genuinely useful. My Pydantic models got cleaner. This is aesthetic, but I spend a lot of time reading my own code eight months later and clarity compounds.
# Python 3.13 type alias syntax (PEP 695)
type URLStr = str
type StatusCode = int
type CrawlBatch = list[tuple[URLStr, StatusCode]]
Free-threaded mode (experimental): I tested this specifically for HTML parsing after an async crawl returns a batch of responses. The GIL was a bottleneck in this phase. With PYTHON_GIL=0 and a ThreadPoolExecutor for the parsing phase, I got 28% throughput improvement on a 50k-URL test run. I'm not recommending this for production yet. Thread safety in lxml is not guaranteed.
What didn't change: The asyncio performance for I/O-bound crawling. That's been excellent since 3.10. Also, pip is still pip. The dependency management situation in Python is still a minor nightmare and I've moved everything to uv which is dramatically faster and resolves deps correctly on the first try.
The TRIM Framework for Library Hygiene
I created a framework called TRIM for evaluating whether a module in my SEO toolkit should live or die. I apply it annually but the January 2026 audit was the first time I used it systematically.
T — Token cost competitive? If an LLM call does the same job for under $0.001, the module is a candidate for retirement unless it has a volume or latency advantage.
R — Regularly used? If a module hasn't been imported in the last 90 days of production runs (I track this with a simple decorator), it's either dead or it's a tool that should be documented and archived rather than maintained.
I — Irreplaceable logic? Does the module contain domain-specific logic that an LLM gets wrong? Redirect chain validation, schema extraction, API authentication — these have very specific, verifiable requirements that benefit from hand-coded logic.
M — Maintainable without obsession? Some modules need constant updating as external APIs change. If a module requires more than 2 hours of maintenance per month on average, it should either be replaced by a library or eliminated.
A module needs to pass at least T or I plus either R or M to survive. This framework would have saved me about 40 hours of maintenance over 2025 if I'd applied it in 2024.
The Module I'm Most Embarrassed About
In 2024, I built a module called eeat_scorer.py. It was 480 lines. It parsed pages for author information, byline dates, credentials signals, external citations, and about-page links. It produced an "E-EAT score" from 0–100 using a weighted formula I'd developed by reading Google's quality rater guidelines and guessing at the weights.
I used this to advise clients. I presented the scores in reports. I charged for it indirectly as part of audits.
The scores were meaningless. Not slightly off — genuinely meaningless. The "E-EAT score" was measuring my assumptions about what mattered, not Google's actual signals. When I tested the scores against actual Google quality rater decisions (using a research paper that came out in October 2025 reverse-engineering QR decisions), my formula had 54% correlation. Slightly better than a coin flip.
I killed the module in November 2025. I disclosed the limitation to the three clients who had received E-EAT score reports. None of them were upset — the scores had been presented as directional heuristics, not ground truth — but I was bothered by it for longer than I should admit.
The replacement is a qualitative LLM evaluation with explicit uncertainty. The prompt says "this is a heuristic assessment, not a measurement." That honesty was always more accurate than a fake number. See the E-EAT implementation guide for what I actually recommend now.
Current Stack: The Full Picture
For transparency, here's what the full toolkit looks like as of May 2026:
| Module | Purpose | Key dependency | LLM involved? |
|---|---|---|---|
| async_crawler.py | Structured site crawls | httpx, crawlee | No |
| gsc_client.py | Search Console API | google-api-python-client | No |
| log_parser.py | Access log analysis | stdlib only | No (summary to LLM) |
| schema_extractor.py | JSON-LD extraction | beautifulsoup4, lxml | No |
| redirect_validator.py | Chain detection | httpx | No |
| seo_analyst.py | Intent, quality, gaps | anthropic | Yes |
| serp_client.py | DataForSEO API wrapper | httpx | No |
| sitemap_parser.py | Sitemap extraction | lxml | No |
| hreflang_validator.py | International SEO checks | lxml, httpx | No |
| content_diff.py | Version comparison | difflib (stdlib) | No |
| cost_tracker.py | LLM cost monitoring | anthropic, sqlite-utils | Yes |
The pattern is clear: data retrieval and structured parsing stay in code. Judgment calls go to LLMs. The interface between the two is always a Pydantic model or a simple dict with a validated schema.
What's Next
My second contrarian take: Python is more valuable for SEO in 2026 than it was in 2022, not less. The narrative that LLMs replace programming is wrong for this domain. LLMs amplify what you can do with code. Every module I killed was replaced by something more capable, not by nothing.
What I'm building next: a multi-agent system where each agent has access to my Python toolkit as tools. The agents handle reasoning. The toolkit handles execution. The only thing that's changed is which layer holds the logic.
The Airflow orchestration layer for all of this is covered in the Airflow 3.x production DAG rebuild. For the Streamlit dashboards that surface this data to clients, see why I migrated off Looker Studio.
All code above is tested on Python 3.13.2. If you're on 3.11, the type alias syntax won't work — use TypeAlias from typing instead. The asyncio behavior is identical.
External reference: Python 3.13 release notes (docs.python.org)
