Keyword clustering is the step where keyword research becomes content strategy. Without clustering, you have a list of keywords. With it, you have a content architecture that maps clearly to URLs, prevents cannibalization, and creates measurable topical authority. The problem is that manual clustering at scale is brutally slow, and most automated tools use surface-level string matching that misses semantic relationships entirely.
This tutorial covers two proven clustering approaches in Python — SERP-similarity clustering (the most accurate method) and embedding-based semantic clustering (the most scalable method) — with complete, runnable code for both.
Why Keyword Clustering Matters for SEO
The core problem clustering solves: you have 500 keyword variants, each slightly different, and you need to answer one question per variant — does this keyword deserve its own page, or should it be handled as part of a broader page? Getting this wrong in either direction has measurable costs:
- Under-splitting (one page tries to rank for too many divergent keywords): Google can't determine what the page is primarily about, ranking signals dilute across intents, and the page ends up ranking for nothing well.
- Over-splitting (every keyword gets its own page): Keyword cannibalization, crawl budget waste, thin-content risk, and internal link equity fragmentation.
The correct split point is where the SERP composition changes. If two keywords return substantially the same top-10 results, they belong on the same page. If they return different results, they represent different intents and deserve different content. This SERP-similarity principle is the foundation of the clustering approaches below.
Clustering Approaches Compared
| Method | Accuracy | Scalability | Cost | Best For |
|---|---|---|---|---|
| Manual clustering | Highest | Very low (< 200 kws) | High (time) | Small, high-stakes keyword sets |
| String similarity (TF-IDF / n-gram) | Low | Very high | Near zero | Initial deduplication pass only |
| SERP-similarity clustering | Very high | Medium (API rate limits) | Medium (API calls) | Gold-standard clustering up to 5,000 kws |
| Embedding-based clustering | High | Very high | Low (OpenAI API or local model) | Large keyword sets, 5,000–100,000+ kws |
| Tool-based (Keyword Insights, etc.) | High | High | Per-keyword subscription cost | Teams without Python capability |
SERP-Similarity Clustering with Python
SERP-similarity clustering works by measuring how many URLs overlap in the top-10 results for two keywords. If keywords A and B share 4 or more top-10 URLs, they trigger substantially the same SERP and should be targeted on the same page. The threshold (typically 3–5 overlapping URLs) is adjustable based on how aggressively you want to split.
Prerequisites
pip install requests pandas scikit-learn python-dotenv networkx
Step 1: Fetch SERP Data
We'll use the ValueSERP API (inexpensive, reliable). Swap for SerpAPI, DataForSEO, or any SERP API you prefer — the data structure will differ slightly.
import requests
import json
import time
import pandas as pd
from collections import defaultdict
VALUESERP_API_KEY = "YOUR_API_KEY"
def fetch_serp(keyword, num_results=10, country="us", language="en"):
"""Fetch top-N organic results for a keyword from ValueSERP."""
params = {
"api_key": VALUESERP_API_KEY,
"q": keyword,
"num": num_results,
"gl": country,
"hl": language,
"output": "json"
}
response = requests.get("https://api.valueserp.com/search", params=params)
if response.status_code == 200:
data = response.json()
results = data.get("organic_results", [])
urls = [r.get("link", "") for r in results[:num_results]]
return urls
return []
def fetch_serps_for_keywords(keywords, delay=1.0):
"""Fetch SERPs for a list of keywords with rate limiting."""
serp_data = {}
for i, kw in enumerate(keywords):
print(f"Fetching SERP {i+1}/{len(keywords)}: {kw}")
serp_data[kw] = fetch_serp(kw)
time.sleep(delay) # Respect API rate limits
return serp_data
Step 2: Calculate SERP Overlap Matrix
def calculate_serp_overlap(serp_data):
"""
Build a similarity matrix based on URL overlap in top-10 SERPs.
Returns a dict of dicts: overlap[kw1][kw2] = number of shared URLs.
"""
keywords = list(serp_data.keys())
overlap = defaultdict(dict)
for i, kw1 in enumerate(keywords):
for j, kw2 in enumerate(keywords):
if i == j:
overlap[kw1][kw2] = 10 # Self-similarity = max
else:
set1 = set(serp_data[kw1])
set2 = set(serp_data[kw2])
shared = len(set1.intersection(set2))
overlap[kw1][kw2] = shared
return overlap, keywords
Step 3: Build Clusters Using Graph-Based Grouping
import networkx as nx
def build_clusters(overlap, keywords, threshold=3):
"""
Build keyword clusters using a graph where edges exist between
keywords sharing >= threshold SERP URLs.
Each connected component becomes a cluster.
"""
G = nx.Graph()
G.add_nodes_from(keywords)
for kw1 in keywords:
for kw2 in keywords:
if kw1 != kw2 and overlap[kw1][kw2] >= threshold:
G.add_edge(kw1, kw2)
clusters = []
for component in nx.connected_components(G):
clusters.append(list(component))
return clusters
def export_clusters_to_csv(clusters, output_path="keyword_clusters.csv"):
"""Export cluster assignments to a CSV file."""
rows = []
for i, cluster in enumerate(clusters):
for kw in cluster:
rows.append({
"cluster_id": i + 1,
"keyword": kw,
"cluster_size": len(cluster)
})
df = pd.DataFrame(rows)
df.sort_values(["cluster_id", "keyword"], inplace=True)
df.to_csv(output_path, index=False)
print(f"Exported {len(rows)} keywords in {len(clusters)} clusters to {output_path}")
return df
Step 4: Run the Full Pipeline
if __name__ == "__main__":
# Your keyword list
keywords = [
"project management software",
"project management tool",
"best project management app",
"task management software",
"team project management",
"project tracking software",
"construction project management software",
"project management for construction",
"construction scheduling software",
]
# Fetch SERPs
serp_data = fetch_serps_for_keywords(keywords, delay=1.5)
# Build overlap matrix
overlap, kw_list = calculate_serp_overlap(serp_data)
# Cluster with threshold=3 (keywords sharing 3+ top-10 URLs go together)
clusters = build_clusters(overlap, kw_list, threshold=3)
# Export results
df = export_clusters_to_csv(clusters)
print(df.groupby("cluster_id").size().reset_index(name="count"))
Expected output: the first 6 keywords likely cluster together (they share the same product-category SERP), while the last 3 cluster separately (construction-specific SERP diverges).
Embedding-Based Semantic Clustering
For keyword sets too large for SERP-API calls (5,000+), embedding-based clustering provides high-quality results without per-query API costs. The approach: convert each keyword to a vector embedding, then use k-means or HDBSCAN to group semantically similar keywords.
Using OpenAI Embeddings
from openai import OpenAI
import numpy as np
from sklearn.cluster import KMeans
from sklearn.preprocessing import normalize
import pandas as pd
client = OpenAI(api_key="YOUR_OPENAI_API_KEY")
def get_embeddings(texts, model="text-embedding-3-small"):
"""Get embeddings for a list of texts using OpenAI API."""
# Batch into groups of 100 to stay within API limits
all_embeddings = []
batch_size = 100
for i in range(0, len(texts), batch_size):
batch = texts[i:i+batch_size]
response = client.embeddings.create(input=batch, model=model)
batch_embeddings = [item.embedding for item in response.data]
all_embeddings.extend(batch_embeddings)
print(f"Processed {min(i+batch_size, len(texts))}/{len(texts)} embeddings")
return np.array(all_embeddings)
def cluster_by_embeddings(keywords, n_clusters=None, auto_n=True):
"""
Cluster keywords using embeddings and K-means.
If auto_n=True, estimate optimal cluster count using elbow method.
"""
embeddings = get_embeddings(keywords)
normalized = normalize(embeddings) # Cosine similarity requires normalisation
if auto_n:
# Elbow method: find n where inertia reduction plateaus
inertias = []
n_range = range(2, min(50, len(keywords) // 5))
for n in n_range:
km = KMeans(n_clusters=n, random_state=42, n_init=10)
km.fit(normalized)
inertias.append(km.inertia_)
# Simple elbow detection: biggest second derivative
diffs = np.diff(inertias)
second_diffs = np.diff(diffs)
optimal_n = n_range[np.argmax(second_diffs) + 2]
n_clusters = optimal_n
print(f"Auto-selected {n_clusters} clusters")
km = KMeans(n_clusters=n_clusters, random_state=42, n_init=10)
labels = km.fit_predict(normalized)
return labels, embeddings
def export_embedding_clusters(keywords, labels, output_path="embedding_clusters.csv"):
df = pd.DataFrame({
"keyword": keywords,
"cluster_id": labels
})
df["cluster_size"] = df.groupby("cluster_id")["cluster_id"].transform("count")
df.sort_values(["cluster_id", "keyword"]).to_csv(output_path, index=False)
print(f"Exported {len(df)} keywords in {df['cluster_id'].nunique()} clusters")
return df
Using a Local Model (No API Cost)
# Alternative: Use sentence-transformers locally (free, no API)
# pip install sentence-transformers
from sentence_transformers import SentenceTransformer
def get_local_embeddings(texts, model_name="all-MiniLM-L6-v2"):
"""Get embeddings using a local sentence transformer model."""
model = SentenceTransformer(model_name)
embeddings = model.encode(texts, show_progress_bar=True)
return embeddings
# Usage: replace get_embeddings() with get_local_embeddings()
# all-MiniLM-L6-v2 is fast and performs well for short texts like keywords
The Hybrid Approach
In practice, the best workflow combines both methods: use embedding clustering for the full keyword set to produce coarse clusters, then run SERP-similarity validation on each cluster's candidate keywords to confirm or split the coarse clusters. This gives you the scalability of embeddings with the accuracy of SERP similarity.
def hybrid_cluster(keywords, serp_data=None):
"""
Step 1: Coarse embedding clustering
Step 2: SERP-similarity validation within each coarse cluster
Returns refined clusters.
"""
# Step 1: Embedding clustering
labels, _ = cluster_by_embeddings(keywords, auto_n=True)
coarse_clusters = {}
for kw, label in zip(keywords, labels):
coarse_clusters.setdefault(label, []).append(kw)
# Step 2: SERP validation within each coarse cluster
if serp_data is None:
print("No SERP data provided — returning coarse clusters only")
return list(coarse_clusters.values())
refined_clusters = []
for cluster_id, cluster_kws in coarse_clusters.items():
if len(cluster_kws) == 1:
refined_clusters.append(cluster_kws)
continue
cluster_serp = {kw: serp_data[kw] for kw in cluster_kws if kw in serp_data}
overlap, _ = calculate_serp_overlap(cluster_serp)
sub_clusters = build_clusters(overlap, list(cluster_serp.keys()), threshold=3)
refined_clusters.extend(sub_clusters)
return refined_clusters
Processing and Using Cluster Output
Raw cluster output is a list of keyword groups. To make it actionable, you need to assign each cluster a primary keyword (the head term), identify its content format, and map it to either an existing URL (for content updates) or a new URL (for content creation).
def label_clusters(clusters, volume_data=None):
"""
Label each cluster with a primary keyword.
Primary = highest volume keyword in cluster (if volume_data provided)
Otherwise = longest keyword (assumes more specific = more representative)
"""
labelled = []
for i, cluster in enumerate(clusters):
if volume_data:
primary = max(cluster, key=lambda kw: volume_data.get(kw, 0))
else:
primary = max(cluster, key=len)
labelled.append({
"cluster_id": i + 1,
"primary_keyword": primary,
"all_keywords": cluster,
"keyword_count": len(cluster)
})
return labelled
Case Study: 4,000-Keyword Cluster in Practice
A B2B software company had accumulated 4,000 keywords from a Semrush export and a competitor gap analysis. Manual clustering was estimated at 40+ hours. We ran the hybrid clustering approach: embedding pass first (OpenAI text-embedding-3-small, cost: ~$0.02 for 4,000 keywords), auto-selected 87 coarse clusters. We then ran SERP validation on the 20 largest coarse clusters (those with 30+ keywords) — 600 SERP API calls at $0.001/call = $0.60.
The process took 4 hours of compute time and produced 112 validated clusters from 4,000 keywords. Manual review of cluster labels took 3 hours. Total cost: <$2 in API fees + 3 hours of human review. The output fed directly into a 6-month content calendar with 112 content briefs, each with a validated primary keyword and supporting keyword list.
The cluster output is then run through our cannibalization detection workflow to identify cases where existing pages already target cluster primary keywords, preventing duplicate content creation. See also the topic clusters and pillar pages guide for how cluster output maps to site architecture.
FAQ
What SERP overlap threshold should I use for clustering?
For most use cases, 3 overlapping top-10 URLs is the right threshold. At 2, you'll over-cluster (too few groups, too many keywords per page). At 4+, you'll under-cluster (too many groups, near-duplicate pages). For highly competitive verticals where the top-10 is stable and consistent, you can raise to 4. For dynamic or niche SERPs, stay at 3 or even 2.
How many keywords can I cluster before SERP API costs become prohibitive?
SERP-similarity clustering requires O(n) API calls — one per keyword. At $0.001–$0.003 per call (ValueSERP, DataForSEO pricing), 1,000 keywords costs $1–$3. At 10,000 keywords: $10–$30. For sets above 5,000 keywords, the hybrid embedding approach reduces SERP API calls to only the validation step, typically 10–20% of the total keyword set.
What embedding model should I use?
For keyword clustering specifically: OpenAI's text-embedding-3-small performs well and is cost-effective. For a completely free, offline solution: sentence-transformers' all-MiniLM-L6-v2 is fast, runs locally, and produces good results for short text inputs like keywords. Avoid using large general-purpose LLMs for embedding generation — they're slow and expensive for this task.
How do I handle clusters that are too large (30+ keywords)?
Large clusters typically mean the threshold is too low, or the topic is genuinely broad enough to warrant a hub-and-spoke structure. Review the largest clusters manually: if they contain clear sub-topic groupings, split at the sub-topic level. If they're all truly targeting the same intent, consider a long-form comprehensive page that covers all the variant keywords as sections.
Can I cluster in a language other than English?
Yes. For SERP-similarity clustering, language doesn't matter — you're comparing URLs regardless of keyword language. For embedding-based clustering, use a multilingual model: OpenAI's text-embedding-3-small is multilingual, and sentence-transformers has excellent multilingual models (paraphrase-multilingual-MiniLM-L12-v2). Always fetch SERPs for the target country and language combination in your API parameters.
How do I validate that my clusters are correct?
Manual spot-check: for each cluster, open the primary keyword in an incognito browser and verify that the SERP matches what you'd expect given the cluster's other keywords. For statistical validation, calculate the average within-cluster SERP overlap and compare to the between-cluster overlap. Good clusters show within-cluster overlap of 4+ and between-cluster overlap of 1 or less.
Key Takeaways
- SERP-similarity clustering is the most accurate method: keywords sharing 3+ top-10 URLs belong on the same page. This is the ground truth for clustering decisions.
- Embedding-based clustering (OpenAI or sentence-transformers) handles large keyword sets cost-effectively and can be run entirely locally for free with all-MiniLM-L6-v2.
- The hybrid approach — embedding clustering for coarse grouping, SERP validation for refinement — gives you the best of both methods at low cost.
- Each cluster needs a primary keyword (highest volume or most representative) that will drive the page's title, H1, and URL slug.
- Total cost for 4,000 keywords using the hybrid approach: under $2 in API fees. Manual review remains the largest time investment.
- Cluster output feeds directly into content brief creation, cannibalization audits, and site architecture planning — it's the bridge between keyword research and content strategy.
- Always validate cluster output against your existing site URLs to identify which clusters already have coverage and which need new content.
Conclusion
Keyword clustering with Python isn't difficult — the code above is a complete implementation you can run today. The harder part is making good decisions about thresholds, cluster sizes, and what to do with the output. Start with the SERP-similarity approach for your most strategically important keyword sets, and use the embedding approach for large exports that need rapid initial organisation.
The payoff for getting clustering right is significant: a properly clustered keyword set produces a content architecture that avoids cannibalization, builds topical authority efficiently, and gives you a clear content calendar rather than an undifferentiated keyword dump.
For the next step after clustering, see our guide on detecting and fixing keyword cannibalization. For how clusters map to pillar pages, see the topic clusters and pillar pages guide.
Reference: Sentence Transformers documentation covers the full range of embedding models available for local deployment.
