Skip to content
DATA & AUTOMATION / FIELD NOTE 014

Keyword Clustering with Python: A Complete Tutorial with Code

Reading map: Why Keyword Clustering Matters for SEO; Clustering Approaches Compared; SERP-Similarity Clustering with Python; Embedding-Based Semantic Clustering
A reading map of this field note. Download SVG ↓

Keyword clustering is the step where keyword research becomes content strategy. Without clustering, you have a list of keywords. With it, you have a content architecture that maps clearly to URLs, prevents cannibalization, and creates measurable topical authority. The problem is that manual clustering at scale is brutally slow, and most automated tools use surface-level string matching that misses semantic relationships entirely.

This tutorial covers two proven clustering approaches in Python — SERP-similarity clustering (the most accurate method) and embedding-based semantic clustering (the most scalable method) — with complete, runnable code for both.

Why Keyword Clustering Matters for SEO

The core problem clustering solves: you have 500 keyword variants, each slightly different, and you need to answer one question per variant — does this keyword deserve its own page, or should it be handled as part of a broader page? Getting this wrong in either direction has measurable costs:

  • Under-splitting (one page tries to rank for too many divergent keywords): Google can't determine what the page is primarily about, ranking signals dilute across intents, and the page ends up ranking for nothing well.
  • Over-splitting (every keyword gets its own page): Keyword cannibalization, crawl budget waste, thin-content risk, and internal link equity fragmentation.

The correct split point is where the SERP composition changes. If two keywords return substantially the same top-10 results, they belong on the same page. If they return different results, they represent different intents and deserve different content. This SERP-similarity principle is the foundation of the clustering approaches below.

Clustering Approaches Compared

Method Accuracy Scalability Cost Best For
Manual clustering Highest Very low (< 200 kws) High (time) Small, high-stakes keyword sets
String similarity (TF-IDF / n-gram) Low Very high Near zero Initial deduplication pass only
SERP-similarity clustering Very high Medium (API rate limits) Medium (API calls) Gold-standard clustering up to 5,000 kws
Embedding-based clustering High Very high Low (OpenAI API or local model) Large keyword sets, 5,000–100,000+ kws
Tool-based (Keyword Insights, etc.) High High Per-keyword subscription cost Teams without Python capability

SERP-Similarity Clustering with Python

SERP-similarity clustering works by measuring how many URLs overlap in the top-10 results for two keywords. If keywords A and B share 4 or more top-10 URLs, they trigger substantially the same SERP and should be targeted on the same page. The threshold (typically 3–5 overlapping URLs) is adjustable based on how aggressively you want to split.

Prerequisites

pip install requests pandas scikit-learn python-dotenv networkx

Step 1: Fetch SERP Data

We'll use the ValueSERP API (inexpensive, reliable). Swap for SerpAPI, DataForSEO, or any SERP API you prefer — the data structure will differ slightly.

import requests
import json
import time
import pandas as pd
from collections import defaultdict

VALUESERP_API_KEY = "YOUR_API_KEY"

def fetch_serp(keyword, num_results=10, country="us", language="en"):
    """Fetch top-N organic results for a keyword from ValueSERP."""
    params = {
        "api_key": VALUESERP_API_KEY,
        "q": keyword,
        "num": num_results,
        "gl": country,
        "hl": language,
        "output": "json"
    }
    response = requests.get("https://api.valueserp.com/search", params=params)
    if response.status_code == 200:
        data = response.json()
        results = data.get("organic_results", [])
        urls = [r.get("link", "") for r in results[:num_results]]
        return urls
    return []

def fetch_serps_for_keywords(keywords, delay=1.0):
    """Fetch SERPs for a list of keywords with rate limiting."""
    serp_data = {}
    for i, kw in enumerate(keywords):
        print(f"Fetching SERP {i+1}/{len(keywords)}: {kw}")
        serp_data[kw] = fetch_serp(kw)
        time.sleep(delay)  # Respect API rate limits
    return serp_data

Step 2: Calculate SERP Overlap Matrix

def calculate_serp_overlap(serp_data):
    """
    Build a similarity matrix based on URL overlap in top-10 SERPs.
    Returns a dict of dicts: overlap[kw1][kw2] = number of shared URLs.
    """
    keywords = list(serp_data.keys())
    overlap = defaultdict(dict)

    for i, kw1 in enumerate(keywords):
        for j, kw2 in enumerate(keywords):
            if i == j:
                overlap[kw1][kw2] = 10  # Self-similarity = max
            else:
                set1 = set(serp_data[kw1])
                set2 = set(serp_data[kw2])
                shared = len(set1.intersection(set2))
                overlap[kw1][kw2] = shared

    return overlap, keywords

Step 3: Build Clusters Using Graph-Based Grouping

import networkx as nx

def build_clusters(overlap, keywords, threshold=3):
    """
    Build keyword clusters using a graph where edges exist between
    keywords sharing >= threshold SERP URLs.
    Each connected component becomes a cluster.
    """
    G = nx.Graph()
    G.add_nodes_from(keywords)

    for kw1 in keywords:
        for kw2 in keywords:
            if kw1 != kw2 and overlap[kw1][kw2] >= threshold:
                G.add_edge(kw1, kw2)

    clusters = []
    for component in nx.connected_components(G):
        clusters.append(list(component))

    return clusters

def export_clusters_to_csv(clusters, output_path="keyword_clusters.csv"):
    """Export cluster assignments to a CSV file."""
    rows = []
    for i, cluster in enumerate(clusters):
        for kw in cluster:
            rows.append({
                "cluster_id": i + 1,
                "keyword": kw,
                "cluster_size": len(cluster)
            })
    df = pd.DataFrame(rows)
    df.sort_values(["cluster_id", "keyword"], inplace=True)
    df.to_csv(output_path, index=False)
    print(f"Exported {len(rows)} keywords in {len(clusters)} clusters to {output_path}")
    return df

Step 4: Run the Full Pipeline

if __name__ == "__main__":
    # Your keyword list
    keywords = [
        "project management software",
        "project management tool",
        "best project management app",
        "task management software",
        "team project management",
        "project tracking software",
        "construction project management software",
        "project management for construction",
        "construction scheduling software",
    ]

    # Fetch SERPs
    serp_data = fetch_serps_for_keywords(keywords, delay=1.5)

    # Build overlap matrix
    overlap, kw_list = calculate_serp_overlap(serp_data)

    # Cluster with threshold=3 (keywords sharing 3+ top-10 URLs go together)
    clusters = build_clusters(overlap, kw_list, threshold=3)

    # Export results
    df = export_clusters_to_csv(clusters)
    print(df.groupby("cluster_id").size().reset_index(name="count"))

Expected output: the first 6 keywords likely cluster together (they share the same product-category SERP), while the last 3 cluster separately (construction-specific SERP diverges).

Embedding-Based Semantic Clustering

For keyword sets too large for SERP-API calls (5,000+), embedding-based clustering provides high-quality results without per-query API costs. The approach: convert each keyword to a vector embedding, then use k-means or HDBSCAN to group semantically similar keywords.

Using OpenAI Embeddings

from openai import OpenAI
import numpy as np
from sklearn.cluster import KMeans
from sklearn.preprocessing import normalize
import pandas as pd

client = OpenAI(api_key="YOUR_OPENAI_API_KEY")

def get_embeddings(texts, model="text-embedding-3-small"):
    """Get embeddings for a list of texts using OpenAI API."""
    # Batch into groups of 100 to stay within API limits
    all_embeddings = []
    batch_size = 100

    for i in range(0, len(texts), batch_size):
        batch = texts[i:i+batch_size]
        response = client.embeddings.create(input=batch, model=model)
        batch_embeddings = [item.embedding for item in response.data]
        all_embeddings.extend(batch_embeddings)
        print(f"Processed {min(i+batch_size, len(texts))}/{len(texts)} embeddings")

    return np.array(all_embeddings)

def cluster_by_embeddings(keywords, n_clusters=None, auto_n=True):
    """
    Cluster keywords using embeddings and K-means.
    If auto_n=True, estimate optimal cluster count using elbow method.
    """
    embeddings = get_embeddings(keywords)
    normalized = normalize(embeddings)  # Cosine similarity requires normalisation

    if auto_n:
        # Elbow method: find n where inertia reduction plateaus
        inertias = []
        n_range = range(2, min(50, len(keywords) // 5))
        for n in n_range:
            km = KMeans(n_clusters=n, random_state=42, n_init=10)
            km.fit(normalized)
            inertias.append(km.inertia_)

        # Simple elbow detection: biggest second derivative
        diffs = np.diff(inertias)
        second_diffs = np.diff(diffs)
        optimal_n = n_range[np.argmax(second_diffs) + 2]
        n_clusters = optimal_n
        print(f"Auto-selected {n_clusters} clusters")

    km = KMeans(n_clusters=n_clusters, random_state=42, n_init=10)
    labels = km.fit_predict(normalized)

    return labels, embeddings

def export_embedding_clusters(keywords, labels, output_path="embedding_clusters.csv"):
    df = pd.DataFrame({
        "keyword": keywords,
        "cluster_id": labels
    })
    df["cluster_size"] = df.groupby("cluster_id")["cluster_id"].transform("count")
    df.sort_values(["cluster_id", "keyword"]).to_csv(output_path, index=False)
    print(f"Exported {len(df)} keywords in {df['cluster_id'].nunique()} clusters")
    return df

Using a Local Model (No API Cost)

# Alternative: Use sentence-transformers locally (free, no API)
# pip install sentence-transformers

from sentence_transformers import SentenceTransformer

def get_local_embeddings(texts, model_name="all-MiniLM-L6-v2"):
    """Get embeddings using a local sentence transformer model."""
    model = SentenceTransformer(model_name)
    embeddings = model.encode(texts, show_progress_bar=True)
    return embeddings

# Usage: replace get_embeddings() with get_local_embeddings()
# all-MiniLM-L6-v2 is fast and performs well for short texts like keywords

The Hybrid Approach

In practice, the best workflow combines both methods: use embedding clustering for the full keyword set to produce coarse clusters, then run SERP-similarity validation on each cluster's candidate keywords to confirm or split the coarse clusters. This gives you the scalability of embeddings with the accuracy of SERP similarity.

def hybrid_cluster(keywords, serp_data=None):
    """
    Step 1: Coarse embedding clustering
    Step 2: SERP-similarity validation within each coarse cluster
    Returns refined clusters.
    """
    # Step 1: Embedding clustering
    labels, _ = cluster_by_embeddings(keywords, auto_n=True)
    coarse_clusters = {}
    for kw, label in zip(keywords, labels):
        coarse_clusters.setdefault(label, []).append(kw)

    # Step 2: SERP validation within each coarse cluster
    if serp_data is None:
        print("No SERP data provided — returning coarse clusters only")
        return list(coarse_clusters.values())

    refined_clusters = []
    for cluster_id, cluster_kws in coarse_clusters.items():
        if len(cluster_kws) == 1:
            refined_clusters.append(cluster_kws)
            continue
        cluster_serp = {kw: serp_data[kw] for kw in cluster_kws if kw in serp_data}
        overlap, _ = calculate_serp_overlap(cluster_serp)
        sub_clusters = build_clusters(overlap, list(cluster_serp.keys()), threshold=3)
        refined_clusters.extend(sub_clusters)

    return refined_clusters

Processing and Using Cluster Output

Raw cluster output is a list of keyword groups. To make it actionable, you need to assign each cluster a primary keyword (the head term), identify its content format, and map it to either an existing URL (for content updates) or a new URL (for content creation).

def label_clusters(clusters, volume_data=None):
    """
    Label each cluster with a primary keyword.
    Primary = highest volume keyword in cluster (if volume_data provided)
    Otherwise = longest keyword (assumes more specific = more representative)
    """
    labelled = []
    for i, cluster in enumerate(clusters):
        if volume_data:
            primary = max(cluster, key=lambda kw: volume_data.get(kw, 0))
        else:
            primary = max(cluster, key=len)

        labelled.append({
            "cluster_id": i + 1,
            "primary_keyword": primary,
            "all_keywords": cluster,
            "keyword_count": len(cluster)
        })
    return labelled

Case Study: 4,000-Keyword Cluster in Practice

A B2B software company had accumulated 4,000 keywords from a Semrush export and a competitor gap analysis. Manual clustering was estimated at 40+ hours. We ran the hybrid clustering approach: embedding pass first (OpenAI text-embedding-3-small, cost: ~$0.02 for 4,000 keywords), auto-selected 87 coarse clusters. We then ran SERP validation on the 20 largest coarse clusters (those with 30+ keywords) — 600 SERP API calls at $0.001/call = $0.60.

The process took 4 hours of compute time and produced 112 validated clusters from 4,000 keywords. Manual review of cluster labels took 3 hours. Total cost: <$2 in API fees + 3 hours of human review. The output fed directly into a 6-month content calendar with 112 content briefs, each with a validated primary keyword and supporting keyword list.

The cluster output is then run through our cannibalization detection workflow to identify cases where existing pages already target cluster primary keywords, preventing duplicate content creation. See also the topic clusters and pillar pages guide for how cluster output maps to site architecture.

FAQ

What SERP overlap threshold should I use for clustering?

For most use cases, 3 overlapping top-10 URLs is the right threshold. At 2, you'll over-cluster (too few groups, too many keywords per page). At 4+, you'll under-cluster (too many groups, near-duplicate pages). For highly competitive verticals where the top-10 is stable and consistent, you can raise to 4. For dynamic or niche SERPs, stay at 3 or even 2.

How many keywords can I cluster before SERP API costs become prohibitive?

SERP-similarity clustering requires O(n) API calls — one per keyword. At $0.001–$0.003 per call (ValueSERP, DataForSEO pricing), 1,000 keywords costs $1–$3. At 10,000 keywords: $10–$30. For sets above 5,000 keywords, the hybrid embedding approach reduces SERP API calls to only the validation step, typically 10–20% of the total keyword set.

What embedding model should I use?

For keyword clustering specifically: OpenAI's text-embedding-3-small performs well and is cost-effective. For a completely free, offline solution: sentence-transformers' all-MiniLM-L6-v2 is fast, runs locally, and produces good results for short text inputs like keywords. Avoid using large general-purpose LLMs for embedding generation — they're slow and expensive for this task.

How do I handle clusters that are too large (30+ keywords)?

Large clusters typically mean the threshold is too low, or the topic is genuinely broad enough to warrant a hub-and-spoke structure. Review the largest clusters manually: if they contain clear sub-topic groupings, split at the sub-topic level. If they're all truly targeting the same intent, consider a long-form comprehensive page that covers all the variant keywords as sections.

Can I cluster in a language other than English?

Yes. For SERP-similarity clustering, language doesn't matter — you're comparing URLs regardless of keyword language. For embedding-based clustering, use a multilingual model: OpenAI's text-embedding-3-small is multilingual, and sentence-transformers has excellent multilingual models (paraphrase-multilingual-MiniLM-L12-v2). Always fetch SERPs for the target country and language combination in your API parameters.

How do I validate that my clusters are correct?

Manual spot-check: for each cluster, open the primary keyword in an incognito browser and verify that the SERP matches what you'd expect given the cluster's other keywords. For statistical validation, calculate the average within-cluster SERP overlap and compare to the between-cluster overlap. Good clusters show within-cluster overlap of 4+ and between-cluster overlap of 1 or less.

Key Takeaways

  • SERP-similarity clustering is the most accurate method: keywords sharing 3+ top-10 URLs belong on the same page. This is the ground truth for clustering decisions.
  • Embedding-based clustering (OpenAI or sentence-transformers) handles large keyword sets cost-effectively and can be run entirely locally for free with all-MiniLM-L6-v2.
  • The hybrid approach — embedding clustering for coarse grouping, SERP validation for refinement — gives you the best of both methods at low cost.
  • Each cluster needs a primary keyword (highest volume or most representative) that will drive the page's title, H1, and URL slug.
  • Total cost for 4,000 keywords using the hybrid approach: under $2 in API fees. Manual review remains the largest time investment.
  • Cluster output feeds directly into content brief creation, cannibalization audits, and site architecture planning — it's the bridge between keyword research and content strategy.
  • Always validate cluster output against your existing site URLs to identify which clusters already have coverage and which need new content.

Conclusion

Keyword clustering with Python isn't difficult — the code above is a complete implementation you can run today. The harder part is making good decisions about thresholds, cluster sizes, and what to do with the output. Start with the SERP-similarity approach for your most strategically important keyword sets, and use the embedding approach for large exports that need rapid initial organisation.

The payoff for getting clustering right is significant: a properly clustered keyword set produces a content architecture that avoids cannibalization, builds topical authority efficiently, and gives you a clear content calendar rather than an undifferentiated keyword dump.

For the next step after clustering, see our guide on detecting and fixing keyword cannibalization. For how clusters map to pillar pages, see the topic clusters and pillar pages guide.

Reference: Sentence Transformers documentation covers the full range of embedding models available for local deployment.

YOUR READING CHECKLIST

Make the ideas stick.

Mark the sections you’ve worked through. Saved in this browser.

0 of 4 reviewed
Andrii Stanetskyi
ABOUT THE AUTHOR

Andrii Stanetskyi

Head of SEO / Technical SEO Lead based in Tallinn, Estonia. Technical architecture, enterprise eCommerce, Python automation, and AI-assisted workflows.

More about Andrii ↗
LET’S FIND THE REAL BOTTLENECK

A clearer picture.
A practical next step.

Get a focused SEO audit or a consultation on your next technical decision. We’ll agree on the scope and fee before any work begins.

01 / Diagnose02 / Prioritize03 / Plan
How can I help?

Scope and fee agreed before any work begins.

Choose your language

Explore SEO services in 26 languages. Journal articles retain their original language.

ENEnglish↗DEDeutsch↗FRFrançais↗ESEspañol↗ITItaliano↗PTPortuguês↗NLNederlands↗PLPolski↗SVSvenska↗DADansk↗FISuomi↗NONorsk↗ETEesti↗LVLatviešu↗LTLietuvių↗CSČeština↗RORomână↗HUMagyar↗ELΕλληνικά↗BGБългарски↗HRHrvatski↗SKSlovenčina↗SLSlovenščina↗RUРусский↗UKУкраїнська↗TRTürkçe↗
LET’S WORK ON YOUR WEBSITE
A CLEAR NEXT STEP

Let’s talk
about your site.

A focused SEO audit or a conversation about a specific challenge. Tell me where you are and what you want to change.

Andrii Stanetskyi
Andrii StanetskyiHead of SEO / Technical SEO Lead
[email protected] ↗
How can I help?

Scope and fee agreed before any work begins.