Skip to content
CONTENT & AUTHORITY / FIELD NOTE 058

Vector Embeddings for SEO: Semantic Similarity Analysis at Scale

Reading map: Embedding Fundamentals for SEO Practitioners; Semantic Gap Analysis: Finding Content Holes; SERP Semantic Similarity: Understanding What Google Wants; Content Clustering and Cannibalization Detection
A reading map of this field note. Download SVG ↓

The shift from keyword matching to semantic understanding in search engines isn't metaphorical — it's a specific architectural change from BM25-style term frequency scoring to dense vector retrieval systems. Understanding how vector embeddings work and how to apply them operationally gives you a 12–18 month analytical advantage over practitioners still operating in keyword frequency space. This article covers embedding models, semantic gap analysis, content clustering, SERP similarity measurement, and the BigQuery + Vertex AI pipeline that makes this actionable at scale.

Embedding Fundamentals for SEO Practitioners

A text embedding is a high-dimensional numerical vector (typically 768–3072 dimensions) that encodes semantic meaning such that semantically similar texts have small cosine distances between their vectors. Google's search system uses multiple embedding models: one for understanding query intent, one for document content, one for entity mentions, and specialized models for specific content types (code, images, structured data).

The SEO relevance: when you optimize a page for a keyword, you're not just optimizing for exact-match retrieval — you're optimizing for semantic proximity to what Google has learned the query vector should map to, based on historical click signals and document quality judgments. Pages that are semantically far from Google's learned query-document mapping, regardless of keyword density, will underperform.

Embedding Model Selection for SEO Applications

Model Dimensions Best For Cost Latency
text-embedding-004 (Google) 768 General semantic similarity, closest to Google's internal $0.000025/1K chars ~200ms
text-embedding-3-large (OpenAI) 3072 Nuanced semantic tasks, multilingual $0.00013/1K tokens ~300ms
all-mpnet-base-v2 (SBERT) 768 Local processing, high volume Self-hosted ~50ms/GPU
nomic-embed-text-v1.5 768 Long documents (8192 token context) Self-hosted or API ~100ms
gemini-embedding-exp (Google) 3072 Closest alignment to Gemini-era search Experimental ~400ms

For SEO analysis, text-embedding-004 is the recommended default — it's Google's own model, making it the closest available proxy to what the search system actually computes. Using OpenAI embeddings for Google SEO analysis is measuring with the wrong instrument.

Semantic Gap Analysis: Finding Content Holes

Semantic gap analysis identifies the difference between the embedding space your content covers and the embedding space your target queries occupy. The practical output: a ranked list of query clusters for which you have no semantically proximate content.

import numpy as np
from google.cloud import aiplatform
from vertexai.language_models import TextEmbeddingModel
import pandas as pd
from sklearn.metrics.pairwise import cosine_similarity

# Initialize Vertex AI
aiplatform.init(project="your-project", location="us-central1")
model = TextEmbeddingModel.from_pretrained("text-embedding-004")

def embed_texts(texts: list[str], batch_size: int = 250) -> np.ndarray:
    """
    Embed a list of texts using Vertex AI text-embedding-004.
    Returns (N, 768) numpy array.
    """
    all_embeddings = []
    for i in range(0, len(texts), batch_size):
        batch = texts[i:i + batch_size]
        embeddings = model.get_embeddings(batch)
        vectors = [e.values for e in embeddings]
        all_embeddings.extend(vectors)
    return np.array(all_embeddings, dtype=np.float32)

def semantic_gap_analysis(
    page_contents: list[dict],  # [{"url": ..., "content": ...}]
    target_queries: list[str],
    similarity_threshold: float = 0.75
) -> pd.DataFrame:
    """
    Find queries with no semantically proximate content page.
    Returns DataFrame of uncovered queries sorted by gap severity.
    """
    # Embed pages
    page_texts = [p["content"][:2000] for p in page_contents]  # First 2000 chars
    page_urls = [p["url"] for p in page_contents]
    page_embeddings = embed_texts(page_texts)

    # Embed queries
    query_embeddings = embed_texts(target_queries)

    # Compute all-pairs similarity
    sim_matrix = cosine_similarity(query_embeddings, page_embeddings)

    # Find max similarity for each query
    max_similarities = sim_matrix.max(axis=1)
    best_matching_page_idx = sim_matrix.argmax(axis=1)

    results = []
    for i, (query, max_sim, best_idx) in enumerate(
        zip(target_queries, max_similarities, best_matching_page_idx)
    ):
        results.append({
            "query": query,
            "max_similarity": round(float(max_sim), 4),
            "best_matching_page": page_urls[best_idx],
            "has_coverage": max_sim >= similarity_threshold,
            "gap_severity": round(float(1 - max_sim), 4),
        })

    df = pd.DataFrame(results)
    return df.sort_values("gap_severity", ascending=False)

# Usage
pages = [{"url": "https://site.com/page1", "content": "...page content..."}]
queries = ["how to optimize knowledge graph", "entity disambiguation SEO", ...]
gaps = semantic_gap_analysis(pages, queries, similarity_threshold=0.72)
print(gaps[~gaps["has_coverage"]].head(20))
See our guide to programmatic content gap analysis using GSC data

SERP Semantic Similarity: Understanding What Google Wants

The most powerful application of embeddings in SEO is computing the semantic centroid of a SERP and measuring your page's distance from it. If the top 10 results for your target query are clustered tightly in embedding space, and your page is an outlier, you have a content alignment problem — not a link problem.

import requests
from bs4 import BeautifulSoup
import numpy as np
from sklearn.metrics.pairwise import cosine_similarity

def scrape_serp_contents(query: str, api_key: str, cx: str, num: int = 10) -> list[dict]:
    """
    Use Custom Search API to get top results for a query.
    Returns list of {url, snippet} dicts.
    """
    url = "https://www.googleapis.com/customsearch/v1"
    params = {
        "key": api_key,
        "cx": cx,
        "q": query,
        "num": num,
    }
    r = requests.get(url, params=params, timeout=10)
    r.raise_for_status()
    data = r.json()

    results = []
    for item in data.get("items", []):
        results.append({
            "url": item["link"],
            "title": item["title"],
            "snippet": item.get("snippet", ""),
            "position": len(results) + 1,
        })
    return results

def analyze_serp_semantic_alignment(
    query: str,
    your_page_content: str,
    serp_results: list[dict],
    embed_func  # callable: list[str] -> np.ndarray
) -> dict:
    """
    Measure your page's semantic alignment with the SERP cluster.
    Returns alignment score, centroid distance, and outlier flag.
    """
    texts = [r["snippet"] for r in serp_results] + [your_page_content[:500]]
    embeddings = embed_func(texts)

    serp_embeddings = embeddings[:-1]
    your_embedding = embeddings[-1:]

    # Compute SERP centroid
    centroid = serp_embeddings.mean(axis=0, keepdims=True)

    # Distances from centroid
    serp_distances = 1 - cosine_similarity(serp_embeddings, centroid).flatten()
    your_distance = float(1 - cosine_similarity(your_embedding, centroid)[0][0])

    avg_serp_distance = float(serp_distances.mean())
    max_serp_distance = float(serp_distances.max())

    return {
        "query": query,
        "your_centroid_distance": round(your_distance, 4),
        "avg_ranking_page_distance": round(avg_serp_distance, 4),
        "max_ranking_page_distance": round(max_serp_distance, 4),
        "alignment_score": round(1 - your_distance, 4),
        "is_semantic_outlier": your_distance > max_serp_distance,
        "gap_vs_serp_avg": round(your_distance - avg_serp_distance, 4),
    }

# A negative gap_vs_serp_avg means you're MORE aligned than average ranking page
# A positive gap means you're semantically farther from the centroid than your competitors

Detecting Content Type Mismatch via Embeddings

Beyond topic alignment, embeddings can detect when your content type doesn't match what's ranking. A query cluster centered on "how-to" guides has a distinctly different embedding signature than one centered on "product comparison" or "definition" content. This is operationally the most valuable insight from SERP analysis — content type mismatches can't be fixed by adding keywords.

# Classify SERP intent using embedding similarity to intent templates
INTENT_TEMPLATES = {
    "informational_how_to": "Step-by-step guide explaining how to accomplish a task with instructions",
    "informational_definition": "Definition and explanation of what a concept means with examples",
    "transactional_product": "Product listing with price, specifications, and purchase options",
    "transactional_service": "Service description with pricing, process, and contact information",
    "navigational": "Homepage or specific brand page for direct navigation",
    "commercial_comparison": "Comparison of multiple options with pros and cons analysis",
}

def classify_serp_intent(serp_snippets: list[str], embed_func) -> dict:
    """Classify aggregate SERP intent using embedding similarity to templates."""
    template_texts = list(INTENT_TEMPLATES.values())
    template_keys = list(INTENT_TEMPLATES.keys())

    # Embed templates and SERP content
    all_texts = template_texts + serp_snippets
    embeddings = embed_func(all_texts)

    template_embeddings = embeddings[:len(template_texts)]
    serp_embeddings = embeddings[len(template_texts):]

    # Average SERP embedding (centroid)
    serp_centroid = serp_embeddings.mean(axis=0, keepdims=True)

    # Similarity to each intent template
    similarities = cosine_similarity(serp_centroid, template_embeddings).flatten()

    intent_scores = dict(zip(template_keys, similarities.tolist()))
    dominant_intent = max(intent_scores, key=intent_scores.get)

    return {
        "dominant_intent": dominant_intent,
        "confidence": round(max(similarities), 4),
        "all_scores": {k: round(v, 4) for k, v in intent_scores.items()},
    }

Content Clustering and Cannibalization Detection

Keyword cannibalization analysis using keyword overlap is inherently limited — two pages can target very different keywords while covering the same semantic territory. Embedding-based cannibalization detection finds pages that are semantically redundant regardless of keyword overlap.

import numpy as np
from sklearn.metrics.pairwise import cosine_similarity
from sklearn.cluster import DBSCAN
import pandas as pd

def detect_semantic_cannibalization(
    pages: list[dict],  # [{"url": ..., "title": ..., "content": ..., "target_keyword": ...}]
    embed_func,
    high_similarity_threshold: float = 0.90,
    cluster_eps: float = 0.15,  # DBSCAN distance threshold
    cluster_min_samples: int = 2
) -> dict:
    """
    Detect semantically cannibalistic pages using DBSCAN clustering.
    Returns clusters and pairwise high-similarity pairs.
    """
    # Combine title + first 1000 chars of content for embedding
    texts = [f"{p['title']} {p['content'][:1000]}" for p in pages]
    urls = [p["url"] for p in pages]

    embeddings = embed_func(texts)

    # Compute pairwise cosine similarity
    sim_matrix = cosine_similarity(embeddings)

    # Find high-similarity pairs
    cannibalization_pairs = []
    n = len(urls)
    for i in range(n):
        for j in range(i + 1, n):
            sim = float(sim_matrix[i][j])
            if sim >= high_similarity_threshold:
                cannibalization_pairs.append({
                    "url_a": urls[i],
                    "url_b": urls[j],
                    "similarity": round(sim, 4),
                    "keyword_a": pages[i].get("target_keyword", ""),
                    "keyword_b": pages[j].get("target_keyword", ""),
                    "severity": "critical" if sim > 0.95 else "high",
                })

    # DBSCAN clustering to identify content silos and thin clusters
    distance_matrix = 1 - sim_matrix
    distance_matrix = np.clip(distance_matrix, 0, None)  # Numerical stability

    clustering = DBSCAN(
        eps=cluster_eps,
        min_samples=cluster_min_samples,
        metric="precomputed"
    ).fit(distance_matrix)

    labels = clustering.labels_
    cluster_map = {}
    for url, label in zip(urls, labels):
        cluster_id = int(label)
        if cluster_id not in cluster_map:
            cluster_map[cluster_id] = []
        cluster_map[cluster_id].append(url)

    return {
        "cannibalization_pairs": sorted(cannibalization_pairs, key=lambda x: -x["similarity"]),
        "clusters": cluster_map,
        "noise_pages": cluster_map.get(-1, []),  # DBSCAN labels outliers as -1
        "total_clusters": len(set(labels)) - (1 if -1 in labels else 0),
    }

Query Intent Classification at Scale

GSC exports give you thousands of queries. Classifying them by intent at scale requires embedding-based classification rather than rule-based approaches (which miss too many cases). The pipeline: embed all queries, embed intent class descriptions, nearest-centroid classify.

-- BigQuery: Store and query embedding-classified GSC keywords
-- Assumes you've run the Python embedding pipeline and stored results

CREATE TABLE IF NOT EXISTS your-project.seo.query_embeddings
(
  query STRING,
  embedding ARRAY,
  dominant_intent STRING,
  intent_confidence FLOAT64,
  embedding_model STRING,
  computed_at TIMESTAMP
)
PARTITION BY DATE(computed_at);

-- Find queries in the same semantic cluster as a seed query
-- Using BigQuery ML's vector search (requires BigQuery ML)
SELECT
  query,
  dominant_intent,
  ML.DISTANCE(
    embedding,
    (SELECT embedding FROM your-project.seo.query_embeddings WHERE query = 'knowledge graph optimization'),
    'COSINE'
  ) AS semantic_distance
FROM your-project.seo.query_embeddings
WHERE
  DATE(computed_at) = CURRENT_DATE()
  AND query != 'knowledge graph optimization'
ORDER BY semantic_distance ASC
LIMIT 50;

-- Identify intent distribution across your GSC query portfolio
SELECT
  dominant_intent,
  COUNT(*) AS query_count,
  AVG(intent_confidence) AS avg_confidence,
  APPROX_QUANTILES(intent_confidence, 4)[OFFSET(2)] AS median_confidence
FROM your-project.seo.query_embeddings
WHERE DATE(computed_at) = CURRENT_DATE()
GROUP BY dominant_intent
ORDER BY query_count DESC;

BigQuery + Vertex AI Production Pipeline

The production architecture for embedding-based SEO analysis runs on BigQuery + Vertex AI. The key insight: BigQuery's remote functions feature lets you call Vertex AI embedding models directly from SQL, enabling you to embed arbitrary text at BigQuery scale without managing a Python service.

-- Step 1: Create BigQuery remote function for Vertex AI embeddings
-- (One-time setup — requires appropriate IAM permissions)

CREATE OR REPLACE FUNCTION your-project.seo.embed_text(text STRING)
RETURNS ARRAY
REMOTE WITH CONNECTION your-project.us-central1.vertex-ai-connection
OPTIONS (
  endpoint = 'https://us-central1-aiplatform.googleapis.com/v1/projects/your-project/locations/us-central1/publishers/google/models/text-embedding-004:predict',
  max_batching_rows = 250
);

-- Step 2: Batch embed all GSC queries from the past 30 days
INSERT INTO your-project.seo.query_embeddings
SELECT
  query,
  your-project.seo.embed_text(query) AS embedding,
  NULL AS dominant_intent,  -- Classify in next step
  NULL AS intent_confidence,
  'text-embedding-004' AS embedding_model,
  CURRENT_TIMESTAMP() AS computed_at
FROM (
  SELECT DISTINCT
    query
  FROM your-project.searchconsole.searchdata_site_impression
  WHERE
    data_date >= DATE_SUB(CURRENT_DATE(), INTERVAL 30 DAY)
    AND impressions >= 10  -- Filter noise
    AND query NOT IN (
      SELECT DISTINCT query
      FROM your-project.seo.query_embeddings
      WHERE DATE(computed_at) >= DATE_SUB(CURRENT_DATE(), INTERVAL 7 DAY)
    )
)
LIMIT 10000;  -- Stay within daily quota

Python: Full Pipeline Orchestration

from google.cloud import bigquery
from vertexai.language_models import TextEmbeddingModel
import vertexai
import numpy as np
import pandas as pd
from datetime import datetime, timezone

def run_seo_embedding_pipeline(
    project_id: str,
    gsc_dataset: str,
    seo_dataset: str,
    days_back: int = 30,
    min_impressions: int = 10,
    batch_size: int = 250
) -> dict:
    """
    Full pipeline: fetch new GSC queries → embed → store → classify intent.
    Returns summary stats.
    """
    bq = bigquery.Client(project=project_id)
    vertexai.init(project=project_id, location="us-central1")
    embed_model = TextEmbeddingModel.from_pretrained("text-embedding-004")

    # Fetch unembedded queries
    fetch_sql = f"""
        SELECT DISTINCT query
        FROM {project_id}.{gsc_dataset}.searchdata_site_impression
        WHERE
            data_date >= DATE_SUB(CURRENT_DATE(), INTERVAL {days_back} DAY)
            AND impressions >= {min_impressions}
            AND query NOT IN (
                SELECT DISTINCT query
                FROM {project_id}.{seo_dataset}.query_embeddings
                WHERE DATE(computed_at) >= DATE_SUB(CURRENT_DATE(), INTERVAL 7 DAY)
            )
        LIMIT 5000
    """
    queries_df = bq.query(fetch_sql).to_dataframe()
    queries = queries_df["query"].tolist()

    if not queries:
        return {"status": "no_new_queries", "embedded": 0}

    # Embed in batches
    rows = []
    for i in range(0, len(queries), batch_size):
        batch = queries[i:i + batch_size]
        try:
            embeddings = embed_model.get_embeddings(batch)
            for query, emb in zip(batch, embeddings):
                rows.append({
                    "query": query,
                    "embedding": emb.values,
                    "embedding_model": "text-embedding-004",
                    "computed_at": datetime.now(timezone.utc).isoformat(),
                })
        except Exception as e:
            print(f"Batch {i//batch_size} failed: {e}")
            continue

    # Insert to BigQuery
    if rows:
        table_id = f"{project_id}.{seo_dataset}.query_embeddings"
        errors = bq.insert_rows_json(table_id, rows)
        if errors:
            raise RuntimeError(f"BigQuery insert errors: {errors}")

    return {
        "status": "success",
        "queries_fetched": len(queries),
        "embedded": len(rows),
        "failed": len(queries) - len(rows),
    }
See our BigQuery GSC data analysis guide for the full data pipeline

FAQ

Q: How similar are commercial embedding models to what Google actually uses for ranking?

Google's text-embedding-004 is the closest available proxy, but it's not the ranking model — it's a general semantic similarity model. Google's actual ranking uses multi-task models trained on search-specific signals (clicks, dwell time, query reformulation patterns). Commercial embeddings are useful for directional analysis and content auditing, not for predicting exact ranking positions. Treat similarity scores as relative signals, not absolute ranking predictors.

Q: What cosine similarity threshold indicates semantic cannibalization?

Above 0.90 is reliably cannibalistic for full-page embeddings. Between 0.80 and 0.90 is "competitive overlap" that warrants investigation. Below 0.75, pages are semantically distinct enough to coexist. These thresholds are model-dependent — calibrate against known cannibalization pairs on your own site before using them operationally. Title-only embeddings behave differently than full-content embeddings.

Q: Should I embed the full page content or just the first N tokens?

For most embedding models, the first 512–1024 tokens (roughly 400–800 words) capture the majority of the topical signal. Models have context windows (text-embedding-004 is 3072 tokens) but empirically, the first third of a well-structured document often produces representative embeddings. For long-form content, consider embedding the intro + conclusion separately and averaging. For product pages, embed title + description + key specifications only.

Q: Can I use vector embeddings to identify which pages to consolidate?

Yes, and this is one of the highest-ROI applications. Run DBSCAN clustering on your full page corpus, identify clusters with 3+ pages in very tight proximity (eps < 0.10), then audit for consolidation. Sort candidates by total cluster impressions from GSC — high-impression clusters with tight semantic overlap are priority consolidation targets. Consolidating semantically duplicate content into single authoritative pages consistently produces ranking gains within 60–90 days.

Q: How do I handle multilingual content in embedding pipelines?

Use multilingual models (text-multilingual-embedding-002 from Google, or paraphrase-multilingual-mpnet-base-v2 from SBERT) when comparing content across languages. Never use monolingual models to compare Spanish and English content — the embedding spaces are incomparable. For cross-lingual gap analysis, embed query intent in one language, translate to another using a translation model, then re-embed in the target language space.

Q: What's the BigQuery cost for running embedding analysis on 100K queries?

BigQuery remote function calls to Vertex AI are billed at Vertex AI rates: text-embedding-004 is $0.000025 per 1K characters. At an average query length of 30 characters, 100K queries ≈ 3M characters ≈ $0.075 in model costs. BigQuery compute adds roughly $0.005 per query job. The real cost is the Vertex AI embedding API rate limit (1,500 requests/min by default) which makes 100K queries take ~67 minutes in serial batches of 250.

Key Takeaways

  • Use text-embedding-004 (Google's own model) for SEO embedding analysis — it's the closest proxy to the search system's semantic understanding.
  • SERP centroid analysis is more actionable than individual competitor comparison: measure your page's distance from the semantic center of gravity of the top 10 results.
  • Embedding-based cannibalization detection finds semantic overlap that keyword analysis misses entirely — run it on any content library with >500 pages.
  • BigQuery remote functions + Vertex AI is the production-grade architecture for embedding SEO data at scale without managing Python services.
  • A cosine similarity above 0.90 between two full-page embeddings is reliable signal for cannibalization; between 0.80–0.90 warrants topical differentiation work.
  • DBSCAN clustering on query embeddings from GSC is the most efficient method for identifying content gaps — clusters with no current content = gap priorities.

Conclusion

Vector embeddings aren't a future technology — they're the current analytical substrate of production search systems. Practitioners who can measure semantic alignment, detect content redundancy, and identify gap opportunities in embedding space are operating on the same conceptual plane as the search systems they're trying to influence. The BigQuery + Vertex AI pipeline described here puts that analysis within reach at any reasonable content scale, with costs that make daily runs economically trivial. The competitive advantage is real and the window to build it before it's commoditized is finite.

Next: How to Use BigQuery to Analyze GSC Data at Scale Google Vertex AI Text Embeddings API Documentation
YOUR READING CHECKLIST

Make the ideas stick.

Mark the sections you’ve worked through. Saved in this browser.

0 of 4 reviewed
Andrii Stanetskyi
ABOUT THE AUTHOR

Andrii Stanetskyi

Head of SEO / Technical SEO Lead based in Tallinn, Estonia. Technical architecture, enterprise eCommerce, Python automation, and AI-assisted workflows.

More about Andrii ↗
LET’S FIND THE REAL BOTTLENECK

A clearer picture.
A practical next step.

Get a focused SEO audit or a consultation on your next technical decision. We’ll agree on the scope and fee before any work begins.

01 / Diagnose02 / Prioritize03 / Plan
How can I help?

Scope and fee agreed before any work begins.

Choose your language

Explore SEO services in 26 languages. Journal articles retain their original language.

ENEnglish↗DEDeutsch↗FRFrançais↗ESEspañol↗ITItaliano↗PTPortuguês↗NLNederlands↗PLPolski↗SVSvenska↗DADansk↗FISuomi↗NONorsk↗ETEesti↗LVLatviešu↗LTLietuvių↗CSČeština↗RORomână↗HUMagyar↗ELΕλληνικά↗BGБългарски↗HRHrvatski↗SKSlovenčina↗SLSlovenščina↗RUРусский↗UKУкраїнська↗TRTürkçe↗
LET’S WORK ON YOUR WEBSITE
A CLEAR NEXT STEP

Let’s talk
about your site.

A focused SEO audit or a conversation about a specific challenge. Tell me where you are and what you want to change.

Andrii Stanetskyi
Andrii StanetskyiHead of SEO / Technical SEO Lead
[email protected] ↗
How can I help?

Scope and fee agreed before any work begins.