The shift from keyword matching to semantic understanding in search engines isn't metaphorical — it's a specific architectural change from BM25-style term frequency scoring to dense vector retrieval systems. Understanding how vector embeddings work and how to apply them operationally gives you a 12–18 month analytical advantage over practitioners still operating in keyword frequency space. This article covers embedding models, semantic gap analysis, content clustering, SERP similarity measurement, and the BigQuery + Vertex AI pipeline that makes this actionable at scale.
Embedding Fundamentals for SEO Practitioners
A text embedding is a high-dimensional numerical vector (typically 768–3072 dimensions) that encodes semantic meaning such that semantically similar texts have small cosine distances between their vectors. Google's search system uses multiple embedding models: one for understanding query intent, one for document content, one for entity mentions, and specialized models for specific content types (code, images, structured data).
The SEO relevance: when you optimize a page for a keyword, you're not just optimizing for exact-match retrieval — you're optimizing for semantic proximity to what Google has learned the query vector should map to, based on historical click signals and document quality judgments. Pages that are semantically far from Google's learned query-document mapping, regardless of keyword density, will underperform.
Embedding Model Selection for SEO Applications
| Model | Dimensions | Best For | Cost | Latency |
|---|---|---|---|---|
| text-embedding-004 (Google) | 768 | General semantic similarity, closest to Google's internal | $0.000025/1K chars | ~200ms |
| text-embedding-3-large (OpenAI) | 3072 | Nuanced semantic tasks, multilingual | $0.00013/1K tokens | ~300ms |
| all-mpnet-base-v2 (SBERT) | 768 | Local processing, high volume | Self-hosted | ~50ms/GPU |
| nomic-embed-text-v1.5 | 768 | Long documents (8192 token context) | Self-hosted or API | ~100ms |
| gemini-embedding-exp (Google) | 3072 | Closest alignment to Gemini-era search | Experimental | ~400ms |
For SEO analysis, text-embedding-004 is the recommended default — it's Google's own model, making it the closest available proxy to what the search system actually computes. Using OpenAI embeddings for Google SEO analysis is measuring with the wrong instrument.
Semantic Gap Analysis: Finding Content Holes
Semantic gap analysis identifies the difference between the embedding space your content covers and the embedding space your target queries occupy. The practical output: a ranked list of query clusters for which you have no semantically proximate content.
import numpy as np
from google.cloud import aiplatform
from vertexai.language_models import TextEmbeddingModel
import pandas as pd
from sklearn.metrics.pairwise import cosine_similarity
# Initialize Vertex AI
aiplatform.init(project="your-project", location="us-central1")
model = TextEmbeddingModel.from_pretrained("text-embedding-004")
def embed_texts(texts: list[str], batch_size: int = 250) -> np.ndarray:
"""
Embed a list of texts using Vertex AI text-embedding-004.
Returns (N, 768) numpy array.
"""
all_embeddings = []
for i in range(0, len(texts), batch_size):
batch = texts[i:i + batch_size]
embeddings = model.get_embeddings(batch)
vectors = [e.values for e in embeddings]
all_embeddings.extend(vectors)
return np.array(all_embeddings, dtype=np.float32)
def semantic_gap_analysis(
page_contents: list[dict], # [{"url": ..., "content": ...}]
target_queries: list[str],
similarity_threshold: float = 0.75
) -> pd.DataFrame:
"""
Find queries with no semantically proximate content page.
Returns DataFrame of uncovered queries sorted by gap severity.
"""
# Embed pages
page_texts = [p["content"][:2000] for p in page_contents] # First 2000 chars
page_urls = [p["url"] for p in page_contents]
page_embeddings = embed_texts(page_texts)
# Embed queries
query_embeddings = embed_texts(target_queries)
# Compute all-pairs similarity
sim_matrix = cosine_similarity(query_embeddings, page_embeddings)
# Find max similarity for each query
max_similarities = sim_matrix.max(axis=1)
best_matching_page_idx = sim_matrix.argmax(axis=1)
results = []
for i, (query, max_sim, best_idx) in enumerate(
zip(target_queries, max_similarities, best_matching_page_idx)
):
results.append({
"query": query,
"max_similarity": round(float(max_sim), 4),
"best_matching_page": page_urls[best_idx],
"has_coverage": max_sim >= similarity_threshold,
"gap_severity": round(float(1 - max_sim), 4),
})
df = pd.DataFrame(results)
return df.sort_values("gap_severity", ascending=False)
# Usage
pages = [{"url": "https://site.com/page1", "content": "...page content..."}]
queries = ["how to optimize knowledge graph", "entity disambiguation SEO", ...]
gaps = semantic_gap_analysis(pages, queries, similarity_threshold=0.72)
print(gaps[~gaps["has_coverage"]].head(20))
See our guide to programmatic content gap analysis using GSC data
SERP Semantic Similarity: Understanding What Google Wants
The most powerful application of embeddings in SEO is computing the semantic centroid of a SERP and measuring your page's distance from it. If the top 10 results for your target query are clustered tightly in embedding space, and your page is an outlier, you have a content alignment problem — not a link problem.
import requests
from bs4 import BeautifulSoup
import numpy as np
from sklearn.metrics.pairwise import cosine_similarity
def scrape_serp_contents(query: str, api_key: str, cx: str, num: int = 10) -> list[dict]:
"""
Use Custom Search API to get top results for a query.
Returns list of {url, snippet} dicts.
"""
url = "https://www.googleapis.com/customsearch/v1"
params = {
"key": api_key,
"cx": cx,
"q": query,
"num": num,
}
r = requests.get(url, params=params, timeout=10)
r.raise_for_status()
data = r.json()
results = []
for item in data.get("items", []):
results.append({
"url": item["link"],
"title": item["title"],
"snippet": item.get("snippet", ""),
"position": len(results) + 1,
})
return results
def analyze_serp_semantic_alignment(
query: str,
your_page_content: str,
serp_results: list[dict],
embed_func # callable: list[str] -> np.ndarray
) -> dict:
"""
Measure your page's semantic alignment with the SERP cluster.
Returns alignment score, centroid distance, and outlier flag.
"""
texts = [r["snippet"] for r in serp_results] + [your_page_content[:500]]
embeddings = embed_func(texts)
serp_embeddings = embeddings[:-1]
your_embedding = embeddings[-1:]
# Compute SERP centroid
centroid = serp_embeddings.mean(axis=0, keepdims=True)
# Distances from centroid
serp_distances = 1 - cosine_similarity(serp_embeddings, centroid).flatten()
your_distance = float(1 - cosine_similarity(your_embedding, centroid)[0][0])
avg_serp_distance = float(serp_distances.mean())
max_serp_distance = float(serp_distances.max())
return {
"query": query,
"your_centroid_distance": round(your_distance, 4),
"avg_ranking_page_distance": round(avg_serp_distance, 4),
"max_ranking_page_distance": round(max_serp_distance, 4),
"alignment_score": round(1 - your_distance, 4),
"is_semantic_outlier": your_distance > max_serp_distance,
"gap_vs_serp_avg": round(your_distance - avg_serp_distance, 4),
}
# A negative gap_vs_serp_avg means you're MORE aligned than average ranking page
# A positive gap means you're semantically farther from the centroid than your competitors
Detecting Content Type Mismatch via Embeddings
Beyond topic alignment, embeddings can detect when your content type doesn't match what's ranking. A query cluster centered on "how-to" guides has a distinctly different embedding signature than one centered on "product comparison" or "definition" content. This is operationally the most valuable insight from SERP analysis — content type mismatches can't be fixed by adding keywords.
# Classify SERP intent using embedding similarity to intent templates
INTENT_TEMPLATES = {
"informational_how_to": "Step-by-step guide explaining how to accomplish a task with instructions",
"informational_definition": "Definition and explanation of what a concept means with examples",
"transactional_product": "Product listing with price, specifications, and purchase options",
"transactional_service": "Service description with pricing, process, and contact information",
"navigational": "Homepage or specific brand page for direct navigation",
"commercial_comparison": "Comparison of multiple options with pros and cons analysis",
}
def classify_serp_intent(serp_snippets: list[str], embed_func) -> dict:
"""Classify aggregate SERP intent using embedding similarity to templates."""
template_texts = list(INTENT_TEMPLATES.values())
template_keys = list(INTENT_TEMPLATES.keys())
# Embed templates and SERP content
all_texts = template_texts + serp_snippets
embeddings = embed_func(all_texts)
template_embeddings = embeddings[:len(template_texts)]
serp_embeddings = embeddings[len(template_texts):]
# Average SERP embedding (centroid)
serp_centroid = serp_embeddings.mean(axis=0, keepdims=True)
# Similarity to each intent template
similarities = cosine_similarity(serp_centroid, template_embeddings).flatten()
intent_scores = dict(zip(template_keys, similarities.tolist()))
dominant_intent = max(intent_scores, key=intent_scores.get)
return {
"dominant_intent": dominant_intent,
"confidence": round(max(similarities), 4),
"all_scores": {k: round(v, 4) for k, v in intent_scores.items()},
}
Content Clustering and Cannibalization Detection
Keyword cannibalization analysis using keyword overlap is inherently limited — two pages can target very different keywords while covering the same semantic territory. Embedding-based cannibalization detection finds pages that are semantically redundant regardless of keyword overlap.
import numpy as np
from sklearn.metrics.pairwise import cosine_similarity
from sklearn.cluster import DBSCAN
import pandas as pd
def detect_semantic_cannibalization(
pages: list[dict], # [{"url": ..., "title": ..., "content": ..., "target_keyword": ...}]
embed_func,
high_similarity_threshold: float = 0.90,
cluster_eps: float = 0.15, # DBSCAN distance threshold
cluster_min_samples: int = 2
) -> dict:
"""
Detect semantically cannibalistic pages using DBSCAN clustering.
Returns clusters and pairwise high-similarity pairs.
"""
# Combine title + first 1000 chars of content for embedding
texts = [f"{p['title']} {p['content'][:1000]}" for p in pages]
urls = [p["url"] for p in pages]
embeddings = embed_func(texts)
# Compute pairwise cosine similarity
sim_matrix = cosine_similarity(embeddings)
# Find high-similarity pairs
cannibalization_pairs = []
n = len(urls)
for i in range(n):
for j in range(i + 1, n):
sim = float(sim_matrix[i][j])
if sim >= high_similarity_threshold:
cannibalization_pairs.append({
"url_a": urls[i],
"url_b": urls[j],
"similarity": round(sim, 4),
"keyword_a": pages[i].get("target_keyword", ""),
"keyword_b": pages[j].get("target_keyword", ""),
"severity": "critical" if sim > 0.95 else "high",
})
# DBSCAN clustering to identify content silos and thin clusters
distance_matrix = 1 - sim_matrix
distance_matrix = np.clip(distance_matrix, 0, None) # Numerical stability
clustering = DBSCAN(
eps=cluster_eps,
min_samples=cluster_min_samples,
metric="precomputed"
).fit(distance_matrix)
labels = clustering.labels_
cluster_map = {}
for url, label in zip(urls, labels):
cluster_id = int(label)
if cluster_id not in cluster_map:
cluster_map[cluster_id] = []
cluster_map[cluster_id].append(url)
return {
"cannibalization_pairs": sorted(cannibalization_pairs, key=lambda x: -x["similarity"]),
"clusters": cluster_map,
"noise_pages": cluster_map.get(-1, []), # DBSCAN labels outliers as -1
"total_clusters": len(set(labels)) - (1 if -1 in labels else 0),
}
Query Intent Classification at Scale
GSC exports give you thousands of queries. Classifying them by intent at scale requires embedding-based classification rather than rule-based approaches (which miss too many cases). The pipeline: embed all queries, embed intent class descriptions, nearest-centroid classify.
-- BigQuery: Store and query embedding-classified GSC keywords
-- Assumes you've run the Python embedding pipeline and stored results
CREATE TABLE IF NOT EXISTS your-project.seo.query_embeddings
(
query STRING,
embedding ARRAY,
dominant_intent STRING,
intent_confidence FLOAT64,
embedding_model STRING,
computed_at TIMESTAMP
)
PARTITION BY DATE(computed_at);
-- Find queries in the same semantic cluster as a seed query
-- Using BigQuery ML's vector search (requires BigQuery ML)
SELECT
query,
dominant_intent,
ML.DISTANCE(
embedding,
(SELECT embedding FROM your-project.seo.query_embeddings WHERE query = 'knowledge graph optimization'),
'COSINE'
) AS semantic_distance
FROM your-project.seo.query_embeddings
WHERE
DATE(computed_at) = CURRENT_DATE()
AND query != 'knowledge graph optimization'
ORDER BY semantic_distance ASC
LIMIT 50;
-- Identify intent distribution across your GSC query portfolio
SELECT
dominant_intent,
COUNT(*) AS query_count,
AVG(intent_confidence) AS avg_confidence,
APPROX_QUANTILES(intent_confidence, 4)[OFFSET(2)] AS median_confidence
FROM your-project.seo.query_embeddings
WHERE DATE(computed_at) = CURRENT_DATE()
GROUP BY dominant_intent
ORDER BY query_count DESC;
BigQuery + Vertex AI Production Pipeline
The production architecture for embedding-based SEO analysis runs on BigQuery + Vertex AI. The key insight: BigQuery's remote functions feature lets you call Vertex AI embedding models directly from SQL, enabling you to embed arbitrary text at BigQuery scale without managing a Python service.
-- Step 1: Create BigQuery remote function for Vertex AI embeddings
-- (One-time setup — requires appropriate IAM permissions)
CREATE OR REPLACE FUNCTION your-project.seo.embed_text(text STRING)
RETURNS ARRAY
REMOTE WITH CONNECTION your-project.us-central1.vertex-ai-connection
OPTIONS (
endpoint = 'https://us-central1-aiplatform.googleapis.com/v1/projects/your-project/locations/us-central1/publishers/google/models/text-embedding-004:predict',
max_batching_rows = 250
);
-- Step 2: Batch embed all GSC queries from the past 30 days
INSERT INTO your-project.seo.query_embeddings
SELECT
query,
your-project.seo.embed_text(query) AS embedding,
NULL AS dominant_intent, -- Classify in next step
NULL AS intent_confidence,
'text-embedding-004' AS embedding_model,
CURRENT_TIMESTAMP() AS computed_at
FROM (
SELECT DISTINCT
query
FROM your-project.searchconsole.searchdata_site_impression
WHERE
data_date >= DATE_SUB(CURRENT_DATE(), INTERVAL 30 DAY)
AND impressions >= 10 -- Filter noise
AND query NOT IN (
SELECT DISTINCT query
FROM your-project.seo.query_embeddings
WHERE DATE(computed_at) >= DATE_SUB(CURRENT_DATE(), INTERVAL 7 DAY)
)
)
LIMIT 10000; -- Stay within daily quota
Python: Full Pipeline Orchestration
from google.cloud import bigquery
from vertexai.language_models import TextEmbeddingModel
import vertexai
import numpy as np
import pandas as pd
from datetime import datetime, timezone
def run_seo_embedding_pipeline(
project_id: str,
gsc_dataset: str,
seo_dataset: str,
days_back: int = 30,
min_impressions: int = 10,
batch_size: int = 250
) -> dict:
"""
Full pipeline: fetch new GSC queries → embed → store → classify intent.
Returns summary stats.
"""
bq = bigquery.Client(project=project_id)
vertexai.init(project=project_id, location="us-central1")
embed_model = TextEmbeddingModel.from_pretrained("text-embedding-004")
# Fetch unembedded queries
fetch_sql = f"""
SELECT DISTINCT query
FROM {project_id}.{gsc_dataset}.searchdata_site_impression
WHERE
data_date >= DATE_SUB(CURRENT_DATE(), INTERVAL {days_back} DAY)
AND impressions >= {min_impressions}
AND query NOT IN (
SELECT DISTINCT query
FROM {project_id}.{seo_dataset}.query_embeddings
WHERE DATE(computed_at) >= DATE_SUB(CURRENT_DATE(), INTERVAL 7 DAY)
)
LIMIT 5000
"""
queries_df = bq.query(fetch_sql).to_dataframe()
queries = queries_df["query"].tolist()
if not queries:
return {"status": "no_new_queries", "embedded": 0}
# Embed in batches
rows = []
for i in range(0, len(queries), batch_size):
batch = queries[i:i + batch_size]
try:
embeddings = embed_model.get_embeddings(batch)
for query, emb in zip(batch, embeddings):
rows.append({
"query": query,
"embedding": emb.values,
"embedding_model": "text-embedding-004",
"computed_at": datetime.now(timezone.utc).isoformat(),
})
except Exception as e:
print(f"Batch {i//batch_size} failed: {e}")
continue
# Insert to BigQuery
if rows:
table_id = f"{project_id}.{seo_dataset}.query_embeddings"
errors = bq.insert_rows_json(table_id, rows)
if errors:
raise RuntimeError(f"BigQuery insert errors: {errors}")
return {
"status": "success",
"queries_fetched": len(queries),
"embedded": len(rows),
"failed": len(queries) - len(rows),
}
See our BigQuery GSC data analysis guide for the full data pipeline
FAQ
Q: How similar are commercial embedding models to what Google actually uses for ranking?
Google's text-embedding-004 is the closest available proxy, but it's not the ranking model — it's a general semantic similarity model. Google's actual ranking uses multi-task models trained on search-specific signals (clicks, dwell time, query reformulation patterns). Commercial embeddings are useful for directional analysis and content auditing, not for predicting exact ranking positions. Treat similarity scores as relative signals, not absolute ranking predictors.
Q: What cosine similarity threshold indicates semantic cannibalization?
Above 0.90 is reliably cannibalistic for full-page embeddings. Between 0.80 and 0.90 is "competitive overlap" that warrants investigation. Below 0.75, pages are semantically distinct enough to coexist. These thresholds are model-dependent — calibrate against known cannibalization pairs on your own site before using them operationally. Title-only embeddings behave differently than full-content embeddings.
Q: Should I embed the full page content or just the first N tokens?
For most embedding models, the first 512–1024 tokens (roughly 400–800 words) capture the majority of the topical signal. Models have context windows (text-embedding-004 is 3072 tokens) but empirically, the first third of a well-structured document often produces representative embeddings. For long-form content, consider embedding the intro + conclusion separately and averaging. For product pages, embed title + description + key specifications only.
Q: Can I use vector embeddings to identify which pages to consolidate?
Yes, and this is one of the highest-ROI applications. Run DBSCAN clustering on your full page corpus, identify clusters with 3+ pages in very tight proximity (eps < 0.10), then audit for consolidation. Sort candidates by total cluster impressions from GSC — high-impression clusters with tight semantic overlap are priority consolidation targets. Consolidating semantically duplicate content into single authoritative pages consistently produces ranking gains within 60–90 days.
Q: How do I handle multilingual content in embedding pipelines?
Use multilingual models (text-multilingual-embedding-002 from Google, or paraphrase-multilingual-mpnet-base-v2 from SBERT) when comparing content across languages. Never use monolingual models to compare Spanish and English content — the embedding spaces are incomparable. For cross-lingual gap analysis, embed query intent in one language, translate to another using a translation model, then re-embed in the target language space.
Q: What's the BigQuery cost for running embedding analysis on 100K queries?
BigQuery remote function calls to Vertex AI are billed at Vertex AI rates: text-embedding-004 is $0.000025 per 1K characters. At an average query length of 30 characters, 100K queries ≈ 3M characters ≈ $0.075 in model costs. BigQuery compute adds roughly $0.005 per query job. The real cost is the Vertex AI embedding API rate limit (1,500 requests/min by default) which makes 100K queries take ~67 minutes in serial batches of 250.
Key Takeaways
- Use
text-embedding-004(Google's own model) for SEO embedding analysis — it's the closest proxy to the search system's semantic understanding. - SERP centroid analysis is more actionable than individual competitor comparison: measure your page's distance from the semantic center of gravity of the top 10 results.
- Embedding-based cannibalization detection finds semantic overlap that keyword analysis misses entirely — run it on any content library with >500 pages.
- BigQuery remote functions + Vertex AI is the production-grade architecture for embedding SEO data at scale without managing Python services.
- A cosine similarity above 0.90 between two full-page embeddings is reliable signal for cannibalization; between 0.80–0.90 warrants topical differentiation work.
- DBSCAN clustering on query embeddings from GSC is the most efficient method for identifying content gaps — clusters with no current content = gap priorities.
Conclusion
Vector embeddings aren't a future technology — they're the current analytical substrate of production search systems. Practitioners who can measure semantic alignment, detect content redundancy, and identify gap opportunities in embedding space are operating on the same conceptual plane as the search systems they're trying to influence. The BigQuery + Vertex AI pipeline described here puts that analysis within reach at any reasonable content scale, with costs that make daily runs economically trivial. The competitive advantage is real and the window to build it before it's commoditized is finite.
Next: How to Use BigQuery to Analyze GSC Data at Scale Google Vertex AI Text Embeddings API Documentation