E-commerce structured data is where the gap between knowing the spec and operating at scale becomes brutally obvious. A catalog of 500,000 SKUs with accurate pricing, availability, ratings, and shipping information — rendered correctly in JSON-LD, validated continuously, and updated within minutes of inventory changes — is an engineering problem as much as an SEO problem. This article covers the architecture for doing it right: from Product graph modeling to Merchant Center feed synchronization to edge-rendered pricing at sub-millisecond latency.
Product Graph Architecture
A Product in Schema.org is not a flat object. It's a graph node with relationships to Offers, Organizations (brand/seller), AggregateRating, Reviews, ProductGroup (for variants), and MerchantReturnPolicy. Each of these sub-entities should have their own @id for proper graph modeling, especially when the same seller entity or return policy applies to thousands of products — you define it once, reference it everywhere.
Complete Product Entity with @graph
{
"@context": "https://schema.org",
"@graph": [
{
"@type": "Product",
"@id": "https://shop.com/products/widget-pro-blue/#product",
"name": "Widget Pro — Blue",
"description": "High-performance widget for professional applications. 3000 RPM max.",
"sku": "WGT-PRO-BLU-001",
"gtin14": "00012345678905",
"mpn": "WGTPRO-B",
"brand": {
"@type": "Brand",
"@id": "https://shop.com/#brand-widgetco",
"name": "WidgetCo"
},
"image": [
{
"@type": "ImageObject",
"@id": "https://shop.com/products/widget-pro-blue/#image-main",
"url": "https://cdn.shop.com/widget-pro-blue-main.jpg",
"width": 1200,
"height": 1200,
"caption": "Widget Pro Blue — front view"
},
{
"@type": "ImageObject",
"@id": "https://shop.com/products/widget-pro-blue/#image-side",
"url": "https://cdn.shop.com/widget-pro-blue-side.jpg",
"width": 1200,
"height": 1200
}
],
"offers": {
"@id": "https://shop.com/products/widget-pro-blue/#offer"
},
"aggregateRating": {
"@id": "https://shop.com/products/widget-pro-blue/#aggregaterating"
},
"isVariantOf": {
"@type": "ProductGroup",
"@id": "https://shop.com/products/widget-pro/#productgroup",
"name": "Widget Pro",
"hasVariant": [
{"@id": "https://shop.com/products/widget-pro-blue/#product"},
{"@id": "https://shop.com/products/widget-pro-red/#product"},
{"@id": "https://shop.com/products/widget-pro-black/#product"}
],
"variesBy": ["color"]
}
},
{
"@type": "Offer",
"@id": "https://shop.com/products/widget-pro-blue/#offer",
"url": "https://shop.com/products/widget-pro-blue/",
"priceCurrency": "USD",
"price": "149.99",
"priceValidUntil": "2026-05-31",
"availability": "https://schema.org/InStock",
"itemCondition": "https://schema.org/NewCondition",
"seller": {
"@type": "Organization",
"@id": "https://shop.com/#organization"
},
"shippingDetails": {
"@id": "https://shop.com/#shipping-standard-us"
},
"hasMerchantReturnPolicy": {
"@id": "https://shop.com/#return-policy-30day"
}
},
{
"@type": "AggregateRating",
"@id": "https://shop.com/products/widget-pro-blue/#aggregaterating",
"ratingValue": "4.7",
"reviewCount": 284,
"bestRating": "5",
"worstRating": "1"
},
{
"@type": "OfferShippingDetails",
"@id": "https://shop.com/#shipping-standard-us",
"shippingRate": {
"@type": "MonetaryAmount",
"value": "0",
"currency": "USD"
},
"shippingDestination": {
"@type": "DefinedRegion",
"addressCountry": "US"
},
"deliveryTime": {
"@type": "ShippingDeliveryTime",
"handlingTime": {
"@type": "QuantitativeValue",
"minValue": 0,
"maxValue": 1,
"unitCode": "DAY"
},
"transitTime": {
"@type": "QuantitativeValue",
"minValue": 3,
"maxValue": 5,
"unitCode": "DAY"
}
}
},
{
"@type": "MerchantReturnPolicy",
"@id": "https://shop.com/#return-policy-30day",
"applicableCountry": "US",
"returnPolicyCategory": "https://schema.org/MerchantReturnFiniteReturnWindow",
"merchantReturnDays": 30,
"returnMethod": "https://schema.org/ReturnByMail",
"returnFees": "https://schema.org/FreeReturn"
}
]
}
Offer Freshness: The Core E-commerce SEO Problem
Google's product rich results are suppressed when price or availability in structured data doesn't match what's on the page or in the Merchant Center feed. The priceValidUntil field is not cosmetic — Google uses it to evaluate data staleness. An offer with a priceValidUntil in the past is treated as stale and may be demoted from Shopping results.
At scale, the only sustainable architecture is real-time structured data generation from your inventory system — not static markup baked into templates. The two viable patterns:
| Pattern | Freshness | Infrastructure Cost | Complexity | Recommended For |
|---|---|---|---|---|
| SSR from inventory DB | Real-time | Medium (origin compute) | Low | Sub-100K SKU catalogs |
| Edge KV cache with TTL invalidation | Near real-time (~1 min) | Low (edge compute) | Medium | 100K–1M SKU catalogs |
| Pregenerated JSON-LD + CDN push on price change | Event-driven (seconds) | Low | High (event pipeline) | High-frequency price-change catalogs |
| Static baked JSON-LD | Stale (hours–days) | Very low | Very low | Never for pricing/availability |
Edge KV Pattern with Cloudflare Workers
// Cloudflare Worker with KV-backed product JSON-LD
// KV key: "jsonld:product:{sku}" → pre-serialized JSON string
// Updated by your inventory webhook → Cloudflare KV API
export default {
async fetch(request, env, ctx) {
const url = new URL(request.url);
// Extract SKU from URL pattern /products/{sku}/
const match = url.pathname.match(/^\/products\/([^/]+)\//);
if (!match) return fetch(request);
const sku = match[1];
const kvKey = <code>jsonld:product:${sku}</code>;
// Fetch pre-built JSON-LD from KV (cached at edge)
const cachedJsonLd = await env.PRODUCT_JSONLD.get(kvKey);
if (!cachedJsonLd) {
// Fall back to origin for uncached products
return fetch(request);
}
// Parse and update time-sensitive fields
const jsonLd = JSON.parse(cachedJsonLd);
const offerNode = jsonLd["@graph"]?.find(n => n["@type"] === "Offer");
if (offerNode) {
// Always set priceValidUntil to 30 days from now at edge
const validUntil = new Date(Date.now() + 30 * 86400 * 1000)
.toISOString().split("T")[0];
offerNode.priceValidUntil = validUntil;
}
const origin = await fetch(request);
return new HTMLRewriter()
.on("head", {
element(el) {
el.append(
<code><script type="application/ld+json">${JSON.stringify(jsonLd)}</script></code>,
{ html: true }
);
},
})
.transform(origin);
},
};
Python: Inventory Webhook → Cloudflare KV Update
import httpx
import json
from datetime import datetime, timedelta
CLOUDFLARE_ACCOUNT_ID = "your_account_id"
CLOUDFLARE_API_TOKEN = "your_api_token"
KV_NAMESPACE_ID = "your_namespace_id"
def build_offer_jsonld(product: dict) -> dict:
"""Build complete product JSON-LD from inventory record."""
price_valid_until = (datetime.now() + timedelta(days=30)).strftime("%Y-%m-%d")
availability = (
"https://schema.org/InStock"
if product["qty"] > 0
else "https://schema.org/OutOfStock"
)
return {
"@context": "https://schema.org",
"@graph": [
{
"@type": "Product",
"@id": f"https://shop.com/products/{product['sku']}/#product",
"name": product["name"],
"sku": product["sku"],
"gtin14": product.get("gtin", ""),
"offers": {"@id": f"https://shop.com/products/{product['sku']}/#offer"},
},
{
"@type": "Offer",
"@id": f"https://shop.com/products/{product['sku']}/#offer",
"price": str(product["price"]),
"priceCurrency": "USD",
"priceValidUntil": price_valid_until,
"availability": availability,
"itemCondition": "https://schema.org/NewCondition",
"url": f"https://shop.com/products/{product['sku']}/",
"seller": {"@id": "https://shop.com/#organization"},
"shippingDetails": {"@id": "https://shop.com/#shipping-standard-us"},
"hasMerchantReturnPolicy": {"@id": "https://shop.com/#return-policy-30day"},
},
],
}
def push_to_cloudflare_kv(sku: str, jsonld: dict) -> bool:
"""Push serialized JSON-LD to Cloudflare KV."""
url = (
f"https://api.cloudflare.com/client/v4/accounts/{CLOUDFLARE_ACCOUNT_ID}"
f"/storage/kv/namespaces/{KV_NAMESPACE_ID}/values/jsonld:product:{sku}"
)
headers = {"Authorization": f"Bearer {CLOUDFLARE_API_TOKEN}"}
r = httpx.put(
url,
headers=headers,
content=json.dumps(jsonld),
timeout=5.0
)
return r.status_code == 200
# Called by your inventory webhook
def on_product_update(product: dict):
jsonld = build_offer_jsonld(product)
success = push_to_cloudflare_kv(product["sku"], jsonld)
if not success:
raise RuntimeError(f"KV push failed for SKU {product['sku']}")
Merchant Center + Structured Data Reconciliation
Google's Merchant Center ingests product data independently of organic structured data. When the two are inconsistent, Google can suppress rich results or flag the listing for manual review. The properties that must match exactly: price, priceCurrency, availability, and the product's canonical URL.
The reconciliation logic Google applies is documented in the Google Merchant Center Help and partially in the Merchant Center API documentation. Key rules: if GMC feed has in stock but structured data shows OutOfStock, the structured data is treated as stale. The GMC feed wins for Shopping results; structured data wins for organic rich results. Divergence causes both to underperform.
Python: Reconcile GMC Feed vs. Structured Data
import csv
import json
import asyncio
import aiohttp
import re
from typing import Optional
JSONLD_PATTERN = re.compile(
r'<script[^>]+type=["\']application/ld\+json["\'][^>]*>(.*?)',
re.DOTALL | re.IGNORECASE
)
def parse_gmc_feed(feed_path: str) -> dict[str, dict]:
"""Parse GMC TSV feed into SKU-keyed dict."""
products = {}
with open(feed_path, newline="", encoding="utf-8") as f:
reader = csv.DictReader(f, delimiter="\t")
for row in reader:
products[row["id"]] = {
"price": row["price"].replace(" USD", "").strip(),
"availability": row["availability"],
"link": row["link"],
}
return products
async def fetch_page_offer(session: aiohttp.ClientSession, url: str) -> Optional[dict]:
"""Extract offer data from structured data on a product page."""
try:
async with session.get(url, timeout=aiohttp.ClientTimeout(total=10)) as resp:
html = await resp.text()
except Exception:
return None
for match in JSONLD_PATTERN.finditer(html):
try:
data = json.loads(match.group(1))
except json.JSONDecodeError:
continue
graph = data.get("@graph", [data])
for node in graph:
if node.get("@type") == "Offer":
return {
"price": str(node.get("price", "")),
"availability": node.get("availability", "").split("/")[-1],
"currency": node.get("priceCurrency", ""),
}
return None
async def reconcile_feed_vs_markup(gmc_products: dict, sample_size: int = 100) -> list[dict]:
"""Cross-check GMC feed against live structured data for a sample."""
skus = list(gmc_products.keys())[:sample_size]
discrepancies = []
async with aiohttp.ClientSession() as session:
tasks = {
sku: fetch_page_offer(session, gmc_products[sku]["link"])
for sku in skus
}
results = await asyncio.gather(*tasks.values())
for sku, page_offer in zip(skus, results):
gmc = gmc_products[sku]
if page_offer is None:
discrepancies.append({"sku": sku, "issue": "structured_data_missing"})
continue
if gmc["price"] != page_offer["price"]:
discrepancies.append({
"sku": sku,
"issue": "price_mismatch",
"gmc_price": gmc["price"],
"page_price": page_offer["price"],
})
gmc_avail = "InStock" if gmc["availability"] == "in stock" else "OutOfStock"
if gmc_avail != page_offer["availability"]:
discrepancies.append({
"sku": sku,
"issue": "availability_mismatch",
"gmc": gmc_avail,
"page": page_offer["availability"],
})
return discrepancies
</script[^>
Product Variant Handling at Scale
Google's guidance on product variants has evolved significantly. The current recommendation uses ProductGroup with hasVariant and variesBy. Each variant should be individually addressable with its own canonical URL and its own Offer — aggregate availability across variants is not acceptable for rich result eligibility.
At scale with thousands of variant combinations (color × size × material), generating individual JSON-LD for each variant page requires a systematic template approach, not manual authoring. The critical property for variant grouping is the productGroupID — it must match across all variants and across your GMC feed's item_group_id.
def generate_variant_jsonld(
base_product: dict,
variant: dict,
all_variants: list[dict]
) -> dict:
"""
Generate complete JSON-LD for a single product variant page.
base_product: shared product data (name, brand, description)
variant: this variant's specific data (color, size, sku, price, qty)
all_variants: all variants for ProductGroup construction
"""
base_url = f"https://shop.com/products/{base_product['slug']}"
variant_url = f"{base_url}-{variant['color'].lower()}-{variant['size'].lower()}"
variant_name = f"{base_product['name']} — {variant['color']}, Size {variant['size']}"
availability = "https://schema.org/InStock" if variant["qty"] > 0 else "https://schema.org/OutOfStock"
graph = [
{
"@type": "ProductGroup",
"@id": f"{base_url}/#productgroup",
"name": base_product["name"],
"productGroupID": base_product["slug"],
"description": base_product["description"],
"brand": {
"@type": "Brand",
"name": base_product["brand"]
},
"variesBy": ["color", "size"],
"hasVariant": [
{"@id": f"{base_url}-{v['color'].lower()}-{v['size'].lower()}/#product"}
for v in all_variants
],
},
{
"@type": "Product",
"@id": f"{variant_url}/#product",
"name": variant_name,
"sku": variant["sku"],
"gtin14": variant.get("gtin", ""),
"color": variant["color"],
"size": variant["size"],
"isVariantOf": {"@id": f"{base_url}/#productgroup"},
"offers": {
"@type": "Offer",
"@id": f"{variant_url}/#offer",
"price": str(variant["price"]),
"priceCurrency": "USD",
"availability": availability,
"itemCondition": "https://schema.org/NewCondition",
"url": f"{variant_url}/",
"seller": {"@id": "https://shop.com/#organization"},
"shippingDetails": {"@id": "https://shop.com/#shipping-standard-us"},
"hasMerchantReturnPolicy": {"@id": "https://shop.com/#return-policy-30day"},
},
}
]
return {"@context": "https://schema.org", "@graph": graph}
Review Aggregation Architecture
AggregateRating data must reflect actual reviews on the page or accessible from the page. Google's policy (updated 2023) explicitly prohibits aggregating ratings across product variants under a single URL unless those specific variant reviews are present at that URL. The common mistake: showing aggregate ratings for a product group on a variant page that has no individual reviews.
-- BigQuery: Calculate per-variant AggregateRating from review database
-- Use this to generate JSON-LD data programmatically
WITH variant_ratings AS (
SELECT
product_sku,
COUNT(*) AS review_count,
ROUND(AVG(rating), 1) AS avg_rating,
MIN(rating) AS min_rating,
MAX(rating) AS max_rating,
COUNTIF(rating >= 4) AS positive_count
FROM <code>your-project.ecommerce.product_reviews</code>
WHERE
status = 'approved'
AND created_at >= DATE_SUB(CURRENT_DATE(), INTERVAL 365 DAY)
GROUP BY product_sku
HAVING review_count >= 3 -- Minimum threshold for AggregateRating eligibility
)
SELECT
vr.product_sku,
p.product_name,
p.page_url,
vr.review_count,
vr.avg_rating AS rating_value,
vr.min_rating AS worst_rating,
vr.max_rating AS best_rating,
-- JSON-LD ready output
TO_JSON_STRING(STRUCT(
'AggregateRating' AS type,
CAST(vr.avg_rating AS STRING) AS ratingValue,
vr.review_count AS reviewCount,
CAST(vr.min_rating AS STRING) AS worstRating,
CAST(vr.max_rating AS STRING) AS bestRating
)) AS aggregate_rating_json
FROM variant_ratings vr
JOIN <code>your-project.ecommerce.products</code> p USING (product_sku)
ORDER BY vr.review_count DESC
The Full Pipeline: Generation, Delivery, Validation
At 500K+ SKU scale, the JSON-LD pipeline is a distributed system. Here's the architecture that works in production:
- Inventory Event Stream → Kafka topic on price/availability changes
- JSON-LD Generator Service → consumes Kafka, builds JSON-LD per SKU, pushes to Redis + Cloudflare KV
- Edge Delivery → Cloudflare Worker reads KV, injects JSON-LD into HTML response
- Validation Job → nightly BigQuery job samples 1% of catalog, extracts JSON-LD, validates against schema, writes failures to monitoring table
- GMC Reconciliation → daily comparison of GMC feed prices vs. JSON-LD prices, alerts on divergence > 0.5%
BigQuery: Structured Data Quality Monitoring Table
-- Create monitoring table
CREATE TABLE IF NOT EXISTS <code>your-project.seo.structured_data_audit</code>
(
audit_date DATE,
page_url STRING,
has_product_schema BOOL,
has_offer_schema BOOL,
has_aggregate_rating BOOL,
price FLOAT64,
availability STRING,
price_valid_until DATE,
is_price_stale BOOL,
gmc_price FLOAT64,
price_discrepancy FLOAT64,
issues ARRAY<string>
)
PARTITION BY audit_date
OPTIONS (require_partition_filter = false);
-- Daily insert from validation pipeline output
INSERT INTO <code>your-project.seo.structured_data_audit</code>
SELECT
CURRENT_DATE() AS audit_date,
page_url,
has_product_schema,
has_offer_schema,
has_aggregate_rating,
CAST(price AS FLOAT64) AS price,
availability,
PARSE_DATE('%Y-%m-%d', price_valid_until) AS price_valid_until,
PARSE_DATE('%Y-%m-%d', price_valid_until) < CURRENT_DATE() AS is_price_stale,
gmc.price AS gmc_price,
ABS(CAST(price AS FLOAT64) - gmc.price) AS price_discrepancy,
ARRAY_CONCAT(
IF(NOT has_product_schema, ['missing_product_schema'], []),
IF(NOT has_offer_schema, ['missing_offer_schema'], []),
IF(PARSE_DATE('%Y-%m-%d', price_valid_until) < CURRENT_DATE(), ['stale_price_valid_until'], []),
IF(ABS(CAST(price AS FLOAT64) - gmc.price) > 0.01, ['price_gmc_mismatch'], [])
) AS issues
FROM <code>your-project.seo.structured_data_raw</code> raw
LEFT JOIN <code>your-project.ecommerce.gmc_feed</code> gmc USING (sku)
WHERE raw.audit_date = CURRENT_DATE()
</string>
FAQ
Q: Should every product variant have its own page and its own JSON-LD?
Yes, for maximum rich result eligibility. Google's product rich results require that the Offer data match the specific URL. A variant page for "Widget Pro — Blue, Size M" must have an Offer for that exact SKU at that URL. Pointing all variants to a parent product page with a single generic Offer suppresses per-variant eligibility, though the ProductGroup entity still provides crawlability value.
Q: How do I handle discontinued products — should I remove JSON-LD?
Set availability to https://schema.org/Discontinued on the Offer, remove priceValidUntil, and keep the Product entity intact. Don't 404 immediately — Google may still rank the page for brand/product name queries. After 6 months with no organic traffic, 301 to the closest alternative product. Removing JSON-LD before removing the page creates a structured-data/content mismatch that can trigger manual review flags.
Q: Can I use ItemList on a category page instead of individual Product markup?
ItemList on category pages is valid and supported but doesn't provide individual product rich results for the listed items. It can help with sitelinks and carousel eligibility. The better pattern is ItemList on the category page plus full Product markup on each individual product page. Don't duplicate Product markup from detail pages onto the category page — that creates canonical confusion.
Q: What's the correct way to express sale pricing in Schema.org?
Use the priceSpecification property with a UnitPriceSpecification that includes validFrom and validThrough. The top-level price property should always reflect the current payable price (the sale price during a sale). Many implementations mistakenly show the regular price in price and the sale in a subproperty — this causes Google to display the wrong price in Shopping results.
Q: How does Google handle dynamic pricing (time-of-day, user-segment pricing)?
Google's crawler sees the price at crawl time, which may not match what users see. For auction-style or highly dynamic pricing, use the canonical "from" price and mark it as MinimumAdvertisedPrice in GMC. Structured data should never show a price that a significant portion of users cannot achieve — this triggers the "price mismatch" policy violation in Merchant Center.
Q: Is there a way to suppress rich results for specific products?
Yes: set robots meta to nosnippet (suppresses all snippets including rich results), or remove the JSON-LD block for that specific product. A more targeted approach: use Google Search Console's URL Inspection tool to check current rich result eligibility, then simply remove the Offer node from the JSON-LD while keeping the Product entity — this maintains index signals without triggering rich result display.
Key Takeaways
- The
priceValidUntilfield is a staleness signal — always set it 30 days forward and update it dynamically. Static baked dates from your last deploy are a rich result suppression time bomb. - GMC feed and on-page structured data must be reconciled; divergence of price or availability by even $0.01 can trigger Merchant Center policy flags.
- Edge injection via Cloudflare Workers KV is the only architecture that delivers real-time pricing JSON-LD without SSR compute costs at million-SKU scale.
- ProductGroup with
hasVariantandvariesByis now required for variant product eligibility — not optional if you want per-variant rich results in Shopping. - AggregateRating must be computed per-variant, not at the product-group level, and must reflect only reviews accessible at that specific variant URL.
Conclusion
E-commerce JSON-LD at scale is a distributed systems problem with SEO implications, not an SEO problem with some engineering involved. The sites winning in Shopping results have real-time inventory pipelines feeding their structured data, automated reconciliation catching divergence before Google does, and edge delivery ensuring zero-latency freshness. The sites losing have static markup that was accurate on deploy day and stale by week two. The infrastructure investment is the moat.
Next: Vector Embeddings for SEO — Semantic Similarity Analysis at Scale Google Merchant Center Product Data Specification