Why Most Migrations Still Fail in 2026
I've run 31 platform migrations since 2019. Shopify to Shopify Plus, Shopify Plus to Hydrogen, WordPress monoliths to Next.js, Magento 2 to custom composable stacks. The median client comes to me after a botched migration with a chart that looks like a cliff edge. They want me to explain what happened. Most of the time I can explain it in a single sentence: the team treated the migration as a deployment problem rather than a search-equity problem.
Deploying a new platform correctly and migrating search equity correctly are two completely different disciplines that happen to share a launch date.
The migration I'm documenting here is specific. In late 2025 I managed the move of a 4.1-million-URL Shopify Plus catalog to Hydrogen (React Server Components, Vite build, Oxygen CDN deployment) for a UK-based health and wellness retailer. Simultaneously, I consulted on a 3.8-million-URL WordPress-to-Next.js migration for a US media property. Between those two engagements I developed what I now call the FREEZE framework, and I'm going to walk you through all 47 steps of it.
The Hydrogen migration landed with a measured organic traffic delta of +0.3% at 90 days post-cutover. Not zero loss, technically, but within normal seasonal variance. The Next.js migration hit -2.1% at 30 days before recovering to +4.7% at 90. Neither client lost meaningful ground. Here's how.
Two Things the Migration Industry Gets Wrong
Contrarian Take 1: Pre-migration crawls are overvalued; pre-migration rendering audits are undervalued
Every migration checklist in existence tells you to crawl the origin site before you move. Run Screaming Frog. Export URLs. Count status codes. Fine. That's table stakes. But in 2026, when the destination platform is Hydrogen or Next.js or any RSC-based framework, a URL inventory is nearly useless without a parallel rendering audit.
Here's the specific failure mode I see constantly: a product detail page on the origin Shopify Plus store renders price, review count, and structured data (Product schema) fully in HTML. The migration team builds the Hydrogen equivalent. The RSC server component renders the shell. A client component hydrates the price and review count. The structured data is injected via a client-side useEffect hook because the developer found a tutorial that did it that way. Googlebot renders the page, sees the schema, but the timing of the client-side injection means the JSON-LD is sometimes present and sometimes not in the rendered DOM snapshot. The rich result fluctuates. Impressions drop 11% before anyone notices.
I mandate a rendering audit before any line of redirect mapping begins. Every template type, every schema type, every dynamic injection point must be verified against a headless Chrome rendering environment before the first redirect fires.
Contrarian Take 2: The "migration window" mindset creates the disasters it's designed to prevent
The industry consensus is to pick a low-traffic day, usually Tuesday or Wednesday, and do the migration in a compressed window. The logic is that you minimize the period of instability. I think this is backwards for large sites.
For catalogs above 500,000 URLs, a compressed migration window means Googlebot encounters millions of 301 redirects in a 24–48 hour burst. The crawl rate spikes, then throttles. The redirect graph takes weeks to process fully. Canonicalization signals conflict during that processing window. You get ranking volatility that a staged rollout would have avoided entirely.
My approach since mid-2025 is a staged URL namespace migration: move product categories in order of crawl priority, verified via Google Search Console crawl stats, over a 3-to-4-week period. The redirects process incrementally. Index updates flow continuously. The volatility curve flattens dramatically. This is slower, more operationally complex, and nearly impossible to sell to a project manager. Do it anyway.
The Migration That Lost 14% (What I Learned)
In Q1 2025 I managed a 920,000-URL Magento 2 to Shopify Plus migration for a European consumer electronics retailer. We hit every checklist item. Redirect map was complete. Sitemaps submitted. GSC property verified. Change of address filed. Canonical tags updated.
We lost 14.3% of non-branded organic traffic at the 30-day mark. At 60 days we were still at -9.1%. We recovered to -3.2% by day 90 and never fully closed the gap in the engagement I was retained for.
The post-mortem identified three compounding failures. First: the site had 847 URL parameters that generated unique canonical URLs in Magento but were handled as faceted navigation in Shopify's URL structure. Those 847 parameter combinations represented 23% of the top-500 landing pages in GSC. We redirected the canonical URLs correctly but missed the parameter variants, which had accumulated link equity independently over six years. Second: the international hreflang annotations referenced the Magento URL structure by absolute URL. We updated them in the sitemap but missed 14 hardcoded references in the theme templates. Third: and this is the one that still bothers me, we ran the soft-launch protocol for 72 hours rather than the 14 days I now require as a minimum. We caught rendering issues in the soft-launch but didn't have enough GSC data to see the hreflang problem before we went hard.
Fourteen percent. Three fixable things. The lesson is not that migrations are hard. The lesson is that the validation window must be longer than feels comfortable, and URL parameter equity deserves its own audit category.
The FREEZE Framework
FREEZE is the personal framework I developed to structure large platform migrations. The acronym maps to the five phases:
- Freeze the origin (inventory, baseline, content lock)
- Redirect architecture (map, test, stage)
- Environment validation (rendering, schema, Core Web Vitals)
- Ease in with soft-launch (staged rollout, monitoring)
- Zero-gap cutover (hard switch, GSC handoff)
- Extended monitoring (30/60/90-day windows)
The rest of this article walks through all 47 steps organized under those five phases, plus the extended monitoring window. Code samples reflect the actual tooling used in the Hydrogen and Next.js migrations.
Phase 1: Inventory and Baseline (Steps 1–11)
Step 1: Export the full crawlable URL set
Do not rely on the CMS sitemap alone. For the Hydrogen migration, the Shopify Plus sitemap covered 2.1 million of 4.1 million URLs. The remaining 2 million were generated by parameterized collection filters, legacy blog redirects, and product variant URLs that Shopify excludes from its default sitemap generation. You need both.
# Crawl origin site and export full URL list with status codes
# Using Scrapy with custom settings for large catalogs
scrapy crawl full_site \
--set CONCURRENT_REQUESTS=16 \
--set DOWNLOAD_DELAY=0.25 \
--set DEPTH_LIMIT=8 \
--set HTTPCACHE_ENABLED=True \
-o origin_urls_$(date +%Y%m%d).jsonl
# Merge with sitemap URLs
python3 merge_url_sources.py \
--crawl origin_urls_20251104.jsonl \
--sitemap sitemap_index_export.xml \
--gsc gsc_top_pages_16months.csv \
--output full_url_inventory_deduplicated.csv
Step 2: Pull 16 months of GSC data
12 months captures seasonality. 16 months gives you a comparison buffer that survives algorithm update volatility. Pull clicks, impressions, average position, and CTR for every URL in the inventory. This becomes your baseline measurement set.
from googleapiclient.discovery import build
from google.oauth2 import service_account
import pandas as pd
from datetime import datetime, timedelta
SCOPES = ['https://www.googleapis.com/auth/webmasters.readonly']
SERVICE_ACCOUNT_FILE = 'gsc-service-account.json'
credentials = service_account.Credentials.from_service_account_file(
SERVICE_ACCOUNT_FILE, scopes=SCOPES)
service = build('searchconsole', 'v1', credentials=credentials)
def fetch_gsc_data(site_url, start_date, end_date, row_limit=25000):
request = {
'startDate': start_date,
'endDate': end_date,
'dimensions': ['page', 'query', 'device', 'country'],
'rowLimit': row_limit,
'dataState': 'final'
}
response = service.searchanalytics().query(
siteUrl=site_url, body=request).execute()
return pd.DataFrame(response.get('rows', []))
# Fetch 16-month window in 3-month chunks to avoid row limits
date_ranges = [
('2024-07-01', '2024-09-30'),
('2024-10-01', '2024-12-31'),
('2025-01-01', '2025-03-31'),
('2025-04-01', '2025-06-30'),
('2025-07-01', '2025-09-30'),
]
frames = [fetch_gsc_data('sc-domain:example.com', s, e) for s, e in date_ranges]
baseline = pd.concat(frames).drop_duplicates()
baseline.to_parquet('gsc_baseline_16mo.parquet', index=False)
Step 3: Identify the top-2000 landing pages by traffic
Every URL gets migrated. But the top-2000 by clicks over 16 months get individual QA treatment. Not sampling. Individual. This list is your migration's source of truth.
Step 4: Tag URL types by template
For the Hydrogen migration, URL types were: homepage, collection (top-level), collection (filtered), product detail, product variant, blog index, blog post, static page, account/authentication (noindex), search results (noindex). Each template type needs its own validation checklist because the rendering stack is different.
Step 5: Run a rendering audit on template types
# Render each template type using headless Chrome
# Compare rendered DOM against raw HTML for content parity
from playwright.async_api import async_playwright
import asyncio
import json
async def render_url(url):
async with async_playwright() as p:
browser = await p.chromium.launch()
page = await browser.new_page(
user_agent='Mozilla/5.0 (compatible; Googlebot/2.1; +http://www.google.com/bot.html)'
)
await page.goto(url, wait_until='networkidle')
content = await page.content()
# Extract JSON-LD blocks
ld_blocks = await page.evaluate('''() => {
const scripts = document.querySelectorAll('script[type="application/ld+json"]');
return Array.from(scripts).map(s => s.textContent);
}''')
await browser.close()
return {
'url': url,
'rendered_html_length': len(content),
'json_ld_count': len(ld_blocks),
'json_ld_types': [json.loads(b).get('@type') for b in ld_blocks if b]
}
async def audit_templates(url_list):
tasks = [render_url(url) for url in url_list]
return await asyncio.gather(*tasks)
sample_urls = [
'https://origin.example.com/',
'https://origin.example.com/collections/supplements',
'https://origin.example.com/products/vitamin-d3-5000iu',
'https://origin.example.com/blogs/news/post-slug',
]
results = asyncio.run(audit_templates(sample_urls))
print(json.dumps(results, indent=2))
Step 6: Capture Core Web Vitals baseline per template
Use CrUX API, not lab data. The origin platform's field data is your benchmark. If the destination platform improves CWV, great. If it regresses, you need to know before launch.
Step 7: Audit all structured data types on origin
Product, Organization, BreadcrumbList, FAQPage, Article, Review, ItemList. Every schema type present on origin must be present and valid on destination. Use the Rich Results Test API for the top-2000 URLs. Automate it.
Step 8: Export internal link graph
You need the internal link graph to validate that high-equity pages on destination maintain the same internal linking depth they had on origin. A page that was 2 clicks from the homepage on Shopify Plus cannot become 5 clicks from the homepage on Hydrogen without planning a compensating internal link injection.
Step 9: Document all canonical configurations
Paginated series, parameter handling, hreflang cross-references, self-referential canonicals. Every variant. Export to a structured CSV you can diff against the destination configuration post-launch.
Step 10: Freeze the origin content
No new content. No URL changes. No structural edits. The origin site is a snapshot from this point. This sounds obvious. It gets violated constantly. Build a documented content freeze policy and get sign-off from content and merchandising teams before you proceed.
Step 11: Establish pre-migration traffic baseline report
Create a Looker Studio dashboard (or your reporting tool of choice) that will serve as the migration health scorecard. Metrics: organic sessions by template type, impressions by URL cluster, average position for the top-200 keywords, crawl coverage in GSC. You'll compare this against post-migration data at 30, 60, and 90 days.
Phase 2: Redirect Architecture (Steps 12–21)
Step 12: Build the redirect map in three layers
Layer 1: exact URL-to-URL matches (most URLs). Layer 2: pattern-based rules for URL families that share a structure transformation. Layer 3: exception rules for URLs that break the pattern. Most migrations fail at Layer 3 because the exceptions are undocumented until someone notices a 404.
# redirect_map_structure.py
# Three-layer redirect architecture
import re
import csv
# Layer 1: Exact matches (loaded from CMS export or manual mapping)
exact_redirects = {
'/old-product-slug': '/products/new-product-slug',
'/collections/old-cat': '/collections/new-cat',
# ... 4.1M entries loaded from CSV
}
# Layer 2: Pattern-based rules
pattern_redirects = [
# Shopify Plus to Hydrogen: collection URL structure unchanged
(r'^/collections/([a-z0-9\-]+)$', r'/collections/\1'),
# Product URLs: old format /p/SKU-123 -> /products/sku-123
(r'^/p/([A-Z0-9\-]+)$', lambda m: f'/products/{m.group(1).lower()}'),
# Blog posts: /blog/YYYY/MM/slug -> /blogs/news/slug
(r'^/blog/\d{4}/\d{2}/(.+)$', r'/blogs/news/\1'),
# Parameterized collection filters -> canonical collection
(r'^/collections/([a-z0-9\-]+)\?.*$', r'/collections/\1'),
]
# Layer 3: Exceptions override patterns
exception_redirects = {
'/collections/sale?sort_by=price-ascending': '/collections/sale',
'/collections/new-in': '/collections/new-arrivals',
}
def resolve_redirect(path):
# Exception check first
if path in exception_redirects:
return exception_redirects[path], 'exception'
# Exact match
if path in exact_redirects:
return exact_redirects[path], 'exact'
# Pattern match
for pattern, replacement in pattern_redirects:
match = re.match(pattern, path)
if match:
if callable(replacement):
return replacement(match), 'pattern'
return re.sub(pattern, replacement, path), 'pattern'
return None, 'unmapped'
Step 13: Identify redirect chains from origin
The origin site has chains. Every site that's been running for more than three years has chains. Find them before you build the destination redirect map or you'll bake the chains into the new platform.
# Detect redirect chains on origin
# Using httpx for async requests
import httpx
import asyncio
async def follow_chain(client, url, max_hops=10):
chain = [url]
current = url
for _ in range(max_hops):
try:
r = await client.get(current, follow_redirects=False)
if r.status_code in (301, 302, 307, 308):
location = r.headers.get('location', '')
chain.append({'status': r.status_code, 'location': location})
current = location
else:
chain.append({'status': r.status_code, 'final': True})
break
except Exception as e:
chain.append({'error': str(e)})
break
return chain
async def audit_chains(urls):
limits = httpx.Limits(max_connections=50)
async with httpx.AsyncClient(limits=limits, timeout=10) as client:
tasks = [follow_chain(client, u) for u in urls]
return await asyncio.gather(*tasks)
# Flag chains longer than 1 hop for flattening
results = asyncio.run(audit_chains(all_origin_urls))
chains = [(u, r) for u, r in zip(all_origin_urls, results) if len(r) > 2]
print(f'Found {len(chains)} chained redirects to flatten')
Step 14: Map parameter equity separately
This is the lesson from the -14% migration. Pull every URL variant with GSC clicks from the 16-month export. Any parameterized URL that received more than 50 clicks gets its own explicit redirect entry, not a catch-all parameter strip.
Step 15: Validate redirect map against GSC top pages
Every URL in your top-2000 GSC landing pages list must resolve to a 200 on the destination, via a maximum of one 301. Script this validation against the staging environment.
Step 16: Implement redirects at the edge
For the Hydrogen migration, redirects were implemented in Oxygen's edge routing layer. For the Next.js migration, in Next.js middleware and Vercel's edge config. Not in the application layer. Not in htaccess. At the edge, before any application logic runs. Speed of redirect resolution affects crawl budget consumption during the migration window.
// next.config.js - redirect configuration for Next.js migration
// Large redirect sets should use Edge Middleware, not next.config.js redirects
// next.config.js is for small sets only
/** @type {import('next').NextConfig} */
const nextConfig = {
async redirects() {
// Only structural/permanent pattern redirects here
return [
{
source: '/blog/:year(\\d{4})/:month(\\d{2})/:slug',
destination: '/articles/:slug',
permanent: true,
},
{
source: '/category/:slug*',
destination: '/topics/:slug*',
permanent: true,
},
{
source: '/author/:name',
destination: '/writers/:name',
permanent: true,
},
]
},
}
module.exports = nextConfig
// middleware.ts - Edge Middleware for bulk URL redirect lookup
// Vercel Edge Config stores up to 512KB of redirect map JSON
import { NextResponse } from 'next/server'
import type { NextRequest } from 'next/server'
import { get } from '@vercel/edge-config'
export async function middleware(request: NextRequest) {
const pathname = request.nextUrl.pathname
// Fast path: skip known non-redirect paths
if (
pathname.startsWith('/_next') ||
pathname.startsWith('/api') ||
pathname.includes('.')
) {
return NextResponse.next()
}
try {
const redirectMap = await get>('redirects')
if (redirectMap && redirectMap[pathname]) {
const destination = redirectMap[pathname]
return NextResponse.redirect(
new URL(destination, request.url),
{ status: 301 }
)
}
} catch (e) {
// Edge Config unavailable - fail open
console.error('Edge Config error:', e)
}
return NextResponse.next()
}
export const config = {
matcher: ['/((?!_next/static|_next/image|favicon.ico).*)'],
}
Step 17: Test redirect map in staging with 5,000-URL sample
Automated. Random sample weighted by GSC traffic. Every URL must return the expected destination with status 301 and correct Location header. Failures block launch.
Step 18: Document the redirect implementation for the dev team
Not a Notion doc. A versioned spec in the repo. The implementation is a contract between SEO and engineering and it needs to be treated like one.
Step 19: Audit destination for self-referential canonicals
Every page on the destination must have a canonical that points to itself, using the destination domain, before any traffic is routed. Canonical pointing to the origin domain after cutover is a common failure mode in headless builds where the canonical is set by a CMS field that still contains the old domain.
Step 20: Verify hreflang on destination (if applicable)
Absolute URLs. Correct domain. Correct locale codes. X-default present. Cross-references valid. The 14% loss migration failed here. Don't fail here.
Step 21: Freeze the redirect map
No changes to the redirect map after this point without documented sign-off. This is a release artifact.
Phase 3: Sitemap Freeze and Thaw Protocol (Steps 22–28)
Step 22: Understand the sitemap freeze concept
Sitemap freezing means: you stop generating dynamic sitemaps on the origin platform and replace them with static XML files that represent the final, canonical URL set you intend to migrate. This prevents the origin sitemap from updating with new URLs during the migration window, which would create targets Googlebot might try to crawl after you've cut over.
Step 23: Generate static XML sitemaps for origin
#!/usr/bin/env python3
# sitemap_freeze.py
# Generates static sitemap XML from URL inventory
import xml.etree.ElementTree as ET
from datetime import datetime
import math
import gzip
import shutil
SITEMAP_MAX_URLS = 48000 # Below 50k limit for safety margin
BASE_URL = 'https://origin.example.com'
def generate_sitemap_chunk(urls, filename, priority_map=None):
urlset = ET.Element('urlset')
urlset.set('xmlns', 'http://www.sitemaps.org/schemas/sitemap/0.9')
for url in urls:
url_el = ET.SubElement(urlset, 'url')
loc = ET.SubElement(url_el, 'loc')
loc.text = url['loc']
lastmod = ET.SubElement(url_el, 'lastmod')
lastmod.text = url.get('lastmod', datetime.now().strftime('%Y-%m-%d'))
priority_val = priority_map.get(url['loc'], '0.5') if priority_map else '0.5'
priority = ET.SubElement(url_el, 'priority')
priority.text = priority_val
tree = ET.ElementTree(urlset)
ET.indent(tree, space=' ')
with open(filename, 'wb') as f:
tree.write(f, encoding='utf-8', xml_declaration=True)
# Gzip for serving
with open(filename, 'rb') as f_in:
with gzip.open(f'{filename}.gz', 'wb') as f_out:
shutil.copyfileobj(f_in, f_out)
def generate_sitemap_index(sitemap_files, index_filename):
sitemapindex = ET.Element('sitemapindex')
sitemapindex.set('xmlns', 'http://www.sitemaps.org/schemas/sitemap/0.9')
for sf in sitemap_files:
sitemap_el = ET.SubElement(sitemapindex, 'sitemap')
loc = ET.SubElement(sitemap_el, 'loc')
loc.text = f'{BASE_URL}/sitemaps/{sf}'
lastmod = ET.SubElement(sitemap_el, 'lastmod')
lastmod.text = datetime.now().strftime('%Y-%m-%d')
tree = ET.ElementTree(sitemapindex)
ET.indent(tree, space=' ')
tree.write(index_filename, encoding='unicode', xml_declaration=True)
# Load full URL inventory
import csv
all_urls = []
with open('full_url_inventory_deduplicated.csv') as f:
reader = csv.DictReader(f)
all_urls = [{'loc': row['url'], 'lastmod': row.get('lastmod', '')}
for row in reader if row['status'] == '200']
print(f'Total URLs to freeze: {len(all_urls):,}')
# Chunk into sitemap files
chunks = [all_urls[i:i+SITEMAP_MAX_URLS]
for i in range(0, len(all_urls), SITEMAP_MAX_URLS)]
sitemap_files = []
for idx, chunk in enumerate(chunks):
filename = f'sitemap-{idx+1:03d}.xml'
generate_sitemap_chunk(chunk, f'frozen_sitemaps/{filename}')
sitemap_files.append(filename)
print(f'Generated {filename}: {len(chunk):,} URLs')
generate_sitemap_index(sitemap_files, 'frozen_sitemaps/sitemap_index_frozen.xml')
print(f'Sitemap index generated with {len(sitemap_files)} sitemaps')
Step 24: Deploy frozen sitemaps to origin
Replace the dynamic sitemap generation with static file serving. On Shopify Plus this means disabling the default sitemap.xml generation and routing /sitemap.xml to a hosted static file. On WordPress this means disabling Yoast or RankMath sitemap generation and serving static files via a redirect rule.
Step 25: Verify frozen sitemaps in GSC
Submit the frozen sitemap index to GSC. Confirm the URLs discovered match the frozen inventory count. Give Googlebot 48 hours to acknowledge the submission before proceeding.
Step 26: Prepare destination sitemaps
The destination sitemaps should be built and ready but NOT submitted to GSC yet. They should be accessible at the correct path on the destination domain, verified as valid XML, and contain only URLs that return 200 on the destination staging environment.
Step 27: Build sitemap thaw automation
#!/bin/bash
# sitemap_thaw.sh
# Run post-cutover to activate destination sitemaps in GSC
SITE_URL="https://www.example.com"
GSC_TOKEN=$(cat gsc_oauth_token.json | python3 -c "import sys,json; print(json.load(sys.stdin)['access_token'])")
# Submit new sitemap index
curl -s -X PUT \
"https://www.googleapis.com/webmasters/v3/sites/$(python3 -c "import urllib.parse; print(urllib.parse.quote('$SITE_URL', safe=''))")/sitemaps/$(python3 -c "import urllib.parse; print(urllib.parse.quote('$SITE_URL/sitemap_index.xml', safe=''))")" \
-H "Authorization: Bearer $GSC_TOKEN" \
-H "Content-Type: application/json"
echo "Sitemap index submitted: $SITE_URL/sitemap_index.xml"
# Verify submission
curl -s \
"https://www.googleapis.com/webmasters/v3/sites/$(python3 -c "import urllib.parse; print(urllib.parse.quote('$SITE_URL', safe=''))")/sitemaps" \
-H "Authorization: Bearer $GSC_TOKEN" | python3 -m json.tool
Step 28: Document the thaw trigger conditions
The sitemap thaw (submitting destination sitemaps, withdrawing origin sitemaps) happens at a specific trigger: hard cutover confirmed, 200 responses verified across template types, redirects validated in production. Not before. Write this down. Get sign-off.
Phase 4: Soft-Launch and Crawl Validation (Steps 29–37)
Step 29: Define soft-launch architecture
Soft-launch means: destination platform is live on a separate subdomain or behind an IP allowlist, receiving no organic traffic, with a robots.txt disallowing all crawlers. You QA everything here before redirects fire.
For the Hydrogen migration, the soft-launch environment was staging.example.com with a restrictive robots.txt and Cloudflare access authentication preventing public access. Organic traffic continued flowing to origin throughout the soft-launch period.
Step 30: Validate all top-2000 URLs on soft-launch
# validate_soft_launch.py
# Automated validation of top-2000 URLs on staging environment
import asyncio
import httpx
import json
from urllib.parse import urlparse, urljoin
STAGING_BASE = 'https://staging.example.com'
ORIGIN_BASE = 'https://www.example.com'
AUTH_HEADERS = {'CF-Access-Client-Id': 'xxx', 'CF-Access-Client-Secret': 'yyy'}
async def validate_url(client, origin_url):
staging_path = urlparse(origin_url).path
staging_url = urljoin(STAGING_BASE, staging_path)
result = {
'origin_url': origin_url,
'staging_url': staging_url,
'checks': {}
}
try:
r = await client.get(staging_url, headers=AUTH_HEADERS, follow_redirects=True)
result['checks']['status_200'] = r.status_code == 200
result['checks']['canonical_self'] = (
f'rel="canonical"' in r.text and
staging_url.replace(STAGING_BASE, ORIGIN_BASE) in r.text
)
result['checks']['has_json_ld'] = 'application/ld+json' in r.text
result['checks']['has_title'] = '' in r.text
result['checks']['no_noindex'] = (
'noindex' not in r.headers.get('x-robots-tag', '') and
'noindex' not in r.text[:2000] # Check meta robots in head
)
result['pass'] = all(result['checks'].values())
except Exception as e:
result['error'] = str(e)
result['pass'] = False
return result
async def run_validation(urls):
limits = httpx.Limits(max_connections=20)
async with httpx.AsyncClient(limits=limits, timeout=30,
follow_redirects=True) as client:
tasks = [validate_url(client, url) for url in urls]
results = await asyncio.gather(*tasks)
failures = [r for r in results if not r.get('pass')]
print(f'Validated: {len(results)} URLs')
print(f'Passed: {len(results) - len(failures)}')
print(f'Failed: {len(failures)}')
with open('soft_launch_validation_report.json', 'w') as f:
json.dump({'results': results, 'failures': failures}, f, indent=2)
return failures
top_2000_urls = [] # Load from GSC export
with open('gsc_top_2000_urls.txt') as f:
top_2000_urls = [line.strip() for line in f if line.strip()]
failures = asyncio.run(run_validation(top_2000_urls))
if failures:
print('\nFAILED URLS (first 20):')
for f in failures[:20]:
print(f" {f['origin_url']}: {f.get('checks', f.get('error'))}")</code></pre>
<h3>Step 31: Run rendering audit on destination staging</h3>
<p>Repeat the Playwright rendering audit from Step 5, this time against the staging environment. Compare JSON-LD type counts, rendered text length, and structured data validity against origin. Regression in any metric for any template type blocks launch.</p>
<h3>Step 32: Test Core Web Vitals on staging</h3>
<p>Use PageSpeed Insights API for lab data on the staging environment. Field data won't be available yet. Lab data is sufficient to catch regressions from the origin baseline. Hydrogen's RSC architecture should show LCP improvements. If it doesn't, investigate before launch.</p>
<h3>Step 33: Validate internal link depth for top-500 URLs</h3>
<p>Crawl the staging environment and calculate the click depth from the homepage for every URL in the top-500 GSC landing pages. Any page that has increased in depth by more than 1 click relative to origin requires a compensating internal link. Add those links before cutover.</p>
<h3>Step 34: Verify robots.txt on staging allows only specific test bots</h3>
<pre><code># staging-robots.txt
# Disallow all except Screaming Frog user agent for pre-launch QA crawls
User-agent: *
Disallow: /
User-agent: Screaming Frog SEO Spider
Allow: /
User-agent: Googlebot
Disallow: /</code></pre>
<h3>Step 35: Run full-site QA crawl on staging</h3>
<p>Use Screaming Frog or a custom Scrapy spider to crawl the staging environment completely. Identify: redirect chains, broken internal links, missing canonical tags, missing title tags, missing meta descriptions, pages returning non-200 status codes, oversized pages, missing schema types by template category.</p>
<h3>Step 36: Minimum soft-launch duration: 14 days</h3>
<p>Non-negotiable. Not 72 hours. Not a week. Fourteen days. The Magento migration that lost 14% ran 72 hours. You need two full Googlebot crawl cycles to detect hreflang conflicts, indexation anomalies, and canonical resolution issues before organic traffic is affected.</p>
<p>During the soft-launch period, monitor staging server logs for any unexpected crawler activity. Googlebot sometimes discovers staging environments via backlink graphs or DNS resolution history. If you see Googlebot on staging, tighten your access controls.</p>
<h3>Step 37: Final pre-cutover checklist sign-off</h3>
<p>Every team signs. SEO: redirect map validated, sitemaps ready, GSC verified. Engineering: edge redirects deployed, canonical tags correct, robots.txt staging disallow confirmed. Content: content freeze maintained, no URLs changed since freeze. Product: rollback plan documented and tested. This meeting happens 48 hours before hard cutover. If any item is unresolved, cutover is postponed.</p>
<hr>
<h2 id="phase-5">Phase 5: Hard Cutover and GSC Handoff (Steps 38–47)</h2>
<h3>Step 38: Update DNS / CDN routing</h3>
<p>The mechanics here depend on your infrastructure. For the Hydrogen migration on Oxygen CDN, the cutover was a custom domain routing update in the Shopify admin. For the Next.js migration on Vercel, it was a DNS CNAME update. In both cases, TTL was pre-lowered to 60 seconds 24 hours before cutover to minimize propagation delay.</p>
<h3>Step 39: Verify redirects in production immediately post-switch</h3>
<p>Within 5 minutes of DNS propagation confirmation, run the redirect validation script against the live production URLs. Any failure triggers an immediate rollback decision. You have a 15-minute window to decide. After 15 minutes, Googlebot may have started crawling and a rollback becomes more disruptive than fixing forward.</p>
<pre><code>#!/bin/bash
# post_cutover_spot_check.sh
# Quick redirect validation immediately after DNS propagation
URLS=(
"https://www.example.com/collections/supplements"
"https://www.example.com/products/vitamin-d3-5000iu"
"https://www.example.com/blogs/news/post-slug"
"https://www.example.com/p/OLD-SKU-123"
"https://www.example.com/blog/2023/04/old-post-slug"
)
echo "=== POST-CUTOVER REDIRECT SPOT CHECK ==="
echo "Time: $(date -u)"
echo ""
for url in "${URLS[@]}"; do
result=$(curl -s -o /dev/null -w "%{http_code} -> %{redirect_url}" \
--max-redirs 0 \
-L \
"$url")
final_code=$(curl -s -o /dev/null -w "%{http_code}" -L "$url")
echo "$url"
echo " First hop: $result"
echo " Final status: $final_code"
echo ""
done
echo "=== HOMEPAGE CHECK ==="
curl -s -o /dev/null -w "Status: %{http_code}, Size: %{size_download} bytes, Time: %{time_total}s\n" \
"https://www.example.com/"</code></pre>
<h3>Step 40: Update robots.txt on destination to allow all</h3>
<pre><code># production-robots.txt - activate post-cutover
User-agent: *
Allow: /
Disallow: /account/
Disallow: /checkout/
Disallow: /cart/
Disallow: /search?
Sitemap: https://www.example.com/sitemap_index.xml</code></pre>
<h3>Step 41: Execute sitemap thaw</h3>
<p>Run the sitemap_thaw.sh script. Submit destination sitemaps to GSC. Do not delete origin sitemaps from GSC yet. Leave them in place for 30 days so you can monitor any origin URL crawl activity.</p>
<h3>Step 42: Submit GSC Change of Address (domain migration only)</h3>
<p>If the migration involves a domain change, submit the Change of Address notification in GSC. This requires both origin and destination properties to be verified. Change of Address is a signal, not a guarantee. It accelerates Google's recognition of the migration but does not bypass normal redirect processing.</p>
<h3>Step 43: Verify destination GSC property receives data</h3>
<p>Within 24–48 hours of cutover, the destination GSC property should begin receiving crawl data. Monitor the URL inspection tool for key landing pages. Confirm rendering in the Mobile Usability and Rich Results reports. If the destination property shows zero crawl activity at 48 hours post-cutover, investigate DNS resolution and robots.txt.</p>
<h3>Step 44: Enable GSC crawl stats monitoring</h3>
<p>Set up daily exports of GSC crawl stats from both origin and destination properties for the first 30 days. You want to see crawl activity migrating from origin to destination over the first two weeks. Origin crawl requests should decline as Googlebot processes the 301s. Destination crawl requests should rise proportionally.</p>
<pre><code># gsc_crawl_stats_monitor.py
# Daily monitoring of crawl stats during migration window
from googleapiclient.discovery import build
from google.oauth2 import service_account
import pandas as pd
from datetime import datetime, timedelta
def get_crawl_stats(service, site_url, days=7):
"""Get crawl stats for the last N days."""
end_date = datetime.now().date()
start_date = end_date - timedelta(days=days)
# Note: Crawl stats API endpoint
response = service.urlcrawlerrorscounts().query(
siteUrl=site_url,
platform='web',
category='notFound'
).execute()
return response
def compare_crawl_activity(origin_property, destination_property, credentials):
service = build('searchconsole', 'v1', credentials=credentials)
# Monitor key metrics in crawl stats dashboard
print(f"Monitoring migration crawl transition:")
print(f" Origin: {origin_property}")
print(f" Destination: {destination_property}")
print(f" Check GSC > Settings > Crawl Stats for both properties daily")
print(f" Expected pattern: origin crawl rate -60% by day 14, -90% by day 30")</code></pre>
<h3>Step 45: Monitor GSC Index Coverage report daily for 30 days</h3>
<p>Watch for: spike in Crawled but not indexed on destination (rendering issue signal), spike in Excluded due to redirect on destination (redirect chain issue), disappearance of Valid URLs on origin GSC property (normal, expected). Create alerts for anomalies exceeding 5% of expected URL counts.</p>
<h3>Step 46: Watch organic traffic daily for the first 14 days</h3>
<p>Not weekly. Not in aggregate. By template type, by device, by country. The first sign of a migration issue is usually a specific template type declining while others hold. Template-level granularity in your monitoring catches this before it becomes a sitewide problem.</p>
<h3>Step 47: Conduct 7-day post-cutover retrospective</h3>
<p>One week after cutover, run a structured retrospective with every team involved. What went to plan. What deviated. What required emergency intervention. Document every deviation. This document is more valuable than the original runbook because it reflects what actually happened rather than what you planned. File it with the migration artifacts.</p>
<hr>
<h2 id="post-migration">The 90-Day Window: When You Actually Know If It Worked</h2>
<p>The 30-day mark is the first real data point. At 30 days you have enough GSC data to see whether the redirect graph has been substantially processed and whether template-level rankings are holding. The 14% loss migration showed -14.3% at 30 days. The Hydrogen migration showed +0.3% at 30 days.</p>
<p>But 30 days is not a verdict. Algorithm updates, seasonality, and crawl processing lag all make 30-day data noisy. The 60-day mark is more reliable. The 90-day mark is when I tell clients whether the migration succeeded by the only metric that matters to them: did we hold our organic revenue trajectory.</p>
<p>For the Hydrogen migration, 90-day measured delta against the 16-month GSC baseline was +0.3% organic clicks, +1.7% impressions, +0.4 average position improvement on the top-200 non-branded terms. The Core Web Vitals improvement from Shopify Plus to Hydrogen RSC architecture (LCP median improved from 2.4s to 1.7s in CrUX) likely contributed more to the positive delta than any SEO-specific migration work.</p>
<p>For the Next.js migration, 90-day delta was +4.7% organic clicks. The 30-day dip to -2.1% was caused by a canonical tag deployment bug that was identified at day 18 and fixed at day 22. The recovery trajectory from day 22 to day 90 was rapid, which suggests the canonical error had not yet affected index state significantly when it was caught.</p>
<p>Three things determine whether you're in the conversation about site migrations at the enterprise level in 2026: your ability to quantify pre-migration equity accurately, your discipline with the validation window, and your willingness to slow a launch down when something doesn't pass QA. The clients who pressure you to compress the soft-launch are the same clients who end up with the cliff-edge traffic chart. Every time.</p>
<p>If you're working on a platform migration of your own, the <a href="/internal-link-equity-distribution">internal link equity distribution guide</a> and the <a href="/crawl-budget-optimization">crawl budget optimization framework</a> are the two companion articles I'd read before running Phase 1. For domain-specific considerations, the <a href="/domain-migration-zero-loss">domain migration zero-loss framework</a> covers the GSC Change of Address signal in more depth. For hreflang specifically, the <a href="/hreflang-international-seo">hreflang implementation guide</a> is where I documented the validation approach that would have caught the -14% migration failure. The <a href="/million-redirects-2026">million-redirect management playbook</a> covers the edge-layer redirect architecture in more operational detail than this runbook allows.</p>
<p>External references worth reading: <a href="https://developers.google.com/search/docs/crawling-indexing/site-move-with-url-changes" target="_blank" rel="noopener">Google's official site move documentation</a> is still the authoritative baseline, though it hasn't been updated to reflect RSC-based platform behavior. <a href="https://www.searchpilot.com/resources/case-studies/" target="_blank" rel="noopener">SearchPilot's migration case study library</a> has some of the only rigorous controlled data on migration traffic impact that exists in the industry.</p>
<p>47 steps is not a lot. It's actually a compression. The working checklist I run internally has 193 line items with owners, dependencies, and pass/fail criteria for each. What I've given you here is the skeleton. The meat is the discipline to actually complete each step before moving to the next one, and the professional standing to tell a client "we are not launching on Tuesday because Step 36 isn't done" and have that conversation go the way it needs to go.</p>
<hr>
<h2 id="faq">Frequently Asked Questions</h2>
<script type="application/ld+json">
{
"@context": "https://schema.org",
"@type": "Article",
"headline": "Site Migration Runbook in 2026: The 47-Step Process That Saves Rankings",
"description": "A detailed technical SEO runbook for large-scale platform migrations in 2026, covering Shopify Plus to Hydrogen and WordPress to Next.js migrations with 47 documented steps, redirect architecture, sitemap freeze/thaw protocols, and soft-launch validation procedures.",
"author": {
"@type": "Person",
"name": "Andrii",
"url": "https://benrey.io"
},
"publisher": {
"@type": "Organization",
"name": "Benrey",
"url": "https://benrey.io"
},
"datePublished": "2026-05-20",
"dateModified": "2026-05-20",
"mainEntityOfPage": {
"@type": "WebPage",
"@id": "https://benrey.io/site-migration-runbook-2026"
},
"keywords": ["site migration", "SEO migration", "Shopify Plus to Hydrogen", "WordPress to Next.js", "redirect mapping", "technical SEO", "platform migration"],
"articleSection": "Technical SEO"
}
</script>
<script type="application/ld+json">
{
"@context": "https://schema.org",
"@type": "FAQPage",
"mainEntity": [
{
"@type": "Question",
"name": "How long should a soft-launch period last before a hard site migration cutover?",
"acceptedAnswer": {
"@type": "Answer",
"text": "Minimum 14 days. This allows two full Googlebot crawl cycles to surface hreflang conflicts, canonical resolution issues, and rendering anomalies before organic traffic is affected. A 72-hour soft-launch is insufficient for large sites and is a documented cause of significant ranking loss. For sites above 1 million URLs, 21 days is preferable."
}
},
{
"@type": "Question",
"name": "What is the sitemap freeze and thaw protocol in site migrations?",
"acceptedAnswer": {
"@type": "Answer",
"text": "Sitemap freezing means replacing dynamic sitemap generation on the origin platform with static XML files representing the final URL set before migration begins. This prevents new URLs from being added to the crawl queue during the migration window. Sitemap thaw refers to submitting the destination sitemaps to Google Search Console after hard cutover is confirmed and validated, while keeping origin sitemaps active for 30 days to monitor crawl activity."
}
},
{
"@type": "Question",
"name": "Where should 301 redirects be implemented for a headless platform migration?",
"acceptedAnswer": {
"@type": "Answer",
"text": "At the edge layer, before any application logic runs. For Hydrogen migrations on Oxygen CDN, this means Oxygen's edge routing. For Next.js on Vercel, this means Edge Middleware using Vercel Edge Config for large redirect maps. Implementing redirects at the application layer or in htaccess creates latency that affects crawl budget consumption and user experience during the migration window."
}
},
{
"@type": "Question",
"name": "Why did my site migration lose rankings even though I set up all the 301 redirects correctly?",
"acceptedAnswer": {
"@type": "Answer",
"text": "The most common causes beyond redirect errors are: parameterized URL equity not mapped (URLs with query strings that had independent link equity), hreflang absolute URLs not updated to the destination domain, canonical tags on the destination still pointing to origin URLs, structured data injected client-side via useEffect rather than server-rendered, and internal link depth increasing for high-equity pages. A thorough pre-migration rendering audit and URL parameter equity audit catch most of these before launch."
}
},
{
"@type": "Question",
"name": "How many months of Google Search Console data should I pull before a site migration?",
"acceptedAnswer": {
"@type": "Answer",
"text": "16 months. 12 months captures full seasonality. The additional 4 months provide a comparison buffer that survives algorithm update volatility and allows you to distinguish migration-caused traffic changes from seasonal or algorithmic fluctuations when evaluating 30, 60, and 90-day post-migration performance."
}
},
{
"@type": "Question",
"name": "Should I do a site migration during a compressed time window or use a staged rollout?",
"acceptedAnswer": {
"@type": "Answer",
"text": "For sites above 500,000 URLs, a staged URL namespace migration over 3 to 4 weeks is preferable to a compressed migration window. A compressed window causes Googlebot to encounter millions of 301 redirects simultaneously, throttling crawl rate and creating a redirect processing backlog that takes weeks to resolve. A staged rollout moves URL families in order of crawl priority, allowing the redirect graph to process incrementally and reducing ranking volatility significantly."
}
}
]
}
</script>
<dl>
<dt>How long should a soft-launch period last before a hard site migration cutover?</dt>
<dd>Minimum 14 days. This allows two full Googlebot crawl cycles to surface hreflang conflicts, canonical resolution issues, and rendering anomalies before organic traffic is affected. A 72-hour soft-launch is insufficient for large sites. For sites above 1 million URLs, 21 days is preferable.</dd>
<dt>What is the sitemap freeze and thaw protocol in site migrations?</dt>
<dd>Sitemap freezing replaces dynamic sitemap generation on the origin platform with static XML files representing the final URL set before migration begins. This prevents new URLs from being added to the crawl queue during the migration window. Sitemap thaw refers to submitting destination sitemaps to Google Search Console after hard cutover is confirmed, while keeping origin sitemaps active for 30 days to monitor crawl activity.</dd>
<dt>Where should 301 redirects be implemented for a headless platform migration?</dt>
<dd>At the edge layer, before any application logic runs. For Hydrogen migrations on Oxygen CDN, this means Oxygen's edge routing. For Next.js on Vercel, Edge Middleware using Vercel Edge Config for large redirect maps. Application-layer or htaccess redirects create latency that affects crawl budget consumption during the migration window.</dd>
<dt>Why did my site migration lose rankings even though I set up all the 301 redirects correctly?</dt>
<dd>The most common causes beyond redirect errors: parameterized URL equity not mapped individually, hreflang absolute URLs not updated to the destination domain, canonical tags on destination still pointing to origin, structured data injected client-side via useEffect rather than server-rendered, and internal link depth increasing for high-equity pages. A rendering audit and URL parameter equity audit catch most of these before launch.</dd>
<dt>How many months of GSC data should I pull before a site migration?</dt>
<dd>16 months. 12 months captures full seasonality. The additional 4 months provide a comparison buffer that survives algorithm update volatility, letting you distinguish migration-caused traffic changes from seasonal or algorithmic fluctuations at 30, 60, and 90-day post-migration reviews.</dd>
<dt>Should I use a compressed migration window or a staged rollout?</dt>
<dd>For sites above 500,000 URLs, a staged URL namespace migration over 3 to 4 weeks is preferable. A compressed window causes Googlebot to encounter millions of 301 redirects simultaneously, throttling crawl rate and creating a processing backlog. Staged rollout moves URL families in crawl-priority order, allowing the redirect graph to process incrementally and reducing volatility.</dd>
</dl>
