Written 19 May 2026. The project that broke my old mental model ran from August through November 2025. I'm still finding edge cases.
The Project That Forced This
In August 2025 I took on an IA audit for a specialty industrial supplies retailer. 783,000 active SKUs across 14 categories. The site had grown via three acquisitions since 2021, each of which brought its own taxonomy. The result was a URL structure that looked coherent in a spreadsheet and was completely incoherent to a crawler.
I had done large-catalog audits before. My previous record was about 340k SKUs for a home improvement client in 2023. That project taught me rules I believed were portable. They weren't.
The specific shock came from log file data. Between July and October 2025, Googlebot crawl frequency on this site dropped 31% year-over-year — while the catalog grew 18%. We had more pages getting crawled less often. I initially attributed this to a server response issue, then to a robots.txt misconfiguration from the most recent acquisition integration. Neither was the cause. What I eventually traced it to was a structural problem: the IA was not telling Googlebot's entity-resolution layer anything coherent, so the crawler kept revisiting old high-authority nodes rather than discovering new ones.
That August-to-November engagement produced what I now call the CLR model — Category, Lens, Resolution. This article explains it.
What Broke: The Flat-Tree Assumption
Most e-commerce IA advice is built around a simple tree: Department > Category > Subcategory > Product. Three or four levels, clean hierarchy, breadcrumbs map 1:1 to the tree. This model worked well from roughly 2012 through 2023.
It breaks at scale for two reasons that have gotten worse in 2025–2026.
First: flat trees cannot express the semantic relationships that Google's entity graph now expects. A product like "3/8-inch stainless steel hex bolt, grade 5, pack of 100" belongs to multiple valid category paths simultaneously. It's a fastener, it's a stainless steel item, it's a grade-5 hardware item, it's a pack-quantity SKU. A flat tree forces one canonical path. That canonical path may not be the one Google's entity graph prefers for the queries that actually drive revenue.
Second: AI Overviews changed how snippet extraction works for commercial queries. When someone searches "best grade-5 stainless hex bolts," Google's AI Overview in 2026 often extracts data directly from category or subcategory pages — not from the product pages themselves. If your subcategory page is thin (a heading, 80 words, and a grid of product tiles), the extraction has nothing to work with. The AI Overview cites someone else who wrote a proper entity-dense subcategory page.
I watched this happen in real-time for the industrial client. Their "Hex Bolts" subcategory page — 12 words of body content, no specifications table, no material callouts — was outranked in AI Overview citations by a competitor whose category page contained a 400-word materials specification block with entity-rich structured content.
The CLR Model Explained
CLR stands for Category, Lens, Resolution. Each layer has a distinct job. The mistake most large catalogs make is blurring these jobs together — usually by treating everything as a Product Detail Page (PDP) variant of the same template.
Layer 1: Categories (Stable Entities)
Categories are the permanent, entity-anchored nodes of the IA. They should correspond to recognizable entities in Google's Knowledge Graph. "Hex Bolts" is an entity. "Fasteners" is an entity. "3/8-inch fasteners" is not yet a stable Knowledge Graph entity — it's a filtered view.
Every Category page needs:
- A stable, keyword-anchored slug that does not change when merchandising changes
- Entity-rich body content: 300–600 words minimum, specification tables, material callouts, common use cases
- JSON-LD markup that explicitly declares the category as a
ProductCollectionwithaboutpointing to the parent entity - Breadcrumbs mapped to the stable taxonomy tree, not to the user's navigation session
Category pages in the CLR model receive the bulk of internal link equity. They are the indexable spine.
Layer 2: Lenses (Justified Filtered Views)
Lenses are filtered views that get their own crawlable URL — but only when they meet a justification threshold. My current threshold is 150+ average monthly searches AND a coherent entity label that doesn't already exist at the Category layer.
A Lens URL looks like:
/hex-bolts/stainless-steel/
/hex-bolts/grade-5/
/hex-bolts/metric/
Not like:
/hex-bolts/?material=stainless&grade=5&size=3-8-inch&pack=100
The first pattern creates indexable entity nodes. The second creates parameter soup that gets crawled, wastes budget, and doesn't earn snippet extractions.
Lenses that don't meet the threshold use rel="canonical" pointing to their parent Category. JavaScript-driven filtering on the front end is fine for the user experience — as long as the crawlable URL surface only exposes justified Lenses.
This is where I made the mistake I'll describe in the contrarian section: I originally set my threshold too low (50 monthly searches) and created 34,000 Lens URLs that diluted crawl budget without contributing meaningfully to organic traffic.
Layer 3: Resolution (SKUs and Variants)
Resolution pages are the Product Detail Pages. At 780k SKUs, even with aggressive canonicalization you likely have 200,000–400,000 indexable PDPs. The question is which ones should be indexed.
My current rule: a PDP should be independently indexed only if it has at least one of the following:
- A unique specification that doesn't appear on any other PDP in the same Lens
- A product name that users actually search for as a discrete query (verifiable in GSC)
- A minimum of 3 inbound links from external sources (manufacturer site, spec sheets, trade publications)
Everything else: canonical to the most appropriate Lens page. Not noindex — canonical. The distinction matters for crawl budget and for how Google handles duplicate content signals.
Variant pages (color, size, pack quantity) almost never need independent indexation at this scale. Canonical all variants to the base SKU. The exception is when a variant has a distinct part number that users search for directly — this is more common in industrial and medical supply than in consumer goods.
AI Overview Snippet Extraction Patterns
I've been tracking AI Overview extraction behavior on commercial queries since the feature expanded in Q4 2024. By May 2026, I have observations across 7 client accounts. The patterns are consistent enough that I'm willing to publish them, with the caveat that Google has iterated the extraction logic several times and will do so again.
For product category queries ("best [category]", "types of [product]", "[product] buying guide"):
- Google extracts from pages with visible specification tables more often than from pages with equivalent content in prose form
- Pages with
ItemListschema showing fewer than 8 items get extracted more often than pages with 40+ items — apparently the AI prefers curated short lists over comprehensive grids - Entity co-occurrence matters: a category page that mentions related entities (brands, materials, standards bodies like ASTM or ISO) gets extracted at higher rates than one that only describes the product
For navigational-intent queries (brand + category type):
- The AI Overview almost never extracts from faceted navigation pages, even well-optimized ones
- It consistently prefers pages with a clear editorial voice and first-person specificity ("we stock grade-5 and grade-8 in stainless and zinc-plated") over anonymous catalog-speak
The implication for IA: your Category pages and top-tier Lens pages need to read like expert editorial content, not like catalog headers. This is uncomfortable for many e-commerce teams whose content workflows treat category descriptions as an afterthought. It was uncomfortable for my industrial supplies client, whose merchandising team owned category content and had no incentive to write 500 words about hex bolt metallurgy.
We solved it by building a structured content template — 7 required fields, each mapped to a specific entity type — that merchandising could fill out without writing prose. The template auto-generated a specification block; a content editor added 2–3 sentences of editorial context. Time per category: 22 minutes. Lift in AI Overview appearances: 41% over 90 days. See my earlier analysis of AI Overview CTR impact for baseline numbers.
Crawl Budget in 2026: The Shrinking Window
This is the part of the story that should concern every large-catalog operator.
Google's own crawl behavior is changing as AI Overviews handle more of the query surface. I've now seen this in log data from 4 separate large e-commerce clients (100k+ SKUs). The pattern: total Googlebot crawl requests dropped 18–34% year-over-year between mid-2024 and mid-2025, despite site growth. The crawl that does happen is increasingly concentrated on a smaller set of high-authority pages.
The hypothesis — and it's a hypothesis, not confirmed by Google — is that when AI Overviews can answer a query from already-indexed content, Googlebot has less incentive to crawl the pages that would have answered it. The entity graph becomes self-reinforcing: pages that are already well-represented in the graph get re-crawled, pages that aren't stay dark.
What this means practically:
- Crawl budget is more precious than it was in 2023. Every wasted crawl on a thin Lens URL or a redundant variant page is a crawl that didn't go to a Category page that needed refreshing.
- Internal linking patterns need to be radically more deliberate. Random related-product widgets that add 40 internal links per PDP are budget drains, not equity distributors.
- XML sitemaps need to be tiered. I now maintain three sitemaps for large catalogs: a Priority sitemap (Category + top-tier Lens pages, submitted first), a Standard sitemap (justified SKU pages), and an Archived sitemap (everything else, not submitted, just available if Googlebot finds it).
One thing I tested in January 2026: removing 62,000 thin Lens URLs from the industrial client's crawlable surface via canonical (not noindex, not robots.txt) and updating the Priority sitemap. Within 6 weeks, Googlebot crawl frequency on Category pages increased 28%. Whether that's correlation is fair to ask — we had done other work simultaneously — but the directional signal is real.
My earlier piece on crawl budget has the technical mechanics of how to audit current crawl distribution from log files.
Entity-Graph Alignment for Product Pages
Google's Knowledge Graph has expanded its product-entity coverage significantly since the 2024 shopping graph integrations. For industrial and technical products, there are now recognized entities for material standards, product grades, and common specifications. Aligning your IA to these entities — rather than to your internal merchandising taxonomy — is the highest-leverage structural change you can make.
What alignment looks like in practice:
Your internal taxonomy might call something "Stainless Fasteners > Hex > Metric." Google's entity graph recognizes "Hex bolt" as an entity (with subtype relationships to metric, inch, grade), "Stainless steel" as a material entity, and ASTM A193 as a standard entity. If your URL structure, H1, and JSON-LD all use your internal naming convention but none of Google's entity labels, you are invisible to entity-graph crawl prioritization.
The fix is not to abandon your internal taxonomy. It's to map your taxonomy nodes to Knowledge Graph entities and surface those entities explicitly in page content, URL slugs, and structured data. A slug like /hex-bolts/stainless-steel/ is better than /fasteners/sstl-hex/ even if your internal SKU system uses the latter.
JSON-LD that helps with entity alignment:
{
"@context": "https://schema.org",
"@type": "CollectionPage",
"name": "Stainless Steel Hex Bolts",
"about": {
"@type": "Product",
"name": "Hex Bolt",
"material": "Stainless Steel",
"additionalProperty": [
{
"@type": "PropertyValue",
"name": "Standard",
"value": "ASTM A193"
}
]
},
"breadcrumb": {
"@type": "BreadcrumbList",
"itemListElement": [
{"@type": "ListItem", "position": 1, "name": "Fasteners", "item": "https://example.com/fasteners/"},
{"@type": "ListItem", "position": 2, "name": "Hex Bolts", "item": "https://example.com/hex-bolts/"},
{"@type": "ListItem", "position": 3, "name": "Stainless Steel Hex Bolts", "item": "https://example.com/hex-bolts/stainless-steel/"}
]
}
}
See my Knowledge Graph optimization deep dive and the JSON-LD at scale piece for more on both topics.
URL Patterns That Still Work
After the CLR model, here's the URL taxonomy I implement for large catalogs:
# Layer 1: Categories
/[department-slug]/[category-slug]/
/fasteners/hex-bolts/
/fasteners/washers/
# Layer 2: Lenses (only if justified)
/[category-slug]/[entity-attribute]/
/hex-bolts/stainless-steel/
/hex-bolts/grade-5/
/hex-bolts/metric/
# Layer 3: Resolution (PDPs)
/[category-slug]/[product-name-slug]/[sku]/
/hex-bolts/hex-bolt-stainless-3-8-inch/HBSS-375-100
# Variants: canonical to base SKU
/hex-bolts/hex-bolt-stainless-3-8-inch/HBSS-375-050
→ canonical: /hex-bolts/hex-bolt-stainless-3-8-inch/HBSS-375-100
Patterns I've moved away from:
# Department in every URL (redundant, wastes slug depth)
/fasteners/hex-bolts/stainless-steel/hex-bolt-stainless-3-8-inch/
# Parameter-based facets (crawl budget drain)
/hex-bolts/?material=stainless&grade=5
# ID-only PDPs (zero entity signal in URL)
/product/78234/
One contrarian position I hold: keeping the department slug in the URL path is not worth it at this scale. The breadcrumb JSON-LD communicates the hierarchy to Google. The URL slug adding /fasteners/ before /hex-bolts/ wastes characters, makes URLs longer in SERPs, and creates redirect complexity when you restructure departments (which happens every 18 months at large retailers). I removed department prefixes from 12,000 category and Lens URLs at the industrial client. Zero ranking loss after 90 days. Some minor ranking gains on mobile where shorter URLs fit the SERP display better.
Two Things I Was Wrong About
I built my original IA framework in 2019–2021 and held onto assumptions I should have tested earlier.
Wrong thing one: Lens thresholds. I used to create Lens URLs for any facet combination with 50+ monthly searches. My reasoning: even modest traffic is worth capturing, and the canonicalization overhead is manageable. I was wrong. On the industrial client, I inherited 34,000 Lens URLs meeting my old 50-search threshold. 31,000 of them had zero clicks in 18 months of GSC data. They were being crawled an average of once every 67 days — rare enough to not help, frequent enough to steal budget from better pages. I raised the threshold to 150 searches, canonicalized everything below it, and crawl budget concentration on meaningful pages improved substantially. The 50-search threshold was comfortable for me because it felt conservative. It wasn't.
Wrong thing two: Category page content length. I used to target 150–250 words for category descriptions. Felt like enough to signal intent without overwhelming users who came to browse, not read. The AI Overview extraction data killed this assumption. Pages with 300+ words and structured specification blocks get extracted at rates roughly 3x higher than pages with 150 words. The content doesn't need to dominate the visual layout — it can sit below the product grid, in an expandable section, or in a sidebar — but it needs to exist in the DOM where Googlebot reads it. I've since updated my category page templates to require 350 words minimum, with a specification table as a required element.
What to Actually Do Next Week
If you operate a catalog over 100k SKUs and haven't audited your IA against the CLR model, here's the fastest diagnostic:
- Pull your Googlebot crawl log for the last 90 days. Segment by URL pattern. What percentage of crawl is going to PDPs vs. Category pages vs. Lens pages?
- Pull GSC clicks for all Lens-type URLs (filtered views, faceted navigation pages). Calculate clicks-per-URL. If median is under 3 clicks/month, your Lens layer is over-extended.
- Check 10 Category pages for content length and structured data. If average is under 200 words and you have no specification tables, your entity-graph alignment is weak.
- Search 5 of your key category terms in Google and check whether AI Overviews appear. If they do, look at what's being cited. If it's not you, read the cited pages and compare to yours.
For external context on how large catalogs are handling this, the SEMrush blog published a useful case study series in March 2026 on faceted navigation behavior post-AI Overviews. Worth reading alongside this. The Google crawl budget documentation updated in February 2026 is also more specific about crawl prioritization signals than earlier versions.
The CLR model isn't magic. It's a forcing function for decisions that most large-catalog teams defer indefinitely: how many filtered views should exist, what content minimum qualifies a page for indexation, and what URL patterns actually communicate entity structure to a crawler. Those decisions don't make themselves. If you don't make them deliberately, your CMS and your developers will make them for you, and the defaults are almost never right.
I'm still finding edge cases in the industrial client's data — most recently around kit and bundle SKUs that don't fit cleanly into any CLR layer. If that's a problem you're dealing with, the ecommerce category page piece has a section on product relationship schemas that partially addresses it.
Filed under: information architecture, e-commerce SEO, crawl budget, AI Overviews, CLR model. Last audited against live client data: May 2026.
