Skip to content
TECHNICAL SEO / FIELD NOTE 163

Taxonomy Design for SEO in 2026: Card-Sorting Studies That Stick After the AI Pivot

Reading map: Why Taxonomy Is Back at the Top of My Audit Stack; What Card Sorting Actually Tells You (and What It Doesn't); Three Mistakes I Made in 2025 Taxonomy Projects; The CREST Framework: How I Evaluate Taxonomy Decisions Post-AI Pivot
A reading map of this field note. Download SVG ↓

By Andrii — Published 19 May 2026

Why Taxonomy Is Back at the Top of My Audit Stack

Taxonomy design spent most of 2022–2024 in a strange limbo. Everyone agreed it mattered, but the discussion was almost entirely UX-framed: card sorting, tree testing, information scent. SEO practitioners treated taxonomy as a given — something you inherited from the CMS team and worked around. I was guilty of that myself.

Two things changed. First, the AI Overviews rollout throughout 2025 made it apparent very fast that topical authority isn't just a content problem. It's a structural problem. You can have excellent articles on adjacent subtopics and still get no AI Overview citations if those subtopics aren't organized into a coherent entity cluster that Google's knowledge graph can resolve. Taxonomy is the primary mechanism for creating those clusters. Second, Googlebot's crawl rate reductions — which I started seeing in log files from clients around September 2025 and confirmed with a client in the industrial parts space by November — mean that the old strategy of "index everything and let Google sort it out" is producing measurable ranking losses for large sites. Taxonomy determines what gets crawled first and what signals get consolidated. That's no longer a secondary concern.

So I went back to doing proper taxonomy work in Q4 2025. This article describes what I learned running three separate card-sorting studies between October 2025 and March 2026, plus the CREST evaluation framework I built to translate user research results into decisions that hold up under SEO constraints.

What Card Sorting Actually Tells You (and What It Doesn't)

Card sorting is a user research method where participants organize items (cards, each labeled with a piece of content or concept) into groups and then name those groups. Open card sort: participants create their own groups. Closed card sort: you give them predefined categories and they assign cards to those. Hybrid: a mix.

What it actually tells you is vocabulary. That's the most underrated output. When 23 out of 30 participants label a group "supplies" instead of "consumables," that's a signal that your taxonomy node name is working against your content's findability. People don't search for the words you use internally. They search for the words they use when thinking about a problem. Card sorting surfaces that gap reliably when you recruit the right participants.

What it does not tell you is anything about crawl efficiency, entity graph structure, or snippet extraction likelihood. Those are engineering and SEO constraints that card sorting is constitutionally unable to surface. A perfectly user-tested taxonomy can be a crawl disaster. I saw this directly.

A Short War Story from November 2025

A mid-sized B2B tools and equipment retailer hired me to redesign their taxonomy after a March 2025 algorithm update wiped out 31% of their organic sessions. Their previous consultant had run a thorough open card sort with 40 participants, implemented a beautifully structured hierarchy, and that's where the project ended. The taxonomy looked excellent on paper. Navigation clarity had improved measurably.

The problem was that the card sort had produced 9 top-level categories covering roughly 68 subcategories and 340 leaf-level nodes. Google was crawling the site's category pages at a rate of about 1,200 per month against a catalog of 48,000 SKUs. The new taxonomy had increased the total number of indexable category pages by 34% without any mechanism to signal priority to Googlebot. Combined with thin content on 60% of the new subcategory pages, the result was predictable in retrospect: crawl dilution, thin-page signals across a wider surface, worse rankings.

The card sort wasn't wrong. The taxonomy design process had simply stopped before asking the second set of questions.

Three Mistakes I Made in 2025 Taxonomy Projects

Might as well be honest about where I got it wrong before describing what I do now.

Mistake 1: recruiting general ecommerce shoppers instead of category-specific buyers. For the tools retailer project, I ran a supplemental card sort in January 2025 and pulled participants from a general consumer panel. The result was vocabulary data that reflected how someone who occasionally buys tools thinks about tool categories — not how a maintenance professional, which was the actual buyer persona, searches for parts. The groups they created were logical for a consumer context and useless for B2B procurement SEO. I had to re-run the study in March with correctly recruited participants and the data looked almost nothing like the first round.

Mistake 2: treating card sort output as taxonomy output. Card sorting produces clusters and vocabularies. Those clusters still have to be translated into a URL structure, a crawl hierarchy, and a set of page-level content requirements. Skipping that translation step — going from card sort data straight to sitemap — produces taxonomies that are semantically coherent but structurally incoherent from a crawl and entity perspective.

Mistake 3: not modeling temporal stability. One of the three 2025 projects was for a news-adjacent publisher covering a fast-moving technology sector. The card sort produced sensible categories, but six months after implementation, three of the top-level categories had either collapsed (two became clearly overlapping once the technology landscape stabilized) or needed splitting (one was actually three distinct buyer intents). I hadn't built any mechanism to revisit the taxonomy on a defined schedule. The resulting instability — multiple category renames, URL structure changes, redirect chains — cost significant crawl budget and rankings.

The CREST Framework: How I Evaluate Taxonomy Decisions Post-AI Pivot

After the third project, I started formalizing what the second evaluation layer should look like. The output is the CREST framework — five criteria applied to each proposed taxonomy node after the user research phase is complete.

CREST: Crawlability, Relevance-signal density, Entity alignment, Snippet surface area, Temporal stability.

CREST at a glance
  • C — Crawlability: Can Googlebot reach and prioritize this node given current crawl budget? Is URL depth ≤3 hops from homepage for priority nodes?
  • R — Relevance-signal density: Will the page assembled for this node have enough content mass and keyword specificity to register a clear topical signal?
  • E — Entity alignment: Does this node map to a named entity or concept in Google's Knowledge Graph? Can it be marked up with schema that asserts that mapping?
  • S — Snippet surface area: Is the node's topic narrow enough that a single page can give Google a definitive answer to the queries it targets, generating AI Overview citation potential?
  • T — Temporal stability: Is the category likely to exist in recognizably the same form in 18–24 months, or is it coupled to a product trend, a technology, or terminology that may shift?

Each node proposed from card sort data gets scored against these five criteria before it goes into the structural design. Nodes that fail Crawlability (too deep, too similar to an existing node) get merged or restructured. Nodes that fail Entity alignment get either sharpened or elevated into a parent concept that does have Knowledge Graph representation. Nodes that fail Temporal stability get flagged for a defined review date and, where possible, structured as facets rather than permanent hierarchy levels.

This is not a magic checklist. It's a forcing function for conversations with clients and developers that would otherwise not happen until something breaks.

Taxonomy Nodes as Entity Containers

The single biggest shift in how I think about taxonomy design over the past 18 months is the move from thinking of taxonomy nodes as navigation helpers to thinking of them as entity containers. The difference matters.

A navigation helper is designed to get a user from point A to point B. Its job is wayfinding. An entity container is designed to accumulate and consolidate all of the relevance signals, schema markup, internal links, and content mass associated with a specific named entity — a product category, a topic, a process, a type of thing — so that Google's entity graph can resolve the node to a Knowledge Graph entry and assign it authority.

# Entity container vs. navigation node: structural comparison

Navigation node (UX-first design):
  URL: /shop/categories/hand-tools/
  Purpose: help users find hand tools
  Content: minimal, possibly just a product grid
  Schema: none, or generic BreadcrumbList
  Internal links: from nav menu, maybe parent category
  Entity signal: weak

Entity container (CREST-pass design):
  URL: /hand-tools/
  Purpose: consolidate all authority signals for "hand tools" as entity
  Content: 300–600 word category editorial + buying guide + spec table
  Schema: ItemList + BreadcrumbList + (where applicable) Product aggregate
  Internal links: from all relevant product pages, from blog posts mentioning hand tools,
                  from comparison pages, from parent taxonomy node
  Entity signal: strong — multiple signal types pointing at one URL

The URL depth difference above is intentional. One of the consistent findings from my log analysis in 2025 was that category pages buried at /shop/categories/level1/level2/ received meaningfully fewer crawl visits per month than equivalent pages at /level1/level2/ on the same domain. The intermediate path segments — /shop/, /categories/ — consume URL depth without contributing entity signal. Strip them where CMS architecture permits.

Schema That Actually Asserts Entity Membership

For taxonomy nodes designed as entity containers, the schema implementation should do more than just describe the page. It should assert the node's membership in a broader entity hierarchy. The sameAs property pointing to a Wikidata or Wikipedia entry is the most direct way to do this for well-defined categories.

{
  "@context": "https://schema.org",
  "@type": "CollectionPage",
  "name": "Hand Tools",
  "description": "Professional and consumer hand tools including wrenches, screwdrivers, pliers, and striking tools.",
  "url": "https://example.com/hand-tools/",
  "sameAs": "https://www.wikidata.org/wiki/Q192451",
  "breadcrumb": {
    "@type": "BreadcrumbList",
    "itemListElement": [
      {
        "@type": "ListItem",
        "position": 1,
        "name": "Tools & Equipment",
        "item": "https://example.com/tools-equipment/"
      },
      {
        "@type": "ListItem",
        "position": 2,
        "name": "Hand Tools",
        "item": "https://example.com/hand-tools/"
      }
    ]
  },
  "hasPart": [
    { "@type": "CollectionPage", "name": "Wrenches", "url": "https://example.com/hand-tools/wrenches/" },
    { "@type": "CollectionPage", "name": "Screwdrivers", "url": "https://example.com/hand-tools/screwdrivers/" }
  ]
}

The hasPart relationship here is doing something important. It asserts, in structured data, the hierarchical relationship between this node and its children — which reinforces the same relationship that internal linking and URL structure assert. Three signals, same relationship. That redundancy is not accidental.

Snippet Surface Area and Why Broad Categories Now Underperform

This is the contrarian take most taxonomy consultants aren't saying yet, so I'll be direct: the broad top-level category page — the one designed to be the authoritative hub for a large, wide topic — is significantly less effective at capturing AI Overview snippet citations than a mid-level category page with a narrow, specific focus.

I started noticing this pattern in GSC data around August 2025 on a home improvement client. Their top-level /flooring/ category had strong authority signals — high PageRank, thousands of internal links, rich schema. But Google's AI Overviews for flooring-category queries were consistently citing a competitor's more specific pages: /hardwood-flooring/, /vinyl-plank-flooring/, not their equivalent of /flooring/. The broad page was authoritative. It just wasn't specific enough to be the definitive answer to anything.

AI Overview extraction appears to favor pages that can be cited as the source for a particular fact or recommendation. A page that covers 12 flooring types at 150 words each cannot be cited as the definitive source on any of those 12 types. A page that covers one flooring type at 800 words with comparison tables, specification data, and a clear recommendation structure can be cited for that type specifically.

The taxonomy implication is that you should design for depth of coverage at the mid-level, not breadth of coverage at the top. Top-level nodes should exist primarily for crawl consolidation and entity hierarchy assertion. The content investment — and the snippet opportunity — sits one level down.

Two Takes Most Taxonomy Literature Gets Wrong

1. "More granular is better for SEO"

Standard SEO advice on taxonomy granularity runs toward more specific being better: more precise categories mean less competition, more targeted content, cleaner intent signals. This is true at the individual page level. It is not true at the structural level.

When you add 200 leaf-level taxonomy nodes to a site that gets 1,800 Googlebot crawl visits per day, you are not adding 200 pages of authority — you are diluting 1,800 crawl visits across 200 additional pages, each of which now gets crawled less frequently, each of which accumulates authority more slowly, and most of which will never have enough content to send a strong entity signal. The granularity question is always: granular relative to what crawl budget?

I run a crawl budget model before finalizing any taxonomy: estimated monthly crawl visits (from log data or GSC coverage reports) divided by target taxonomy node count gives a rough crawl visits per node per month. Anything under 15 visits per node per month is a signal to consolidate. On the industrial parts client from the CLR model project (described in the large ecommerce IA article), we eliminated 340 leaf-level category pages that were getting fewer than 8 crawl visits per month. Three months later, crawl visits to the remaining pages increased by 41% and the eliminated pages' traffic had consolidated, not disappeared.

2. "Card sorting should drive taxonomy, full stop"

The UX discipline position is that user mental models should define taxonomy. If users think of something as belonging in category X, put it in category X. Don't override user research with internal assumptions.

I have a lot of respect for this principle. I also think it's incomplete in a way that causes real harm when applied to SEO taxonomy design without modification.

User mental models are often at odds with search query vocabulary. A card sort might reveal that your target audience groups "safety glasses" under "personal protective equipment." But if the majority of relevant search queries use "safety glasses" as a primary term and "personal protective equipment" is a secondary or regulatory term, building /personal-protective-equipment/safety-glasses/ into your URL structure is actively working against your relevance signals. The URL communicates taxonomy position. Taxonomy position communicates entity context. Entity context influences which queries the page ranks for.

What I do now: run the card sort to get vocabulary data and group structure. Then run keyword volume and entity graph research against the card sort outputs before building the URL hierarchy. Where there's conflict between user vocabulary and search vocabulary, I resolve it at the page level with content strategy — the page speaks to both vocabularies in its copy — and at the URL level with the search vocabulary, because the URL is the signal that matters to the crawler.

Temporal Stability: The Dimension Nobody Models

Every taxonomy degrades. That's not a failure of design — it's a consequence of operating in a domain that changes over time. Products are discontinued, terminology shifts, new categories emerge, old categories merge. The question isn't whether your taxonomy will need revision. It's whether you have a process for managing revision without destroying the URL architecture you've spent years building authority into.

What I missed in the 2025 publisher project was that some taxonomy nodes I created were implicitly coupled to a specific moment in a technology's development. The technology moved. The nodes didn't. Within eight months, two category names were search-vocabulary relics and one category had become two distinct buyer intents that were fragmenting the page's ability to rank for either.

The CREST Temporal Stability criterion asks explicitly: is this node coupled to a terminology trend or a product-moment that may not persist? If yes, the node design should account for that. Options:

  • Design the node around the stable underlying entity (the technology, the material, the process) rather than the current terminology for it, and use content on the page to connect to current terminology.
  • Implement the volatile concept as a facet or filter rather than a permanent hierarchy node, so URL structure doesn't encode the assumption that this concept is stable.
  • Document the review trigger explicitly: "revisit this node if search volume for [term X] drops below [threshold Y] for three consecutive months."

The third option sounds trivially obvious. Almost nobody does it. There's no operational process for taxonomy maintenance on most sites I audit. The taxonomy gets designed, implemented, and then revisited only when something breaks badly enough to create a visible traffic loss.

The Process I Run Now, Step by Step

This isn't a methodology pitch. It's what I actually do, in order, with the caveats noted.

Step 1: Audit the existing taxonomy against CREST before touching user research. No point running card sorts until you know what's actually broken. Pull Screaming Frog or Sitebulb data, segment by taxonomy level, check crawl data from logs or GSC coverage, identify which nodes are failing which CREST criteria. This takes 4–8 hours for a site under 50,000 URLs.

Step 2: Pull keyword and entity data for the content domain. Before recruiting card sort participants, I want to know the search vocabulary for this domain: what terms are high-volume, what entities appear in Google's People Also Ask and related searches for core terms, what the Knowledge Graph returns for the top 10–15 parent concepts. This shapes which cards I put in the card sort deck and helps me interpret the output.

Step 3: Run open card sort with domain-specific participants. 25–35 participants. Recruitment criterion: they must have searched for and purchased (or evaluated) products or content in this domain in the past 90 days. Remote unmoderated is fine for most content types. For complex B2B domains, I prefer moderated sessions because the think-aloud protocol surfaces vocabulary that participants don't encode in their group labels.

Step 4: Analyze for vocabulary consensus, not structural consensus. Card sort analysis typically focuses on which items ended up together — agreement on co-location. I'm equally interested in the label vocabulary. What words did people use to name their groups? What words appeared across multiple participants' labels? That's the vocabulary that belongs in URLs and page titles.

Step 5: Translate card sort output through CREST. For each proposed node, score Crawlability, Relevance-signal density, Entity alignment, Snippet surface area, Temporal stability. Adjust structure to maximize passes. Document nodes that fail and the specific adjustment made.

Step 6: Run closed card sort to validate the adjusted structure. Show 15–20 participants the proposed structure and have them place a sample of content items into it. Measure agreement. Items with agreement under 60% usually indicate either a label problem (fix the name) or a structural problem (the node boundaries don't match mental models). Distinguish between them before assuming the fix is editorial.

Step 7: Build the URL hierarchy and content specification document simultaneously. Every taxonomy node should have a minimum content spec before it goes to development: minimum word count, required schema types, required content elements (comparison table, specification list, editorial section, whatever is appropriate). Nodes without content specs become thin pages. Thin pages fail Relevance-signal density and usually fail Snippet surface area. They are not worth indexing.

Step 8: Define review cadence and triggers. Document which nodes have been flagged for Temporal Stability risk and what the review trigger is. Calendar the review. This is operational process, not design, but it's the part that prevents the taxonomy from degrading invisibly.

# Taxonomy node specification template

Node name: [Name as it appears to users]
URL: /[path-segment]/
CREST scores: C: pass/fail | R: pass/fail | E: pass/fail | S: pass/fail | T: pass/review
Entity mapping: [Wikidata entity ID or "no direct mapping — parent concept: X"]
Schema type: [CollectionPage / ItemList / FAQPage / etc.]
sameAs: [URL if applicable]
Min content spec:
  - Editorial introduction: 200 words min
  - Required elements: [specification table / comparison grid / buying guide section / etc.]
  - Required structured data: [list]
Review trigger: [condition] — [review date or "stable"]
Internal link requirement: [minimum inbound links from what page types]

This document goes to the content team, the developer, and the client before a single page is built. The conversation it creates — usually about whether the content spec is realistic given resource constraints — is the most valuable part of the process. Better to have that argument before the taxonomy is implemented than after.

What This Costs in Real Terms

The honest answer: doing taxonomy design this way costs roughly 3–4x more time than the UX-only approach. A proper card sorting study with correct recruitment, analysis, and CREST evaluation for a site with 50–80 taxonomy nodes takes me about 6 weeks of part-time work. A UX team doing just the card sort and structural design would do it in 2–3 weeks.

That cost calculation changes when you compare it to the cost of rebuilding a taxonomy that was done without the SEO evaluation layer. The B2B tools client's restructure — triggered by taxonomy-related traffic losses — took 4 months and cost considerably more than a properly scoped initial project would have.

There's an external reference worth sitting with here: Nielsen Norman Group's card sorting guidance is excellent on methodology but is explicitly UX-framed and doesn't address crawl or entity-graph considerations. Read it for method. Apply CREST after.

The other thing that's changed since AI Overviews became a significant traffic factor: the cost of getting taxonomy wrong is higher than it was. A badly structured taxonomy in 2022 might cost you rankings on competitive head terms. A badly structured taxonomy in 2026 can cost you AI Overview citations across hundreds of mid-tail queries simultaneously, because the entity-graph alignment that drives snippet extraction is a structural property of the whole site, not an individual page property. You fix it by fixing the taxonomy. That's not quick work.

If you're auditing a site where AI Overview appearances have dropped since mid-2025 without a corresponding quality change in the content, check the taxonomy before you check the content. In my experience — three projects where this was the presenting problem — the structural explanation was more often correct than the content explanation.

For related structural work, the faceted navigation piece covers the overlap between taxonomy design and faceted filtering in detail. The crawl budget article has the log analysis methodology referenced in the crawl budget modeling step above. And if you're thinking about how taxonomy connects to site-wide information architecture, the large ecommerce IA piece and the media site IA piece cover the structural contexts where these taxonomy decisions play out at scale.


Filed under: taxonomy design, information architecture, card sorting, CREST framework, AI Overviews, entity graph. Last updated against live project data: May 2026.

YOUR READING CHECKLIST

Make the ideas stick.

Mark the sections you’ve worked through. Saved in this browser.

0 of 4 reviewed
Andrii Stanetskyi
ABOUT THE AUTHOR

Andrii Stanetskyi

Head of SEO / Technical SEO Lead based in Tallinn, Estonia. Technical architecture, enterprise eCommerce, Python automation, and AI-assisted workflows.

More about Andrii ↗
LET’S FIND THE REAL BOTTLENECK

A clearer picture.
A practical next step.

Get a focused SEO audit or a consultation on your next technical decision. We’ll agree on the scope and fee before any work begins.

01 / Diagnose02 / Prioritize03 / Plan
How can I help?

Scope and fee agreed before any work begins.

Choose your language

Explore SEO services in 26 languages. Journal articles retain their original language.

ENEnglish↗DEDeutsch↗FRFrançais↗ESEspañol↗ITItaliano↗PTPortuguês↗NLNederlands↗PLPolski↗SVSvenska↗DADansk↗FISuomi↗NONorsk↗ETEesti↗LVLatviešu↗LTLietuvių↗CSČeština↗RORomână↗HUMagyar↗ELΕλληνικά↗BGБългарски↗HRHrvatski↗SKSlovenčina↗SLSlovenščina↗RUРусский↗UKУкраїнська↗TRTürkçe↗
LET’S WORK ON YOUR WEBSITE
A CLEAR NEXT STEP

Let’s talk
about your site.

A focused SEO audit or a conversation about a specific challenge. Tell me where you are and what you want to change.

Andrii Stanetskyi
Andrii StanetskyiHead of SEO / Technical SEO Lead
[email protected] ↗
How can I help?

Scope and fee agreed before any work begins.