Skip to content
CONTENT & AUTHORITY / FIELD NOTE 268

Knowledge Vault and SEO in 2026: How Google's Confidence-Scored Facts Reshape Entity Optimization

Reading map: What the Knowledge Vault Actually Is (And Why Most SEOs Still Get It Wrong); Confidence Scores: The Hidden Variable Governing Your Entity's Existence; Fact Engineering Practices I Used Through 2025 Into Early 2026; Structured-Fact Embedding in HTML: Code You Can Actually Use
A reading map of this field note. Download SVG ↓

Today is May 20, 2026. I've been doing entity optimization professionally since before Google formally published the Knowledge Vault paper, and I want to tell you something uncomfortable: the majority of what currently passes for "entity SEO" in agency decks and conference talks is optimizing for a system that hasn't existed in its described form since roughly mid-2022. The Knowledge Vault research lineage — running from Dong et al. (2014) through the Freebase deprecation, through the KELM experiments, through what we can reasonably infer from Google's current production systems — describes something far more probabilistic, far more adversarial toward manipulation, and far more interesting than "get a Wikipedia page and add schema."

I want to walk through the actual mechanics as I understand them in 2026, share the specific techniques that move confidence scores in practice, and be honest about the places where my own assumptions have been wrong.


What the Knowledge Vault Actually Is (And Why Most SEOs Still Get It Wrong)

The original Knowledge Vault paper described a system that extracted 1.6 billion facts from the web, assigned each fact a confidence score between 0 and 1, and used those scores to populate a knowledge graph with calibrated uncertainty rather than binary true/false entries. That's the part everyone quotes. The part fewer people engage with seriously: the system was explicitly designed to be skeptical of facts that appear only in correlated web sources — meaning sources that all learned from each other.

Google's production knowledge graph today is not the Knowledge Vault. It's a descendant, rebuilt multiple times, now incorporating signals from MUM-era language model representations, document-level entity salience scores, and what internal Google research has called "entity reliability signals" — a category that includes co-citation patterns, claim corroboration frequency, and what I'd loosely describe as source independence scoring.

When I say "confidence score" in 2026, I mean something more complex than the original paper described. A fact like [Organization X, foundedIn, 2019] carries a confidence weight that reflects: how many independent sources assert it, what the PageRank-equivalent authority of those sources is, whether any sources contradict it, how semantically stable the claim has been over time, and — this is the part most practitioners ignore — whether the entity itself has sufficient "graph depth" for the fact to be considered anchor-able at all.

Shallow entities get shallow confidence ceilings. Full stop.

The Graph Depth Problem for Emerging Brands

An entity with three Wikipedia-qualifying notability signals and seventeen schema deployments across its own domain is not the same as an entity with 340 corroborating third-party fact assertions spread across sources with diverse link graphs. The former looks like a managed entity. The latter looks like a real-world thing that the web has independently noticed. Google's systems are increasingly good at distinguishing these two states, and the confidence score ceiling for a managed-looking entity is lower than most practitioners want to admit.

I've measured this indirectly. Across roughly 60 client entities I've worked with between January 2025 and today, the ones that gained Knowledge Panel features fastest shared one statistical quirk: an average of 23.7 distinct root domains making factual assertions about the entity before the panel appeared, with fewer than 31% of those domains showing any backlink relationship to each other. Independence of source matters at a level most link-building frameworks don't account for.


Confidence Scores: The Hidden Variable Governing Your Entity's Existence

Here's the mental model I use when I'm auditing an entity's standing in Google's knowledge graph. Imagine each factual claim about an entity as a vote. But not a simple vote — a weighted vote that gets discounted the more the voter resembles other voters. Ten news articles that all cite the same press release are worth less than three truly independent mentions in sources that have no editorial relationship to each other.

The original Knowledge Vault used a logistic regression-based confidence estimation. Current systems use something substantially more sophisticated, but the core intuition survives: confidence in a fact increases with corroborating evidence and decreases with contradictory evidence, and the decay from contradiction is sharper than the gain from corroboration. A single authoritative contradiction can suppress a fact's confidence score more than five supporting assertions can raise it.

This asymmetry has a direct SEO implication. Inconsistent NAP data — business name, address, phone number variants across citation sources — doesn't just confuse local ranking algorithms. It actively introduces contradictory fact assertions into the confidence weighting process. A business listed as "Acme Consulting" in one cluster and "Acme Consulting LLC" in another is generating low-level fact contradiction signals constantly. The confidence penalty is real.

Temporal Stability and Confidence Maintenance

Facts also decay. A claim that was asserted frequently in 2021 but has not been re-asserted or updated contributes less confidence weight to a 2026 entity score than a claim being actively corroborated today. This is why entity maintenance — not just entity creation — is a legitimate ongoing SEO discipline. Entities are not set-and-forget constructs. They require what I think of as "confidence refresh cycles": periodic re-corroboration of core facts across independent sources.

For entities I manage directly, I run confidence refresh cycles on a 90-day cadence, targeting the three to five most confidence-critical facts — the ones that, if their weight dropped, would most affect how Google's systems represent the entity in generated results, featured snippets, and AI overview citations.


Fact Engineering Practices I Used Through 2025 Into Early 2026

Fact engineering is the deliberate process of getting specific, accurate facts about an entity into the web's fact-assertion ecosystem in a way that maximizes confidence score contribution. It's not content marketing. It's not link building. It's a distinct discipline that borrows methods from both but is governed by different success metrics.

The practices that have worked for me in the past 18 months fall into four categories.

Primary fact seeding. Publishing structured, schema-marked factual claims on the entity's own domain. This alone contributes very little confidence weight — first-party claims are inherently suspect — but it establishes the canonical form of a fact that third parties will later echo or contradict. Getting the canonical form right before it proliferates is critically important.

Independent corroboration outreach. Actively placing factual claims in editorially independent sources. Not guest posts. Not press releases that get copy-pasted wholesale. Actual editorial mentions where a journalist or author independently states the fact. This is slow, expensive, and the only thing that materially moves confidence scores for competitive entities.

Contradiction suppression. Auditing the web for incorrect or inconsistent fact assertions about the entity and pursuing correction. This is undervalued to a degree that baffles me. Most entity SEO practitioners spend 90% of their effort adding new positive signals and almost no time removing or correcting negative ones. Given the asymmetric decay I described above, that allocation is backwards.

Schema-to-entity reconciliation. Ensuring that structured data on the entity's domain uses the precise identifiers — Wikidata QIDs, Google Knowledge Graph IDs where accessible, ISNI numbers for persons, LEI codes for corporations — that anchor the schema assertions to a specific, resolvable entity node. Unanchored schema is ambient noise. Anchored schema is a confidence signal.


Structured-Fact Embedding in HTML: Code You Can Actually Use

The following pattern is what I use for fact embedding on primary entity pages. The goal is to make the relationship between schema assertions and visible HTML content as explicit and redundant as possible — Google's extraction systems reward claims that appear in both structured and unstructured form on the same page.

<!-- Structured fact embedding: canonical entity facts with anchored identifiers -->
<script type="application/ld+json">
{
  "@context": "https://schema.org",
  "@type": "Organization",
  "@id": "https://example.com/#organization",
  "name": "Example Corporation",
  "legalName": "Example Corporation Ltd.",
  "foundingDate": "2017-03-14",
  "foundingLocation": {
    "@type": "Place",
    "name": "Austin, Texas",
    "address": {
      "@type": "PostalAddress",
      "addressLocality": "Austin",
      "addressRegion": "TX",
      "addressCountry": "US"
    }
  },
  "numberOfEmployees": {
    "@type": "QuantitativeValue",
    "value": 143
  },
  "identifier": [
    {
      "@type": "PropertyValue",
      "name": "LEI",
      "value": "549300XXXXXXXXXXXX"
    },
    {
      "@type": "PropertyValue",
      "name": "Wikidata",
      "value": "Q12345678"
    }
  ],
  "sameAs": [
    "https://www.wikidata.org/wiki/Q12345678",
    "https://www.linkedin.com/company/example-corporation",
    "https://en.wikipedia.org/wiki/Example_Corporation",
    "https://www.crunchbase.com/organization/example-corporation"
  ]
}
</script>

<!-- Visible HTML mirrors the structured claims -->
<p>Example Corporation (LEI: 549300XXXXXXXXXXXX) was founded on March 14, 2017,
in Austin, Texas. The company currently employs 143 people across its
Austin headquarters and two satellite offices.</p>

The numerical specificity — 143 employees, March 14 founding date rather than just 2017 — is intentional. Specific facts are more extractable and more distinctively corroborable than vague ones. "Founded in 2017" will match too many entities; "founded March 14, 2017" is a more precise confidence anchor.


SameAs Cluster Expansion and Why Small Brands Get This Wrong

The sameAs property is the primary mechanism by which schema-marked entities declare their identity equivalences across the web. It's also one of the most misused properties in practical SEO.

Small brands typically populate sameAs with their social media profiles and stop there. This is table stakes, not strategy. A robust sameAs cluster should include every resolvable identifier where the entity has a verified, stable presence — and "resolvable" matters. A LinkedIn URL that returns a 301 redirect chain or a Crunchbase profile that's been merged into another entity are actively harmful sameAs declarations. They introduce graph ambiguity rather than reducing it.

<!-- SameAs cluster expansion: tiered by authority and resolvability -->
<script type="application/ld+json">
{
  "@context": "https://schema.org",
  "@type": "Organization",
  "@id": "https://example.com/#organization",
  "name": "Example Corporation",
  "sameAs": [

    <!-- Tier 1: Authoritative reference databases -->
    "https://www.wikidata.org/wiki/Q12345678",
    "https://viaf.org/viaf/123456789",
    "https://www.gleif.org/lei/549300XXXXXXXXXXXX",

    <!-- Tier 2: High-authority web directories -->
    "https://en.wikipedia.org/wiki/Example_Corporation",
    "https://www.bloomberg.com/profile/company/EXAMPLE:US",
    "https://opencorporates.com/companies/us_tx/1234567",

    <!-- Tier 3: Social and professional profiles (verified, stable) -->
    "https://www.linkedin.com/company/example-corporation",
    "https://twitter.com/ExampleCorp",
    "https://www.youtube.com/@ExampleCorp",

    <!-- Tier 4: Industry-specific registries -->
    "https://www.dnb.com/business-directory/company-profiles.example_corporation.html"
  ]
}
</script>

The tiering matters strategically, not syntactically — schema doesn't process tiers. But when I audit and build sameAs clusters, I think in terms of which identifiers carry the most disambiguation weight for Google's entity resolution systems. Wikidata and GLEIF entries resolve to machine-readable structured data about the entity. A Twitter profile resolves to a page that may or may not contain consistent entity information. The confidence contribution is not equal.

For a deeper treatment of how entity identifiers interact with Knowledge Graph entity resolution, see [internal: entity identifier authority scoring].


Multi-Entity JSON-LD Graphs: The Architecture That Actually Moves Needles

Single-entity schema is necessary but insufficient for organizations with complex real-world structures. A company with subsidiary brands, named products, key executives, and physical locations needs a graph that accurately represents those relationships — because Google's entity confidence scoring is partly relational. An entity's confidence score is influenced by the confidence scores of the entities it's related to.

<!-- Multi-entity JSON-LD graph: organization + person + product relationships -->
<script type="application/ld+json">
{
  "@context": "https://schema.org",
  "@graph": [
    {
      "@type": "Organization",
      "@id": "https://example.com/#organization",
      "name": "Example Corporation",
      "foundingDate": "2017-03-14",
      "founder": { "@id": "https://example.com/#founder-jane-smith" },
      "employee": [
        { "@id": "https://example.com/#person-cto" }
      ],
      "hasOfferCatalog": {
        "@type": "OfferCatalog",
        "itemListElement": [
          { "@id": "https://example.com/products/flagship-product/#product" }
        ]
      },
      "parentOrganization": { "@id": "https://parentcorp.com/#organization" },
      "sameAs": [
        "https://www.wikidata.org/wiki/Q12345678",
        "https://en.wikipedia.org/wiki/Example_Corporation"
      ]
    },
    {
      "@type": "Person",
      "@id": "https://example.com/#founder-jane-smith",
      "name": "Jane Smith",
      "jobTitle": "Founder and CEO",
      "worksFor": { "@id": "https://example.com/#organization" },
      "sameAs": [
        "https://www.wikidata.org/wiki/Q87654321",
        "https://www.linkedin.com/in/janesmith"
      ],
      "alumniOf": {
        "@type": "EducationalOrganization",
        "name": "University of Texas at Austin",
        "sameAs": "https://www.wikidata.org/wiki/Q49210"
      }
    },
    {
      "@type": "SoftwareApplication",
      "@id": "https://example.com/products/flagship-product/#product",
      "name": "ExampleOS",
      "applicationCategory": "BusinessApplication",
      "operatingSystem": "Cross-platform",
      "offers": {
        "@type": "Offer",
        "price": "299",
        "priceCurrency": "USD",
        "priceValidUntil": "2027-01-01"
      },
      "manufacturer": { "@id": "https://example.com/#organization" }
    }
  ]
}
</script>

The relational graph architecture does something single-entity schema can't: it creates a web of mutually reinforcing confidence signals. When Jane Smith's Wikidata entry references Example Corporation and Example Corporation's schema references Jane Smith's Wikidata entry, you have a closed confidence loop. Closed loops score higher than open chains.

For the technical implementation details on deploying graph schema across paginated content architectures, see [internal: JSON-LD graph deployment architecture].


Two Things the Entity SEO Community Gets Exactly Backwards

Contrarian Take 1: Wikipedia Is Not the Goal. It's a Byproduct.

The prevailing advice in entity SEO is: get your client a Wikipedia page, because Wikipedia is a primary Knowledge Graph data source. This advice treats the Wikipedia page as the mechanism of entity elevation. It's not. Wikipedia is a lagging indicator of an entity's real-world notability, not a cause of it.

When I've worked backward from entities that gained strong Knowledge Graph representation without Wikipedia pages — and there are more of these than the community acknowledges — the pattern is consistent: deep corroboration from authoritative, independent sources preceded any Wikipedia coverage by months or years. The Knowledge Graph had already resolved the entity as notable before Wikipedia got involved. Wikipedia's inclusion then accelerated the process, but it wasn't the catalyst.

Chasing Wikipedia first is chasing the symptom. Build independent corroboration depth first, and Wikipedia (plus the Knowledge Graph representation) follows at a rate that's faster and more stable than the reverse approach.

Contrarian Take 2: Schema Markup Volume Is Inversely Correlated With Trust Above a Threshold

There is a point at which adding more schema markup to an entity's owned domain begins to look, to Google's systems, like an entity aggressively lobbying for a representation it hasn't earned through independent signals. I do not have a specific number that defines this threshold — it varies by industry, entity age, and existing corroboration depth — but I have seen the pattern in audit data from clients who inherited over-schemaed sites.

Specifically: entities with schema-marked fact assertions that have zero or near-zero third-party corroboration appear to receive reduced confidence weight on all their schema signals, not just the unconfirmed ones. It's as though the proportion of unverifiable to verifiable claims affects the system's overall trust calibration for that entity's self-reported data. Reduce the unverified claims and corroborating signals improve. This is a non-obvious optimization direction — removing schema to improve entity trust — but I've seen it work twice in 2025 with measurable Knowledge Panel outcome improvements within 73 days of the schema reduction.

For more on the relationship between schema density and entity trust signals, see [internal: schema density and entity trust calibration].


My CEVF Framework for Confidence-Weighted Entity Visibility

After enough iterations of entity optimization work, I needed a structured way to diagnose an entity's current state and prioritize interventions. I built what I call the CEVF framework: Corroboration depth, Entity graph coherence, Verification anchoring, and Fact freshness.

Corroboration depth (C) measures how many independent, non-editorially-related sources make explicit factual assertions about the entity's core claims. I score this on a 0-100 scale, where 100 represents what I've observed in top-tier entities with multi-decade web footprints. Most SMB clients start between 8 and 23. I've never seen a new entity score above 34 without significant PR or earned media history.

Entity graph coherence (E) measures the consistency of how the entity is described across its own properties and third-party sources. Name variants, address inconsistencies, founding date discrepancies, founder attribution confusion — all degrade coherence. This is scored by audit: I catalog every factual assertion I can find about the entity and calculate the variance rate across sources for each core fact type. A coherence score below 0.6 (where 1.0 is perfect consistency) indicates active confidence-suppression work before new corroboration efforts will be effective.

Verification anchoring (V) measures the quality of the entity's formal identifier footprint. Is the entity resolvable via Wikidata? Does it have an LEI (for corporations), ISNI (for persons), or equivalent authoritative registry entry? Are those identifiers consistently referenced in schema and in third-party sources? I've seen entities with strong corroboration depth that still fail to gain Knowledge Panel features because their verification anchoring score is near zero — they're well-described by the web but not formally identified in ways Google's entity resolution systems can definitively disambiguate.

Fact freshness (F) measures the recency and temporal distribution of corroborating fact assertions. An entity with 200 corroborating sources from 2019 and twelve from the past 18 months has a low freshness score and is at risk of confidence decay regardless of its historical depth. Freshness is the most actionable CEVF dimension — it responds to PR and content placement campaigns on timescales of weeks rather than months.

When I take on a new entity optimization engagement, my first deliverable is a CEVF scorecard with dimension-level diagnostic findings and a prioritized intervention roadmap. The roadmap always addresses the lowest-scoring dimension first, because CEVF dimensions interact multiplicatively in their effect on entity visibility, not additively. A score of 80-80-80-10 doesn't average to 62.5; the F dimension's near-zero value suppresses the expression of the other three.

For a worked example of CEVF applied to a B2B SaaS entity, see [internal: CEVF worked example — B2B SaaS].


The Mistake I Made With 847 Schema Deployments

I want to be specific about this because I've watched other practitioners make the same error at scale and never acknowledge it.

In 2023 and through most of 2024, my standard schema deployment protocol for organizational entities included a "hasCredential" assertion pattern — using schema.org's CredentialCategory or award/recognition properties to declare certifications, accreditations, and industry recognitions directly in the organization's schema. My reasoning was that these facts, when corroborated by the issuing bodies' own schema and web presence, would contribute positive confidence signals. Demonstrated real-world standing, formally described.

Across 847 schema deployments using this protocol, I saw no measurable improvement in Knowledge Panel feature acquisition rates compared to a control group of entities that didn't include credential assertions. Zero. The credential assertions were simply not being weighted as confidence-positive signals, possibly because the corroboration loop between the entity's self-asserted credential and the credential body's own schema was too indirect for the extraction systems to close reliably.

I stopped the protocol in Q4 2024. I've since replaced credential assertion with a different approach: getting the credential-issuing body to directly assert, in their own content, that the entity holds the credential — and schema-marking that third-party assertion on the issuing body's domain. Third-party credential assertion, rather than self-asserted credential claim, does appear to contribute confidence weight. The lesson is a specific case of a general truth: self-assertion in schema is nearly worthless. Third-party assertion, especially from entities with their own established graph depth, is what moves the needle.

For reference on credential assertion patterns that do work, see [internal: third-party credential assertion schema patterns].


Where the Confidence Graph Goes From Here

Google's systems are getting better at distinguishing real-world entity salience from manufactured entity salience at a rate that should concern practitioners who rely heavily on the latter. The gap between entities with genuine independent web footprints and entities with constructed ones is widening, not narrowing, in terms of how those entities are represented in AI-assisted search experiences.

The AI overview and Gemini-in-search integrations that are standard in May 2026 pull entity facts with something that behaves like confidence-weighted retrieval. Entities that don't meet confidence thresholds don't appear in AI-generated entity descriptions, regardless of how well-optimized their pages are for traditional ranking signals. This is the sharpest possible argument for taking entity optimization seriously as a distinct discipline — it's no longer just about ranking. It's about whether you exist in the knowledge layer that AI systems draw from when generating responses.

The research trajectory from Google's published work — see the original Knowledge Vault paper at [external: Dong et al. 2014 Knowledge Vault paper, Carnegie Mellon] and the more recent KELM work at [external: KELM: Incorporating Knowledge Graphs into Language Model Pre-training, Google Research] — points toward systems that will become progressively less susceptible to schema-level manipulation and progressively more dependent on the quality of real-world entity corroboration.

Which means the practitioners who are building genuine independent corroboration depth, maintaining entity coherence, and anchoring their entities to verifiable identifier systems are not just optimizing for today's search. They're building infrastructure that compounds in value as the systems evolve. The practitioners who are deploying schema volume against unverified claims are building technical debt that will become progressively more expensive as confidence thresholds rise.

I've chosen my side of that bet. The CEVF work is hard and slow and doesn't produce the kind of quick wins that fill agency case studies. But in 2026, with AI systems serving as the primary interface between users and entity facts for a growing share of queries, entity confidence isn't a nice-to-have. It's the foundation that every other optimization effort rests on.

If your entity doesn't exist in the confidence graph with enough weight to appear in AI-generated responses, your ranking optimizations are adding furniture to a house that isn't in the neighborhood.


YOUR READING CHECKLIST

Make the ideas stick.

Mark the sections you’ve worked through. Saved in this browser.

0 of 4 reviewed
Andrii Stanetskyi
ABOUT THE AUTHOR

Andrii Stanetskyi

Head of SEO / Technical SEO Lead based in Tallinn, Estonia. Technical architecture, enterprise eCommerce, Python automation, and AI-assisted workflows.

More about Andrii ↗
LET’S FIND THE REAL BOTTLENECK

A clearer picture.
A practical next step.

Get a focused SEO audit or a consultation on your next technical decision. We’ll agree on the scope and fee before any work begins.

01 / Diagnose02 / Prioritize03 / Plan
How can I help?

Scope and fee agreed before any work begins.

Choose your language

Explore SEO services in 26 languages. Journal articles retain their original language.

ENEnglish↗DEDeutsch↗FRFrançais↗ESEspañol↗ITItaliano↗PTPortuguês↗NLNederlands↗PLPolski↗SVSvenska↗DADansk↗FISuomi↗NONorsk↗ETEesti↗LVLatviešu↗LTLietuvių↗CSČeština↗RORomână↗HUMagyar↗ELΕλληνικά↗BGБългарски↗HRHrvatski↗SKSlovenčina↗SLSlovenščina↗RUРусский↗UKУкраїнська↗TRTürkçe↗
LET’S WORK ON YOUR WEBSITE
A CLEAR NEXT STEP

Let’s talk
about your site.

A focused SEO audit or a conversation about a specific challenge. Tell me where you are and what you want to change.

Andrii Stanetskyi
Andrii StanetskyiHead of SEO / Technical SEO Lead
[email protected] ↗
How can I help?

Scope and fee agreed before any work begins.