Two Years In: Separating Signal from SEO Twitter
The May 2024 leak of Google's internal API documentation is background noise now. Not ancient history — background noise. The kind of thing you reference in a client kickoff call without explaining what it was, the way you'd mention Penguin or Panda without a footnote. Two years out, the initial wave of hot takes has mostly settled. What's left is more interesting: a set of real behavioral changes, a handful of confirmed-by-behavior attributes that actually map to auditable signals, and at least one mistake I made publicly that I've had to walk back.
I want to write this as a practitioner account, not a recap. You can find recaps. What you can't find easily is someone who ran audits across a range of sites in 2025 and early 2026 and actually connected what they found to the attributes that surfaced in the leak. That's what this is.
Quick orientation for anyone who needs it: in May 2024, roughly 2,500 pages of internal Google API documentation were leaked on GitHub. The documents described modules, attributes, and data types that appeared to be part of Google's content warehouse and ranking systems. Attributes like siteAuthority, navBoost, siteDiversity, chromeInTotal, and dozens more. Google confirmed the documents were authentic but disputed specific interpretations. The SEO community collectively lost its mind for about six weeks, then largely moved on.
Most moved on to the wrong conclusions. Either "nothing matters, we can't know" (fatalistic) or "now we know everything, optimize for all of it" (delusional). The useful space is narrower: certain attributes correlate well enough with auditable, actionable signals that they've changed how I scope and prioritize work. That's the only claim I'm willing to defend.
siteAuthority and siteDiversity — What I Actually See in the Data
siteAuthority was the one that generated the most initial excitement. It appeared to be a site-level authority score, which immediately activated every Domain Authority–brained SEO in the industry. The reaction was predictable: "Google has a site-level authority score, therefore DA is validated, therefore link building to boost DA is good." That chain of reasoning is wrong at almost every step, but it's sticky because it ends where people wanted to end up.
What I've come to believe — and what the audit data from a mid-sized health information site (roughly 3.2 million pages, audited in Q3 2025) reinforced — is that siteAuthority is more likely a weighting factor applied contextually than a stable numeric score you can move with a link campaign. The site in question had a strong backlink profile, a high Ahrefs DR, and was still experiencing selective ranking suppression on its newer subsections. Old content ranked well. New content on adjacent topics, even when technically better, was getting stuck in the 15–40 SERP range for months before any movement.
The pattern: established site sections with long clickstream histories ranked; newer sections without that history didn't, regardless of backlink profiles. That points toward siteAuthority being evaluated per section or per content cluster, not sitewide. Which has real audit implications.
siteDiversity is less talked about and more interesting. The attribute in the leaked documents appears to be something Google uses to prevent a single domain from dominating results — a deduplication or cap mechanism. I've seen this play out in SERPs for some competitive finance queries where the same domain appears twice and never three times, even when their third-ranking page would otherwise outperform competitors. That's not new behavior, but the attribute name gives us a more precise thing to observe. For sites trying to rank multiple pages for the same broad topic, siteDiversity probably acts as a ceiling. Knowing that, the audit question shifts: instead of "why isn't this page ranking?" the question becomes "which page in this cluster is getting the site's one allowed slot, and is it the right one?"
navBoost: The Attribute Everyone Misread
The navBoost attribute is where I see the most persistent misreading, two years on. The majority interpretation in current SEO content is that navBoost is a signal tied to navigational queries — branded searches, people looking for a specific site. Therefore, the argument goes, you should build branded search volume to boost your navBoost score.
That interpretation has a surface plausibility that makes it resilient. But the leaked documents actually describe navBoost in the context of click and navigation data more broadly. It's tied to how often users navigate to a URL after a search interaction — not exclusively branded search. Long-click behavior. Return visits. The difference matters because it means navBoost is downstream of content quality and user satisfaction, not upstream of it. You cannot run a brand awareness campaign and expect navBoost to move. You make content worth returning to, and navigation follows.
I audited a B2B SaaS site in January 2026 — 47,000 pages, recently migrated from a legacy CMS — and one of the things we tracked was the delta between pages with strong branded click patterns in GSC versus pages that showed repeat-visit behavior in their GA4 data without branded search driving it. The pages with organic repeat-visit behavior (users coming back via direct or organic non-branded) outranked their technically equivalent counterparts by a margin that couldn't be explained by backlinks or on-page optimization alone. navBoost, or something very close to what that attribute describes, is real. And it's earned, not engineered.
My DPAF Framework for Post-Leak Auditing
After spending most of 2025 trying to reconcile what the leaked attributes described with what I was actually seeing in GSC, GA4, and crawl data, I formalized the way I now scope and prioritize technical and content audits. I call it DPAF: Diversity-Positioning-Authority-Flow.
D — Diversity. Before recommending new content, map which cluster owns the site's one or two SERP slots per topic area. Stop creating new pages that compete internally for the same slot. Start identifying which page deserves the slot and pointing equity there.
P — Positioning. Review actual SERP positions with click-through curves overlaid. A page at position 4 with a CTR matching a position-2 average is a navBoost candidate — something in user behavior already signals it should rank higher. These pages get priority for refresh, not creation.
A — Authority. Map authority at the section level, not the domain level. Which sections of the site have the longest content history, the densest internal link structures, and the most stable ranking patterns? Those sections are likely benefiting from elevated section-level siteAuthority. Expansion efforts should grow from those sections outward, not parachute into new territory.
F — Flow. Track link equity flow at a granular level — not just "is the homepage linking to this section?" but whether topically adjacent pages are cross-linking in ways that reinforce cluster coherence. Thin internal link flow into new content is one of the most consistent predictors of slow indexing and ranking lag I see in large site audits right now.
DPAF is not a magic framework. It's a forcing function for sequencing work in an order the client environment can actually support. It's helped me stop writing audit recommendations that read like a wishlist and start writing ones that have a clear operational sequence.
The Chrome Data Question Is Settled, and I Was Wrong
Here's the mistake. When the leak dropped, I wrote a fairly confident thread arguing that chromeInTotal and related Chrome usage attributes were probably legacy or internal-only signals — not active ranking factors. My reasoning was that using browser data at scale would create obvious GDPR and privacy issues that Google would avoid. I thought the attributes described aspirational infrastructure rather than live systems.
I was wrong. Not catastrophically wrong, but wrong enough to have corrected the record with the people I work with.
The behavioral evidence from 2025 is fairly clear: pages with strong Chrome bookmark and revisit patterns — observable as a proxy through direct traffic spikes in GA4 — do show ranking advantages that persist even through core updates. A travel publisher I worked with through Q4 2025 had a subset of destination guides that consistently retained positions through three significant algorithm updates that hammered most of their category. Those pages had anomalously high direct traffic ratios. They weren't the most-linked pages. They weren't the most recently updated. But users kept coming back to them directly. Whether that's Chrome data specifically or a cluster of correlated signals, the behavioral fingerprint matches what chromeInTotal and its sibling attributes describe.
The correction I'd make to my original position: the privacy concern is real, but Google has mechanisms — aggregated, de-identified, consented-through-Chrome-sync data — that let them use browser behavior signals without individual tracking. I underweighted that. Don't make the same mistake in your audits.
Weird Numbers from 2025 Audits
A few data points that don't fit neatly into a clean narrative but are relevant to how I'm thinking about the leak's impact on practice.
On a 14,000-page e-commerce site audited in March 2025: pages with an average session duration above 3 minutes ranked 2.4 positions higher on average than topically equivalent pages on the same site with sessions under 90 seconds. That's not a controlled experiment. But the within-site comparison controls for domain authority, and the magnitude of the difference is large enough to matter. Something in user engagement is being detected and weighted.
On a content site in the personal finance space, audited in July 2025: after a significant internal link restructuring that reduced the average number of outbound internal links per page from 22 to 8 (while increasing topical relevance of those links), the pages we restructured saw a median ranking improvement of 3.1 positions over 90 days. Pages we didn't touch moved 0.4 positions on average over the same period. The restructuring was designed explicitly to concentrate link equity flow in ways that align with what siteAuthority at the section level would suggest — routing equity into the pages with the established history and letting that authority gradient work.
One more: a regional news site where I tracked the velocity of indexation for new articles before and after implementing a denser internal linking protocol from the homepage and section fronts. Before: median time from publication to indexed in GSC was 4.8 hours. After: 1.9 hours. The change was in link structure only. No XML sitemap changes, no IndexNow, no manual fetch requests. The crawl priority signal that siteAuthority and related attributes likely feed into seems to affect not just rankings but crawl scheduling.
Two Takes That Go Against What Most Are Publishing Right Now
The Link-Building-as-authority-restoration thesis is mostly wrong
The mainstream 2026 SEO consensus, at least as I read it across the major publications and conference talks, is that the leak validated the importance of link building by confirming that site-level authority signals exist. The logic: Google has a site authority attribute, therefore links that build authority matter, therefore your link building program is justified.
I think this is wrong in a specific and important way. The leaked attributes suggest authority is evaluated at a granular level — per section, possibly per cluster — and that it's built through behavioral and historical signals, not link acquisition alone. A link campaign that raises your DR from 52 to 61 is not moving a section-level authority signal. The section-level signal is built by producing content in that section over time, earning repeat visits in that section, and having that section's content linked and navigated to internally and externally in a consistent pattern over months and years. You cannot buy your way into that with a 90-day link sprint.
The clients who are spending significant budget on link acquisition in 2026 to recover from HCU or core update drops — most of them are misdiagnosing the problem. The drop is usually a siteAuthority or navBoost deficit, both of which respond to content behavior, not to backlink volume.
The "optimize for AI Overviews" push is cannibalizing real ranking work
The second contrarian take is about resource allocation. The current dominant conversation in SEO is GEO — Generative Engine Optimization, optimizing to be cited in AI Overviews and other LLM-driven results. I don't think that's wrong as a long-term consideration. But I do think it's pulling audit and strategy resources away from the core organic ranking work that the leaked attributes point toward as still fundamental.
AI Overview citations are heavily weighted toward sites that already rank well in traditional organic results. Chasing AI citations without fixing the underlying organic signals is a loop that doesn't close. The leaked documentation doesn't describe a world where content optimized for snippet extraction outranks content with strong behavioral engagement signals. It describes the opposite: behavioral signals feeding into authority calculations that then determine what gets surfaced in any format, AI or otherwise.
The sites I see gaining in 2026 are the ones that spent 2025 improving content quality, internal link coherence, and page-level engagement — not the ones that rewrote everything for structured snippet formats. Correlation, not causation, but the pattern is consistent enough to push back on the rush toward AI-first optimization.
Detecting Proxy Signals: SQL and Python I Actually Run
The leaked attributes aren't queryable directly. But proxy signals for most of them are auditable with data you already have in BigQuery (GA4 + GSC export) or with Python crawl analysis. Here's what I actually run.
Detecting navBoost proxy signals via BigQuery
-- Identify pages with high return-visit ratio (navBoost proxy)
-- Requires GA4 BigQuery export linked to GSC performance data
WITH page_sessions AS (
SELECT
page_location,
COUNT(*) AS total_sessions,
COUNTIF(session_number > 1) AS return_sessions,
COUNTIF(medium = 'organic') AS organic_sessions,
AVG(session_duration) AS avg_duration
FROM your_project.analytics_XXXXXXX.events_*
WHERE _TABLE_SUFFIX BETWEEN '20250101' AND '20260430'
AND event_name = 'session_start'
GROUP BY page_location
),
gsc_performance AS (
SELECT
url,
AVG(position) AS avg_position,
SUM(clicks) AS total_clicks,
SUM(impressions) AS total_impressions
FROM your_project.searchconsole.searchdata_url_impressions
WHERE data_date BETWEEN '2025-01-01' AND '2026-04-30'
GROUP BY url
)
SELECT
p.page_location,
p.total_sessions,
ROUND(p.return_sessions / NULLIF(p.total_sessions, 0), 3) AS return_visit_ratio,
p.avg_duration,
g.avg_position,
g.total_clicks
FROM page_sessions p
JOIN gsc_performance g ON p.page_location = g.url
WHERE p.total_sessions > 100
AND g.avg_position BETWEEN 1 AND 20
ORDER BY return_visit_ratio DESC
LIMIT 200;
Pages with a return_visit_ratio above 0.35 and avg_position above 8 are candidates for a navBoost gap — user behavior signals they should rank higher than they do. These get priority in content refresh queues.
Detecting siteAuthority gradient by section — Python
import pandas as pd
import numpy as np
from urllib.parse import urlparse
# Load GSC data export (url, position, clicks, impressions, date)
gsc_df = pd.read_csv('gsc_export_2025_2026.csv', parse_dates=['date'])
# Extract section from URL path (first segment after domain)
def extract_section(url):
path = urlparse(url).path
parts = [p for p in path.split('/') if p]
return parts[0] if parts else 'root'
gsc_df['section'] = gsc_df['url'].apply(extract_section)
gsc_df['year_month'] = gsc_df['date'].dt.to_period('M')
# Calculate section-level authority proxy:
# stable_rank_score = inverse of position variance over time
# (low variance = stable positions = likely strong siteAuthority at section level)
section_stats = gsc_df.groupby(['section', 'url']).agg(
avg_position=('position', 'mean'),
position_std=('position', 'std'),
total_clicks=('clicks', 'sum'),
months_active=('year_month', 'nunique')
).reset_index()
section_stats['stability_score'] = (
1 / (section_stats['position_std'].fillna(10) + 0.01)
) * section_stats['months_active']
section_authority = section_stats.groupby('section').agg(
avg_stability=('stability_score', 'mean'),
avg_position=('avg_position', 'mean'),
total_clicks=('total_clicks', 'sum'),
url_count=('url', 'count')
).reset_index().sort_values('avg_stability', ascending=False)
print(section_authority.head(20))
# Sections with high avg_stability and low avg_position are
# your established authority clusters (siteAuthority proxy = high)
# Sections with low avg_stability despite decent avg_position are
# vulnerable — likely floating on thin authority, at risk in core updates
The stability_score calculation is a rough proxy. What I'm looking for is sections where position variance is low over a long window — that's the behavioral signature of what siteAuthority likely captures at the section level. Compare those stable sections to sections with high position volatility and you can prioritize where new content is likely to take hold versus where you're going to fight for months.
Detecting siteDiversity cap in SERPs — Python with SERP data
import requests
from collections import defaultdict
# Example using any SERP API that returns full result sets
# Replace with your preferred provider endpoint
def analyze_domain_cap(keyword_list, api_key, domain):
domain_appearances = defaultdict(list)
for keyword in keyword_list:
response = requests.get(
'https://your-serp-api.example.com/search',
params={'q': keyword, 'num': 20, 'api_key': api_key}
)
results = response.json().get('organic_results', [])
positions = [
r['position'] for r in results
if domain in r.get('link', '')
]
domain_appearances[keyword] = positions
# Identify keywords where domain appears exactly once or twice
# even when you have multiple strong candidate URLs
capped_keywords = {
kw: pos for kw, pos in domain_appearances.items()
if len(pos) <= 2
}
return capped_keywords
# siteDiversity cap behavior: if your domain consistently
# appears once per SERP across a topic cluster, and you have
# 3+ strong candidate pages internally, you are hitting the cap.
# Solution: consolidate rather than expand.
What Changed in My Actual Practice
The concrete list is shorter than people expect, given how much noise the leak generated.
I now always include a section-level authority map in technical audits. Before the leak, I was doing domain-level authority assessment and page-level technical checks. The gap between those two levels — the section or cluster level — was something I was treating informally. Now it's a structured audit component. I look at ranking stability per section, cross-link density per section, and crawl frequency per section from log data when I can get it.
I've shifted how I frame content recommendations. Instead of "create X new pages on this topic," I'm more often saying "promote this existing page within the cluster, reduce competing pages, and create behavioral conditions for it to earn navBoost." That's a different kind of recommendation and a harder one to sell to clients who want to see deliverables in the form of new content. But it's more often correct.
I deprioritized link acquisition recommendations for recovery situations. This is the most contested change I've made, and the one that's generated the most pushback from clients and other consultants. When a site drops in a core update, the temptation — reinforced by years of industry conditioning — is to diagnose a link deficit. In most cases I'm seeing now, the diagnosis is behavioral and historical, not link-based. The remedy is different. Slower to show results. Harder to measure. But more often durable.
I've also started tracking what I call "authority orphans" — pages on strong domains that sit in sections or subsections with thin historical content and weak cross-linking. These pages have the domain's reputation but none of the section-level authority signal. They're the ones that bounce between page 2 and page 3 despite strong external backlinks. Identifying them, consolidating them into better-positioned sections, or building out the section around them — that's become a routine audit action item that didn't exist in my process before the leak.
The leak didn't give anyone a cheat code. What it gave practitioners who looked carefully was a vocabulary for patterns that were already visible in the data. The patterns were always there. The names just help me explain them to clients without sounding like I'm making things up.
If you want to go deeper on the technical side of section-level authority mapping, the BigQuery SEO patterns guide has the full dataset preparation steps. For how this intersects with crawl budget and log analysis, see the log analysis workflow. The behavioral signals discussed here connect directly to how HCU recovery works in practice — the section authority framing is particularly relevant there. On the content side, this informs topic authority as a ranking currency. And for how navBoost behavior ties into engagement-driven architecture decisions, the SEO-CRO integration piece gets into the overlap.
For primary source reading on the leaked attributes, the archived analysis at iPullRank's original breakdown remains the most thorough treatment of the raw documents, even if some interpretations have aged differently. The Search Engine Land contemporaneous coverage is useful for the historical record of what was claimed in May 2024.
Two years out, the leak is part of the baseline. Not a revelation anymore. Just another input into how the work gets done — which is exactly where it should be.
Frequently Asked Questions
What was actually confirmed from the 2024 Google API leak?
Google confirmed the leaked API documentation was authentic but disputed specific interpretations. The documents described internal modules and attributes including siteAuthority, navBoost, siteDiversity, and chromeInTotal. What those attributes do precisely in live systems has not been officially clarified, but behavioral evidence from 2025–2026 audits is consistent with several of the interpretations that emerged from careful reading of the documents.
Does Google use Chrome browsing data as a ranking signal?
The evidence from audit data in 2025 suggests yes, in some form. Pages with strong direct and return-visit traffic patterns — a proxy for what Chrome usage data would capture — show ranking advantages that persist through core updates. Google likely uses aggregated, de-identified, consent-based Chrome sync data rather than individual user tracking.
What is navBoost and how do you optimize for it?
navBoost appears to be a signal tied to how often users navigate to a URL after a search interaction, including long-click behavior and return visits. You cannot directly optimize for navBoost. It is earned by producing content that users find genuinely useful and return to. Proxy signals include return visit ratio in GA4, average session duration, and low bounce rates relative to topic competitors on the same domain.
Is siteAuthority the same as Domain Authority?
No. siteAuthority as described in the leaked documents appears to operate at a more granular level than a single domain-wide score. Audit evidence suggests it is evaluated per site section or content cluster, built through content history, behavioral engagement, and crawl patterns over time. Third-party domain authority metrics like Ahrefs DR or Moz DA do not map directly to this signal.
Should I change my link building strategy based on the leak?
If you are running a link campaign primarily to recover from a core update or HCU drop, the audit evidence from 2025–2026 suggests you are likely misdiagnosing the problem. Most drops of this kind reflect behavioral and historical authority deficits at the section level, not link deficits. Link acquisition does not directly move those signals. Content behavior, internal link coherence, and user engagement do.
