Skip to content
TECHNICAL SEO / FIELD NOTE 053

Reverse-Engineering Google's PageRank: What Still Applies in 2026

Reading map: The Original Algorithm: What the Math Actually Says; The Damping Factor and Its SEO Implications; Topic-Sensitive PageRank and TrustRank; The Reasonable Surfer Model Patent
A reading map of this field note. Download SVG ↓

The original PageRank paper by Brin and Page (1998) remains one of the most cited computer science papers in history. What is less understood is how far Google has moved from that original formulation—and equally, which mathematical properties of the original algorithm are so fundamental that they survive in some form regardless of what Google calls the current implementation. This article does not speculate about Google's internal systems. It examines what is mathematically verifiable, what is supported by Google patents (many annotated by the late Bill Slawski), and what practitioners have validated through large-scale experiments.

The Original Algorithm: What the Math Actually Says

The core equation from Brin and Page's 1998 paper "The Anatomy of a Large-Scale Hypertextual Web Search Engine" is:

PR(A) = (1 - d) + d * Σ (PR(T_i) / C(T_i))

Where PR(A) is the PageRank of page A, d is the damping factor (originally set to 0.85), T_i are pages linking to A, and C(T_i) is the count of outbound links from page T_i. The formula encodes the random surfer model: a user who randomly follows links, occasionally teleporting to a random page. The teleportation probability is (1-d), ensuring convergence and preventing rank sinks.

Two mathematically significant properties that are often glossed over:

Conservation of rank: In the simplified model, total PageRank in the system is conserved. A page passing rank through N outbound links passes 1/N of its rank through each link. This is not a Google policy choice—it is a mathematical consequence of the Markov chain formulation. Any link graph analysis tool that violates this conservation is wrong.

Power iteration convergence: PageRank is computed iteratively until convergence (change below threshold δ). For a web graph of N nodes, convergence requires O(log N) iterations typically. The order in which nodes are updated does not affect the final result—only the convergence rate. This means Google can compute partial PageRank updates on changed subgraphs without recomputing the entire web graph, which has practical crawl frequency implications.


# Reference PageRank implementation — exact formulation from Brin & Page 1998
import numpy as np
from scipy.sparse import csr_matrix

def compute_pagerank(adjacency_matrix, damping=0.85, max_iter=100, tol=1e-6):
    """
    Compute PageRank using power iteration.
    adjacency_matrix: sparse NxN matrix where M[i,j]=1 means j links to i
    """
    N = adjacency_matrix.shape[0]

    # Normalize columns to create column-stochastic matrix
    # Each column sums to 1 (outbound link weight distribution)
    col_sums = np.array(adjacency_matrix.sum(axis=0)).flatten()
    col_sums[col_sums == 0] = 1  # Dangling nodes — avoid division by zero
    norm_matrix = adjacency_matrix / col_sums

    # Handle dangling nodes (pages with no outbound links)
    # Dangling nodes teleport uniformly to all pages
    dangling_nodes = np.where(col_sums == 0)[0]
    dangling_vector = np.zeros(N)
    dangling_vector[dangling_nodes] = 1.0 / N

    # Initialize rank vector uniformly
    rank = np.ones(N) / N

    for iteration in range(max_iter):
        prev_rank = rank.copy()

        # Standard PageRank update
        rank = (damping * (norm_matrix @ rank + dangling_vector @ prev_rank)
                + (1 - damping) / N)

        # Check convergence (L1 norm)
        delta = np.abs(rank - prev_rank).sum()
        if delta < tol:
            print(f"Converged at iteration {iteration + 1}, delta={delta:.2e}")
            break

    return rank / rank.sum()  # Normalize to sum to 1


# Practical usage: crawl link graph → NetworkX → adjacency matrix → PageRank
import networkx as nx

def build_link_graph_from_crawl(crawl_items):
    """Build directed link graph from crawler output."""
    G = nx.DiGraph()

    for item in crawl_items:
        source = item['url']
        G.add_node(source)
        for link in item.get('internal_links', []):
            if not link.get('nofollow'):  # Respect nofollow
                G.add_edge(source, link['url'])

    return G


def analyze_pagerank_distribution(G, damping=0.85):
    """Compute and analyze PageRank distribution."""
    pr = nx.pagerank(G, alpha=damping, max_iter=200, tol=1e-8)

    # Sort by PageRank descending
    sorted_pr = sorted(pr.items(), key=lambda x: x[1], reverse=True)

    # Compute distribution statistics
    values = list(pr.values())
    print(f"Total pages: {len(values)}")
    print(f"Max PR: {max(values):.6f} ({sorted_pr[0][0]})")
    print(f"Mean PR: {np.mean(values):.6f}")
    print(f"Median PR: {np.median(values):.6f}")
    print(f"Top 10% of pages hold {sum(sorted(values, reverse=True)[:len(values)//10]):.2%} of total PR")

    return sorted_pr

The Damping Factor and Its SEO Implications

The damping factor d=0.85 was chosen empirically by Brin and Page. The mathematical interpretation: the probability that a random surfer continues clicking links rather than teleporting. At d=0.85, a page 6 clicks deep from the homepage retains approximately 0.85^6 ≈ 37.7% of the rank it would have at depth 0. At d=0.9 (which some researchers argue better fits observed crawl behavior), that becomes 0.9^6 ≈ 53.1%.

The practical implication is logarithmic rank decay with depth. Pages at crawl depth 7+ receive negligibly small PageRank shares from any realistic link graph, which is the mathematical explanation for why deep content gets crawled and ranked poorly—it is not a separate algorithmic choice by Google, it is an emergent property of the random surfer model.

PageRank Retention at Depth (d=0.85)
DepthRetention Factor% of Surface PR
10.8585.0%
20.722572.3%
30.614161.4%
40.522052.2%
50.443744.4%
60.377137.7%
70.320632.1%
100.196919.7%

Topic-Sensitive PageRank and TrustRank

Haveliwala (2002) introduced Topic-Sensitive PageRank, computing 16 separate rank vectors—one per ODP (Open Directory Project) top-level category—and combining them at query time based on query-topic affinity. Google filed related patents and has referenced topical authority in Quality Rater Guidelines. The key insight: a link from a topically relevant page passes more effective rank than a link from an unrelated page, even if their raw PageRank scores are equal.

TrustRank (Gyöngyi et al., 2004) inverts the PageRank formulation: instead of computing rank from all incoming links, compute it from a seed set of manually verified trusted pages (universities, government sites, established news outlets). Pages that receive link flow from trusted seeds accumulate TrustRank; spam pages, which typically cannot receive links from legitimate trusted sites, score low. Google's implementation of spam-resistant ranking (referenced in several patents) follows this architecture.


# Topic-Sensitive PageRank: compute rank vectors per topic seed set
def compute_topic_sensitive_pr(G, topic_seeds, damping=0.85):
    """
    topic_seeds: dict of {topic_name: [seed_urls]}
    Returns: dict of {url: {topic: rank_value}}
    """
    nodes = list(G.nodes())
    N = len(nodes)
    node_idx = {url: i for i, url in enumerate(nodes)}

    # Build adjacency matrix
    adj = np.zeros((N, N))
    for source, target in G.edges():
        if source in node_idx and target in node_idx:
            adj[node_idx[target]][node_idx[source]] = 1.0

    # Normalize columns
    col_sums = adj.sum(axis=0)
    col_sums[col_sums == 0] = 1
    adj = adj / col_sums

    topic_ranks = {}

    for topic, seeds in topic_seeds.items():
        # Teleportation vector: uniform over seed set instead of all pages
        teleport = np.zeros(N)
        valid_seeds = [s for s in seeds if s in node_idx]
        if not valid_seeds:
            continue

        for seed in valid_seeds:
            teleport[node_idx[seed]] = 1.0 / len(valid_seeds)

        # Power iteration with topic-biased teleportation
        rank = np.ones(N) / N
        for _ in range(200):
            new_rank = damping * (adj @ rank) + (1 - damping) * teleport
            if np.abs(new_rank - rank).sum() < 1e-8:
                break
            rank = new_rank

        topic_ranks[topic] = {nodes[i]: rank[i] for i in range(N)}

    return topic_ranks

The Reasonable Surfer Model Patent

Google Patent US7716225B1, "Generating and using document signatures" and more relevantly the "Reasonable Surfer" patent (US8117209B2, inventors Lawrence Page et al.) describe a modification where not all links on a page pass equal rank. The probability that a surfer clicks a link depends on its features: position on page, anchor text relevance, link type (navigation vs. in-content), and visual prominence.

This means a link in the main body copy of an article passes more effective PageRank than a footer link, even on the same page. This is a significant departure from the original formulation which treats all links equally. The Reasonable Surfer model can be approximated by weighting links in your internal link analysis:


# Reasonable Surfer weight estimation for internal links
def estimate_link_weight(link_context):
    """
    Estimate relative PageRank passing weight based on link context.
    Based on signals described in Google's Reasonable Surfer patent.
    """
    weight = 1.0

    # Position-based multipliers
    position = link_context.get('position', 'body')
    position_weights = {
        'body': 1.0,       # Main content area
        'navigation': 0.7, # Navigation menus — lower probability of click
        'footer': 0.3,     # Footer links — lowest click probability
        'sidebar': 0.6,    # Sidebar
        'header': 0.8,     # Header (above fold)
    }
    weight *= position_weights.get(position, 1.0)

    # Anchor text quality multiplier
    anchor = link_context.get('anchor_text', '').strip()
    if len(anchor) < 3:
        weight *= 0.5  # Generic anchors ('click here', 'read more')
    elif len(anchor) > 100:
        weight *= 0.7  # Overly long anchor text

    # NoFollow reduces effective weight (not to zero in current Google interpretation)
    if link_context.get('nofollow'):
        weight *= 0.0  # Google claims not to pass rank through nofollow

    # Image-only links — no anchor text signal
    if link_context.get('is_image_only'):
        weight *= 0.8

    return weight


def compute_weighted_pagerank(G, link_contexts, damping=0.85, max_iter=200):
    """PageRank with Reasonable Surfer link weights."""
    nodes = list(G.nodes())
    N = len(nodes)
    node_idx = {url: i for i, url in enumerate(nodes)}

    # Build weighted adjacency matrix
    adj = np.zeros((N, N))
    for source, target, data in G.edges(data=True):
        if source in node_idx and target in node_idx:
            context_key = f"{source}:{target}"
            weight = estimate_link_weight(link_contexts.get(context_key, {}))
            adj[node_idx[target]][node_idx[source]] = weight

    # Normalize columns (weighted)
    col_sums = adj.sum(axis=0)
    col_sums[col_sums == 0] = 1
    adj = adj / col_sums

    rank = np.ones(N) / N
    for _ in range(max_iter):
        new_rank = damping * (adj @ rank) + (1 - damping) / N
        if np.abs(new_rank - rank).sum() < 1e-8:
            break
        rank = new_rank

    return {nodes[i]: rank[i] for i in range(N)}

# Full pipeline: crawl data → link graph → PageRank → authority report
import pandas as pd
import networkx as nx
from collections import defaultdict

def generate_pagerank_report(crawl_df):
    """
    crawl_df: DataFrame with columns [url, internal_links (list of dicts)]
    Returns: DataFrame with PageRank scores, link counts, authority classification
    """
    G = nx.DiGraph()

    # Add all crawled URLs as nodes
    for _, row in crawl_df.iterrows():
        G.add_node(row['url'])

    # Add edges from internal links (skip nofollow)
    for _, row in crawl_df.iterrows():
        for link in row.get('internal_links', []):
            if not link.get('nofollow') and link['url'] in G.nodes:
                G.add_edge(row['url'], link['url'])

    # Compute PageRank
    pr = nx.pagerank(G, alpha=0.85, max_iter=500)

    # Compute additional graph metrics
    in_degree = dict(G.in_degree())
    out_degree = dict(G.out_degree())

    # Identify orphan pages (no incoming links)
    orphans = [n for n in G.nodes() if in_degree.get(n, 0) == 0]

    # Identify rank sinks (no outgoing links)
    sinks = [n for n in G.nodes() if out_degree.get(n, 0) == 0]

    # Identify link clusters (strongly connected components)
    sccs = list(nx.strongly_connected_components(G))
    node_to_scc = {}
    for i, scc in enumerate(sccs):
        for node in scc:
            node_to_scc[node] = i

    # Build report DataFrame
    report_rows = []
    for url in G.nodes():
        rank_score = pr.get(url, 0)
        report_rows.append({
            'url': url,
            'pagerank': rank_score,
            'pagerank_percentile': None,  # Filled below
            'inbound_links': in_degree.get(url, 0),
            'outbound_links': out_degree.get(url, 0),
            'is_orphan': url in orphans,
            'is_sink': url in sinks,
            'scc_id': node_to_scc.get(url, -1),
        })

    report_df = pd.DataFrame(report_rows)
    report_df['pagerank_percentile'] = report_df['pagerank'].rank(pct=True)
    report_df['authority_tier'] = pd.cut(
        report_df['pagerank_percentile'],
        bins=[0, 0.5, 0.8, 0.95, 1.0],
        labels=['low', 'medium', 'high', 'top']
    )

    return report_df.sort_values('pagerank', ascending=False)

Approximating PageRank Distribution via BigQuery


-- Approximate PageRank authority tiers from GSC + crawl data in BigQuery
-- Uses inbound link count as a PageRank proxy for large graphs

WITH crawl_links AS (
  -- From your custom crawler output
  SELECT
    url AS target_url,
    COUNT(DISTINCT referrer_url) AS inbound_link_count,
    SUM(CASE WHEN crawl_depth <= 2 THEN 1 ELSE 0 END) AS shallow_page_links
  FROM your-project.seo_crawl.pages
  WHERE crawl_run_id = @latest_run_id
    AND status_code = 200
  GROUP BY url
),

gsc_performance AS (
  SELECT
    url,
    SUM(clicks) AS total_clicks,
    SUM(impressions) AS total_impressions,
    AVG(position) AS avg_position
  FROM your-project.gsc_export.searchdata_url_impression
  WHERE data_date >= DATE_SUB(CURRENT_DATE(), INTERVAL 90 DAY)
  GROUP BY url
),

authority_model AS (
  SELECT
    c.target_url,
    c.inbound_link_count,
    c.shallow_page_links,
    g.total_clicks,
    g.avg_position,
    -- Composite authority score: links + GSC signal
    LOG(c.inbound_link_count + 1) * 0.6
    + LOG(GREATEST(g.total_clicks, 1)) * 0.3
    + (50 - LEAST(g.avg_position, 50)) / 50 * 0.1 AS authority_score,
    -- Rank within site
    PERCENT_RANK() OVER (
      ORDER BY c.inbound_link_count ASC
    ) AS link_percentile
  FROM crawl_links c
  LEFT JOIN gsc_performance g ON c.target_url = g.url
)

SELECT
  target_url,
  inbound_link_count,
  total_clicks,
  ROUND(avg_position, 1) AS avg_position,
  ROUND(authority_score, 4) AS authority_score,
  ROUND(link_percentile, 3) AS link_percentile,
  CASE
    WHEN link_percentile >= 0.95 THEN 'T1_authority'
    WHEN link_percentile >= 0.80 THEN 'T2_high'
    WHEN link_percentile >= 0.50 THEN 'T3_medium'
    ELSE 'T4_low'
  END AS authority_tier
FROM authority_model
ORDER BY authority_score DESC
LIMIT 1000

What Has Provably Changed Since 1998

Several patent-supported and experiment-validated changes are worth documenting precisely:

NoFollow and link attributes: The original PageRank passes rank through all links. Google introduced rel="nofollow" in 2005 as a zero-rank-passing directive. In 2019, Google changed NoFollow to a "hint"—rank may or may not pass. rel="ugc" and rel="sponsored" were added as semantic classifiers. The practical implication: do not assume zero rank flow through nofollow links, but do not assume full flow either.

Link velocity and freshness: Multiple Google patents describe temporal aspects of link acquisition. A sudden spike in links from a narrow set of IP ranges or domain registrars may be algorithmically discounted or flagged for manual review. Historically, this has been the primary mechanism for Penguin-style spam detection.

Passage indexing: Announced in 2020, passage indexing allows Google to rank individual passages within a page independently. This implies that internal PageRank reaching a page may not uniformly benefit all content on that page—sections can be ranked independently of the page-level authority signal.

EEAT as a PageRank complement: Experience, Expertise, Authoritativeness, and Trustworthiness (EEAT) in the Search Quality Evaluator Guidelines describes human-rater signals that complement algorithmic link analysis. EEAT signals include authorship attribution, credential verification, entity recognition, and citation patterns in Google's Knowledge Graph. These are not replacements for PageRank—they are orthogonal signals weighted at ranking time.

Practical Architecture: Internal PageRank Sculpting

PageRank sculpting—the practice of directing link equity through internal linking structure—remains valid. What changed: rel="nofollow" no longer concentrates rank (Google ignores it for sculpting purposes since 2009). What works: controlling which pages receive inbound links via site architecture, contextual linking, and crawl budget management.

StrategyMechanismEffectRisk
Hub page architectureHigh-equity pages link to priority targetsConcentrates rank on revenue pagesLow if natural
Orphan page rescueAdd internal links to zero-inbound pagesPrevents rank waste on uncrawled pagesMinimal
Pagination handlingSelf-canonical on facets, rel=next removedConsolidates rank to root categoryLow
Breadcrumb optimizationConsistent parent-child links in breadcrumbsReinforces site hierarchyNone
Link depth reductionFlatten architecture to max 3 clicks from rootEnsures damping factor retentionLow
Sidebar/footer link auditReduce sitewide nav to essential links onlyImproves Reasonable Surfer weight on body linksLow

For a deep dive on internal link orchestration see the schema and knowledge graph article. For monitoring how link equity flows through your crawl data, see the BigQuery analysis article.

FAQ

Is the public PageRank toolbar metric useful in 2026?

Google discontinued the public PageRank toolbar in 2016. Third-party metrics (Moz DA, Ahrefs DR, Majestic TF) are correlation-based approximations trained on ranking outcomes, not actual PageRank computations. They are useful for comparative analysis (site A vs. site B) and trend monitoring, but do not confuse them with internal Google signals. Treat them as probabilistic proxies with known biases toward link count over link quality.

Does PageRank still matter relative to content quality signals?

Empirically yes, especially for competitive head terms. Google has invested in content quality signals (BERT, MUM, EEAT), but link authority remains a strong prior for ranking competitive queries where many pages have adequate content quality. PageRank does not help a low-quality page; content quality does not overcome a severe authority deficit for competitive terms. They are complementary signals, not substitutes.

How does Google handle link spam at scale in 2026?

The SpamBrain neural network (described in Google's blog in 2021 and evolved significantly since) classifies links as spam/not-spam using features including: link neighborhood topology, anchor text distribution, temporal acquisition patterns, and cross-referencing with known spam domains. Algorithmically spammy links are not counted (neither positive nor negative in most cases) rather than penalized. Manual penalties apply to sites actively buying or selling links.

What is the practical click depth limit for PageRank retention?

At d=0.85, pages at depth 7+ retain less than 32% of surface-page rank. In practice, pages beyond 5 clicks from the homepage see dramatically reduced crawl frequency and link equity. The industry heuristic of "max 3 clicks from homepage" for important pages is mathematically justified: at depth 3, retention is ~61%, which is sufficient. At depth 5, it drops to ~44%.

Does PageRank distribute through redirect chains?

Google has confirmed that PageRank passes through 301 redirects, though historically there has been a small transfer penalty. 302 redirects are more ambiguous—Google generally follows them but the canonical assignment can be unstable. Redirect chains (A→B→C→D) likely incur incremental loss at each hop. Best practice: resolve chains to single-hop 301s. The loss per hop is not published but is commonly estimated at 10–15% by practitioners who have tested this experimentally.

Can JavaScript-rendered links pass PageRank?

Yes, once Google renders the page. Googlebot renders JavaScript using a deferred second-wave rendering process—initial HTML is crawled first, rendering queued. Links only visible in the rendered DOM pass PageRank only after rendering. This creates a crawl latency: newly added JavaScript-only links may take days to weeks to be discovered and counted for link equity purposes. Critical navigation links should be in the initial HTML response.

Key Takeaways

  • The core PageRank conservation property—a page passes 1/N of its equity through each outbound link—remains mathematically fundamental regardless of implementation changes.
  • At d=0.85, rank retention at depth 5 is 44%; at depth 7 it is 32%. Flatten site architecture to 3 clicks from root for priority pages.
  • The Reasonable Surfer patent (US8117209B2) describes non-uniform link weight by position, anchor quality, and link type—body links pass more effective rank than navigation or footer links.
  • Topic-Sensitive PageRank means topical relevance of the linking page matters, not just its absolute authority—a link from a relevant mid-authority page may outperform a link from an unrelated high-authority page.
  • NoFollow is a "hint" since 2019—do not rely on it for rank concentration; use site architecture and crawl budget management instead.
  • JavaScript-only links delay PageRank flow due to deferred rendering; critical internal links must be in initial HTML.

Conclusion

The original PageRank algorithm is 28 years old. Its mathematical foundations—Markov chain modeling of random web traversal, rank conservation, power iteration convergence—are not obsolete. They are embedded in any link-based ranking system, regardless of what engineers build on top. Understanding the math is what separates practitioners who know why link equity behaves the way it does from those who apply cargo-cult rules. The superstructure has changed enormously: Reasonable Surfer, Topic-Sensitive PR, SpamBrain, passage indexing, EEAT. But the foundation still holds. Build internal link architecture that respects the math—depth, conservation, link weight distribution—and the specifics of Google's current implementation matter less than you might think.

YOUR READING CHECKLIST

Make the ideas stick.

Mark the sections you’ve worked through. Saved in this browser.

0 of 4 reviewed
Andrii Stanetskyi
ABOUT THE AUTHOR

Andrii Stanetskyi

Head of SEO / Technical SEO Lead based in Tallinn, Estonia. Technical architecture, enterprise eCommerce, Python automation, and AI-assisted workflows.

More about Andrii ↗
LET’S FIND THE REAL BOTTLENECK

A clearer picture.
A practical next step.

Get a focused SEO audit or a consultation on your next technical decision. We’ll agree on the scope and fee before any work begins.

01 / Diagnose02 / Prioritize03 / Plan
How can I help?

Scope and fee agreed before any work begins.

Choose your language

Explore SEO services in 26 languages. Journal articles retain their original language.

ENEnglish↗DEDeutsch↗FRFrançais↗ESEspañol↗ITItaliano↗PTPortuguês↗NLNederlands↗PLPolski↗SVSvenska↗DADansk↗FISuomi↗NONorsk↗ETEesti↗LVLatviešu↗LTLietuvių↗CSČeština↗RORomână↗HUMagyar↗ELΕλληνικά↗BGБългарски↗HRHrvatski↗SKSlovenčina↗SLSlovenščina↗RUРусский↗UKУкраїнська↗TRTürkçe↗
LET’S WORK ON YOUR WEBSITE
A CLEAR NEXT STEP

Let’s talk
about your site.

A focused SEO audit or a conversation about a specific challenge. Tell me where you are and what you want to change.

Andrii Stanetskyi
Andrii StanetskyiHead of SEO / Technical SEO Lead
[email protected] ↗
How can I help?

Scope and fee agreed before any work begins.