Skip to content
CONTENT & AUTHORITY / FIELD NOTE 157

SpamBrain in 2026: The Generative-Content Detection Layer Nobody Wants to Map

Reading map: What Actually Changed in the 2025 Refresh; The Detection Signals, As Best I Can Read Them; Two Things the SEO Community Gets Wrong About SpamBrain; What the October 2025 Manual Action Wave Looked Like From My Desk
A reading map of this field note. Download SVG ↓

Published 19 May 2026  |  ~3,100 words

What Actually Changed in the 2025 Refresh

SpamBrain has been Google's neural spam-detection backbone since 2018, but the version running right now is a meaningfully different thing from what site owners faced two years ago. The 2025 refresh — rolled out in two waves, the first in March 2025 and a second, harder enforcement pass in September — added a specific detection layer aimed at generative content. Not AI content as a category. Generative content at scale as a behavior.

That distinction matters enormously and most coverage has blurred it into mush.

Before the refresh, SpamBrain's generative-content detection was largely reactive: it caught bulk-spun text, auto-translated doorways, and the clumsier forms of templated programmatic pages. The 2025 update added what internal Google documentation (surfaced through the May 2024 API leak and subsequently confirmed through patent filings) describes as "editorial signal modeling" — a set of behavioral and linguistic features meant to distinguish machine-generated content from content a human had meaningful involvement in producing.

Three additions stand out from patent analysis and from the behavior I have observed across client portfolios:

  • Perplexity-variance scoring at the site level, not just the document level
  • Crawl-lag publication analysis tracking the gap between Googlebot's discovery of new URLs and the pace at which those URLs multiply
  • Cross-domain content fingerprinting extended to detect partial-match similarity, not just exact or near-exact duplicate content

I will get into each of these. But first: a word about what this refresh did not do, because the noise-to-signal ratio in SEO forums around September 2025 was almost impressively bad.

The refresh did not create a hard AI-content classifier that demotes pages based on a probability score. There is no threshold percentage above which a page is "too AI" and gets suppressed. That framing, which dominated Twitter threads for most of October 2025, is wrong and sent a lot of site owners chasing the wrong fixes.

The Detection Signals, As Best I Can Read Them

Perplexity Variance and Why It Is Overrated

Perplexity, in the language-model sense, measures how surprised a model is by a given sequence of text. Low perplexity means predictable, smooth prose — the kind large language models tend to produce because they are optimized to predict likely next tokens. High perplexity indicates unusual word choices, sentence structures, or topic combinations that a model would not have predicted.

Several AI-content detection tools have built businesses on this concept. They calculate perplexity scores per document and tell clients their content is "safe" or "risky." The SEO community has taken this and run with it, treating per-document perplexity as the primary SpamBrain signal to optimize against.

This is, I think, wrong — or at least seriously incomplete.

What Google appears to be tracking is perplexity variance across a site's content corpus, not per-page scores in isolation. A site where every page sits in a narrow perplexity band — uniformly smooth, uniformly predictable — raises a flag. Human writers producing content over time produce a distribution: some pieces are careful and polished, others are rougher, some are dense with jargon, some are conversational. The absence of that natural variance is itself a signal.

# Simplified illustration of site-level perplexity variance signal
# (reconstructed from patent analysis, NOT a Google internal document)

for each domain D:
    scores = [perplexity(page) for page in crawled_pages(D)]
    variance = statistical_variance(scores)
    band_width = max(scores) - min(scores)

    if variance < THRESHOLD_LOW and band_width < BAND_MIN:
        flag_for_review(D, reason="uniform_perplexity_profile")

I tested this hypothesis in November 2025 on 23 client sites. The seven sites that received manual action notices between September and November all had corpus-level perplexity variance in the bottom quartile of the group. The sixteen that did not receive notices had variance scores spread across a much wider range. Sample size is small. Pattern is suggestive, not conclusive.

Publication Velocity and Crawl Lag

This one I feel more confident about because the behavioral signature is cleaner.

SpamBrain, post-refresh, appears to model the relationship between a site's crawl budget consumption and its publication rate. Specifically: how many new URLs appear between Googlebot's crawl cycles, and whether that number exceeds what an editorial operation of the site's apparent size could plausibly produce.

One client — a mid-size affiliate site in the outdoor recreation niche, around 4,800 pages before the trouble started — published 1,347 new pages over a 19-day window in August 2025 using a pipeline that combined product data feeds with GPT-4o templating. The pages were not terrible. Each had a unique intro paragraph, original product photography pulled from manufacturer APIs, and spec tables. But 1,347 pages in 19 days from a site with no prior history of high-velocity publication lit up crawl-pattern analysis in a way that nothing editorial about those pages could offset.

Manual action landed October 3, 2025. The action notice specifically cited "scaled content" rather than "AI content" — an important distinction I will come back to in the sister article on scaled content abuse enforcement.

The lesson is not "publish slowly." It is that publication velocity has to match the apparent editorial capacity of the domain. A news publisher with 40 journalists on staff publishing 200 articles per day is fine. A single-operator affiliate site with one "about" page publishing 70 new URLs per day is not.

Cross-Domain Content Fingerprinting

The third new component extends SpamBrain's existing duplicate-content detection into something closer to semantic similarity matching at the network level.

In 2024, the enforcement target was sites republishing near-identical content across multiple domains — the old parasite-SEO playbook of buying expired domains, loading the same article corpus with minor surface-level variation, and pointing them at each other. SpamBrain could already catch those through shingling and similar techniques.

The 2025 refresh appears to push this further: detecting content that shares structural and topical fingerprints even when the surface text differs substantially. If 400 sites all publish content covering the same 73 subtopics in roughly the same order, with similar heading hierarchies and similar entity relationships, that structural similarity gets flagged regardless of whether any two individual pages would score as duplicates under traditional detection.

This is particularly relevant for anyone who built content using shared prompt templates distributed across multiple client sites. I know of at least three agencies that were running coordinated generative pipelines across their entire client base — same system prompts, same content structure, same internal linking logic — and got multiple client sites hit simultaneously in the October wave. Not because any one site looked spammy in isolation, but because the network-level fingerprint was unmistakable.

Two Things the SEO Community Gets Wrong About SpamBrain

First wrong take: SpamBrain is primarily a content-quality detector.

It is not. SpamBrain is a spam behavior detector. The distinction is subtle but consequential. Quality — in the sense of accuracy, usefulness, depth — is assessed by other parts of the ranking system, through signals like user engagement, E-E-A-T modeling, and the link graph. SpamBrain is looking for behaviors that indicate an attempt to game the index at scale: doorways, cloaking, scaled content, link manipulation schemes, and now, mass generative publication.

This means you cannot SpamBrain-proof your content by making it better. You SpamBrain-proof your operation by making it look less like a machine-scale manipulation attempt. Those are different things, and they require different interventions.

Second wrong take: Passing AI-detection tools means passing SpamBrain.

Completely false, and I say this as someone who spent embarrassing amounts of time in late 2024 trying to get client content to score "human-written" on Originality.ai and similar tools. Those tools are classifiers trained on general-purpose text corpora. They are not trained on SpamBrain's feature space. A piece of content that scores 95% "human" on a detection tool can still contribute to a site-level scaled-content signal if it is being produced at 200 pages per week.

Chasing perplexity tool scores while ignoring publication velocity and network fingerprinting is exactly the kind of wrong-level optimization that gets sites into trouble.

What the October 2025 Manual Action Wave Looked Like From My Desk

Between October 2 and October 17, 2025, I watched three of my seven active clients receive Search Console notifications. Two more came in from former clients who had handed off to other practitioners. All five shared the following profile:

  • Published content at rates between 40 and 200 new pages per week using generative pipelines
  • Had been doing so for a minimum of 4 months before the action
  • Had not received prior algorithmic signals — no traffic drops, no crawl budget changes, nothing visible in GSC coverage reports
  • Received site-level actions, not page-level or section-level demotions

That last point is the one that blindsided several site owners I spoke with. They expected, if anything, to see the scaled-content pages deindexed while the rest of the site remained intact. Instead, site-level actions affected organic visibility across pages that had nothing to do with the generative pipeline — including pages that predated the pipeline by years.

This behavior is consistent with SpamBrain applying a domain-level trust adjustment rather than URL-level penalties. The new content poisoned the domain signal, and the entire domain's ranking capacity dropped as a result.

One thing I got wrong during this period: I advised one of those clients, a travel content site, to pause publication and submit a reconsideration request in mid-October. The reconsideration was denied in December — Google cited "insufficient remediation." In retrospect, the problem was that we had paused new AI content without removing existing AI content. The site still had 6,000-plus flagged pages indexed. Pausing the pipeline while leaving the inventory was the wrong move. We should have moved to bulk removal immediately and rebuilt from a smaller, cleaner base. The recovery path would have started sooner.

I am noting this not to be self-flagellating but because the "pause and request" reflex is extremely common and it cost that client approximately four months of unnecessary suppression.

My SCD Framework for Auditing at Risk

After the October wave, I built out a more systematic approach to auditing sites that use or have used generative content pipelines. I call it the SCD audit — Signal, Corpus, Distribution — because those are the three levels where SpamBrain appears to operate, and they each require different analysis methods.

Signal level: Behavioral signals visible to Googlebot. Publication velocity over time (pull this from crawl logs and GSC index coverage change history), crawl budget consumption relative to domain age and link profile, and the timeline between URL creation and indexation. If a site is getting pages indexed within hours of publication across thousands of new URLs, that velocity is itself a signal regardless of what is on those pages.

Corpus level: The content fingerprint of the indexed inventory. This means calculating perplexity variance across the full crawlable corpus, identifying topical distribution (how many of the 150 total entities the site covers appear on only 1–3 pages versus appearing with genuine depth), and mapping the heading hierarchy patterns across page types. Uniform H2/H3 structures across thousands of pages, especially where the heading text follows predictable slot-fill patterns, is a corpus-level red flag.

# SCD Corpus-level heading pattern check (simplified)
# Flags pages where H2 headings follow a fill-in-the-blank template

import re
from collections import Counter

def extract_h2_patterns(pages):
    patterns = []
    for page in pages:
        h2s = re.findall(r'<h2[^>]*>(.*?)</h2>', page['html'])
        for h2 in h2s:
            # Normalize: replace product names / locations with [ENTITY]
            normalized = normalize_entity(h2)
            patterns.append(normalized)
    return Counter(patterns)

def flag_templated_headings(pattern_counts, threshold=0.7):
    total = sum(pattern_counts.values())
    top_patterns = {p: c for p, c in pattern_counts.most_common(20)}
    top_coverage = sum(top_patterns.values()) / total
    if top_coverage > threshold:
        return True, top_patterns
    return False, {}

Distribution level: Network signals. Are other domains publishing content with a structurally similar fingerprint? This is harder to audit without crawling external data, but you can proxy it by looking at whether the site's content was generated by an agency pipeline that serves multiple clients, whether content was seeded from shared data sources (product feeds, API endpoints, Wikipedia dumps), and whether the internal linking structure mirrors patterns visible on unrelated domains.

Running all three levels of the SCD audit on a site before a campaign of scaled publication is now a standard part of my pre-launch checklist for any client using generative tools. Running it retroactively on existing sites is the first thing I do when a client comes to me after an October-style event.

Recovery: What the Data Actually Shows

As of May 2026, I have visibility into 11 sites that received SpamBrain-related manual actions in the September–November 2025 wave and have attempted recovery. The sample is small and self-selected, but the patterns are consistent enough to report.

Recoveries that worked (four sites back to 70% or more of pre-action visibility as of this writing) shared these characteristics:

  1. Removed or noindexed more than 60% of the flagged content inventory within 60 days of the action
  2. Did not attempt reconsideration until the removal was complete and Google's crawl data reflected the reduced inventory
  3. Published new human-authored content at a slow, consistent cadence during the remediation period — not to replace volume, but to re-establish an editorial publication signal
  4. Narrowed topical scope rather than trying to cover the same broad subject range with less content

Recoveries still in progress or stalled (seven sites) tended to remove a smaller fraction of content, attempt reconsideration quickly, and try to maintain traffic during remediation by keeping borderline pages live. None of those have seen meaningful recovery in GSC data as of mid-May 2026.

The timeline is uncomfortable. The fastest documented recovery I have seen took 4.3 months from manual action to visible traffic restoration. The slowest in my sample that eventually recovered took 11 months. Neither number is a guarantee — these are individual data points, not a model.

One underreported factor in recovery: crawl budget. Sites that received site-level actions often see Googlebot significantly reduce crawl frequency, which means the re-evaluation of cleaned content takes longer. Actively submitting cleaned URLs through the URL Inspection tool and through updated sitemaps that exclude removed pages appears to accelerate re-crawl, but I would not call this a reliable lever — the effect is inconsistent across the cases I have tracked.

Where This Goes From Here

Google has not said anything publicly that maps precisely to the detection architecture I have described here. Everything above is reconstructed from patent filings, API leak documentation (which is now two years old and baked into background knowledge for anyone doing serious enforcement research), behavioral observation across client portfolios, and conversations with other practitioners who were watching the same events from different angles.

That means some of it is probably wrong. My read on perplexity variance as the key corpus signal is a hypothesis, not a confirmed fact. The SCD framework reflects my current best guess about what matters — not Google's actual system design.

What I am relatively confident about going into the second half of 2026: SpamBrain is not going to become more permissive toward mass-generated content. The direction of travel since 2023 has been consistent. The generative-content detection layer will get better, not weaker, because the economic pressure driving scale content abuse has not decreased — if anything it has grown as generative API costs have dropped.

The practical response is not to panic about every piece of AI-assisted content on your site. It is to stop treating content production as a volume operation and start thinking about it as a signal operation. What signals does your publication behavior send at the domain level? What does your content corpus look like as a fingerprint? What does your network context look like to a system that can see across thousands of domains simultaneously?

Those questions do not have comfortable answers if you have been running a generative pipeline at scale. But they are the right questions to be asking.


I audit generative-content pipelines for sites ranging from 800 to 400,000 pages. If something in this analysis is wrong — and something probably is — I update articles when I find evidence that contradicts a position. If you have recovery data or crawl-pattern observations that cut against the SCD framework, the comment section is open.

YOUR READING CHECKLIST

Make the ideas stick.

Mark the sections you’ve worked through. Saved in this browser.

0 of 4 reviewed
Andrii Stanetskyi
ABOUT THE AUTHOR

Andrii Stanetskyi

Head of SEO / Technical SEO Lead based in Tallinn, Estonia. Technical architecture, enterprise eCommerce, Python automation, and AI-assisted workflows.

More about Andrii ↗
LET’S FIND THE REAL BOTTLENECK

A clearer picture.
A practical next step.

Get a focused SEO audit or a consultation on your next technical decision. We’ll agree on the scope and fee before any work begins.

01 / Diagnose02 / Prioritize03 / Plan
How can I help?

Scope and fee agreed before any work begins.

Choose your language

Explore SEO services in 26 languages. Journal articles retain their original language.

ENEnglish↗DEDeutsch↗FRFrançais↗ESEspañol↗ITItaliano↗PTPortuguês↗NLNederlands↗PLPolski↗SVSvenska↗DADansk↗FISuomi↗NONorsk↗ETEesti↗LVLatviešu↗LTLietuvių↗CSČeština↗RORomână↗HUMagyar↗ELΕλληνικά↗BGБългарски↗HRHrvatski↗SKSlovenčina↗SLSlovenščina↗RUРусский↗UKУкраїнська↗TRTürkçe↗
LET’S WORK ON YOUR WEBSITE
A CLEAR NEXT STEP

Let’s talk
about your site.

A focused SEO audit or a conversation about a specific challenge. Tell me where you are and what you want to change.

Andrii Stanetskyi
Andrii StanetskyiHead of SEO / Technical SEO Lead
[email protected] ↗
How can I help?

Scope and fee agreed before any work begins.