Skip to content
CONTENT & AUTHORITY / FIELD NOTE 246

Sentiment Analysis and SEO in 2026: The Tone Audit That Recovered a Penalty Page

Reading map: The Page That Wouldn't Rank; What Sentiment Analysis Actually Measures (and What SEOs Get Wrong); Building the Tone Audit: The SIGNAL Framework; The Python Stack I Use Every Week
A reading map of this field note. Download SVG ↓

The Page That Wouldn't Rank

Eleven months ago I inherited a client page that had been stuck between positions 18 and 23 for the better part of a year. The topic was competitive but not brutal. Domain authority was fine. Backlinks were fine. Technical health was clean. Page speed: fast. Schema: present. Everything the standard checklist told us to care about was in order, and yet the page sat there, inert, while a thinner competitor piece held position 4 with fewer words and worse structured data.

The first thing most consultants do in that situation is reach for links. I almost did. Instead, on a Wednesday afternoon when I was procrastinating on a proposal, I ran the page through a compound sentiment pipeline and got a score that stopped me: a VADER compound of −0.14. Mildly negative. Not angry, not alarming, just... persistently bleak. The kind of tone you get when a subject-matter expert writes about risks and caveats and "what to avoid" without ever articulating what success actually feels like.

That number became the beginning of a six-week project. What came out the other side was a page sitting at position 5, a 34% lift in organic CTR, and a personal obsession with sentiment auditing that has since become a core part of how I evaluate content before I touch anything else.

This is that story, and the replicable framework I built from it.

What Sentiment Analysis Actually Measures (and What SEOs Get Wrong)

Let's get the basics right before we go further, because the field is full of people using "sentiment" to mean wildly different things.

In NLP, sentiment analysis assigns a polarity to text: positive, negative, or neutral, sometimes with a confidence score. VADER (Valence Aware Dictionary and sEntiment Reasoner) is the classic lexicon-based tool, fast and interpretable, built originally for social media. Transformer-based models like those available through HuggingFace's pipeline("sentiment-analysis") are slower and heavier but pick up on subtler constructions that VADER misses entirely.

Neither of these tools was designed for SEO. That's the first thing to internalize.

What they measure is surface-level affective tone, which is not the same as user satisfaction, trust, or purchase intent. A detailed warning about medication interactions is correctly scored as negative. It should be negative. That page deserves to rank. So the naive application of "make everything positive" is actively harmful and I've seen agencies pursue it with disastrous results.

The Actual Hypothesis

The hypothesis I'm working with is narrower and, I think, more defensible: for informational and commercial-investigation queries, pages that carry sustained negative sentiment across their primary value sections perform measurably worse on behavioral signals (dwell time, scroll depth, return-to-SERP rate) than pages with balanced or moderately positive tone in those same sections. And Google has had enough behavioral data for long enough that those signals are influencing rankings in ways that technical auditors routinely miss.

That's not the same as "be cheerful." It's saying: be calibrated. Threat without resolution reads as anxiety-inducing. Most readers don't consciously notice this, but their behavior reflects it. They leave. They come back to the SERP. Google notices.

There's also a second axis that almost nobody talks about, which is consistency. A page that oscillates violently between fear-based language and overpromising creates cognitive whiplash. Even if the average sentiment lands at neutral, the variance in per-paragraph scores can correlate with poor engagement. I've started tracking standard deviation alongside mean sentiment scores for exactly this reason.

Building the Tone Audit: The SIGNAL Framework

After running sentiment audits on 40-something pages across eight client verticals over the past year, I settled on a structured approach I call SIGNAL.

  • S — Segment the page into logical sections (intro, value body, proof sections, CTA zones). Don't score the whole page as a unit.
  • I — Instrument each segment with at least two model types (lexicon-based + transformer). Single-model scores lie too often.
  • G — Grade each segment against a reference corpus of top-5 ranking pages for the same query.
  • N — Note variance, not just mean. Flag any segment where paragraph-level standard deviation exceeds 0.3 on a normalized −1 to +1 scale.
  • A — Align problem segments with behavioral data. Pull GSC click-through data and, where available, heatmap or session recording data. Sentiment alone doesn't justify a rewrite.
  • L — Log before/after scores and tie them to ranking and traffic outcomes over a minimum 60-day window.

The logging step is the one people skip. It's also the only step that lets you build a feedback loop instead of just vibing your way through rewrites.

What Counts as a "Problem" Score

I use these thresholds as starting points, not laws:

  • VADER compound below −0.05 in an intro section: investigate.
  • VADER compound below −0.20 sustained across three or more consecutive value-body paragraphs: rewrite candidate.
  • Transformer model returning >60% negative label confidence across a section: rewrite candidate regardless of VADER score.
  • Paragraph-level standard deviation above 0.35 across the full page body: flag for structural review.

These numbers came from my own dataset. Your verticals may calibrate differently. Finance and healthcare pages run naturally more negative than lifestyle or productivity content. The reference corpus comparison in step G is what makes the threshold meaningful rather than arbitrary.

The Python Stack I Use Every Week

Two scripts handle most of the heavy lifting. The first is a quick VADER pass that I run on any page I'm asked to evaluate. The second is a transformer pipeline I pull out when VADER scores look ambiguous or when the page is in a domain where irony and hedged language are common.


# sentiment_audit_vader.py
# Requires: vaderSentiment, requests, beautifulsoup4

import requests
from bs4 import BeautifulSoup
from vaderSentiment.vaderSentiment import SentimentIntensityAnalyzer
import statistics
import json

def scrape_paragraphs(url):
    resp = requests.get(url, timeout=10, headers={"User-Agent": "Mozilla/5.0"})
    soup = BeautifulSoup(resp.text, "html.parser")
    # Target main content area — adjust selector for your CMS
    content = soup.select_one("article, main, .post-content, #content")
    if not content:
        content = soup
    paragraphs = [p.get_text(strip=True) for p in content.find_all("p") if len(p.get_text(strip=True)) > 60]
    return paragraphs

def run_vader_audit(url):
    analyzer = SentimentIntensityAnalyzer()
    paragraphs = scrape_paragraphs(url)
    scores = []
    results = []

    for i, para in enumerate(paragraphs):
        vs = analyzer.polarity_scores(para)
        compound = vs["compound"]
        scores.append(compound)
        results.append({
            "paragraph_index": i + 1,
            "text_preview": para[:80] + "...",
            "compound": compound,
            "label": "positive" if compound >= 0.05 else "negative" if compound <= -0.05 else "neutral"
        })

    summary = {
        "url": url,
        "paragraph_count": len(scores),
        "mean_compound": round(statistics.mean(scores), 4) if scores else None,
        "stdev_compound": round(statistics.stdev(scores), 4) if len(scores) > 1 else None,
        "min_compound": round(min(scores), 4) if scores else None,
        "max_compound": round(max(scores), 4) if scores else None,
        "paragraphs": results
    }

    print(json.dumps(summary, indent=2))
    return summary

if __name__ == "__main__":
    import sys
    run_vader_audit(sys.argv[1])

When I need more nuance, I layer in a transformer model. HuggingFace's distilbert-base-uncased-finetuned-sst-2-english is fast enough for paragraph-level analysis without requiring a GPU on most modern laptops.


# sentiment_audit_transformers.py
# Requires: transformers, torch, requests, beautifulsoup4

from transformers import pipeline
import requests
from bs4 import BeautifulSoup
import statistics
import json

def scrape_paragraphs(url, min_length=60):
    resp = requests.get(url, timeout=10, headers={"User-Agent": "Mozilla/5.0"})
    soup = BeautifulSoup(resp.text, "html.parser")
    content = soup.select_one("article, main, .post-content, #content") or soup
    return [p.get_text(strip=True) for p in content.find_all("p") if len(p.get_text(strip=True)) > min_length]

def run_transformer_audit(url, model_name="distilbert-base-uncased-finetuned-sst-2-english"):
    classifier = pipeline("sentiment-analysis", model=model_name, truncation=True, max_length=512)
    paragraphs = scrape_paragraphs(url)

    results = []
    normalized_scores = []

    for i, para in enumerate(paragraphs):
        output = classifier(para)[0]
        label = output["label"]    # "POSITIVE" or "NEGATIVE"
        score = output["score"]    # confidence 0–1
        # Normalize to −1 to +1
        normalized = score if label == "POSITIVE" else -score
        normalized_scores.append(normalized)

        results.append({
            "paragraph_index": i + 1,
            "text_preview": para[:80] + "...",
            "label": label,
            "confidence": round(score, 4),
            "normalized": round(normalized, 4)
        })

    summary = {
        "url": url,
        "model": model_name,
        "paragraph_count": len(normalized_scores),
        "mean_normalized": round(statistics.mean(normalized_scores), 4) if normalized_scores else None,
        "stdev_normalized": round(statistics.stdev(normalized_scores), 4) if len(normalized_scores) > 1 else None,
        "paragraphs": results
    }

    print(json.dumps(summary, indent=2))
    return summary

if __name__ == "__main__":
    import sys
    run_transformer_audit(sys.argv[1])

Running both on the same URL and comparing outputs takes less than two minutes. When they agree, I'm confident. When they diverge significantly, that divergence is itself a signal worth investigating manually.

Aggregating Sentiment Signals at Scale with BigQuery

For clients with large content operations (500+ pages), running individual Python scripts is impractical. I've moved the aggregation layer into BigQuery, pulling pre-computed sentiment scores from a Cloud Function that runs on a weekly crawl schedule and deposits results into a partitioned table.

The query below assumes you have a sentiment_scores table with columns for url, crawl_date, mean_compound, stdev_compound, paragraph_count, and gsc_avg_position (joined from a Search Console export).


-- BigQuery: Identify pages with negative sentiment AND poor rank position
-- for triage prioritization

WITH ranked_pages AS (
  SELECT
    url,
    crawl_date,
    mean_compound,
    stdev_compound,
    paragraph_count,
    gsc_avg_position,
    gsc_impressions_28d,
    -- Risk score: lower is worse
    ROUND(
      (mean_compound * 0.5)                         -- penalize negative mean
      - (stdev_compound * 0.3)                      -- penalize high variance
      + (1.0 / NULLIF(gsc_avg_position, 0) * 0.2), -- reward high rank
    4) AS sentiment_risk_score
  FROM
    your_project.seo_data.sentiment_scores
  WHERE
    crawl_date = (SELECT MAX(crawl_date) FROM your_project.seo_data.sentiment_scores)
    AND paragraph_count >= 8        -- ignore stub pages
    AND gsc_impressions_28d >= 100  -- ignore invisible pages
)

SELECT
  url,
  mean_compound,
  stdev_compound,
  gsc_avg_position,
  gsc_impressions_28d,
  sentiment_risk_score
FROM ranked_pages
WHERE mean_compound < -0.05
ORDER BY sentiment_risk_score ASC
LIMIT 50;

The sentiment_risk_score formula is deliberately simple. It exists to sort a triage list, not to predict exact ranking outcomes. The 50 pages at the bottom of that list are the candidates for manual SIGNAL audits. From there, human judgment takes over.

One thing I've found: pages with stdev_compound above 0.40 and a neutral mean are often more problematic than pages with consistently negative scores. The volatile ones tend to be listicles or "pros and cons" articles that alternate between enthusiastic claims and bleak warnings without structural resolution. Readers experience those as confused and untrustworthy even when they can't articulate why.

Two Things Most SEO Consultants Won't Tell You

Contrarian Take 1: Positive Sentiment Doesn't Correlate With Rankings in YMYL Categories

I've run this analysis across health, legal, and personal finance verticals and the data doesn't support the "be positive, rank higher" story that some content marketers are now pushing. In YMYL categories, the top-ranking pages are often meticulous about articulating risk, limitation, and uncertainty. They should be. A financial planning article that presents only upside scenarios reads as promotional and is probably less trustworthy to both readers and quality raters.

In those verticals, what I actually find is that the top pages have moderate negative sentiment with low variance and consistently strong resolution language at the end of each major section. They take you into the problem clearly, and then they take you out with equal clarity. That's a different behavioral target than raw positivity.

The agencies that run blanket positivity rewrites on YMYL content are not helping their clients. In some cases they're actively reducing the perceived expertise of the page by stripping out the honest acknowledgment of complexity. I've seen one client recover rankings after we restored cautionary language that a previous agency had edited out in the name of "positive user experience."

Contrarian Take 2: Your Competitors' Sentiment Scores Matter More Than Your Own

This one surprises people when I present it in workshops. The sentiment score of your page is only meaningful in relation to what Google is already rewarding for that query. If the top 5 pages for your target keyword average a VADER compound of +0.18, and your page sits at +0.22, you do not have a sentiment problem. You might have a hundred other problems, but tone is not one of them.

Conversely, if the top 5 average −0.08 (which happens in competitive, high-stakes industries) and your page is at +0.31, you may read as shallower or less rigorous than what's already ranking. The query intent encodes an expected emotional register, and departing too far from it in either direction creates friction.

Reference corpus comparison is non-optional. Without it, you're optimizing against an imaginary benchmark.

The Mistake I Made

I owe an admission here: in the early months of building this framework, I relied exclusively on VADER for all my scoring. I was confident in it because it's fast, well-documented, and returns intuitive scores. Then a client asked me to audit a series of pages in the cybersecurity vertical. VADER returned mildly negative scores on most of them. The transformer pipeline returned strongly negative scores on the same sections. When I read the text manually, the transformer was right. The content used technical hedging language and conditional threat framing that VADER's lexicon didn't adequately capture.

I had recommended against rewrites on two of those pages based on VADER alone. When we finally ran the transformer audit and made targeted changes, both pages moved up measurably within 45 days. I don't know exactly what I cost that client by the delay, but it was probably several months of compounding ranking improvement. I now treat dual-model scoring as mandatory, not optional. That's the process change I should have made earlier.

The Rewrite: From −0.14 to +0.42

Back to the original page. Here's what the SIGNAL audit found.

The intro section (paragraphs 1–4) had a VADER compound of −0.21. The writer had opened with a catalog of failure modes: things that go wrong, mistakes people make, reasons the topic is harder than it looks. Technically accurate. Completely appropriate for a knowledgeable author who wanted to establish credibility by not oversimplifying. But as an opening register, it was a cold room. Readers came in from a SERP where the competing snippets read as warm and approachable, and they landed in something that felt more like a warning label.

The value body (paragraphs 5–19) showed the variance problem I described earlier. Mean compound: −0.03. Standard deviation: 0.41. The page lurched. A section on best practices scored +0.38, immediately followed by a risk section scoring −0.44, followed by an example scoring +0.29. No section transitions, no framing that prepared readers for the emotional shifts. The page read as competent but jittery.

The CTA zone (final 3 paragraphs) scored −0.06. The writer had hedged the recommendation into near-meaninglessness. "Depending on your situation, you might consider exploring whether this approach could potentially be worth investigating." Not a sentence that drives action.

The rewrite targeted three specific interventions:

  1. Intro reframe. Lead with what success looks like, then introduce the complications. Same information, different sequencing. New intro compound: +0.17.
  2. Transition bridging. Added a single sentence between high-variance sections to signal the emotional shift to the reader. "Here's where it gets complicated" is unsophisticated, but it works. Variance dropped from 0.41 to 0.22.
  3. CTA commitment language. Stripped the hedge stack. Made a direct recommendation with a clearly bounded scope. New CTA compound: +0.38.

We did not change the factual content. We did not add word count. We did not build links. The word count actually decreased by 180 words because removing the hedge stack is inherently deflationary.

Page-level compound moved from −0.14 to +0.42. Forty-three days later, the page was at position 5. CTR improved by 34% against the same impression base. Average session duration from organic traffic increased by 1 minute 12 seconds. I can't isolate sentiment as the sole variable; no SEO experiment can fully isolate variables. But the intervention was targeted, the change was substantial, and the direction of improvement was consistent across every behavioral metric I had access to.

That's enough for me to keep doing it.

Where This Goes from Here

The integration of LLM-based content generation into publishing pipelines has created a new class of sentiment problem: machine-generated flatness. LLMs trained on RLHF objectives tend to produce text that scores consistently neutral to slightly positive on sentiment metrics, with low variance and a distinctive cadence of hedged optimism. That sounds harmless. In practice it reads as lifeless to both readers and, increasingly, to automated quality systems that can distinguish between the emotional texture of lived expertise and the smooth output of a fine-tuned model trying not to offend anyone.

The most interesting work I'm seeing in early 2026 is in using sentiment auditing as a detection layer for AI-generated homogenization rather than purely as a content quality signal. Pages that are too consistent, too moderate, too emotionally uniform may be raising flags in ways that aggregate behavioral signals will eventually reflect. Variance, in this reading, becomes a feature rather than a flaw. Real expertise has moments of strong concern, genuine enthusiasm, and honest uncertainty. Flattening those out doesn't produce trustworthiness; it produces blandness.

For SEOs who want to get ahead of this, the practical implication is clear. Stop auditing just for presence of sentiment problems. Start auditing for absence of appropriate emotional range. A page that never registers above +0.15 or below −0.05 across its entire body may be as much in need of intervention as a page that's dragging a −0.20 average.

The tools described in this piece are operational right now. The Python scripts run in any environment with the relevant packages. The BigQuery query adapts to your schema with minimal changes. The SIGNAL framework is a checklist, not proprietary software. You can begin a sentiment audit on your most underperforming informational pages this week, with no budget, and have actionable data within a few hours of crawl time.

What you do with that data is where judgment comes in. Numbers tell you where to look. They don't write the fix. That part is still a craft skill, and in a world increasingly full of content that optimizes for metric conformity, craft may end up being the actual differentiator.

I'll keep publishing what I find. The dataset grows every month, the framework evolves, and some of what I've written here will probably look naive in eighteen months. That's the point. SEO that doesn't update is just nostalgia.

If you want to dig deeper into behavioral signals and content quality, the behavioral signals guide covers the GSC data layer in more detail. For the technical setup on automated content crawling, see content crawl automation. The SERP feature optimization piece connects sentiment calibration to featured snippet eligibility in ways that surprised me when I ran the numbers. And if you're working in a YMYL vertical specifically, the YMYL content quality framework is a better starting point than this article for establishing your baseline approach.

For the underlying research on behavioral signals and search quality, this ACM paper on user engagement and search quality remains one of the more rigorous public treatments of the topic. The RoBERTa paper is worth reading if you want to understand why transformer-based sentiment models outperform lexicon approaches on complex domain text.

YOUR READING CHECKLIST

Make the ideas stick.

Mark the sections you’ve worked through. Saved in this browser.

0 of 4 reviewed
Andrii Stanetskyi
ABOUT THE AUTHOR

Andrii Stanetskyi

Head of SEO / Technical SEO Lead based in Tallinn, Estonia. Technical architecture, enterprise eCommerce, Python automation, and AI-assisted workflows.

More about Andrii ↗
LET’S FIND THE REAL BOTTLENECK

A clearer picture.
A practical next step.

Get a focused SEO audit or a consultation on your next technical decision. We’ll agree on the scope and fee before any work begins.

01 / Diagnose02 / Prioritize03 / Plan
How can I help?

Scope and fee agreed before any work begins.

Choose your language

Explore SEO services in 26 languages. Journal articles retain their original language.

ENEnglish↗DEDeutsch↗FRFrançais↗ESEspañol↗ITItaliano↗PTPortuguês↗NLNederlands↗PLPolski↗SVSvenska↗DADansk↗FISuomi↗NONorsk↗ETEesti↗LVLatviešu↗LTLietuvių↗CSČeština↗RORomână↗HUMagyar↗ELΕλληνικά↗BGБългарски↗HRHrvatski↗SKSlovenčina↗SLSlovenščina↗RUРусский↗UKУкраїнська↗TRTürkçe↗
LET’S WORK ON YOUR WEBSITE
A CLEAR NEXT STEP

Let’s talk
about your site.

A focused SEO audit or a conversation about a specific challenge. Tell me where you are and what you want to change.

Andrii Stanetskyi
Andrii StanetskyiHead of SEO / Technical SEO Lead
[email protected] ↗
How can I help?

Scope and fee agreed before any work begins.