Skip to content
DATA & AUTOMATION / FIELD NOTE 258

First-Party Data and SEO in 2026: The CDP-to-Content Loop That Lifted RPM by 4.2x

Reading map: Why Third-Party Death Matters Less Than You Think; Building the CDP Layer: Segment, RudderStack, and the Pipes Nobody Talks About; BigQuery First-Party Joins: The Query That Changed Everything; Content Personalization Patterns That Actually Move RPM
A reading map of this field note. Download SVG ↓

Why Third-Party Death Matters Less Than You Think

Everybody spent 2023 panicking about cookie deprecation. Then 2024. Then the first half of 2025. At some point you have to stop mourning the infrastructure you lost and start building the one you actually need.

I run content operations across three mid-sized publishing properties. Combined, they pull somewhere around 4.1 million organic sessions a month. In January 2025, our blended RPM across those properties sat at $9.40. By November 2025, after an 11-month CDP-to-content integration project, it reached $39.48. That's the 4.2x figure I keep citing, and I want to be precise about what actually caused it versus what I just happened to do simultaneously.

The short answer: we stopped treating SEO and first-party data as separate disciplines. The longer answer is what this article is about.

Third-party cookies were a crutch that let publishers be lazy about their own user relationships. Their deprecation forced a reckoning that, honestly, most SEO teams were not structurally equipped to have. SEO lives in content and technical infrastructure. First-party data lives in CDP tooling and data engineering. The org chart kept these worlds apart, and the Google algorithm didn't care about org charts.

What changed for us was a single realization: the signals that predict content resonance are sitting in our own event streams, not in a third-party audience segment we rented. When someone reads 73% of an article about variable annuity riders, spends 4 minutes 12 seconds on it, then bounces to a Google search for "indexed universal life vs variable annuity," that behavioral sequence is enormously valuable editorial signal. And we were throwing it away.

Building the CDP Layer: Segment, RudderStack, and the Pipes Nobody Talks About

Why We Started with Segment

In February 2025, we instrumented our primary property with Segment. Not because it was the cheapest option (it wasn't) but because the destination ecosystem was mature enough to reduce custom engineering. We needed data flowing to BigQuery, to our ad server, and to a nascent personalization layer in under six weeks. Segment's prebuilt BigQuery destination cut maybe three weeks off that timeline.

Here's the destination config we settled on after two rounds of iteration:

{
  "name": "BigQuery - Editorial Behavior",
  "type": "BIGQUERY",
  "config": {
    "projectId": "your-gcp-project-id",
    "datasetId": "segment_editorial_raw",
    "tableSuffix": "_events",
    "uploadInterval": 60,
    "enablePartitionDecorator": true,
    "partitionField": "received_at",
    "mergeStateFilePath": "gs://your-bucket/segment-merge-state/",
    "transformations": [
      {
        "name": "strip_pii",
        "type": "REMOVE_FIELD",
        "fields": ["email", "phone", "ip"]
      }
    ],
    "eventFilters": {
      "allow": [
        "Article Viewed",
        "Article Scrolled",
        "Outbound Link Clicked",
        "Search Query Submitted",
        "Session Started",
        "Ad Impression Recorded",
        "Ad Click Recorded"
      ]
    }
  }
}

The uploadInterval of 60 seconds matters for near-real-time editorial feedback loops. We experimented with 300 seconds and found it created too much latency in our content performance dashboards. The PII stripping transformation was non-negotiable from a compliance standpoint; we had legal review the config before any data moved.

The RudderStack Migration on Properties Two and Three

For properties two and three, we chose RudderStack over Segment. Mostly a cost decision once we understood our event volume (around 47 million events per month across both), but also because RudderStack's open-source core gave us more flexibility for custom transformations without paying per-transformation fees.

The workflow that powers our content signal pipeline looks like this:

# RudderStack Transformation: Content Engagement Scoring
export function transformEvent(event, metadata) {
  if (event.type !== 'track') return event;

  const engagementEvents = [
    'Article Scrolled',
    'Article Viewed',
    'Time on Page Recorded'
  ];

  if (!engagementEvents.includes(event.event)) return event;

  const scrollDepth = event.properties?.scroll_depth || 0;
  const timeOnPage = event.properties?.time_on_page_seconds || 0;
  const wordCount = event.properties?.article_word_count || 1000;

  // Normalize time-on-page against estimated read time
  const estimatedReadTime = (wordCount / 238) * 60; // 238 wpm average
  const readTimeRatio = Math.min(timeOnPage / estimatedReadTime, 1.5);

  // Composite engagement score (0–100)
  const engagementScore = Math.round(
    (scrollDepth * 0.4) +
    (readTimeRatio * 100 * 0.45) +
    (event.properties?.internal_link_clicked ? 15 : 0)
  );

  event.properties.engagement_score = Math.min(engagementScore, 100);
  event.properties.content_tier = engagementScore >= 70 ? 'high' :
    engagementScore >= 40 ? 'medium' : 'low';

  return event;
}

This transformation runs in-stream before data reaches BigQuery. The engagement score it produces became the central variable in our content optimization loop. I'll come back to that.

One thing nobody warns you about with RudderStack transformations: error handling is brutal in production. If your transformation throws on an unexpected event shape, you get silent drops. We lost 6 days of clean data in May 2025 because a downstream property pushed a CMS update that changed the article_word_count field name to wordCount (camelCase). The transformation divided by undefined, scored everything zero, and we didn't catch it for nearly a week. I'll revisit that mistake in more detail later.

BigQuery First-Party Joins: The Query That Changed Everything

Getting data into BigQuery is table stakes. What you do with it is where the SEO story actually begins.

The query that fundamentally changed how we approach content planning runs weekly. It joins three datasets: the Segment/RudderStack behavioral events, our Google Search Console data (pulled via API into BigQuery), and our CMS metadata table. The goal is to identify mismatches between what users engage with deeply and what we're ranking for.

-- Weekly Content-Signal Mismatch Report
WITH high_engagement_content AS (
  SELECT
    e.properties.page_path AS page_path,
    e.properties.article_id AS article_id,
    AVG(e.properties.engagement_score) AS avg_engagement,
    COUNT(DISTINCT e.anonymous_id) AS unique_readers,
    APPROX_QUANTILES(e.properties.scroll_depth, 100)[OFFSET(50)] AS median_scroll_depth
  FROM
    your-project.segment_editorial_raw.Article_Scrolled_events e
  WHERE
    e.received_at >= TIMESTAMP_SUB(CURRENT_TIMESTAMP(), INTERVAL 30 DAY)
    AND e.properties.engagement_score IS NOT NULL
  GROUP BY 1, 2
  HAVING unique_readers >= 50
),
search_performance AS (
  SELECT
    page,
    SUM(impressions) AS total_impressions,
    SUM(clicks) AS total_clicks,
    AVG(position) AS avg_position,
    SUM(clicks) / NULLIF(SUM(impressions), 0) AS ctr
  FROM
    your-project.search_console.searchdata_url_impression
  WHERE
    data_date >= DATE_SUB(CURRENT_DATE(), INTERVAL 30 DAY)
  GROUP BY 1
),
cms_metadata AS (
  SELECT
    url_path,
    article_id,
    primary_topic,
    word_count,
    published_date,
    last_updated_date
  FROM
    your-project.cms.articles
)
SELECT
  h.page_path,
  h.avg_engagement,
  h.unique_readers,
  h.median_scroll_depth,
  s.total_impressions,
  s.avg_position,
  s.ctr,
  m.primary_topic,
  m.word_count,
  -- The signal mismatch: high engagement, low search visibility
  CASE
    WHEN h.avg_engagement >= 65 AND s.avg_position > 20 THEN 'HIGH_PRIORITY_REFRESH'
    WHEN h.avg_engagement >= 65 AND s.total_impressions < 500 THEN 'INDEXATION_REVIEW'
    WHEN h.avg_engagement < 35 AND s.avg_position <= 5 THEN 'CONTENT_QUALITY_RISK'
    ELSE 'MONITOR'
  END AS content_signal_flag
FROM high_engagement_content h
LEFT JOIN search_performance s
  ON CONCAT('https://your-domain.com', h.page_path) = s.page
LEFT JOIN cms_metadata m
  ON h.article_id = m.article_id
ORDER BY h.avg_engagement DESC, s.avg_position ASC;

The HIGH_PRIORITY_REFRESH flag is where we find our best opportunities. These are articles that real users love, demonstrated by behavioral signals, but that haven't achieved search visibility. The working theory (validated over nine months of testing) is that deep engagement predicts topical authority signals, which predict ranking improvement after a refresh. We've seen average position improvement of 6.3 positions within 60 days of refreshing a HIGH_PRIORITY_REFRESH article. That's not universal, but it's consistent enough that it drives our editorial calendar now.

The CONTENT_QUALITY_RISK flag is equally important and more uncomfortable to act on. Articles ranking in position 1-5 but generating engagement scores below 35 are a liability. They're drawing clicks they can't satisfy. We've preemptively deprecated or substantially rewritten 23 such articles since September 2025. In 19 of those cases, the page either held its ranking or improved after the rewrite. In 4 cases, we lost the ranking temporarily. Worth it every time.

See also: [Internal: Content Decay Detection Framework] for how we operationalize the CONTENT_QUALITY_RISK bucket.

Content Personalization Patterns That Actually Move RPM

The Personalization Misconception

Most publishers hear "personalization" and imagine Netflix-style recommendation engines. That's not what we built. We built something much narrower and, I'd argue, more defensible from an SEO perspective.

The pattern we use is segment-aware content assembly, not dynamic content replacement. Meaning: the page a user lands on from search is static, fully crawlable, indexable without condition. What changes is the contextual module set that loads after the initial render, based on behavioral cohort data passed through our CDP. Google sees the canonical content. The user sees an enhanced experience layered on top.

// Content personalization cohort loader
// Reads Hightouch-synced cohort from edge KV store
async function loadContentCohort(anonymousId) {
  const cohortKey = cohort:${anonymousId};

  try {
    const cohortData = await KV_STORE.get(cohortKey, { type: 'json' });

    if (!cohortData) {
      return { cohort: 'default', interests: [], engagementTier: 'unknown' };
    }

    return {
      cohort: cohortData.primary_cohort,
      interests: cohortData.topic_affinities || [],
      engagementTier: cohortData.engagement_tier,
      lastSeenTopics: cohortData.recent_topics?.slice(0, 5) || [],
      adTargetingSegments: cohortData.ad_segments || []
    };
  } catch (err) {
    console.error('Cohort load failed:', err);
    return { cohort: 'default', interests: [], engagementTier: 'unknown' };
  }
}

// Module selection based on cohort
function selectContextualModules(cohortData, articleMeta) {
  const modules = ['related-articles']; // always shown

  if (cohortData.engagementTier === 'high') {
    modules.push('deep-dive-cta');
    modules.push('newsletter-specialist');
  }

  if (cohortData.interests.includes(articleMeta.primaryTopic)) {
    modules.push('topic-series-nav');
  }

  if (cohortData.adTargetingSegments.length > 0) {
    modules.push('contextual-ad-enhanced');
  } else {
    modules.push('contextual-ad-default');
  }

  return modules;
}

The contextual-ad-enhanced module is where RPM lifts materially. When we can pass first-party audience segment data to our ad server at render time, CPMs on those impressions are consistently 2.1x to 3.8x higher than the default contextual rate. Across high-engagement pages (where cohort data is most likely to be present), this compounds quickly.

Related reading on the ad stack side: [Internal: Programmatic Architecture for Content Publishers, 2026].

What the Numbers Actually Look Like

In October 2025, we ran an A/B test across property one. The control group received no cohort-aware module loading. The test group received the full personalization stack. Results over 30 days:

  • RPM: $31.20 control vs. $44.70 test (43.3% lift)
  • Pages per session: 1.84 control vs. 2.31 test (25.5% lift)
  • Return visitor rate (14-day window): 11.2% control vs. 17.8% test (58.9% lift)
  • Avg session duration: 2m 14s control vs. 3m 02s test (35.8% lift)

The return visitor rate is the number I keep coming back to. Not because it's the biggest lift, but because return visitors generate richer CDP profiles, which makes the personalization more accurate over time, which improves engagement scores, which feeds the content signal mismatch query at the top of the stack. The loop compounds.

The Mistake I Made in Q3 2025 (And What It Cost)

I mentioned the RudderStack transformation bug briefly. Let me be more specific about what happened and what it cost, because I think transparency here is useful for anyone building similar systems.

Between May 14 and May 20, 2025, our engagement scoring transformation was silently failing for property two and property three due to the field name mismatch (article_word_count vs. wordCount). Every article scored zero. Every article got flagged as content_tier: 'low'.

We had a weekly Hightouch sync that pushed those content tier signals into our ad server targeting logic. So for one full week, our ad server was treating all content on two properties as low-tier, which depressed CPMs by an average of 41% across those properties for that period.

Estimated revenue impact: approximately $23,400 across the two properties over that six-day window. Real money.

The fix was embarrassingly simple: a null-check and a field alias fallback in the transformation function. The lesson was systemic: we had no alerting on downstream data quality. We were monitoring for event volume (events are arriving, system is running) but not for score distribution (are the scores within expected ranges?).

We now run a daily BigQuery job that alerts if the mean engagement score for any property drops below 30 or above 90. Both are anomaly signals. That check has caught two subsequent data quality issues before they compounded.

If you are building first-party data pipelines for editorial use, instrument your outputs as carefully as your inputs. Event volume tells you the pipe is flowing. Score distribution tells you what's actually in the pipe.

The FOCAL Framework: My Personal System for CDP-to-Content Loops

After running this project for 14 months across three properties, I've distilled the operational model into a framework I call FOCAL. It's not designed to be published in a marketing deck. It's designed to survive contact with an actual editorial team that also has to publish 40 articles a week.

F — Flow. Data must move continuously from user behavior to analytical layer. Not in batches that accumulate errors, not in daily dumps that are stale by the time anyone acts on them. Sub-five-minute latency from event to BigQuery is the target we built toward. The Segment 60-second upload interval was part of this.

O — Observation. The BigQuery signal mismatch query is the observation layer. This is where you watch what users do versus what search surfaces. Most SEO teams do keyword research. Observation replaces the keyword research process with behavioral evidence. You still check search volume, but user behavior overrides it when they conflict.

C — Cohort. Users must be grouped into actionable behavioral cohorts, not demographic guesses. Our cohorts are built on topic affinity (derived from what someone reads, not what they say they are), engagement tier (derived from scroll depth and time-on-page patterns), and recency (how recently they exhibited those behaviors). Hightouch syncs these cohorts to every destination that needs them, on a schedule we control.

A — Activate. Cohorts do nothing until activated at an endpoint. For us, activation means: the ad server gets first-party segment identifiers at ad request time, the content module selector gets engagement tier at page render time, and the editorial team gets the signal mismatch report at planning time. Three activation points, three measurable outcomes.

L — Loop. The loop closes when activation improves the behavioral signals that feed Flow. Higher-quality ad experiences improve user tolerance for advertising, which reduces rage-clicks and improves session quality, which improves engagement scores. Better content refreshes improve ranking positions, which bring higher-intent organic visitors, who generate richer behavioral signals. The loop doesn't compound overnight. It compounded over nine months for us.

See also: [Internal: Implementing FOCAL in Content Operations] for the operational runbook.

Two Things Everyone Gets Wrong About First-Party SEO

Contrarian Take One: Engagement Metrics Are Not Ranking Factors, and That's Not the Point

Google has been clear (and I believe them, which itself might be the contrarian position) that user engagement metrics from your site are not direct ranking factors. They don't have access to your scroll depth. They're not reading your session duration. The FOCAL framework is not predicated on fooling Google into thinking your content is better.

The point is different and more interesting: the behavioral signals that predict high engagement also predict the on-page quality characteristics that do influence rankings. Deep engagement correlates with topical comprehensiveness, with accurate answers to implicit sub-questions, with appropriate reading level matching. When you use first-party data to find your highest-engagement content and use it as a template for what to refresh, you're not gaming signals. You're reverse-engineering what quality actually looks like in your specific niche with your specific audience.

The RPM lift comes from the ad side, where first-party data is a direct, documented, well-understood commercial input. The SEO lift comes from making better editorial decisions faster. These are separate mechanisms that stack.

Contrarian Take Two: You Probably Have Enough Data Already

The CDP consulting industry has a vested interest in making first-party data collection seem complicated and incomplete. "You need more events. You need richer profiles. You need a data warehouse and a reverse ETL layer and a real-time feature store." Maybe eventually. Not to start.

We generated meaningful signal mismatch insights with three event types: Article Viewed, Article Scrolled, and Session Started. If your CMS logs a pageview and you have a JavaScript listener on scroll, you have what you need to run a version of that BigQuery query. The fancy transformation logic and the cohort activation layers came later and added incremental value. They were not prerequisites.

The minimum viable first-party SEO setup is: capture scroll depth with an event, get it into a queryable store, join it against Search Console data once a week. That's it. You can build from there. Don't let the vision of the complete system stop you from running the simplest version today.

For external context on what "sufficient" first-party data means for publisher monetization, the IAB's guidance on first-party data for publishers is the most grounded reference I've found. They're not trying to sell you a CDP.

Hightouch is the piece of the stack I'm most reluctant to write about, because it's also the piece most dependent on your specific ad and CMS architecture. But the conceptual pattern is worth explaining.

We use Hightouch to sync cohort membership data from BigQuery to four destinations: our ad server (Xandr), our email platform, our CMS's personalization layer, and a Slack channel for the editorial team. That last destination is not a joke. Hightouch's Slack destination is how our editors see, every Monday morning, which content categories generated the highest engagement scores in the prior week. It's the most-read Slack message in our organization.

The Hightouch model that powers the ad-server sync is straightforward:

-- Hightouch Model: Ad Server Audience Sync
SELECT
  u.anonymous_id,
  u.primary_cohort,
  u.engagement_tier,
  ARRAY_TO_STRING(u.topic_affinities, ',') AS topic_affinity_string,
  u.ad_segments,
  u.last_seen_at
FROM (
  SELECT
    e.anonymous_id,
    -- Primary cohort: the topic with highest cumulative engagement
    ARRAY_AGG(
      m.primary_topic ORDER BY SUM(e.properties.engagement_score) DESC
      LIMIT 1
    )[OFFSET(0)] AS primary_cohort,
    -- Engagement tier based on median score in last 30 days
    CASE
      WHEN AVG(e.properties.engagement_score) >= 65 THEN 'high'
      WHEN AVG(e.properties.engagement_score) >= 35 THEN 'medium'
      ELSE 'low'
    END AS engagement_tier,
    ARRAY_AGG(DISTINCT m.primary_topic ORDER BY m.primary_topic) AS topic_affinities,
    MAX(e.received_at) AS last_seen_at
  FROM
    your-project.segment_editorial_raw.Article_Scrolled_events e
  JOIN
    your-project.cms.articles m
    ON e.properties.article_id = m.article_id
  WHERE
    e.received_at >= TIMESTAMP_SUB(CURRENT_TIMESTAMP(), INTERVAL 30 DAY)
    AND e.properties.engagement_score >= 40
  GROUP BY e.anonymous_id
  HAVING COUNT(DISTINCT DATE(e.received_at)) >= 2 -- at least 2 distinct visit days
) u
CROSS JOIN UNNEST(
  CASE
    WHEN u.engagement_tier = 'high' AND 'personal-finance' IN UNNEST(u.topic_affinities)
      THEN ['seg_higheng_pf', 'seg_premium']
    WHEN u.engagement_tier = 'high'
      THEN ['seg_higheng_general']
    ELSE ['seg_standard']
  END
) AS ad_segment
GROUP BY u.anonymous_id, u.primary_cohort, u.engagement_tier, u.topic_affinities, u.last_seen_at

The HAVING COUNT(DISTINCT DATE(e.received_at)) >= 2 clause is important. It filters out drive-by readers and ensures only return visitors with demonstrated cross-session interest are synced to ad targeting. This keeps our audience quality high, which keeps CPMs high. Volume of audience matters less than signal quality.

Additional background on reverse ETL patterns for publishers: [Internal: Reverse ETL for Editorial Properties: A Practical Guide].

For authoritative external reading on CDP-to-ad-server integration patterns, Segment's destination documentation remains the most complete reference available publicly, despite being written by a vendor with obvious interests.

What This Means in May 2026

The 4.2x RPM lift is real and documented. I've been careful throughout this article to separate the mechanisms: ad revenue lifts from first-party data activation, SEO improvement from better content decisions informed by behavioral evidence. Both are happening. They're related but not identical.

What I believe as of today, May 20, 2026: the publisher properties that survive the next algorithmic cycle are not the ones with the most content or the fastest publishing velocity. They're the ones that have built feedback loops between what users do and what editors create. First-party data is the substrate that makes that loop possible at scale.

The tools (Segment, RudderStack, Hightouch, BigQuery) are available to anyone. The integrations are documented. The hard part is organizational: convincing an editorial team that behavioral data is editorial signal, not just an analytics afterthought. That took us about four months to accomplish and was more important than any technical configuration we wrote.

RPM of $39.48 is not our ceiling. The cohort quality improves every month as the dataset matures. The content refresh pipeline gets smarter as we accumulate more matched pairs of "high engagement score + search performance delta." The loop is still tightening.

If you're starting from zero today, don't start with the complete FOCAL stack. Start with scroll depth events and a BigQuery table. Run the signal mismatch query against Search Console. Find one article with high engagement and poor rankings. Refresh it. Measure what happens. That's the first iteration. Everything else in this article is what comes after you believe the loop is real.

YOUR READING CHECKLIST

Make the ideas stick.

Mark the sections you’ve worked through. Saved in this browser.

0 of 4 reviewed
Andrii Stanetskyi
ABOUT THE AUTHOR

Andrii Stanetskyi

Head of SEO / Technical SEO Lead based in Tallinn, Estonia. Technical architecture, enterprise eCommerce, Python automation, and AI-assisted workflows.

More about Andrii ↗
LET’S FIND THE REAL BOTTLENECK

A clearer picture.
A practical next step.

Get a focused SEO audit or a consultation on your next technical decision. We’ll agree on the scope and fee before any work begins.

01 / Diagnose02 / Prioritize03 / Plan
How can I help?

Scope and fee agreed before any work begins.

Choose your language

Explore SEO services in 26 languages. Journal articles retain their original language.

ENEnglish↗DEDeutsch↗FRFrançais↗ESEspañol↗ITItaliano↗PTPortuguês↗NLNederlands↗PLPolski↗SVSvenska↗DADansk↗FISuomi↗NONorsk↗ETEesti↗LVLatviešu↗LTLietuvių↗CSČeština↗RORomână↗HUMagyar↗ELΕλληνικά↗BGБългарски↗HRHrvatski↗SKSlovenčina↗SLSlovenščina↗RUРусский↗UKУкраїнська↗TRTürkçe↗
LET’S WORK ON YOUR WEBSITE
A CLEAR NEXT STEP

Let’s talk
about your site.

A focused SEO audit or a conversation about a specific challenge. Tell me where you are and what you want to change.

Andrii Stanetskyi
Andrii StanetskyiHead of SEO / Technical SEO Lead
[email protected] ↗
How can I help?

Scope and fee agreed before any work begins.