Skip to content
AI & SEARCH / FIELD NOTE 182

Tracking 14,000 AI Citations a Week in 2026: My Stack After GA4 Stayed Useless

Reading map: Why GA4 Is Still the Wrong Tool for This; What "Tracking a Citation" Actually Means; The VELA Measurement Model; The SaaS Layer: Profound and AthenaHQ
A reading map of this field note. Download SVG ↓

Published 19 May 2026. All numbers come from my own monitoring setup across 43 sites unless otherwise noted. I am not paid by Profound, AthenaHQ, or any vendor mentioned here.

Why GA4 Is Still the Wrong Tool for This

I want to say this plainly because I keep seeing people report "AI traffic from GA4" as if it is a meaningful number: it is not. Not in any reliable sense.

Here is what GA4 sees in May 2026. ChatGPT.com referrals: real, but partial. When a user clicks a citation link from ChatGPT's Browse-enabled web interface, GA4 captures it as a referral from chatgpt.com. That is good. But: users who copy-paste a URL from a ChatGPT response into a new tab? Direct traffic. Users on mobile apps for ChatGPT or Claude? Partial or no referrer. Users who follow a Perplexity citation link? Often shows as perplexity.ai referral, but only for the web interface — the mobile app is noisier.

Gemini's AI Overviews in Google Search don't even have an obvious referrer. The user is already on Google. The click goes to your site from the SERP, indistinguishable in GA4 from an organic click. Google Search Console might differentiate it as "AI Overview impression" but the attribution is still opaque at the session level.

I ran a test in November 2025 across seven sites where I manually tracked 40 citation events (I prompted the AI tools myself, clicked through, verified the sessions in real time) and compared to what GA4 reported. GA4 captured 17 of the 40. Missed 14 entirely (direct attribution). Miscategorized 9 (appeared as organic or direct). This is not a GA4-is-broken story — it is a structural problem with how AI interfaces pass referrer data, and GA4 cannot fix what the browser does not send.

The answer is a parallel monitoring system. Not a replacement for GA4. A parallel layer that tracks citation presence on the AI platform side, independent of what traffic eventually appears in your analytics.

What "Tracking a Citation" Actually Means

Before I describe the stack, a definitions problem worth addressing. A "citation" in my tracking data is not a single thing. I track four distinct event types, which I first described in the context of the PACE framework piece:

  1. Named attribution: AI response mentions your brand name or URL in prose. "According to Benrey..." Highest value. Easiest to track.
  2. Source listing: Your URL appears in the footnote/source list of a Perplexity or ChatGPT Browse response without named prose attribution. Trackable. Medium value.
  3. Paraphrase absorption: The AI uses your content structure or specific phrasing without attribution. Detectable only via careful response parsing and comparison to your source content.
  4. Training-data influence: The model's priors reflect your content from training, not retrieval. Auditable only through the training-data exposure audit services — a real commercial category as of early 2026, covered in depth in the training-data audit piece.

My tracking stack covers types 1 and 2 reliably, type 3 partially, and type 4 not at all in real time. Anyone claiming to track type 4 in real time is either lying or working with a very narrow definition of "track."

The VELA Measurement Model

Raw citation counts are almost meaningless without context. A brand mentioned 200 times in AI responses — but always buried in a list of eight alternatives, never first — is different from being mentioned 47 times as the primary recommended source. I needed a quality-weighted model, not just a count.

VELA is what I use. Built it in August 2025 and have not changed it significantly since, which I take as a sign it is reasonably well-calibrated.

V — Volume. Raw citation event count per week, per AI system, per query cluster. Useful for trend detection. Meaningless in isolation.

E — Exactness. Is the citation a direct named attribution (score: 3), a source listing (score: 2), a paraphrase detection (score: 1), or an implied stylistic match (score: 0.5)? I weight volume by exactness to get a "quality-adjusted citation volume" per week.

L — Location. Where in the AI response does the citation appear? First sentence or lead recommendation (score: 3), middle of response (score: 2), end of response or footnote (score: 1). A client cited 80 times in Perplexity footnotes has a worse L-score than a client cited 30 times as the lead source. The VELA composite accounts for this.

A — Attribution completeness. Does the citation include your brand name? Your URL? Just a paraphrase? I score: brand + URL (score: 3), brand only (score: 2), URL only (score: 1.5), none (score: 0). This matters because brand conditioning — the effect where users who see your brand in AI responses arrive at your site with higher trust — requires the brand name to appear, not just the URL.

The VELA composite per week for a given client: sum(V × E_weight × L_weight × A_weight) across all tracked prompts. I normalize it to a 0–100 scale relative to the best-performing domain in that vertical within my dataset. Not a publishable metric. A practical dashboard number.

The SaaS Layer: Profound and AthenaHQ

I run both tools. They are not redundant — they cover different ground.

Profound

Profound's core feature is prompt-replay at scale. You give it a library of queries relevant to your brand and vertical. It sends those queries to ChatGPT (GPT-4o) and Gemini on a daily or weekly schedule and logs whether your brand or URLs appear in the response, where they appear, and what the surrounding text says. The reporting UI shows citation rate trends over time and — this is the feature I use most — competitor share-of-voice, meaning how often your competitors appear in response to the same prompts.

Things Profound does well: trend detection over time, competitor comparison, Gemini coverage. Things it does less well: Perplexity tracking (the source-URL parsing for Perplexity is less robust than AthenaHQ's), real-time alerting, and any attempt to capture paraphrase-level absorption (type 3 above).

Pricing: starts around $500/month for moderate prompt libraries. Enterprise tiers go much higher and involve custom prompt-library builds. I use it for all clients at or above $3K/month in SEO retainer value because the time savings in reporting justify the cost. Below that, I use DIY.

AthenaHQ

AthenaHQ is purpose-built around Perplexity citation tracking. The reason this matters: Perplexity's interface shows inline source citations with URLs, which is structurally more parseable than ChatGPT's prose attribution. AthenaHQ has built its indexing around that parseability, and the Perplexity coverage is noticeably more accurate than Profound's.

I use AthenaHQ as the Perplexity-specific layer and Profound as the ChatGPT/Gemini layer. The overlap is intentional — I compare their outputs for the same prompts as a quality check. When they disagree on whether my client appeared in a Gemini response, I re-run the prompt manually and audit.

A Note on This Tool Category

AI citation tracking as a distinct software category did not exist in any serious commercial form before Q4 2024. By Q1 2026 it is a recognizable category with at least six vendors I am aware of and three that are well-funded enough to have customer support teams. This is still early. The tools are imperfect. The category will consolidate. Do not build your reporting infrastructure around any single vendor's data model right now.

The DIY Scraper: Playwright + Python

For clients where SaaS cost is not justified, or for high-volume prompt testing I don't want to pay per-prompt API fees for, I run a Playwright-based scraper against Perplexity's web interface. It is fragile and requires maintenance every four to six weeks as the site interface changes, but it works well enough for trend data.


import asyncio
import json
import re
from datetime import datetime
from playwright.async_api import async_playwright

PROMPTS = [
    "best practices for AI-generated content in 2026",
    "how to optimize content for Perplexity citations",
    "GEO generative engine optimization explained",
    # add your target queries here
]

TARGET_DOMAINS = [
    "benrey.io",
    "yourdomain.com",
]

async def query_perplexity(page, prompt):
    await page.goto("https://www.perplexity.ai/", wait_until="networkidle")
    await page.wait_for_selector('textarea[placeholder*="Ask"]', timeout=15000)
    await page.fill('textarea[placeholder*="Ask"]', prompt)
    await page.keyboard.press("Enter")
    await page.wait_for_timeout(8000)  # wait for response

    # grab source URLs from the citation list
    try:
        source_elements = await page.query_selector_all('[data-testid="source-item"] a, .citation-url a')
        sources = []
        for el in source_elements:
            href = await el.get_attribute("href")
            if href:
                sources.append(href)
    except Exception:
        sources = []

    # grab the main response text
    try:
        response_el = await page.query_selector('.prose, [class*="answer"]')
        response_text = await response_el.inner_text() if response_el else ""
    except Exception:
        response_text = ""

    return {"prompt": prompt, "sources": sources, "response_text": response_text}

def score_citation(result, targets):
    citations = []
    for domain in targets:
        in_sources = any(domain in url for url in result["sources"])
        in_prose = domain in result["response_text"]
        if in_sources or in_prose:
            citations.append({
                "domain": domain,
                "in_sources": in_sources,
                "in_prose": in_prose,
                "source_position": next(
                    (i+1 for i, url in enumerate(result["sources"]) if domain in url),
                    None
                )
            })
    return citations

async def run_tracking_session(prompts, targets):
    results = []
    async with async_playwright() as p:
        browser = await p.chromium.launch(headless=True)
        context = await browser.new_context(
            user_agent="Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) AppleWebKit/537.36"
        )
        page = await context.new_page()
        for prompt in prompts:
            print(f"Querying: {prompt[:60]}...")
            result = await query_perplexity(page, prompt)
            result["citations"] = score_citation(result, targets)
            result["timestamp"] = datetime.utcnow().isoformat()
            results.append(result)
            await asyncio.sleep(4)  # rate limiting
        await browser.close()
    return results

if __name__ == "__main__":
    results = asyncio.run(run_tracking_session(PROMPTS, TARGET_DOMAINS))
    with open(f"citations_{datetime.utcnow().strftime('%Y%m%d')}.json", "w") as f:
        json.dump(results, f, indent=2)
    cited = sum(1 for r in results if r["citations"])
    print(f"\nCited in {cited}/{len(results)} prompts ({100*cited//len(results)}%)")

A few practical notes on this: Perplexity will occasionally serve a CAPTCHA or rate-limit aggressively if your scraper runs too many queries in sequence. I add randomized delays between 3–7 seconds and rotate sessions across a weekly batch. I also maintain a headful (non-headless) mode for debugging when the selector paths break after a site update — which happens.

The script above is a simplified version. My production version includes: delta detection (alerting when citation rate drops more than 12% week-over-week for any prompt cluster), response-text diffing to catch when paraphrase patterns match my clients' content, and a SQLite backend for multi-week trend storage.

Correlating Citation Data with Actual Traffic

This is where it gets genuinely interesting and also genuinely murky.

Across my tracked sites, the correlation between VELA composite score and traffic from AI-attributed referral sessions (the portion GA4 does successfully capture, primarily chatgpt.com and perplexity.ai referrals) is 0.54 over a 16-week window. That is meaningful but not tight. The R-squared of 0.29 means citation volume explains less than a third of the variance in AI-referred traffic.

What explains the rest? Three things, based on digging into the outliers:

First, query intent. Some prompt clusters I track have high citation rates but low click-through because the AI's response is fully satisfying — the user does not need to click. Informational queries with a simple factual answer. Other clusters have lower citation rates but much higher click-through per citation because the user needs the full resource. "What is X" prompts generate citations that rarely get clicked. "How do I implement X" prompts generate citations that often do.

Second, UI differences by platform. Perplexity's interface shows source citations prominently with the domain name and a snippet; click-through per citation is higher than ChatGPT, where citations in Browse mode often appear at the bottom of a long response. Same citation event, very different click probability.

Third, the brand conditioning effect I mentioned in the GEO advanced playbook. Some of the traffic lift from high VELA scores shows up in direct and branded search, not in referral sessions. GA4 attributes it to "direct" or "organic branded." The true attribution chain is: AI citation → brand familiarity → later branded search → conversion. Tracking that chain requires multi-touch attribution work that most clients are not set up for.

The Decision I Reversed After Three Months

In October 2025, I decided to consolidate citation tracking entirely into Profound and shut down my DIY scraper, reasoning that the SaaS tool was "good enough" and the maintenance burden of the scraper was not worth the incremental signal.

Three months later I turned the scraper back on. The reason: Profound's prompt library is configured by me, but the prompt weighting and response parsing logic is Profound's black box. When I noticed that Perplexity citation rates for one client appeared flat in Profound while my manual spot-checks showed improvement, I dug in and found that Profound's Perplexity parsing was missing citations in a specific UI state — Perplexity had shipped a UI change that moved source URLs to a collapsed "show sources" widget, and Profound had not updated their parser yet. I only caught this because I had been spot-checking manually.

The lesson: never fully outsource measurement to a tool you cannot audit. The DIY scraper is my audit layer. It runs a smaller prompt set — 60 prompts per week versus Profound's 340 — but it gives me ground truth I can verify down to the HTML.

Contrarian Positions on Citation Tracking

Contrarian Take 1: Citation Rate Is a Vanity Metric Without Revenue Linkage

I have seen agencies report citation rate improvements to clients who then renewed contracts based on those numbers while their actual revenue from AI-sourced traffic was functionally zero. Citation rate is a leading indicator, not an outcome. If you cannot draw a line from citation events to brand impressions to eventual conversion — even a probabilistic, multi-touch line — you are measuring something that feels like progress but might not be.

I still track it. But I track it alongside GA4 AI-referral revenue, branded search volume trends, and conversion rate comparisons across segments with high vs. low AI citation exposure. The citation rate alone is not the number I report to clients. The VELA composite correlated against revenue segments is.

Contrarian Take 2: The Tools Are All Behind the AI Interfaces They're Tracking

Profound, AthenaHQ, and every other citation tracking tool I know of are measuring AI outputs as they exist at the moment of measurement. But AI interfaces change rapidly — new UI states, new retrieval logic, new model versions — and the measurement tools lag. The Perplexity parsing issue I described above is one example. ChatGPT's Browse mode has changed response formats at least four times since January 2025. Every format change potentially breaks parser logic that tool vendors have to chase.

This is not a reason to avoid the tools. It is a reason to treat the data as directional trend information rather than precise measurement. When the trend line looks wrong, it might be the AI changed. Or it might be your content lost favor. You need to know the difference.

Operational Reality: What This Actually Costs

For a client at a $4,000/month SEO retainer, my citation tracking overhead is roughly:

  • Profound subscription: $500/month (shared across multiple clients at a firm tier)
  • AthenaHQ: $400/month (similar shared tier)
  • Infrastructure for DIY scraper (a small VPS): $18/month
  • My time for prompt library management and anomaly investigation: ~3 hours/month per client

That is not trivial for a small client. At sub-$2K/month retainers, I run DIY only and skip the SaaS tools. The tracking quality drops but the cost is appropriate to the engagement size.

I also bill a one-time setup fee for citation tracking infrastructure — typically $800–1,200 — to cover the initial prompt library build, which is the most time-intensive part. Ongoing maintenance is lower.

What I Know Now That I Didn't in October 2024

October 2024 is when I started taking GEO seriously as a measurable discipline rather than a theoretical one. Eighteen months of data later:

I know that citation rates move slowly. Median lag from content publishing to first citation: 23 days in my dataset. Expecting weekly or even biweekly citation improvements from content changes is too optimistic. Monthly trend windows are the minimum meaningful measurement period.

I know that the prompt library design is most of the work. A poorly designed prompt library — one that asks generic questions your clients' content happens to rank for, rather than questions that reflect actual user intent in AI searches — will produce artificially high citation rates that do not translate to traffic or brand value. I have rebuilt three clients' prompt libraries from scratch after discovering this.

I know that prompt-injection in competitor content is now a real concern in my tracking workflow. I started noticing in Q1 2026 that several competitor pages for one client's vertical were embedding text designed to redirect AI responses toward their own URLs. The patterns are specific and detectable. I document them in the prompt-injection defense piece.

And I know — this is the uncomfortable one — that a significant fraction of what appears to be "AI citation success" in any given month is statistical noise in small prompt sets. The way I manage this: I never report a single-week citation spike as a win. I report 4-week rolling averages with a minimum of 100 prompt events in the measurement window. Below that threshold, the signal-to-noise ratio is too poor to act on.


Related reading: GEO Advanced Playbook (PACE framework) | Prompt-Injection Defense for Content Teams | Training-Data Exposure Audits

External references: Profound (AI visibility platform) | AthenaHQ (AI citation tracking)

YOUR READING CHECKLIST

Make the ideas stick.

Mark the sections you’ve worked through. Saved in this browser.

0 of 4 reviewed
Andrii Stanetskyi
ABOUT THE AUTHOR

Andrii Stanetskyi

Head of SEO / Technical SEO Lead based in Tallinn, Estonia. Technical architecture, enterprise eCommerce, Python automation, and AI-assisted workflows.

More about Andrii ↗
LET’S FIND THE REAL BOTTLENECK

A clearer picture.
A practical next step.

Get a focused SEO audit or a consultation on your next technical decision. We’ll agree on the scope and fee before any work begins.

01 / Diagnose02 / Prioritize03 / Plan
How can I help?

Scope and fee agreed before any work begins.

Choose your language

Explore SEO services in 26 languages. Journal articles retain their original language.

ENEnglish↗DEDeutsch↗FRFrançais↗ESEspañol↗ITItaliano↗PTPortuguês↗NLNederlands↗PLPolski↗SVSvenska↗DADansk↗FISuomi↗NONorsk↗ETEesti↗LVLatviešu↗LTLietuvių↗CSČeština↗RORomână↗HUMagyar↗ELΕλληνικά↗BGБългарски↗HRHrvatski↗SKSlovenčina↗SLSlovenščina↗RUРусский↗UKУкраїнська↗TRTürkçe↗
LET’S WORK ON YOUR WEBSITE
A CLEAR NEXT STEP

Let’s talk
about your site.

A focused SEO audit or a conversation about a specific challenge. Tell me where you are and what you want to change.

Andrii Stanetskyi
Andrii StanetskyiHead of SEO / Technical SEO Lead
[email protected] ↗
How can I help?

Scope and fee agreed before any work begins.