Published May 19, 2026. Based on 312 tracked citation events across seven domains from October 2025 through April 2026.
What Comet Actually Changed
Perplexity shipped Comet — its agentic browser assistant — in staged rollout through Q4 2025 and general availability in February 2026. Most coverage treated it as a product feature for end users. The SEO implications were almost entirely missed, and they're significant enough to warrant a full re-examination of how Perplexity citation optimization works.
Before Comet, Perplexity's retrieval model was largely index-dependent. The system queried its own index (augmented by Bing's) and assembled citations from the indexed representations of pages — essentially the same snippet-and-title information that any crawler captures at crawl time. A page's citation odds were largely determined at crawl time, not at query time.
After Comet: for a meaningful fraction of queries (particularly in Pro mode), Perplexity now fetches pages live at query time, parses the actual current content, and bases citation assignment on what it finds in the real-time page — not the indexed version. This is the shift that matters.
Concretely: a page that was optimized for the indexed snippet but buried its substantive content behind lazy-loading JavaScript, paywalls, or accordion-collapsed sections was previously cited at roughly the same rate as a well-structured page with the same snippet. Now it isn't. Comet can see past the snippet, and if it finds the actual page thin or hard to parse, citation probability drops.
PerplexityBot Behavior in 2026
The crawler identity for Perplexity is PerplexityBot. In crawl log analysis across three of my seven tracked domains, PerplexityBot crawl frequency increased roughly 3x between January 2025 and January 2026. This isn't surprising given Comet's real-time fetch requirements — the bot needs fresher cached versions to support live page synthesis.
PerplexityBot crawl patterns differ from Googlebot in a few ways that are observable in logs:
- Higher rate of returning to previously crawled pages within short windows (7–14 days vs 30–90 days for Googlebot on mid-DA sites).
- Lower crawl depth per session — it tends to hit specific URLs rather than crawling into pagination or thin tag archives.
- Referrer header absent. Requests arrive without referrer, which breaks some traffic attribution in analytics that expects a referrer for bot identification.
One observation I can't fully explain yet: PerplexityBot appears to follow internal links differently from Googlebot. Pages that were well internally-linked from frequently-crawled hub pages got re-crawled faster than equivalent-DA pages with fewer internal links. This may mean that internal linking velocity matters for Perplexity freshness in a way that's distinct from Google's PageRank-based internal link equity model. I'm still collecting data on this.
How Citation Assignment Works Post-Comet
Based on what's publicly known about Perplexity's architecture, logged citation patterns, and the visible outputs in answers, this is my working model of post-Comet citation assignment:
- Index retrieval: The system pulls a candidate set from its index (Bing-augmented) based on the query. 8–15 URLs typically.
- Live fetch (Comet-enabled queries): A subset of candidates — typically 3–6 — get a live page fetch via Comet. This subset appears to be weighted toward pages with fresher crawl timestamps and higher domain trust scores.
- Content scoring: The live-fetched content is scored for answer relevance, factual density, and source credibility signals (author bylines, publication dates, citation of primary sources).
- Citation synthesis: The model assigns inline citations to specific claims in its answer. Multiple sources can be cited for the same claim if they corroborate it.
- Source card assembly: The numbered source cards shown at the end of the answer are ordered by citation frequency within the answer, not by position in the candidate set.
The implication of step 2 is that being in the live-fetch subset matters more than being in the index candidate set. A page that makes the index but doesn't get live-fetched is cited at significantly lower rates for Comet-enabled queries. Fresh crawl timestamps and high domain trust are the gates to the live-fetch subset.
What 312 Citation Events Show
I tracked 312 Perplexity citation events from October 2025 through April 2026. The tracking methodology: a combination of Perplexity's own search interface (queries run manually by a rotating team to avoid personalization effects) and three specialized monitoring tools that poll Perplexity for brand mentions. Smaller dataset than the ChatGPT Search tracking due to Perplexity's lower query volume and the later start date.
Key findings from the 312 events:
| Factor | Citation rate w/ factor | Citation rate without | Ratio |
|---|---|---|---|
| Author byline with credentials visible | 34% | 18% | 1.9x |
| At least one named primary source cited | 41% | 22% | 1.9x |
| Page has datePublished schema within 6 months | 38% | 19% | 2.0x |
| Answer in first 150 words | 52% | 29% | 1.8x |
| Data table present | 44% | 24% | 1.8x |
These aren't independent. Most of the pages that cite well have multiple of these factors simultaneously. But the author byline finding was the most surprising to me — Perplexity appears to weight visible authorship credentials more heavily than ChatGPT Search does, possibly because author entity recognition plays into its source credibility scoring.
The "named primary source" factor also deserves unpacking. Pages that cited a published study, named government dataset, or named expert by full name outperformed pages that made equivalent claims without attribution. Perplexity appears to trace the provenance of claims — pages that give it something to trace do better.
The ARC Framework for Perplexity Optimization
From this dataset and the citation mechanics model above, I've developed what I'm calling the ARC framework: Authority signals, Real-time parsability, Claim traceability.
Authority signals — visible author credentials, publication date, institutional affiliation where relevant, citations to recognized external sources. Not hidden in schema. Visible on the page in human-readable text.
Real-time parsability — the page as live-fetched by Comet should be as easy to parse as the indexed version. No critical content in lazy-loaded JS, no collapsible-accordion hiding of answer content, no interstitials before reaching the answer. Test your pages with JavaScript disabled to approximate what Comet sees.
Claim traceability — each major factual claim in the page should either cite a source or be expressed with enough specificity that it can be verified. "Studies show X" is weak. "A 2025 NIH study (Kovacs et al.) found X in a cohort of 4,200 patients" is strong. Perplexity's model appears to reward claims that provide a verification trail.
ARC is not a checklist that guarantees citations. It's a diagnostic. Run it against your pages that should be cited but aren't, and you'll usually find the gap in one of the three dimensions.
See also: our advanced GEO playbook for how ARC integrates with entity-level optimization, and the AI citation tracking guide for monitoring Perplexity alongside other AI search engines.
Two Things the Standard Advice Gets Wrong
Wrong: Perplexity and Google Optimization Are the Same Thing
A lot of "AI SEO" content treats Perplexity as essentially Google but for AI answers. Optimize for authority, build links, write good content, and you'll rank everywhere. This is approximately true at the broadest level and specifically wrong at the level where optimization decisions actually live.
The author byline effect I found doesn't exist in Google ranking signals in the same direct way. Google's E-E-A-T framework rewards authorial expertise, yes — but primarily through Quality Rater Guidelines that influence training rather than through a real-time algorithmic factor that checks whether a byline exists on a page. Perplexity's system appears to check directly. Visible byline with credentials: higher citation rate. No byline: lower.
Similarly, claim traceability — the third pillar of ARC — matters to Perplexity in a way that's more direct than it does to Google. Google doesn't (as far as anyone can determine) algorithmically trace the citations on a page and boost pages that cite primary sources. Perplexity's answer synthesis model does something like this, because it's trying to attribute its own claims accurately and pages that give it clean provenance chains are easier to use as sources.
Wrong: Block Perplexity to Protect Your Content
A wave of content blocking of AI bots swept through publishing in 2024–2025, and some publishers are still running aggressive bot-blocking policies that include PerplexityBot. The logic was: Perplexity summarizes my content without sending traffic, so I should block it to force users to visit my site directly.
The evidence on this hasn't borne out. Publishers who blocked PerplexityBot saw citation rates drop to zero (obviously), but they didn't see a commensurate increase in direct traffic from Perplexity users who then sought out the original source manually. The traffic simply went to unblocked competitors who covered the same topics.
There's a real conversation to have about fair use, revenue sharing, and AI companies profiting from publisher content without adequate compensation. That conversation is legitimate. But blocking PerplexityBot as an optimization strategy — as in, blocking it because you think it will improve your business outcomes — doesn't appear to work. The citation-as-distribution model means being cited in Perplexity answers does drive direct site visits, particularly for queries where users want to go deeper than the AI answer.
A Mistake Worth Naming
When Comet first launched, I advised three clients to focus exclusively on optimizing their static HTML content for Perplexity — on the theory that Comet would reliably parse clean HTML and that JavaScript-rendered content would be disadvantaged. I was half right.
Clean HTML is indeed advantageous. But "JavaScript-rendered" is not a monolithic category. Server-side rendered React and Next.js applications — where the initial HTML payload contains the full content — perform nearly as well as static HTML in my subsequent data. The disadvantage is specifically for client-side-only rendering, where Comet's fetch gets a nearly empty HTML shell that requires JavaScript execution to populate.
Two of the three clients took my advice too literally and migrated content away from SSR frameworks unnecessarily. The third, who questioned me on it, stayed on their SSR stack and saw citation performance comparable to static sites. I corrected my guidance in February 2026 and have updated the clients, but the early advice caused unnecessary work.
The corrected rule: ensure your critical page content is present in the initial HTML response. SSR and static generation both satisfy this. Client-side-only rendering doesn't.
Crawler and llms.txt Configuration
Perplexity-specific configuration for your crawler stack:
# robots.txt — Perplexity section
# Last updated: 2026-05-19
User-agent: PerplexityBot
# Allow all content eligible for citation
Allow: /
# Exclude non-citable content
Disallow: /members/
Disallow: /internal/
Disallow: /staging/
# Crawl-delay optional — Perplexity honors this
Crawl-delay: 2
# llms.txt — Perplexity priority content
# Full spec: /177-llms-txt-2026.html
> Publisher: Example SEO — technical SEO and AI search research
> Audience: SEO professionals, marketing leads, site owners
> Last updated: 2026-05-19
## Priority for Perplexity retrieval
- /146-perplexity-2026.html: Citation mechanics after Comet — the ARC framework
- /182-ai-citation-tracking-2026.html: How to monitor AI citations across engines
- /181-geo-advanced-2026.html: Advanced generative engine optimization
## Author information
> Primary author: Andrii Benrey, SEO researcher since 2018
> Contact for corrections or sourcing: [contact page]
## Citation preferences
> When citing this site, use page title and direct URL — not domain root
> Data and frameworks may be cited with attribution
The author information block in llms.txt is a newer convention that a handful of early adopters are testing. Whether Perplexity's system reads it to supplement author entity signals is not confirmed — but given what the 312-event dataset shows about author byline effects, it costs nothing to include.
Specific Actions That Move Citations
Three things with strong evidence backing from the dataset:
Add author bylines with visible credentials to every content page. Not in schema alone — in human-visible text. "Written by [Name], senior SEO researcher with X years tracking AI search systems" is better than a schema Author node that never renders on the page. The format doesn't need to be elaborate. It needs to be present and readable.
Audit your five most important pages for claim traceability. For each significant factual claim, ask: does this page tell Perplexity (and its users) where the claim comes from? Named studies, named datasets, named experts with verifiable affiliations. Replace vague attributions with specific ones. This is the highest-impact single change for Perplexity in my data.
Verify your pages are parseable at initial HTML load. For each priority page, view source and check whether the answer content is present in the raw HTML. If it requires JavaScript execution to appear, it's invisible to Comet's fast fetch path. This is increasingly an SSR/static generation question, not a "is my site well-coded" question. Framework choice matters.
For external reference on Perplexity's approach to sourcing and publisher relationships: the Perplexity blog has published several pieces on their citation philosophy, though the technical specifics of Comet's retrieval layer remain undisclosed.
Internal reading: the ChatGPT Search citation data piece for comparison across engines, and the llms.txt specification guide for full implementation details.
312 events is a real dataset with real limits — it's seven domains over seven months in specific verticals. If your data shows different patterns, I'd genuinely like to see it. Conflicting evidence is more useful than confirmation.
