Where We Started: 11,847 Deindexed Pages and a Hard Lesson
In late 2024, I launched what I thought was a clean programmatic build. We had 18,000 token pages live across three domains, solid internal linking, a Cloudflare-hosted edge rendering setup, and data piped in from CoinGecko's API every four hours. By January 2025, roughly 11,847 of those pages had been deindexed during a quiet spam sweep that most people in the crypto SEO space only noticed in their Search Console email on a Tuesday afternoon.
I'm going to walk through exactly what went wrong, what we rebuilt, and how the rebuild eventually produced 4,847 ranked token pages generating a 31.4% year-over-year lift in organic sessions by Q1 2026. But I need to start with the mistake because it's the only honest way to frame everything that follows.
The mistake was structural, not technical. We had optimized our templates for data completeness rather than user utility. Every page had price, market cap, volume, circulating supply, and a three-paragraph boilerplate "what is [Token Name]?" section that was algorithmically varied but substantively identical. Google's classifiers, which have gotten extremely good at distinguishing information density from information variety, flagged the whole cluster. The deindex was not a penalty. It was a quality filter. That distinction matters enormously for how you respond to it.
What made this particularly stinging is that I had read the signals before they became a verdict. In November 2024, crawl coverage dropped across two of the three domains. Not dramatically (maybe 12% fewer pages in the coverage report week over week) but the pattern was there. I attributed it to crawl budget fluctuation rather than a quality signal. That was wrong, and the January sweep confirmed it. If you're seeing coverage decline on a programmatic build, treat it as a five-alarm fire, not seasonal noise. Google does not quietly test deindexing your pages for fun.
The recovery took four months of rebuilding from the data layer up. We didn't start from the template. We started from the question: what does a person searching for a specific token's name actually need that nobody else is giving them? The answer varied by token category, by market cap tier, by whether the token had a functioning product. That's where the architecture that eventually worked was born.
The YMYL Reality Nobody Talks About Honestly
Your Money Your Life classification for cryptocurrency pages is applied inconsistently, and that inconsistency is deliberate. Google's quality guidelines define YMYL as content where poor information could "directly impact the health, financial stability, safety, or welfare of people." Crypto price pages sit in a gray zone that most SEO practitioners either overreact to or completely ignore.
Here is what I've observed across multiple builds in this vertical. Pages that present historical price data, market statistics, and contextual information about a token's utility without offering investment advice or price predictions typically receive standard quality evaluation. Pages that include phrases like "could reach," "analysts predict," or "potential upside" get flagged for heightened E-E-A-T scrutiny immediately. The model appears to be semantic proximity to advisory language, not the category of the page itself.
So the first rule of building crypto programmatic pages at scale: never let your templates drift into forward-looking language. Lock that out at the template level, not the editorial level. The moment you add a "price prediction" section to your token template because a competitor has one and it's getting traffic, you've changed your YMYL exposure for every single page in your cluster simultaneously. That's thousands of pages shifting risk profile in one deploy.
I learned to treat the template as a regulatory document. Every variable, every conditionally rendered block, every fallback string went through a "does this read as advice?" review before merge.
There's a second dimension to YMYL exposure in crypto that almost nobody talks about: the difference between active and abandoned tokens. A page for a token that launched in 2019, hit its ATH in 2021, lost 98% of its value, and hasn't seen a meaningful on-chain transaction in 14 months is a fundamentally different content problem than a page for an active DeFi protocol. Both might technically meet your data thresholds. But the abandoned token page reads, to a classifier and to a human reviewer, like outdated financial information being kept live for SEO purposes. That's exactly the behavior Google's quality guidelines call out. We added a "last meaningful on-chain activity" date check to our publication filter in April 2025 and unpublished 3,100 effectively dead token pages in one sweep. It was the right call.
The Inconsistency Is Also an Opportunity
CoinMarketCap and CoinGecko are treated as authoritative sources despite being largely programmatic at their core. So the bar isn't "is this programmatic?" The question is: "does this programmatic content demonstrate sufficient expertise and serve a real user need?" Build to that standard and the YMYL concern recedes. The pages that survived our 2025 rebuild were the ones that answered a specific, answerable question about a specific token in a way that required knowing something about that token in particular. Not boilerplate. Not generic crypto education with the token name swapped in.
The practical implication: stop trying to compete with CoinGecko on breadth and start competing on depth in categories where CoinGecko is weakest. DeFi protocol metrics. Layer-2 bridge activity. NFT collection token economics. Real-world asset tokenization data. These are areas where programmatic pages built on on-chain data sources can genuinely offer more useful information than the major aggregators, and where YMYL exposure is lower because the content is more analytical than purely price-focused.
On-Chain Data Ingestion That Actually Scales
The data layer is where most teams cut corners and pay for it later. Our current architecture uses a three-source pipeline: CoinGecko's v3 Pro API for market data, a self-hosted Alchemy node for Ethereum and EVM-compatible on-chain reads, and Flipside Crypto's SQL interface for historical analytics. Each source feeds into a Postgres database that our Next.js site reads at build time for static generation and at request time for live price widgets.
# On-chain data ingestion — Python pipeline
# Fetches token metadata + on-chain metrics for page generation
import asyncio
import httpx
import psycopg2
from datetime import datetime, timezone
COINGECKO_BASE = "https://pro-api.coingecko.com/api/v3"
ALCHEMY_RPC = "https://eth-mainnet.g.alchemy.com/v2/{API_KEY}"
async def fetch_token_batch(session: httpx.AsyncClient, token_ids: list[str]) -> dict:
"""
Fetch market data for a batch of tokens from CoinGecko.
Rate limit: 500 calls/min on Pro tier. Batch up to 250 IDs per call.
"""
ids_param = ",".join(token_ids)
resp = await session.get(
f"{COINGECKO_BASE}/coins/markets",
params={
"vs_currency": "usd",
"ids": ids_param,
"order": "market_cap_desc",
"per_page": 250,
"page": 1,
"sparkline": True,
"price_change_percentage": "1h,24h,7d,30d,1y",
"locale": "en",
"precision": 8,
},
headers={"x-cg-pro-api-key": CG_API_KEY},
timeout=30.0,
)
resp.raise_for_status()
return resp.json()
async def fetch_onchain_supply(w3_session: httpx.AsyncClient, contract_address: str) -> dict:
"""
Call ERC-20 totalSupply() and read decimals directly from chain.
Avoids relying on aggregator data for supply figures,
which can lag 24-48h on low-cap tokens.
"""
payload = {
"jsonrpc": "2.0",
"method": "eth_call",
"params": [
{
"to": contract_address,
# totalSupply() selector
"data": "0x18160ddd",
},
"latest",
],
"id": 1,
}
resp = await w3_session.post(ALCHEMY_RPC, json=payload, timeout=10.0)
raw_hex = resp.json().get("result", "0x0")
total_supply_raw = int(raw_hex, 16)
# Fetch decimals
dec_payload = {**payload, "params": [{"to": contract_address, "data": "0x313ce567"}, "latest"]}
dec_resp = await w3_session.post(ALCHEMY_RPC, json=dec_payload, timeout=10.0)
decimals = int(dec_resp.json().get("result", "0x12"), 16)
return {
"total_supply": total_supply_raw / (10 ** decimals),
"decimals": decimals,
"fetched_at": datetime.now(timezone.utc).isoformat(),
}
def upsert_token_record(conn, token_data: dict, onchain: dict) -> None:
"""
Write enriched token record to Postgres.
Uses ON CONFLICT to handle re-runs without duplicates.
"""
with conn.cursor() as cur:
cur.execute(
"""
INSERT INTO token_pages (
coingecko_id, symbol, name, slug,
price_usd, market_cap_usd, volume_24h_usd,
price_change_24h_pct, price_change_7d_pct,
circulating_supply, total_supply_onchain,
ath_usd, atl_usd, ath_date, atl_date,
contract_address, chain,
sparkline_7d, updated_at
) VALUES (
%(id)s, %(symbol)s, %(name)s, %(slug)s,
%(current_price)s, %(market_cap)s, %(total_volume)s,
%(price_change_percentage_24h)s, %(price_change_percentage_7d_in_currency)s,
%(circulating_supply)s, %(total_supply_onchain)s,
%(ath)s, %(atl)s, %(ath_date)s, %(atl_date)s,
%(contract_address)s, %(chain)s,
%(sparkline_in_7d)s::jsonb, NOW()
)
ON CONFLICT (coingecko_id) DO UPDATE SET
price_usd = EXCLUDED.price_usd,
market_cap_usd = EXCLUDED.market_cap_usd,
volume_24h_usd = EXCLUDED.volume_24h_usd,
price_change_24h_pct = EXCLUDED.price_change_24h_pct,
price_change_7d_pct = EXCLUDED.price_change_7d_pct,
circulating_supply = EXCLUDED.circulating_supply,
total_supply_onchain = EXCLUDED.total_supply_onchain,
sparkline_7d = EXCLUDED.sparkline_7d,
updated_at = NOW()
""",
{**token_data, **onchain},
)
conn.commit()
A few things worth highlighting in that pattern. We read total supply directly from the contract rather than relying on CoinGecko's reported figure, because for tokens below the top-500 by market cap, the aggregator data on circulating vs. total supply lags considerably. That lag produces factual errors on your pages, which is the worst possible outcome in YMYL territory. One wrong number undermines every other accurate number on the page.
Data Freshness as a Ranking Signal
Google's crawlers have gotten good at detecting stale financial data through comparison signals. If your page shows a price that differs substantially from what Googlebot sees in structured data vs. rendered HTML vs. what competitor pages show for the same asset, that's a freshness trust signal problem. We run full rebuilds every six hours for the top 5,000 tokens by market cap. Long-tail tokens below rank 5,000 get rebuilt daily, and tokens below rank 20,000 get rebuilt on content-change triggers only.
The build system for this is worth describing briefly. We use a priority queue backed by Redis. When market data ingestion runs, it calculates a staleness score for each token based on how much prices have moved since the last build and where that token sits in our traffic-weighted priority ranking. High-traffic tokens that have moved more than 3% since last build get bumped to the front of the rebuild queue. This means that in volatile market conditions, the pages people are actually searching for get rebuilt first, not the pages that happen to be alphabetically early in our token list. That detail alone probably accounted for several percentage points of the crawl coverage improvement we saw in Q3 2025.
One more thing about the data architecture: error states need to be explicit, not silent. When the CoinGecko API returns a null for a field we expect to be populated, we don't render a zero or an empty cell. We render a "data unavailable" state with a timestamp of the last successful read. This is not just good UX. It's a signal to Googlebot that you're aware of your data quality and managing it actively, rather than serving garbage data that happens to fill a template slot.
The Token Page Template That Ranked 4,847 Pages
The template architecture that survived Google's scrutiny is not what most programmatic SEO guides recommend. The conventional wisdom says: pick a template, apply it uniformly, differentiate with data. That works for most verticals. For crypto under YMYL scrutiny, uniform templates are a liability because they make the programmatic nature of your content immediately legible to classifiers.
Our current template uses five content blocks, three of which are conditionally rendered based on data availability thresholds.
{/* Token page template — Next.js / TypeScript (simplified) */}
{/* File: app/tokens/[slug]/page.tsx */}
import { getTokenData } from "@/lib/db/tokens";
import { PriceHero } from "@/components/token/PriceHero";
import { MarketMetrics } from "@/components/token/MarketMetrics";
import { OnChainActivity } from "@/components/token/OnChainActivity";
import { ExchangeListings } from "@/components/token/ExchangeListings";
import { TokenContext } from "@/components/token/TokenContext";
import { generateTokenJsonLd } from "@/lib/schema/financialProduct";
export async function generateMetadata({ params }) {
const token = await getTokenData(params.slug);
if (!token) return { title: "Token Not Found" };
// Price-in-title only for top 1000 by market cap
// Avoids stale prices in SERP snippets for long-tail tokens
const priceClause =
token.market_cap_rank <= 1000
? — $${formatPrice(token.price_usd)} Today
: "";
return {
title: ${token.name} (${token.symbol.toUpperCase()}) Price, Market Cap & On-Chain Data${priceClause},
description: ${token.name} (${token.symbol}) live price is $${formatPrice(token.price_usd)} with a market cap of $${formatMarketCap(token.market_cap_usd)}. View on-chain supply, exchange listings, and historical data.,
alternates: {
canonical: https://example.com/tokens/${token.slug},
},
};
}
export default async function TokenPage({ params }) {
const token = await getTokenData(params.slug);
if (!token) notFound();
// Conditional rendering thresholds
const hasRichOnChainData = token.total_supply_onchain !== null && token.decimals !== null;
const hasExchangeListings = token.exchange_count >= 2;
const hasHistoricalContext = token.ath_usd !== null && token.atl_usd !== null;
const jsonLd = generateTokenJsonLd(token);
return (
<>
{/* Always rendered */}
{/* Conditional: only when on-chain data is fresh (<6h) */}
{hasRichOnChainData && (
)}
{/* Conditional: only when listed on 2+ exchanges */}
{hasExchangeListings && (
)}
{/* Always rendered — editorial context block, NOT boilerplate */}
);
}
The TokenContext component is the critical differentiator. It doesn't pull boilerplate. It renders one of eight distinct editorial frames based on token attributes: layer-1 infrastructure, DeFi protocol, stablecoin, wrapped asset, meme token with documented community origin, gaming/metaverse utility token, oracle network, or cross-chain bridge. Each frame asks different questions, surfaces different data points, and uses different language. A stablecoin page talks about peg mechanisms and reserve audits. A layer-1 page discusses consensus mechanism and transaction throughput. An oracle network page references data provider count and heartbeat intervals.
This is not AI content generation. These are hand-edited editorial templates per category. The data injection is programmatic. The editorial logic is human-authored and reviewed quarterly.
The metadata strategy also deserves its own explanation because it took three iterations to get right. Our first approach put the live price directly in the title tag for every token, regardless of market cap rank. This worked fine for the top 200 tokens where we rebuilt frequently enough to keep prices reasonably current. For tokens ranked 2,000-10,000, Googlebot would crawl the page, cache the title with a price from that crawl visit, and then not recrawl for days or weeks. Users searching and seeing a stale price in the SERP snippet then clicking through to find a different current price: that's a trust gap. The fix was simple once we understood the problem: we only include live price in the title tag for tokens where we have enough traffic signal to know Google is recrawling frequently. For everything else, the title focuses on the token's utility and category, not its price.
The Category Hub Architecture
Individual token pages don't live in isolation. They're organized into category hub pages (one page per token category) that serve as the primary internal navigation layer and as topical authority anchors for each cluster. The DeFi hub page links to every DeFi protocol token in our index, ordered by market cap, with a summary card for each. The hub page itself ranks for category-level queries like "defi tokens by market cap 2026" and passes relevance signals down to every token in the cluster.
This hub-and-spoke structure is not novel in programmatic SEO. What's different in our implementation is that hub pages are genuinely editorial, not just paginated lists. Each hub page has an opening section written by someone who actively uses tokens in that category, explaining what distinguishes this category from adjacent ones, what metrics matter most for evaluating tokens in it, and what the current landscape looks like as of our last editorial review. That makes the hub page a ranking asset in its own right, not just a navigation wrapper. Our full programmatic SEO strategy guide covers the hub architecture in more detail.
The PACE Framework: My Personal Architecture for Crypto pSEO
After running three separate programmatic crypto builds through 2024 and 2025, I developed a decision framework I now call PACE. It stands for Provenance, Accuracy threshold, Context category, and Evergreen anchor. I apply it before any page type goes to production.
The reason I formalized this into a named framework rather than leaving it as a mental checklist is that programmatic builds involve so many moving parts (data sources, rendering systems, deployment pipelines, editorial templates, sitemap management) that without a common vocabulary, quality decisions don't travel well through a team. When I say "this page fails the Provenance check," everyone working on the build understands exactly what that means and what needs to happen to fix it. Named frameworks are not for conferences. They're for keeping a team aligned when you're maintaining 5,000 live pages across multiple data sources and something breaks at 11pm.
Provenance means every data point on a page has a traceable source. Not "CoinGecko" as a vague attribution, but a specific API endpoint, a timestamp, and a fallback behavior when the source is unavailable. If a data source goes down, pages should display cached data with a visible "last updated" timestamp rather than silently displaying stale numbers. Google's crawlers, and your users, both notice when financial figures are obviously wrong.
Accuracy threshold governs which pages get published at all. We don't publish a token page if critical fields fall below completeness thresholds: price (required), market cap (required), at least one exchange listing (required), contract address for EVM tokens (required for on-chain block to render), and a non-null description from CoinGecko's coin detail endpoint (required for the context block). Tokens that don't meet threshold get a placeholder record in the database but no live page. This kept approximately 29,000 low-quality token records from ever reaching production.
Context category determines which editorial frame renders, as described above. The category is assigned during data ingestion using CoinGecko's category tags plus a manual override table we maintain. The override table currently has 847 entries correcting misclassifications.
Evergreen anchor is the section of every token page that doesn't change with price data. It describes what the protocol does, when it launched, who created it (with sourced attribution), and what distinguishes it from similar tokens. This section is written once per token for significant tokens. For long-tail tokens, it renders from the category editorial template. But the existence of an evergreen section means the page has substantive value even when it's crawled between price updates.
Running PACE as an actual checklist before any page type ships has caught problems I would otherwise have discovered only after a deindex. The Provenance check caught a situation where we were pulling liquidity pool TVL from a third-party aggregator that had a known 48-hour data lag (we didn't know until we ran the trace). The Accuracy threshold check caught a batch of 340 tokens where contract addresses had been submitted incorrectly in the CoinGecko community edit system, meaning our on-chain supply reads were returning data for completely different contracts. The Context category check identified a cluster of 89 tokens tagged as "stablecoins" in CoinGecko's API that were actually algorithmic stablecoins that had already de-pegged; they needed a different editorial frame entirely, one that accurately represented their status rather than implying they were functioning pegged assets.
Frameworks only matter if you actually run them. We gate deploys behind a PACE check that runs as part of our CI pipeline. If any token in a batch fails the Accuracy threshold, the batch doesn't deploy. This has delayed launches twice. Both times it was the right call.
JSON-LD for FinancialProduct: Getting It Right
The schema.org FinancialProduct type is technically applicable to cryptocurrency tokens but is implemented correctly on fewer than 3% of crypto pages in the wild, based on my manual audits of roughly 200 crypto sites over the past year. Most implementations use a generic WebPage or Article type, which leaves structured data opportunity on the table.
// JSON-LD generator for token pages
// File: lib/schema/financialProduct.ts
export function generateTokenJsonLd(token: TokenRecord) {
const baseUrl = "https://example.com";
return {
"@context": "https://schema.org",
"@type": "FinancialProduct",
"name": ${token.name} (${token.symbol.toUpperCase()}),
"description": token.description_short || ${token.name} is a cryptocurrency token traded on decentralized and centralized exchanges.,
"url": ${baseUrl}/tokens/${token.slug},
"image": token.image_url || ${baseUrl}/images/tokens/${token.slug}.png,
"identifier": {
"@type": "PropertyValue",
"propertyID": "CoinGeckoID",
"value": token.coingecko_id
},
"additionalProperty": [
{
"@type": "PropertyValue",
"name": "Contract Address",
"value": token.contract_address
},
{
"@type": "PropertyValue",
"name": "Blockchain",
"value": token.chain
},
{
"@type": "PropertyValue",
"name": "Market Cap Rank",
"value": token.market_cap_rank?.toString()
}
],
"feesAndCommissionsSpecification": "Trading fees vary by exchange. See individual exchange listings.",
"provider": {
"@type": "Organization",
"name": "Example Crypto Data",
"url": baseUrl,
"sameAs": [
"https://twitter.com/example",
"https://linkedin.com/company/example"
]
},
"dateModified": token.updated_at,
"mainEntityOfPage": {
"@type": "WebPage",
"@id": ${baseUrl}/tokens/${token.slug}
},
// Breadcrumb helps with SERP presentation for long-tail tokens
"breadcrumb": {
"@type": "BreadcrumbList",
"itemListElement": [
{
"@type": "ListItem",
"position": 1,
"name": "Tokens",
"item": ${baseUrl}/tokens
},
{
"@type": "ListItem",
"position": 2,
"name": token.category_label,
"item": ${baseUrl}/tokens/category/${token.category_slug}
},
{
"@type": "ListItem",
"position": 3,
"name": token.name,
"item": ${baseUrl}/tokens/${token.slug}
}
]
}
};
}
Three things that are non-negotiable in this implementation. The dateModified field must reflect actual content changes, not just price refreshes. We only update it when the evergreen anchor or editorial context changes. Price data is dynamic content, not modified content in the schema sense. Second, the feesAndCommissionsSpecification field must be present for FinancialProduct; omitting it triggers validation warnings in Rich Results Test. Third, the breadcrumb embedded here is redundant with your page-level BreadcrumbList. You can remove it if you're already implementing breadcrumb schema separately, but if you're not, it's doing real work for your SERP presentation on tokens that don't have strong brand recognition.
One issue I ran into during validation: Google's Rich Results Test flags FinancialProduct pages for missing offers property in certain evaluation contexts, even though offers is optional in the schema.org spec for this type. The workaround is to include a minimal offers node that describes trading availability rather than a specific price offer. We use a generic offers block pointing to a general "available via cryptocurrency exchanges" description. It clears the validation warning without implying we're offering to sell the token ourselves, which would be both legally problematic and factually wrong.
The schema implementation I care most about in this vertical is actually not on token pages at all. It's the Organization schema on the site root and the WebSite schema that enables the sitelinks search box for branded queries. When someone searches your site's brand name and a sitelinks box appears offering to search within your token database, that's a conversion funnel improvement that compounds as your brand recognition grows. Too many crypto programmatic builds implement schema only on content pages and miss the structural schema opportunities entirely. See the Google structured data documentation for the exact implementation pattern.
CoinGecko and CoinMarketCap API Patterns at Scale
Running a 47,000-token page database against third-party APIs requires careful rate management. Here is the actual pattern we use, including the exponential backoff logic that saved us from a 48-hour data outage in March 2025 when CoinGecko's Pro API had an incident.
# CoinGecko + CoinMarketCap API orchestration
# File: pipeline/api_orchestrator.py
import asyncio
import logging
import random
from datetime import datetime, timedelta
from typing import AsyncGenerator
import httpx
from tenacity import (
retry,
stop_after_attempt,
wait_exponential,
retry_if_exception_type,
before_sleep_log,
)
logger = logging.getLogger(__name__)
class RateLimitedClient:
"""
Wraps httpx.AsyncClient with sliding-window rate limiting.
CoinGecko Pro: 500 req/min. CMC Professional: 333 req/min.
Uses 85% of limit to leave headroom for parallel workers.
"""
def __init__(self, calls_per_minute: int, base_url: str, api_key: str, api_key_header: str):
self.window_calls: list[datetime] = []
self.limit = int(calls_per_minute * 0.85)
self.base_url = base_url
self.client = httpx.AsyncClient(
base_url=base_url,
headers={api_key_header: api_key, "Accept-Encoding": "gzip"},
http2=True,
timeout=httpx.Timeout(connect=5.0, read=30.0, write=5.0, pool=2.0),
)
async def _wait_for_slot(self):
now = datetime.utcnow()
cutoff = now - timedelta(minutes=1)
self.window_calls = [t for t in self.window_calls if t > cutoff]
if len(self.window_calls) >= self.limit:
oldest = self.window_calls[0]
wait_seconds = (oldest + timedelta(minutes=1) - now).total_seconds()
wait_seconds += random.uniform(0.1, 0.5) # jitter
logger.debug(f"Rate limit reached, waiting {wait_seconds:.2f}s")
await asyncio.sleep(max(0, wait_seconds))
self.window_calls.append(datetime.utcnow())
@retry(
retry=retry_if_exception_type((httpx.HTTPStatusError, httpx.TimeoutException)),
wait=wait_exponential(multiplier=1, min=2, max=60),
stop=stop_after_attempt(5),
before_sleep=before_sleep_log(logger, logging.WARNING),
reraise=True,
)
async def get(self, path: str, **kwargs) -> dict:
await self._wait_for_slot()
resp = await self.client.get(path, **kwargs)
if resp.status_code == 429:
retry_after = int(resp.headers.get("Retry-After", 60))
logger.warning(f"429 received, sleeping {retry_after}s")
await asyncio.sleep(retry_after + random.uniform(1, 5))
resp.raise_for_status()
resp.raise_for_status()
return resp.json()
async def paginate_cmc_listings(client: RateLimitedClient) -> AsyncGenerator[list[dict], None]:
"""
CoinMarketCap /v1/cryptocurrency/listings/latest paginates differently than CoinGecko.
CMC uses start + limit offsets. Max 5000 per call on Professional plan.
We use it as a secondary source for tokens not in CoinGecko's database.
"""
start = 1
limit = 5000
while True:
data = await client.get(
"/v1/cryptocurrency/listings/latest",
params={
"start": start,
"limit": limit,
"sort": "market_cap",
"cryptocurrency_type": "all",
"aux": "circulating_supply,total_supply,max_supply,platform,tags",
},
)
batch = data.get("data", [])
if not batch:
break
yield batch
if len(batch) < limit:
break
start += limit
await asyncio.sleep(0.5)
The dual-source setup matters more than it might seem. CoinGecko covers approximately 14,200 coins as of May 2026. CoinMarketCap covers closer to 22,800. The delta is where your long-tail token strategy lives. Most of the tokens in the 14,000-22,000 range are genuinely low-quality projects with thin on-chain activity, but a meaningful subset are legitimate projects that CoinGecko simply hasn't onboarded yet. We pull from CMC for those tokens and cross-reference against our quality thresholds. About 2,100 of our ranked pages use CMC as the primary data source.
The specific failure mode to watch for with CMC as a primary source: CMC's category taxonomy is coarser than CoinGecko's and more frequently wrong on edge cases. We've seen wrapped tokens misclassified as native assets, cross-chain bridge tokens tagged as generic DeFi, and several real-world asset tokens listed under "stablecoins" because they target price stability rather than because they use a traditional peg mechanism. Running CMC-sourced tokens through our manual override table before they hit the context category selection is non-negotiable. The automated category assignment from CMC tags without human review produces demonstrably wrong editorial frames on roughly 8% of tokens we've checked.
Rate limiting behavior also differs between the two APIs in ways that affect your pipeline design. CoinGecko returns a Retry-After header on 429 responses with a specific cooldown period; respect it exactly. CoinMarketCap's 429 responses sometimes include a Retry-After and sometimes don't, in which case you need to back off with exponential delay starting from 60 seconds. Both APIs will silently return stale cached data rather than erroring when they're under load, which is harder to detect. We log a checksum of the response body for each batch request. If the same checksum appears across three consecutive fetches spanning more than two hours, we flag it as a potential silent cache return and trigger a manual review of those token records before they're published.
Two Things the SEO Industry Gets Wrong About Crypto Pages
Contrarian Take One: More Tokens Is Not a Better Strategy
The prevailing advice in programmatic SEO circles is to go wide. Cover every asset, every token, every obscure altcoin. The volume argument is intuitively appealing: if you have 50,000 token pages and even 5% rank for something, that's 2,500 ranking pages. But in crypto, the quality floor matters in a way that doesn't apply in, say, real estate programmatic builds. Financial data about non-existent or abandoned tokens is actively harmful to users. A page about a token that was rugged in 2022 and still shows a "current price" is misinformation. Google's quality classifiers surface this more aggressively in YMYL categories.
Our build deliberately excludes any token with fewer than 30 days of trading history, fewer than three exchange listings, or no verifiable contract address. That filter removes approximately 18,000 tokens from consideration. Our ranked page count is lower than it could be if we went wide, but our deindex rate has been zero since the February 2025 rebuild. Zero. That's the trade I'd make again.
The practical question this raises is: who do you think is served by a page for an abandoned token? Not the person who holds it, who probably knows it's abandoned. Not the person considering buying it, who will be misled by stale data. Not the search engine, which penalizes the domain for the quality failure. The only stakeholder served is a traffic spreadsheet. Programmatic SEO has a habit of optimizing for traffic spreadsheets instead of users, and in YMYL verticals, that habit gets you deindexed. The crypto version of programmatic restraint is understanding that a database of 47,000 tokens and a live site of 5,400 token pages are not contradictions. They are a quality filter in action.
Contrarian Take Two: Schema Markup Is Overweighted as a Differentiator Here
I know this is uncomfortable given that I've spent 400 words on JSON-LD implementation. But the honest truth is that structured data for crypto token pages provides minimal ranking lift directly. What it does provide is better click-through rates on tokens where Google elects to show rich snippets (which is less than 8% of the token pages in our index), and it provides a quality signal that the publisher understands structured data conventions. That second benefit is meaningful in aggregate, but it's not why your token pages rank or don't rank. They rank because they have fresh, accurate, useful data and because other authoritative pages link to them. Schema is table stakes, not a differentiator.
I've seen builds obsess over FinancialProduct schema implementation while running stale price data and thin editorial context. Get the data layer right first. Schema second.
The more useful differentiator for crypto programmatic pages in 2026 is topical authority at the category level, not technical schema implementation at the page level. If your site has 600 DeFi protocol token pages, a DeFi glossary, three well-linked DeFi protocol comparison articles, and incoming links from DeFi-focused publications, those 600 pages rank better than they would if you simply perfected the schema on each one. Authority aggregates at the topic cluster level. Schema helps Google interpret pages correctly. But it doesn't manufacture authority that isn't there.
A word on the validator tools that get used obsessively in this space. Google's Rich Results Test is useful for catching structural errors in your JSON-LD, but it's not a proxy for ranking quality. Schema.org's validator is good for ensuring you're not violating the type definition. Neither tells you whether your schema is helping you rank. The only measure of that is index coverage and click-through data in Search Console, correlated against your schema deploy dates. We saw a 6.2% improvement in rich snippet eligibility after fixing our FinancialProduct implementations, but our ranked page count didn't shift meaningfully until we addressed the data freshness issues. Schema validation is a hygiene task. Treat it as such.
Quality Signals That Google Actually Rewards in This Vertical
The honest version of this section requires acknowledging that I don't have a controlled experiment. I have a before-and-after comparison between two very different builds on similar domains, with a rebuild period in between. Correlation is not causation and all of that. But when you rebuild a site from scratch and can compare the pages that ranked against the pages that didn't, across thousands of data points, patterns emerge that are hard to dismiss as coincidence. These five factors are the ones that appear consistently enough that I'm now building toward them deliberately rather than discovering them after the fact.
Based on correlating our 4,847 ranked pages against the 11,847 deindexed ones, I can identify five factors that distinguish ranked pages consistently.
Exchange listing count. Pages for tokens listed on five or more exchanges rank at dramatically higher rates than tokens listed on one or two DEXs. This isn't just a proxy for token legitimacy, though it is that too. It's because exchange listing pages link back to your token pages when you get included in their ecosystems. Those links are high-quality topical references. Tokens in our top-100 by page-level organic traffic average 11.3 exchange listings each. The exchange listing signal is one of the few factors in our dataset that correlates strongly with ranking even when controlling for market cap, meaning mid-cap tokens with broad exchange presence outrank high-cap tokens with concentrated trading on a single venue. Building your exchange listing data into the page in a structured, machine-readable way (not just text) accelerates how quickly Google can verify the exchange relationship and treat it as a quality indicator.
Data source transparency. Visibly attributing your price data to a specific source, with a timestamp, reduces the trust gap that vague financial data creates. We added a "Data source: CoinGecko Pro API, updated [timestamp]" line to every page in February 2025. There's no way to isolate its contribution, but it coincided with a 14% improvement in crawl coverage for previously under-crawled long-tail tokens in our index. The transparency also creates a natural anchor for the data attribution link. Linking your "updated [timestamp]" text to a static page explaining your data methodology gives Google a way to understand your sourcing practices at the site level, not just at the page level. That site-level trust signal compounds across the index.
Internal linking depth. Token pages that are three or more clicks from the homepage in our crawl graph rank at lower rates than those that are two clicks or fewer. We solved this with category hub pages. Every token belongs to a category. The category page links directly to all tokens in that category. The category page is linked from the homepage navigation. So every token is two clicks from home: home > category > token. Browse our full token database to see how this manifests structurally.
Historical context richness. Pages that include ATH and ATL data with dates, not just current price, rank measurably better for navigational queries like "[token name] all-time high." This is low-hanging fruit. If you're pulling this data from CoinGecko (it's in the markets endpoint), surface it prominently.
Related token linking. Each token page links to five related tokens based on category and market cap tier. This sounds like a standard "related content" pattern but the implementation matters. We link to tokens in the same category that are adjacent in market cap rank, so a token ranked #847 links to tokens ranked #820-870 in the same category. This creates coherent topical neighborhoods in the link graph rather than random internal links. See our internal linking methodology for more on this approach.
What doesn't work as a quality signal, in my experience: keyword density in the evergreen anchor text, synonym variation in programmatic descriptions, and structured data completeness scores. I've seen consultants advise adding more synonyms for the token name, populating more optional schema fields, and running the page copy through readability scoring tools. None of those correlated with ranking improvement in our dataset. The only signals that tracked were data freshness, source transparency, exchange listing count, depth of historical data, and internal link graph cohesion. Build toward those and let the rest follow.
A Note on Content Velocity
We don't add new token pages at a constant rate. We batch publish based on market conditions. When a token surges into the top 200 by market cap, we accelerate its page to production within hours. When a new chain launches and onboards 50 new tokens in a week, we batch-publish the whole cohort together. This mimics organic publishing patterns and creates natural crawl demand signals. Publishing 3,000 pages in a single sitemap ping looks automated. Publishing 30-80 pages in waves driven by market events looks like responsive editorial work, because it is.
The sitemap management piece is worth a full article on its own, but the essential point is this: segment your sitemap by token category, not alphabetically or by market cap. Category-segmented sitemaps let Google understand your content taxonomy from the first crawl. They also make it trivially easy to deprioritize or exclude entire categories if you need to manage crawl budget; you can set lastmod on a category sitemap to signal nothing has changed, without touching individual token sitemaps. We maintain eleven category sitemaps plus a priority sitemap for the top-500 tokens that gets a lastmod update on every rebuild cycle. That priority sitemap gets crawled most aggressively and keeps our top-revenue token pages current in the index.
What This Looks Like in Six Months
The build that's currently live will look different by November 2026 in a few specific ways I'm already planning for.
AI Overviews are claiming featured position for generic token queries ("what is ethereum") at a rate that's making the top-of-funnel traffic for major tokens essentially irretrievable via traditional blue-link results. We're pivoting our template architecture for top-500 tokens toward comparison and contextual queries where AI Overviews are less dominant. The CTR data on AI Overview displacement in financial verticals is stark and the trend is not reversing. What matters for our template adaptation is understanding where AI Overviews don't appear: queries that require real-time data, queries with navigational intent, and queries with enough specificity that Google can't synthesize a reliable answer from training data alone. "[token] vs [token] 2026" comparison queries, "[token] contract address on [chain]" address lookup queries, and "[token] trading volume last 30 days" data queries all fall into the AI Overview gap and represent defensible blue-link real estate for well-structured programmatic pages.
On-chain data is getting richer and more accessible. Dune Analytics now exposes an API that makes protocol-specific metrics queryable without writing custom SQL for each chain. We're integrating Dune as a fourth data source for DeFi protocol tokens specifically, which will unlock TVL, unique active wallets, and transaction count as page-level data points. Those are metrics that CoinGecko doesn't surface and that meaningfully differentiate our DeFi token pages from commodity price-data pages.
The other major shift I'm planning around is Google's continued reduction of blue-link real estate for crypto queries in favor of conversational AI Mode responses. Our traffic data from Q1 2026 shows that informational queries ("what is [token] used for," "how does [protocol] work") have declined more sharply than navigational queries in our organic sessions. Navigational queries ("[token] price," "[token] contract address," "[token] all-time high") have held relatively stable. That distribution is shifting our content strategy away from informational evergreen anchor text and toward richer data presentation for navigational intent. The user who searches "[token] price" and lands on our page is increasingly the core user we're building for, not the user who searches "is [token] a good project." The latter question is being answered by AI Mode. The former still needs a page with up-to-date data.
The 31.4% lift we saw year-over-year is going to be harder to sustain. The easy gains in this space came from fixing the quality problems that got 11,847 pages deindexed. The next phase of growth requires original data, stronger topical authority, and probably some editorial investment in the highest-traffic token categories that goes beyond what any template system can produce. I'm already budgeting for three token category deep-dives per quarter, written by people who actually use the protocols in question. E-E-A-T in crypto isn't just about credentials. It's about demonstrable use of the technology you're writing about.
The hardest part of programmatic SEO for crypto isn't the engineering. The pipeline, the templates, the schema: all of that is solvable with enough time and the right architecture. The hard part is maintaining editorial judgment at scale. Deciding which tokens deserve human attention, which categories need original research, and when a template is doing real work versus hiding a thin-content problem behind good data formatting. Those decisions don't automate. They require someone who understands both search and the crypto market well enough to know the difference between a page that serves a user and a page that merely exists.
The builds that are thriving right now are the ones that treated the 2025 enforcement wave not as a setback but as a forcing function. It forced specificity. It forced honesty about which token pages were genuinely useful versus which ones were just incrementing a page count. If you're building in this space right now and you haven't gone through that kind of forced specificity yet, do it voluntarily before Google does it to you. Audit your lowest-traffic token pages ruthlessly. Unpublish the ones that don't meet your own quality bar. Strengthen the ones that do. The index ratio improvement you'll see from that exercise will do more for your domain's crawl health than any technical optimization you could run in the same time.
We're 4,847 ranked pages in. The next milestone is sustaining that while the search landscape continues to shift under every site building at this scale. That's not a solved problem. But it's the right problem to be working on.
One thing I want to be direct about before signing off: programmatic SEO for crypto is harder in 2026 than it was in 2022, and harder in 2026 than it will be in 2028. The current period is the most demanding because Google's quality enforcement has tightened substantially, AI Overviews are claiming top-of-funnel queries, and the technical complexity of on-chain data ingestion has increased as the number of relevant chains has expanded from a handful to dozens. Anyone telling you this is easy is either working in a different vertical or hasn't done it at real scale. The 11,847 pages I had deindexed were not built carelessly. They were built with the prevailing best practices of 2024. Those best practices were not enough.
The reward for getting it right is also real. Crypto users search with high frequency and high intent. A token in the top 1,000 by market cap generates thousands of monthly searches across price queries, comparison queries, "how to buy" queries, and technical information queries. Pages that rank for even a fraction of that search demand generate meaningful, recurring organic traffic from an audience that is actively engaged with the asset. That's a valuable SEO asset worth building carefully, worth maintaining rigorously, and worth protecting fiercely from the quality failures that get clusters deindexed.
For questions about implementation specifics or if you're running a similar build and want to compare notes, reach out directly. I respond to every message from practitioners who are doing actual work in this space.
Frequently Asked Questions
- Does programmatic SEO for crypto trigger YMYL penalties?
- Programmatic crypto pages are evaluated under YMYL standards, but this doesn't mean automatic penalties. The key distinction is between presenting factual financial data with clear sourcing and timestamps versus offering investment advice or price predictions. Pages that stay in the former category and demonstrate E-E-A-T through data accuracy, source attribution, and editorial context typically receive standard quality evaluation. The risk arises when templates drift into advisory language or when data quality is poor enough to mislead users.
- What is the right schema type for cryptocurrency token pages?
- The most technically accurate schema.org type for a cryptocurrency token page is FinancialProduct. It supports relevant properties including
feesAndCommissionsSpecification(required for validation),identifier,provider, andadditionalPropertyfor chain-specific metadata like contract address and blockchain. Pair it with BreadcrumbList and, for editorial content sections, Article schema. Avoid using WebPage as the primary type. - How do you prevent stale price data from appearing in Google's index?
- Three mechanisms work together. Implement ISR with a revalidation window matching your update frequency. Add a visible "last updated" timestamp to every page in rendered HTML. Exclude live price from your meta description for tokens below the top 1,000 by market cap, since meta descriptions are cached by Google and a stale price in the snippet is a trust signal problem.
- Is CoinGecko or CoinMarketCap better as a primary data source?
- For tokens in the top 15,000 by market cap, CoinGecko Pro is preferable: better API design, more reliable uptime historically, and richer metadata including category tags and developer data. For tokens below that threshold, CoinMarketCap covers a larger universe. Run both in parallel and reconcile discrepancies, which are common for newer tokens on supply figures.
- How many token pages is too many for a single domain?
- The constraint isn't page count but the ratio of indexed to published pages. If that ratio drops below 60%, you have a quality or crawl budget problem that adding more pages will worsen. Our current build maintains an 89% index rate across approximately 5,400 live token pages.
