Published 19 May 2026. Writing from the practitioner side — I run SEO for a dozen sites ranging from a mid-size media brand to two SaaS companies. I've been watching llms.txt since the week it dropped.
Where We Actually Are, Two and a Half Years In
The llms.txt specification was released in September 2023 by Jeremy Howard and the Answer.AI team. That means we're roughly 2.5 years into what was, at the time, a modest proposal: a plain-text file, sitting at the root of your domain, that told large language models what your site was, what pages mattered, and what they could or couldn't use. Structured like a simplified markdown document. Parseable by anything. Intentionally minimal.
I remember reading the original proposal on a Tuesday and thinking it was elegant but probably doomed. My exact prediction, which I wrote in a Slack channel I'll never recover: "Robots.txt took a decade to get real compliance. This will be dead in eighteen months." I was wrong. I'll get to that.
Right now, as of mid-May 2026, there are somewhere between 180,000 and 220,000 public domains serving an llms.txt file — depending on which crawl dataset you pull from. That's a real number. Not transformative, but real. For context, it's roughly the same penetration that XML sitemaps had around 2006, two years after Google started supporting them. The analogy isn't perfect, but the growth curve shape is similar: slow, then a burst, then plateau, then another burst when something external forced adoption.
The external forcing event here was the wave of publisher licensing deals in 2024 and early 2025. When the New York Times settled its litigation with OpenAI (details of that settlement remain partially under NDA, but the broad strokes leaked — a structured data access agreement plus a revenue share on derivative training), other publishers started asking: what exactly are we giving away, and is there a machine-readable way to say "here but not there"? llms.txt became part of that answer, at least symbolically.
The Spec Drift Problem Nobody Talks About
Here's what keeps me up at night about llms.txt in 2026: the spec is no longer one thing.
The original specification was clean. A single llms.txt file at the root. An optional llms-full.txt for expanded context. Markdown headers identifying the site. Bullet links to key pages. A minimal set of directives about allowed usage. That was it.
What we have now is a mess of informal extensions that emerged organically as publishers tried to express more complex licensing positions. Some implementations I've audited in the last four months:
- A major news organization using a custom
X-LLM-Licensefield they invented internally, with no documentation - Three SaaS companies I know of that have added
training: prohibitedas a key-value field, a directive the spec doesn't define - One e-commerce site that serves different llms.txt content based on the User-Agent header of the requester — so GPTBot sees one thing, ClaudeBot sees another
- Publishers who've added
licensing_contactfields pointing to API endpoints for automated licensing negotiation
None of this is in the spec. None of it is formally wrong, either, because the spec deliberately left room for extension. But the result is that "llms.txt compliant" has become a near-meaningless phrase. Compliant with which version? Whose extensions? Which parser?
This is spec drift, and it's a genuine problem — not because it breaks anything immediately, but because it makes the file harder to trust as a signaling mechanism. When I read a client's llms.txt and see fields I don't recognize, I can't tell if they're intentional, cargo-culted from someone else's implementation, or just wrong.
The Full llms.txt Specification (Current)
Below is the current baseline spec as I understand it, synthesized from the original Answer.AI documentation, the community wiki that's emerged around it, and my own reading of how major crawlers seem to interpret the file. I've formatted this as a working template.
# llms.txt — Full Specification Template (May 2026 baseline)
# Based on original Answer.AI spec + community extensions
# Place at: https://yourdomain.com/llms.txt
# Optional extended version at: https://yourdomain.com/llms-full.txt
# ── SECTION 1: SITE IDENTITY ─────────────────────────────────────
# H1 = site name (required)
# Blockquote = short site description (recommended)
# Plain paragraph = extended description (optional)
# Example:
# # Your Site Name
# > One-sentence description of what this site covers.
#
# Longer description of the site, its purpose, audience, and
# the nature of content it publishes. Two to four sentences.
# ── SECTION 2: KEY PAGES ─────────────────────────────────────────
# H2 headers group related links
# Bullet list items are the pages LLMs should prioritize
## Docs
# - [Page Title](https://yourdomain.com/path/): Brief description of page
## Blog
# - [Article Title](https://yourdomain.com/blog/slug/): Brief description
## API Reference
# - [Endpoint Docs](https://yourdomain.com/api/): Brief description
# ── SECTION 3: OPTIONAL / BLOCKED CONTENT ────────────────────────
# H2 "Optional" = lower-priority content
# H2 "Blocked" = content LLMs should not include in responses
## Optional
# - [Less-important page](https://yourdomain.com/misc/): Description
## Blocked
# - [Terms of Service](https://yourdomain.com/tos/): Legal content
# - [Internal search results](https://yourdomain.com/search?q=): Dynamic
# ── SECTION 4: USAGE DIRECTIVES ──────────────────────────────────
# These are NOT in the original spec but have become common practice.
# No crawler is obligated to honor these. Use robots.txt for enforcement.
# training: allowed | prohibited | contact-required
# retrieval: allowed | allowed-with-attribution | prohibited
# licensing_contact: [email protected]
# last_updated: 2026-05-19
# ── SECTION 5: llms-full.txt CONVENTION ──────────────────────────
# If you serve llms-full.txt, it should contain:
# - Full text of all linked pages (or significant excerpts)
# - Same markdown structure as llms.txt but with content inline
# - No size limit defined in spec; practical max is ~2MB for parser compat
# ── ROBOTS.TXT RELATIONSHIP ──────────────────────────────────────
# llms.txt does NOT interact with robots.txt directly.
# robots.txt controls crawl access (enforceable by compliant bots).
# llms.txt signals intent (advisory only, no enforcement mechanism).
# Both can and should coexist. They serve different functions.
A real implementation for a media brand looks something like this:
# The Broadsheet
> Independent journalism covering technology policy and digital rights since 2019.
We publish original reporting, analysis, and interviews. All content is written by
human journalists. Reproduction requires licensing. Contact [email protected].
## Featured Coverage
- [AI Governance Tracker](https://broadsheet.example.com/tracker/): Live database of AI regulation bills globally
- [Big Tech Antitrust Archive](https://broadsheet.example.com/antitrust/): Case filings and analysis since 2021
## Recent Analysis
- [The EU AI Act at One Year](https://broadsheet.example.com/eu-ai-act-year-one/): Implementation failures and wins
- [Who Owns Training Data](https://broadsheet.example.com/training-data-ownership/): Legal landscape as of Q1 2026
## Optional
- [About Us](https://broadsheet.example.com/about/): Masthead and editorial policy
- [Newsletter Archive](https://broadsheet.example.com/newsletter/): Weekly digest back-issues
## Blocked
- [Subscriber-only content](https://broadsheet.example.com/members/): Paywalled, licensing required
- [Search](https://broadsheet.example.com/search/): Dynamic results only
training: prohibited
retrieval: allowed-with-attribution
licensing_contact: [email protected]
last_updated: 2026-05-19
Who Is Actually Parsing It
I've spent time trying to verify this, and verification is harder than it should be. The honest answer: Perplexity is the most confirmed parser. They've stated publicly that their retrieval pipeline reads llms.txt to prioritize which pages to pull context from. That's meaningful — Perplexity has real traffic, and I've seen log file evidence of their bot fetching the file separately from content crawls.
Several smaller RAG (retrieval-augmented generation) pipeline tools — the kind used by enterprise internal chatbots, not consumer products — explicitly support llms.txt as a site configuration mechanism. Jina AI's Reader product respects it. Some of the open-source LLM indexing tools do too.
OpenAI's GPTBot and Anthropic's ClaudeBot: I have no public confirmation that either reads llms.txt for anything other than discovering pages to crawl. My log file analysis on four domains where I control server access shows these bots fetching llms.txt at roughly the same frequency as they fetch robots.txt — which suggests they at least retrieve it. Whether they do anything with the usage directives is opaque.
Google has been notably quiet. Google-Extended (their AI training crawler) does not appear to have any public documentation referencing llms.txt. This is important because if the largest search engine and one of the largest AI companies doesn't parse it, the file's direct influence is limited to a smaller ecosystem.
The Adoption Curve: Real Numbers, Not Projections
I pulled data from three public crawl datasets to triangulate adoption. The numbers are messy but directionally consistent.
As of April 2026: roughly 0.7% of the top 250,000 domains by traffic serve an llms.txt file. Among the top 10,000 domains, that figure jumps to about 4.1%. Among developer-focused domains (SaaS, documentation sites, technical blogs), I've seen figures as high as 11–14% in some samples.
The variance matters. llms.txt adoption is highly concentrated in tech-adjacent publishing. A traditional newspaper in a mid-sized US city? Almost certainly doesn't have one. A DevRel blog for a cloud infrastructure company? High probability.
The growth rate has been roughly doubling year-over-year since late 2024, which sounds impressive until you realize it's doubling from a small base. The plateau risk is real. Adoption probably stalls unless either (a) a major AI system creates a visible, trackable benefit for sites that use it, or (b) content licensing negotiations create a contractual requirement to maintain one.
That second scenario is already happening in isolated cases. I know of two publisher licensing agreements signed in the last six months where the AI company required the publisher to maintain a valid llms.txt as part of the deal structure — essentially as a machine-readable manifest of what's being licensed.
The SAR Framework: My Client Advice Process
I developed a simple three-step framework I call SAR — Signal, Audit, Restrict — for deciding how to approach llms.txt with any client. It keeps the conversation from getting tangled in hypotheticals.
Signal: Start with the file as a disclosure and identity layer. Even if no AI system perfectly parses it, having a well-structured llms.txt gives you a documented, timestamped record of your site's content structure and intent. This has legal utility. Several publishers have used their llms.txt (and its commit history in Git) as evidence in licensing negotiations — "this is what we made available, this is when, here's what we marked as blocked."
Audit: Before publishing anything, audit what you're including in the key pages section. I've seen clients accidentally list pages they'd rather not surface — internal product roadmaps published on a public URL but not prominently linked, old press releases with outdated claims. llms.txt can be a useful forcing function for a content audit.
Restrict: Pair the llms.txt with real enforcement mechanisms. robots.txt for crawler blocking. Cloudflare or Fastly rules for tarpit and rate limiting. Paywall infrastructure for monetized content. llms.txt alone restricts nothing — it's a polite note, not a lock.
The SAR framework is deliberately low-ceremony. I can walk through it with a client in 45 minutes and leave them with a working file and a sensible robots.txt update. See also my piece on managing 40+ AI crawlers in 2026 for the robots.txt side of this.
Two Things I Believe That Most SEOs Don't
First: llms.txt is more valuable for its legal signaling than for its technical function. The SEO community talks about it as a crawler-guidance mechanism, but its biggest practical use right now is as a legal artifact. When publishers negotiate licensing, they need a machine-readable, timestamped record of what they published and what they considered proprietary. llms.txt is evolving into that record — and that's worth more than any parser compliance rate.
Second: the push for a stable, well-governed spec is actually counterproductive. I know that's an odd position. But the spec's informal nature is what allowed it to spread. A formal W3C or IETF standardization process would add two to four years of committee latency and probably produce something nobody actually implements. The messy ecosystem of informal extensions is frustrating, but it's also evidence that real people with real needs are bending the format to fit their situations. Standardization would freeze that.
The Mistake I Made and Would Make Differently
In early 2024 I advised a mid-size SaaS client to add training: prohibited to their llms.txt and treat it as meaningful enforcement. I framed it as protection. It wasn't. Three months later, we found GPTBot had indexed their documentation thoroughly — the llms.txt directive had no effect on crawl behavior. The client felt misled, rightly. I had overstated what the file could do.
What I should have said: "This directive expresses your intent, creates a legal record that you explicitly did not consent to training use, and may influence a small number of compliant retrieval systems. It will not stop GPTBot." The legal-record framing would have been accurate. The enforcement framing was not.
I've corrected my standard client briefing to be clear about this distinction. Related: see the anti-AI scraping article for what actually enforces access restrictions.
What I Actually Tell Clients Right Now
The advice I give in May 2026 comes down to four points.
One: add llms.txt. It's a two-hour implementation for most sites. The cost is negligible. The upside — legal clarity, Perplexity retrieval optimization, forward-compatibility with systems that will parse it better in 12 months — is real even if hard to measure.
Two: keep it accurate. I've audited llms.txt files that listed pages returning 404. I've seen files that linked to content that no longer existed. A stale llms.txt is potentially worse than none — it signals inattention and gives bad data to retrieval systems. Add it to your quarterly content audit checklist.
Three: don't put anything in the file you haven't thought about. The "Blocked" section, in particular, draws attention to the existence of content you might prefer to leave un-highlighted. I've seen sites list member-only content in their Blocked section in a way that essentially advertises it. Think about what you're disclosing.
Four: if you're in a licensing conversation with an AI company, use your llms.txt history as a negotiating document. Export your Git history for the file. Show when specific pages were added or removed. This creates a timeline of your content availability and intent that has real value in contract negotiations.
For the robots.txt and active blocking side of this, see Managing 40+ AI Crawlers in 2026. For licensing strategy more broadly, Content Licensing for AI in 2026 goes deep on deal structure.
External reference: the original llmstxt.org spec remains the canonical starting point, though the community wiki now supplements it significantly.
Where This Goes From Here
llms.txt is not going to become the robots.txt of AI — not at its current adoption rate, not with the spec drift problem unresolved, not without stronger signals from the major AI labs that they take it seriously. But it's also not dead. It's found a niche: developer-adjacent publishers who want a low-friction way to signal their intent, and increasingly, publishers who need a machine-readable artifact for licensing negotiations.
The most interesting version of this story is the one where AI licensing deals become standard enough that llms.txt maintenance becomes contractually required by the buying companies — essentially forcing adoption from the demand side rather than relying on publisher initiative. That would change the adoption curve completely.
Whether that happens in the next 18 months depends heavily on how the broader content licensing market matures. The deals are getting signed. The question is whether the infrastructure around them — including files like this one — gets standardized or stays fragmented.
I'll update this piece when something material changes. If your llms.txt is in a state you're not sure about, the SOW templates piece has a section on scoping AI-related technical work that might help frame what you actually need done.
Filed under: AI Crawlers, Technical SEO, Content Strategy | Last verified: May 2026 | Reading time: ~14 minutes
