How This Started (And Why I Almost Turned Down the First Client)
The email came in October 2025. A head of content at a mid-market SaaS company — project management tooling, boring in the best way — wanted to know whether OpenAI had trained on their documentation. Specifically, whether competitors could now prompt ChatGPT and receive the company's proprietary methodology language, verbatim, without paying for a license or a seat.
My first instinct was to say no. Not "no, I won't help," but "no, that's not really something you audit." I had been doing technical SEO for nine years. Crawl budget analysis, Core Web Vitals, log file interpretation. This felt adjacent but strange, like being asked to assess a building's structural integrity because you once hung drywall.
I took the call anyway. By the end of it, I had quoted $19,500 for a six-week engagement. They said yes without negotiating.
That was audit number one. By February 2026, I had completed eleven of them. Fees ranged from $14,000 for a lean two-week sprint on a single content vertical to $47,000 for a full-site exposure assessment covering 23,000 indexed URLs across three brands. This is what I learned.
What a Training-Data Exposure Audit Actually Is
The term "training-data exposure" sounds more alarming than it usually is in practice. What we are actually measuring is the degree to which a website's content has been ingested by the large language models that power AI assistants, search summaries, and API products. "Exposure" is not inherently bad. The question is what type of content was exposed, to which models, under what data-handling terms, and whether that exposure creates legal, competitive, or brand risk.
A training-data exposure audit (TDEA) has four legs:
- Surface-area mapping. Which URLs were publicly accessible during known training windows? What proportion of those contained proprietary methodology, trademarked terminology, or content that was placed behind a soft paywall after the fact?
- Regurgitation testing. Can we extract verbatim or near-verbatim content from major models using probe prompts? If yes, how much, at what confidence level?
- Crawler signal analysis. What does the robots.txt history show about when (and whether) AI crawler blocks were implemented? What does server log data reveal about CCBot, GPTBot, and ClaudeBot activity?
- Risk stratification. Given the above, what is the legal exposure tier, the competitive exposure tier, and the recommended remediation priority?
None of this existed as a formal service category two years ago. The legal pressure that made it necessary accelerated faster than anyone anticipated.
The SCOPE Framework: My Named Audit Process
After the first three audits, I realized I was reinventing the methodology each time. Different clients asked different questions, I answered them in different orders, and the final reports looked nothing alike. Client four got a better product than client two purely because I had more practice. That is not a situation you can maintain if you plan to scale or bring in contractors.
So I codified it. I call it SCOPE:
- S — Surface-Area Mapping. Enumerate every URL the site has served publicly during the training windows of GPT-4 (cutoff approximately September 2021), GPT-4o (approximately April 2023), Claude 3 (approximately early 2023), and Llama 3 and 3.1 (approximately December 2023). Use Wayback CDX API queries, Search Console historical data, and crawl logs to reconstruct what was exposed and when.
- C — Crawler Signal Review. Pull server logs or CDN logs for the 18 months prior to the audit. Identify known AI crawler user-agent strings. Cross-reference against robots.txt versioning via Wayback Machine. Determine whether blocks were in place before or after major training windows closed.
- O — Output Testing. Run structured probe prompts against GPT-4o, Claude 3 Opus, Gemini 1.5 Pro, and where API access allows, Llama 3.1. Score outputs by verbatim match rate, paraphrase density, and citation behavior.
- P — Proprietary Content Classification. Working from the surface-area map, classify content by type: commodity (freely available elsewhere), differentiated (specific to this site), and sensitive (methodology, pricing, internal process documentation that was accidentally indexed).
- E — Exposure Risk Scoring. Assign each content class a risk tier across three dimensions: legal, competitive, and brand. Output a prioritized remediation list with specific technical recommendations.
SCOPE is not a proprietary technology. It is a documented process. The value is in how you interpret the outputs, not in the framework itself. I tell clients this upfront. They pay for judgment, not for the acronym.
Eleven Clients, Eleven Different Problems
The pattern I expected going in was that larger sites would have larger exposure. That turned out to be completely wrong.
The SaaS company from October 2025 had 4,200 indexed URLs at the time of their audit. Their exposure risk score was 7.1 out of 10. A regional law firm I audited in December 2025 had 310 indexed pages. Their risk score was 8.4. The difference was content type, not volume. The law firm had published detailed procedural guides — step-by-step walkthroughs of estate planning strategies specific to one state's statutes — that were indexed and crawlable for over two years before they implemented any AI crawler restrictions. Those guides regurgitated cleanly under probe prompts at rates above 60% verbatim match.
Client three was a B2B data company. Their concern was competitive: they worried that a rival could use an AI assistant to approximate their methodology without licensing their platform. The audit confirmed the worry. Eleven out of seventeen methodology documents published between 2021 and 2023 showed verbatim or near-verbatim regurgitation in at least one major model. The surface-area metrics for that audit: 17,400 total URLs mapped across three domains, 2,100 classified as differentiated, 340 classified as sensitive. The 340 sensitive URLs are the number the VP of Product keeps in a slide deck.
Client seven — a healthcare content publisher — had the opposite problem. They were worried about HIPAA-adjacent exposure from patient-facing articles. The audit showed almost no verbatim regurgitation, but a significant paraphrase density problem: the models had clearly ingested enough of the site's content to reproduce its distinctive explanatory structure, its preferred analogies, its tonal register. Hard to litigate. Easy to feel.
Client nine is the one I am most proud of. A specialty retailer with a 14-year-old product knowledge base. The audit discovered 680 URLs that had been publicly accessible between 2019 and 2022 but were subsequently moved behind a login gate in a site migration. Those URLs were archived by the Wayback Machine and almost certainly crawled by AI training pipelines before the migration. The client had no idea those pages had ever been public. That finding alone justified the fee.
What varies wildly across the eleven engagements is not the framework — SCOPE applies cleanly every time — but the remediation priorities. Some clients want to block future crawling immediately, which is straightforward. Some want to pursue legal remediation, which is not my lane but which I can inform. Some want to restructure their content strategy to reduce future exposure of proprietary material. That last category is the one where the SEO background actually helps.
The Probe Prompts We Use
The regurgitation testing phase is the most technically interesting part of the audit, and also the part where I get the most questions from other practitioners. I am publishing a subset of the probe prompt templates here because secrecy about methodology does not serve anyone. The interesting work is in interpreting the outputs, not in constructing the prompts.
The basic structure is a three-pass approach:
Pass 1: Direct citation probe. Ask the model to quote or summarize content from the target site. Look for verbatim strings of 15 or more tokens.
Pass 2: Completion probe. Provide the opening phrase of a proprietary paragraph and ask the model to complete it. Measure how closely the completion matches the original.
Pass 3: Paraphrase probe. Ask the model to explain a concept or methodology in its own words, using the same topical frame as the target content. Score for structural and tonal similarity.
## PROBE PROMPT TEMPLATES — TDEA REGURGITATION TESTING
## Use with: GPT-4o, Claude 3 Opus, Gemini 1.5 Pro
## Replace [CLIENT_CONCEPT] and [CLIENT_URL_SLUG] with target content identifiers
--- PASS 1: DIRECT CITATION PROBE ---
Prompt A (attribution trigger):
"What does [CLIENT_DOMAIN] say about [CLIENT_CONCEPT]?
Quote their explanation directly if you have it."
Prompt B (recall trigger):
"I'm looking for the exact wording used by [CLIENT_DOMAIN]
to define [CLIENT_PROPRIETARY_TERM].
Can you reproduce it?"
--- PASS 2: COMPLETION PROBE ---
Prompt C (sentence completion):
"Please complete this passage: '[FIRST_15_WORDS_OF_TARGET_PARAGRAPH]...'"
Prompt D (structural continuation):
"The following is the beginning of an article.
Continue it in the same style and with the same specific claims:
'[FIRST_SENTENCE_OF_TARGET_SECTION]'"
--- PASS 3: PARAPHRASE PROBE ---
Prompt E (concept elicitation):
"Explain [CLIENT_CONCEPT] the way [CLIENT_DOMAIN]
would explain it to a [TARGET_AUDIENCE] audience."
Prompt F (structural similarity):
"Write a [CONTENT_TYPE, e.g., 'step-by-step guide']
about [CLIENT_TOPIC] at roughly the same depth and length
as what you'd find at [CLIENT_URL]."
--- SCORING ---
Verbatim match: Levenshtein distance < 0.15 vs. original = VERBATIM
Near-verbatim: 0.15–0.30 = NEAR-VERBATIM
Structural similarity: cosine similarity of TF-IDF vectors > 0.72 = PARAPHRASE FLAG
Scoring is automated using a small Python script that calls the Wayback Machine CDX API to retrieve the canonical original, then compares against model outputs. I do not publish that script here, but the methodology is reproducible with standard NLP libraries.
One note on Pass 3: the paraphrase probe is the most contentious in legal discussions. A high structural similarity score does not necessarily constitute infringement. It is evidence of training influence, not proof of copying. I am careful in the audit reports to distinguish between "this model reproduces your content" and "this model has been shaped by your content." The legal teams who read these reports care deeply about that distinction.
Blocking the Crawlers That Feed the Models
The robots.txt situation in 2026 is messier than it should be. The major AI labs have published user-agent strings for their training crawlers. Compliance with robots.txt exclusions is not legally mandated in most jurisdictions, but it is the practical standard, and departing from it creates significant PR and legal exposure for the labs. In practice, the major crawlers honor the blocks.
The complication is history. Training runs happen on snapshots. If your robots.txt blocked GPTBot starting in August 2023 but the training data snapshot was taken in April 2023, the block did not help you. This is the core forensic challenge in the crawler signal review phase of SCOPE.
Here is a representative robots.txt block pattern covering the major known AI training crawlers as of May 2026:
## AI TRAINING CRAWLER BLOCK — robots.txt
## Verified agent strings as of May 2026
## Review regularly; new crawlers appear without announcement
User-agent: GPTBot
Disallow: /
User-agent: ChatGPT-User
Disallow: /
User-agent: CCBot
Disallow: /
User-agent: anthropic-ai
Disallow: /
User-agent: ClaudeBot
Disallow: /
User-agent: Claude-Web
Disallow: /
User-agent: Google-Extended
Disallow: /
User-agent: Googlebot-Extended
Disallow: /
User-agent: Omgilibot
Disallow: /
User-agent: FacebookBot
Disallow: /
User-agent: Diffbot
Disallow: /
User-agent: Bytespider
Disallow: /
User-agent: Applebot-Extended
Disallow: /
User-agent: cohere-ai
Disallow: /
User-agent: PerplexityBot
Disallow: /
User-agent: YouBot
Disallow: /
## Selective allow example: block training, allow product search
## Use only if you have a specific reason to permit a crawler
## User-agent: PerplexityBot
## Allow: /blog/
## Disallow: /methodology/
## Disallow: /platform/
Three things to note about this block list. First, "Disallow: /" blocks training crawls but may also affect AI-generated search summaries and citation features. Some clients want that. Others do not. The decision is business strategy, not pure technical hygiene. Second, the selective allow example at the bottom is something I actively debate with every client. Blocking PerplexityBot entirely means your content does not appear in Perplexity citations. For some sites, that traffic is meaningful. For others, the training risk outweighs the visibility benefit. Third, this list will be outdated by the time you read it. The landscape of crawler agents changes faster than any static document can track. Audit your own server logs for unknown user agents on a quarterly basis.
For sites using Cloudflare, there is now a one-click AI crawler block in the dashboard that handles most of these programmatically. It is not a substitute for a well-maintained robots.txt, but it is a reasonable starting point for smaller sites without dedicated technical teams. See our Cloudflare AI blocking walkthrough for implementation details.
What the Deliverable Actually Looks Like
Clients pay for a report. The report has to be useful to three different audiences at the same company: the technical team, the legal team, and the executive team. Getting all three right in a single document is genuinely hard. I have failed at it more than once.
Here is the skeleton I use for every audit deliverable:
## TRAINING-DATA EXPOSURE AUDIT — REPORT TEMPLATE
## [CLIENT NAME] | Audit Period: [DATE RANGE] | Completed: [DATE]
## Prepared by: [AUDITOR] | SCOPE Framework v2.1
---
EXECUTIVE SUMMARY (1 page)
- Risk tier: [LOW / MEDIUM / HIGH / CRITICAL]
- Surface area: [X] total URLs mapped; [Y] differentiated; [Z] sensitive
- Regurgitation findings: [X]% of tested sensitive URLs showed verbatim match
in at least one major model
- Crawler block status: [IMPLEMENTED / PARTIAL / NOT IMPLEMENTED]
as of [DATE]; retroactive coverage assessment: [DESCRIPTION]
- Top 3 recommended actions with timeline and estimated effort
---
SECTION 1: SURFACE-AREA MAPPING
1.1 URL inventory methodology
1.2 Training window coverage by model
- GPT-4 window (est. cutoff Sept 2021): [X] URLs exposed
- GPT-4o window (est. cutoff April 2023): [X] URLs exposed
- Claude 3 window (est. cutoff early 2023): [X] URLs exposed
- Llama 3/3.1 window (est. cutoff Dec 2023): [X] URLs exposed
1.3 Content classification breakdown
- Commodity: [X] URLs ([X]% of total)
- Differentiated: [X] URLs ([X]% of total)
- Sensitive: [X] URLs ([X]% of total)
1.4 Historical accessibility analysis (Wayback CDX findings)
1.5 Post-migration exposure gaps (previously public, now gated)
---
SECTION 2: CRAWLER SIGNAL REVIEW
2.1 Robots.txt version history
2.2 Known AI crawler activity in server/CDN logs
- GPTBot first observed: [DATE]; volume: [X] requests
- CCBot first observed: [DATE]; volume: [X] requests
- ClaudeBot first observed: [DATE]; volume: [X] requests
- [Other crawlers as found]
2.3 Block implementation timeline vs. training window timeline
2.4 Gap analysis: windows where crawling occurred without restriction
---
SECTION 3: REGURGITATION TESTING
3.1 Models tested and versions
3.2 Probe prompt methodology (SCOPE Pass 1–3)
3.3 Verbatim match findings by content class
3.4 Near-verbatim and paraphrase findings
3.5 Notable specific instances (with redacted examples for legal use)
---
SECTION 4: RISK STRATIFICATION
4.1 Legal exposure tier and rationale
4.2 Competitive exposure tier and rationale
4.3 Brand exposure tier and rationale
4.4 Priority remediation matrix (effort vs. impact)
---
SECTION 5: RECOMMENDATIONS
5.1 Immediate actions (0–30 days)
5.2 Medium-term actions (30–90 days)
5.3 Ongoing monitoring protocol
---
APPENDICES
A: Full URL inventory with classification
B: Probe prompt transcripts (full)
C: Server log extracts
D: Wayback CDX query results
E: Scoring methodology notes
The part that generates the most back-and-forth with legal teams is Section 3.5, the specific instances appendix. Lawyers want verbatim examples for potential litigation. They also want those examples handled carefully, because putting a model's verbatim regurgitation into a document that might be discoverable is itself a legal consideration. I now deliver Section 3 appendices as a separate encrypted file with restricted distribution. That decision came after client five's general counsel raised the issue on review. I should have thought of it myself.
Why Legal Teams Are Driving This Now
The 2024 and 2025 training-data lawsuits changed the dynamic. Not because they all succeeded — most are still in various stages of discovery or appeal as of this writing — but because they established that this was a litigable category. The question shifted from "could we theoretically have a claim?" to "do we actually have evidence to support one?"
That is where the audit comes in. General counsel at a company that suspects its content was trained on cannot file anything useful without evidence. A well-constructed TDEA provides the evidentiary foundation: timestamped surface-area data, crawler activity logs, regurgitation test results with scoring methodology. It does not constitute legal advice. I say this repeatedly in every engagement letter and at least once verbally in every kickoff call.
The shift I have noticed across the eleven clients is that the inbound is no longer coming from content or SEO teams. Eight of the eleven engagements were initiated by legal or compliance, who then pulled in the technical teams. Two years ago, this question would have come from a content strategist curious about AI. Now it comes from lawyers who have read the complaint filed in The New York Times Co. v. Microsoft Corp. and want to know if their company has a comparable case, or alternatively, if they are exposed as a defendant in a similar action from another direction.
For more background on the legal framework developing around training data, Cornell's Legal Information Institute copyright overview is a useful starting point, though the specific AI training doctrine is still being written by courts in real time. See also our breakdown of the 2025 training-data cases.
Two Things Nobody Wants to Hear
Most companies have no meaningful recourse, and the audit is still worth doing
The first contrarian take: for the majority of websites, a training-data exposure audit will confirm that yes, your content was almost certainly ingested by one or more models, and no, there is probably not much you can do about it retroactively from a legal standpoint. The copyright framework around training data is genuinely unsettled. Even the cases with the strongest fact patterns — involving clear verbatim reproduction of identifiable copyrighted works — are years from resolution.
I tell clients this before they sign. I tell them that the audit is not primarily a tool for building a lawsuit. It is a tool for understanding what happened, documenting it formally, adjusting future behavior, and making an informed business decision about whether the evidence warrants engaging litigation counsel. Most of the eleven clients would tell you the audit was worth it even though none of them have filed anything. Knowing is different from suspecting.
Blocking AI crawlers is not always the right call
The second contrarian take: the reflexive advice to block all AI crawlers immediately is often wrong for businesses that depend on discovery. A publisher whose revenue model includes being cited in AI search summaries is not obviously better off with a blanket GPTBot block. A SaaS company whose blog ranks because it gets cited in AI Overviews should think carefully before removing that signal.
This is not a popular position among people who frame AI training exposure as purely adversarial. The reality is that the calculus is different for every site, and anyone selling a one-size-fits-all "block everything" solution is selling a reflex, not a strategy. I have had two clients where I explicitly recommended against full crawler blocking after the audit, because the competitive risk of reduced AI visibility outweighed the training-exposure risk of their specific content profile. See when not to block AI crawlers for the longer version of this argument.
The Mistake I Made on Audit #4
Client four was a fintech company. Mid-size, Series C, compliance-heavy vertical. They wanted a surface-area assessment focused on their API documentation and developer guides, because they suspected proprietary implementation details had been ingested.
I ran the surface-area mapping using Search Console historical data and Wayback CDX queries. I classified 412 URLs as sensitive. I ran Pass 1 and Pass 2 probe prompts against all of them. Found verbatim matches on 37 URLs. Wrote a clean report. Delivered it on time.
Three weeks later, their head of engineering emailed me. He had found a subdomain — docs-legacy.theirdomain.com — that I had not mapped. It had been publicly accessible from 2020 through early 2023 and contained the exact implementation documentation they were most worried about. It was not in Search Console because it had never been submitted. It did not show up in my initial Wayback CDX query because I had queried the primary domain only.
I went back and audited the subdomain at no additional charge. Found significant verbatim regurgitation on 23 of the 89 sensitive URLs there. Revised the risk tier upward. The final report was materially different from what I had delivered.
The lesson: surface-area mapping has to include subdomain enumeration as a first pass. I now run DNS brute-forcing and certificate transparency log queries on every engagement before I begin any other phase of the audit. It added roughly four hours to the process. I do not charge extra for it. I probably should, but I feel too much obligation about the mistake to make it a revenue line.
For subdomain enumeration tooling, crt.sh certificate transparency search is the fastest starting point for most engagements.
What We Charge and Why
The fee range across eleven audits: $14,000 to $47,000. Here is how I think about the spread.
The $14,000 engagements are limited-scope sprints: single brand, known URL inventory under 2,000 pages, no server log access (we rely on CDN reports or Search Console only), and a condensed report format. Two weeks of work, mostly mine, with some automated tooling.
The $47,000 engagement was the multi-brand one: 23,000 indexed URLs across three domains, full server log analysis covering 18 months, subdomain enumeration on all three root domains (found six previously unknown subdomains with crawlable content across the three brands), regurgitation testing on 400+ specific URLs across five models, and a report structured for simultaneous delivery to three separate legal teams. That was six weeks, me plus one contractor, plus two rounds of legal review of the methodology section.
The middle of the range, $22,000–$31,000, covers most single-brand enterprise audits with server log access and a standard SCOPE deliverable. That is the sweet spot. It is also where I have the most clients: seven of eleven fall in that band.
I charge a flat project fee rather than hourly. Clients who have been through hourly billing on technical SEO audits have been burned before — scope creep on an hourly engagement is everyone's nightmare. The flat fee aligns my incentives with efficiency and their incentives with predictability. I scope carefully up front, I write detailed statements of work, and I do not take on projects where I cannot define the scope clearly enough to quote flat.
Read how we structure technical SEO audit pricing if you want to understand the broader framework.
Where This Is Heading
Twelve months from now, I expect training-data exposure audits to be a standard line item on enterprise content audits, the way Core Web Vitals assessments became standard after the 2021 Page Experience update. Not because everyone will need the full forensic treatment, but because the question will be expected. Legal teams will ask for it. Insurance carriers will ask for it. In M&A due diligence, it will become part of the IP and content asset review.
That mainstreaming will compress fees. The $47,000 engagement will not exist in two years at that price point, because the tools will automate more of it and competition will drive margins down. What will survive is the interpretive layer: the judgment about what the findings mean for a specific business, specific competitive context, specific legal posture. That part does not automate.
I am also watching the opt-in licensing frameworks that some AI labs have proposed in early 2026. If a viable mechanism develops for content publishers to audit, license, and receive compensation for training use, the nature of the audit service will shift from forensic investigation to ongoing compliance management. Recurring revenue instead of project fees. A different business, but one I would rather have.
For now, though, the forensic work is real, the demand is real, and the clients who have gone through the process — all eleven of them — have better information than they started with. That is a sufficient argument for the service line. It does not need to be more than that to be worth building.
