The Scale Problem Nobody Warned Me About
In the back half of 2025, I was brought in to help a mid-size travel brand retrofit alt text onto roughly 340,000 images — most of them AI-generated illustrations their content team had been producing with Midjourney and Flux since late 2024. The brief sounded tidy: run Vision API over the asset library, generate descriptive captions, inject them into the CMS, done. Three weeks, maybe four. Ship it and move on.
It took five months. And by the end I had a entirely different theory about how Google Lens citation works, a framework I now apply to every visual-heavy project, and one admitted mistake that cost the brand a non-trivial amount of image-search traffic for about six weeks.
This piece is my attempt to put all of it in one place — for practitioners working at scale with AI-generated images right now, in May 2026, when the disclosure requirements are real, the schema vocabulary has expanded in ways most SEOs haven't caught up to, and Google Lens is behaving in ways I can only describe as structurally interesting.
If you just want the schema snippets, scroll to the Schema Layer section. If you want the full picture, start here.
What Google Lens Actually Cites (And Why the Numbers Are Weird)
Let me start with the anomaly that sparked everything.
About eight weeks into the travel brand project, we had alt text deployed on approximately 60,000 images — a subset, all AI-generated illustrations of destinations. I was checking Google Search Console image data, which at this point feeds into a separate "Visual Search" performance tab in GSC (rolled out Q4 2025 for most accounts). The impressions for those newly-captioned images climbed 340% in four weeks. Expected. Normal. Alt text works.
What was not expected: the click-through rate on Lens-attributed traffic was 0.4%. Not 4%. Zero-point-four.
For context, our photographic images — real photography, not AI-generated — were pulling somewhere between 2.1% and 3.8% CTR from Lens. The AI illustrations, despite now having descriptive, accurate, keyword-conscious alt text, were essentially invisible in terms of downstream engagement. Google was showing them. Users weren't clicking.
At first I assumed it was a quality signal. Maybe Lens deprioritizes AI imagery in the results ranking, so the impressions are coming from low-visibility placements — bottom of the panel, below the fold. Plausible. But when I pulled position data, the AI images were averaging position 4.2 versus 5.6 for photos. They were ranking higher. They were getting seen. They were just not converting.
I ran this by two other practitioners working on visual search in early 2026 and both had noticed similar patterns across different verticals. One had a furniture brand with AI-rendered product lifestyle shots — same story. High Lens impressions, anomalously low CTR. The other was in publishing, editorial illustrations. Same gap.
There is no clean explanation for this yet. I have a theory, and I'll get to it. But the numbers are genuinely weird, and anyone telling you they've fully figured out Lens CTR for AI imagery is probably selling something.
The Position-CTR Inversion
One more data point before moving on. In March 2026, after we had schema deployed on the full 340K image set (more on that below), I ran a cohort comparison: AI images with complete ImageObject schema including creator and contentLocation versus AI images with only alt text and no structured data. The schema cohort had 23% higher CTR from Lens. Still dramatically below photo CTR. But schema was clearly doing something.
This is why I now treat ImageObject markup as non-negotiable for AI-generated content, even when the performance delta feels underwhelming in absolute terms.
Contrarian Take 1: Mass Auto-Captioning Is Quietly Wrecking Your Visual Authority
Here's something I believe that is not widely said: running Vision API over 300,000 images and bulk-deploying the output as alt text is one of the better ways to damage your site's visual search standing over the medium term.
I did this. I am telling you not to.
The problem is not that machine-generated captions are wrong — they're usually directionally accurate. The problem is that they are homogeneous at scale. Vision models describe images through a relatively narrow set of syntactic patterns. "A person sitting at a desk with a laptop." "A mountain landscape at sunset with clouds in the background." "An aerial view of a coastal city with blue water." These descriptions are fine, individually. Deployed across 50,000 images in the same site, they create an alt text corpus that reads like it was generated by a single voice in a single moment — because it was.
Google's quality signals for image search are not purely about the presence of alt text. They appear to weight distinctiveness, specificity, and what I'd call contextual coherence — the degree to which the alt text fits the surrounding editorial context rather than just describing the image's visual contents. A Vision API caption tells you what's in the frame. Good editorial alt text tells you why that image is there, what it contributes to the piece, and occasionally what the reader should notice.
When I audited our alt text output about three months in, I found that 18% of our captions were essentially duplicates or near-duplicates — different images, same description. "A scenic view of a tropical beach with palm trees and clear blue water" appeared 847 times. That is not an alt text library. That is a canonicalization problem waiting to happen.
The deeper issue is that mass auto-captioning is frictionless, and frictionless processes tend to get applied without editorial judgment. Teams use it because it clears the accessibility backlog, it satisfies the literal requirement for alt text, and it requires no specialist skill. All of those things are true. What is also true is that it produces a visual content layer that is structurally thin and algorithmically legible as such.
The Schema Layer: ImageObject, C2PA, and What Actually Matters
Let's get into the markup. This is where I spend most of my time now, and where I think the real leverage is for anyone trying to build durable visual search presence with AI-generated imagery.
The core schema you need is ImageObject, and in 2026 there are three properties that have become genuinely important for AI-generated content specifically: creator, contentLocation, and acquireLicensePage. The last one matters because Google is increasingly surfacing licensing information in Lens panels, and sites that declare it cleanly get different treatment in commercial image search contexts.
Here's the baseline ImageObject pattern I use for AI-generated travel imagery:
{
"@context": "https://schema.org",
"@type": "ImageObject",
"name": "Clifftop café overlooking the Aegean at golden hour, Santorini",
"description": "AI-generated illustration of a whitewashed café terrace perched above a volcanic caldera. Created to visualize a dining experience described in the accompanying editorial. Colors emphasize the blue-dome architecture characteristic of Oia.",
"contentUrl": "https://example.com/images/santorini-clifftop-cafe-ai-001.webp",
"encodingFormat": "image/webp",
"width": "1200",
"height": "800",
"creator": {
"@type": "SoftwareApplication",
"name": "Flux 1.1 Pro",
"url": "https://blackforestlabs.ai"
},
"contentLocation": {
"@type": "Place",
"name": "Oia, Santorini",
"geo": {
"@type": "GeoCoordinates",
"latitude": "36.4618",
"longitude": "25.3753"
}
},
"acquireLicensePage": "https://example.com/image-licensing",
"license": "https://creativecommons.org/licenses/by/4.0/",
"creditText": "AI-generated illustration by Example Media. Prompt-engineered and art-directed by the editorial team.",
"copyrightNotice": "2026 Example Media Ltd",
"dateCreated": "2025-11-14",
"isFamilyFriendly": true
}
The creator property taking a SoftwareApplication type is the disclosure-forward approach. Some practitioners use Organization and point to their own company as the creator, which isn't wrong from a schema validity standpoint but is arguably less precise. When the image was made by a model, saying the model made it is more accurate and — I suspect, though I can't prove this — better-received by Google's quality systems.
Embedding Images in Article Context
The ImageObject should not live in isolation. For editorial content, it should appear as part of the parent Article's image property, and the alt text in your HTML img tag should align with the name or description in the schema — not copy it verbatim, but echo the same specifics. Consistency between the markup layer and the visible HTML layer is something I've seen flagged in manual reviews.
<!-- HTML -->
<img
src="/images/santorini-clifftop-cafe-ai-001.webp"
alt="AI-generated illustration of a whitewashed Santorini café terrace above the caldera at golden hour"
width="1200"
height="800"
loading="lazy"
decoding="async"
/>
Notice that the alt text explicitly says "AI-generated illustration." This is intentional, and I'll return to why.
The contentLocation Field Is Doing More Work Than You Think
For travel, real estate, local business, and any geo-sensitive vertical, contentLocation with proper GeoCoordinates is the bridge between image search and Google Maps integration. As of Q1 2026, Lens results for location queries increasingly pull from entities with structured location data, not just visual similarity. An AI-rendered image of Santorini that declares its location in schema can now appear in Lens panels triggered by "Santorini cafes" even if it looks nothing like any real photograph of Santorini. The image is being understood at the semantic layer, not just the pixel layer.
This is genuinely new behavior. It changes the game for illustrated content in geo-sensitive contexts. See our visual search SEO deep dive for more on how Lens entity resolution is evolving.
My Vision API Workflow — And the Mistake I'm Not Proud Of
The workflow I landed on after five months looks like this:
# Vision API batch processing workflow (Python pseudocode)
import anthropic
import json
from pathlib import Path
client = anthropic.Anthropic()
def generate_alt_text(image_path: str, editorial_context: str, location_hint: str = "") -> dict:
"""
Generate structured alt text + schema metadata for an AI-generated image.
editorial_context: the article title and surrounding paragraph text
location_hint: place name if known, for contentLocation
"""
with open(image_path, "rb") as f:
image_data = base64.b64encode(f.read()).decode("utf-8")
prompt = f"""You are reviewing an AI-generated image for a travel editorial.
Editorial context: {editorial_context}
Location (if applicable): {location_hint}
Generate:
1. alt_text: 10-20 words, descriptive, includes "AI-generated" prefix, mentions location if given
2. schema_name: concise title for ImageObject.name (under 100 chars)
3. schema_description: 2-3 sentence description for ImageObject.description — includes what the image depicts,
its editorial purpose, and any notable visual details worth surfacing
4. content_location: the most specific place name inferable from context and image
Return valid JSON only."""
response = client.messages.create(
model="claude-sonnet-4-6",
max_tokens=600,
messages=[{
"role": "user",
"content": [
{
"type": "image",
"source": {
"type": "base64",
"media_type": "image/webp",
"data": image_data
}
},
{
"type": "text",
"text": prompt
}
]
}]
)
return json.loads(response.content[0].text)
def build_image_object_schema(metadata: dict, image_url: str, model_name: str, date_created: str) -> dict:
"""Assemble full ImageObject JSON-LD from Vision API output."""
schema = {
"@context": "https://schema.org",
"@type": "ImageObject",
"name": metadata["schema_name"],
"description": metadata["schema_description"],
"contentUrl": image_url,
"encodingFormat": "image/webp",
"creator": {
"@type": "SoftwareApplication",
"name": model_name
},
"dateCreated": date_created,
"isFamilyFriendly": True
}
if metadata.get("content_location"):
schema["contentLocation"] = {
"@type": "Place",
"name": metadata["content_location"]
}
return schema
The key difference from a naive Vision API pipeline: I pass editorial context alongside the image. The model isn't just describing pixels — it's understanding why the image is there, which produces meaningfully more useful alt text for both accessibility and search.
The Mistake
In October 2025, about six weeks into the project, I made a batch processing decision I'm still annoyed about. We had a backlog of about 90,000 images and the client was pushing for faster throughput. I dropped the editorial context from the prompt — no article titles, no surrounding text, just raw image description. Call it a throughput optimization. It got us through the backlog in two weeks instead of six.
Those 90,000 images performed measurably worse in Lens for the following six weeks. Not dramatically worse, but the click-through rates were lower, the average position was slightly higher (meaning lower rank), and the GSC visual search report showed fewer rich result enhancements applied. We eventually re-ran the full workflow on that cohort. The metrics recovered. But it cost time and demonstrated something I should have known: context is not optional overhead. It's the actual value-add.
If you're running Vision API at scale and skipping the surrounding text to save tokens, you're optimizing the wrong variable. See our n8n SEO automation guide for pipeline architectures that pass editorial context efficiently without blowing out API costs.
The ACAT Framework for AI-Captioned Imagery
After too many late nights puzzling over image search data, I needed a mental model that was simple enough to actually use during client briefings. I call it ACAT — not because the name is especially clever, but because it sticks.
A — Accuracy at the Pixel Level. The description must be factually correct about what is visually present. This sounds obvious, but Vision models hallucinate image details with surprising frequency. A QA step where a human spot-checks 5-10% of outputs is non-negotiable. For AI-generated images specifically, accuracy includes acknowledging the artificial nature of the image.
C — Context from the Editorial Layer. The alt text and schema description should reflect the image's role in the piece, not just its visual contents. "A coastal city at dusk" is a pixel-level description. "AI-generated aerial illustration of Dubrovnik's old town used to open a piece about off-season Adriatic travel" is a contextual description. The second one is roughly three times more useful to a search engine trying to understand relevance.
A — Attribution via Schema. Every AI-generated image needs explicit creator attribution in ImageObject markup. The model, the date, the content URL. Not because it's legally required everywhere (though in some jurisdictions it increasingly is), but because structured attribution is how you participate in the provenance ecosystem that search engines are building right now. See the schema additions guide for the full vocabulary changes from 2025-2026.
T — Topical Alignment. The image's alt text and schema should use vocabulary consistent with the page's topical focus. This is just standard on-page coherence applied to images, but it's consistently overlooked in bulk caption workflows. A Vision API default output uses its own vocabulary. Your site has a vocabulary. They need to align.
ACAT is not a checklist so much as a sequence. You do them in order because each step builds on the last. Accuracy gives you raw material. Context shapes it. Attribution structures it. Topical alignment connects it to the rest of the page.
Contrarian Take 2: Handcrafted Alt Text Is the Last Defensible Moat
This contradicts the premise of the last 2,000 words, and I'm saying it anyway.
If you have a content team capable of writing specific, expert, editorially-grounded alt text for images — human beings who understand the subject matter, the audience, and the site's topical authority — that labor is currently worth more per image than any automated pipeline, including mine.
The reason is signal differentiation. When every major publisher is running Vision API variants over their AI image libraries, the outputs are going to converge. The vocabulary, the sentence structures, the level of specificity — all of it will cluster around the mean of what a vision model produces when instructed to describe an image. This is already happening. I can sometimes tell, when auditing a competitor site's image metadata, that their captions came from the same model generation process as ours. They feel the same.
Handcrafted alt text, written by a subject matter expert, does not converge to that mean. A volcanologist writing alt text for an image of a lava field writes differently than Claude or GPT-4o does. A professional chef captioning food photography uses different vocabulary than any general-purpose vision model. That difference is not just stylistic — it carries topical authority signals that automated systems cannot reliably replicate.
The practical argument against this is cost. At 340,000 images, there is no realistic budget for human-written alt text. Agreed. The moat argument applies to focused, high-authority content — feature articles, flagship guides, anything that needs to rank against serious competition. For bulk asset libraries, automate with ACAT. For your most strategically important content, write it yourself or hire someone who actually knows the subject.
This is not a comfortable position to hold simultaneously with advocating for automated workflows. I hold it anyway because the data supports it: our highest-performing images in Lens, measured by both CTR and the quality of the downstream traffic, are the ones where a human wrote the alt text with the article in mind from the start. Not the ones where we reverse-engineered a caption after the fact.
C2PA Disclosure in the Wild: Where Schema Meets Legal Reality
The Coalition for Content Provenance and Authenticity (C2PA) standard has gone from something you read about in trade press to something publishers with significant AI image libraries are actually implementing. Not everywhere, not even most places, but it's real now in a way it wasn't eighteen months ago.
For SEO purposes, the relevant question is whether C2PA metadata interacts with Google's systems in a way that affects search treatment. The honest answer is: probably, partially, and we can't measure it cleanly yet.
What we can observe: Google has stated publicly that C2PA provenance data is one of the signals it uses to label AI-generated images in Search and Lens. Sites that have C2PA credentials embedded in their images — the actual binary metadata, not just schema declarations — are getting "AI-generated" labels in Lens panels more consistently. This matters for expectation-setting and, relatedly, for CTR. Users who see an AI-generated label before clicking may be more likely to click if the image is clearly serving an illustrative rather than documentary purpose.
Here's what a C2PA-aligned disclosure looks like in schema form, combining ImageObject with AI-generation disclosure vocabulary:
{
"@context": "https://schema.org",
"@type": "ImageObject",
"name": "AI-generated illustration of Lisbon's Alfama district at sunrise",
"contentUrl": "https://example.com/images/lisbon-alfama-ai-dawn.webp",
"creator": {
"@type": "SoftwareApplication",
"name": "Midjourney v7",
"url": "https://www.midjourney.com"
},
"copyrightNotice": "2026 Example Media Ltd. AI-generated image.",
"creditText": "AI-generated by Midjourney v7 under direction of Example Media editorial team",
"acquireLicensePage": "https://example.com/image-licensing",
"usageInfo": "https://example.com/ai-content-policy",
"conditionsOfAccess": "Free for editorial use with attribution. Commercial use requires license.",
"encodingFormat": "image/webp"
}
The usageInfo property pointing to an AI content policy page is something I've added to all client implementations since January 2026. It's a weak signal on its own, but it contributes to the overall provenance picture that Google and Bing are assembling. Think of it as a consistency signal: if your image declaration, your on-page label, your content policy, and your C2PA binary metadata all say the same thing, you are building a coherent provenance record. If any of those contradict each other, you have a trust problem that will eventually manifest in search behavior.
The EU AI Act disclosure requirements, now in their enforcement phase for high-volume publishers, also create regulatory context that makes this more than an SEO decision. Not the place for legal advice, but: the schema implementation and the legal compliance implementation are the same implementation. Do them together.
External reference on C2PA implementation specifics: the C2PA 2.1 specification is the authoritative source. Don't rely on summaries.
Lens Citation Patterns: Three Things I Cannot Fully Explain
I promised you weird numbers. Here they are.
Pattern 1: The Freshness Cliff
AI-generated images on our travel site see a sharp drop in Lens impressions at approximately 90 days post-indexing, then recover to a lower baseline. Photographic images do not show this pattern — they either grow steadily or plateau. The 90-day cliff for AI imagery is consistent across three different image cohorts I've tracked. I don't know what causes it. My best guess is some kind of recency weighting in the visual similarity index that treats AI images differently from photographs, but that's speculation.
The practical implication: if you're tracking AI image Lens performance, don't compare month-one data to month-four data without accounting for this pattern. It will look like a penalty or a drop when it may just be the normal trajectory.
Pattern 2: Alt Text Length Inversely Correlates with Lens CTR Above a Threshold
Images with alt text under 120 characters have higher Lens CTR than images with alt text between 120 and 200 characters. Images with alt text over 200 characters perform similarly to the under-120 cohort. This U-shaped relationship doesn't appear in any guidance I've seen, and I'm not sure what to make of it. My current hypothesis is that mid-length alt text often tries to be comprehensive and ends up being neither punchy enough to trigger a specific query nor rich enough to match a broad one. But I'm holding this loosely.
Pattern 3: Images With contentLocation Schema Outperform in Lens, But Only for Non-Branded Queries
When I split Lens traffic by branded versus non-branded query type (imperfect but possible in GSC with enough filter work), the schema with contentLocation boosts non-branded performance substantially — about 31% more impressions — while having no measurable effect on branded queries. This makes intuitive sense: for a branded query, users already know who you are. For "best cafes Santorini" or "Alfama district Lisbon," location-structured imagery competes in a semantic space where schema gives you an edge.
This is actually actionable. If your AI image library is primarily supporting non-branded discovery, contentLocation is not optional. If you're primarily trying to defend brand SERP territory, deprioritize it and focus on acquireLicensePage and creditText for commercial contexts. See our Google Lens 2026 guide for the broader query-type framework.
The second external reference: Google's own documentation on image license structured data is the baseline, though it lags actual behavior by several months in my experience.
Where This Goes From Here
The thing about working at the intersection of AI-generated content, schema markup, and visual search is that the ground moves fast and the data lags by months. Everything I've described above reflects conditions as of May 2026. Some of it will be wrong by year-end.
What I'm reasonably confident will remain true: the provenance layer matters more over time, not less. The gap between sites that have coherent AI disclosure infrastructure — schema, binary metadata, editorial labeling, policy pages — and sites that have none is going to widen as Google and the EU enforcement apparatus mature. The sites treating C2PA and ImageObject as checkbox compliance rather than strategic infrastructure are going to look like the sites that had no mobile version in 2014.
The Lens CTR gap between AI imagery and photography is also unlikely to close on its own. My current belief is that it reflects genuine user behavior, not just algorithmic treatment. People are more likely to click a photo than an illustration when the query is informational. The solution isn't better alt text — it's being strategic about when you use AI imagery at all, and when you invest in photography for the content types that drive Lens-based discovery.
As for the weird numbers — the 90-day cliff, the U-shaped alt text curve, the branded versus non-branded location split — I'll keep tracking them. If any of it resolves into something I can explain cleanly, I'll publish an update. Right now they're anomalies, and anomalies are where the interesting work lives.
The ACAT framework, the ImageObject patterns, the C2PA disclosure approach: these are what I'm deploying today. Not because they're perfect, but because they're the most coherent response I have to conditions that are genuinely still in formation. Check the AI citation tracking piece for how this connects to the broader GEO picture — because Lens citation and LLM citation are starting to share more infrastructure than most people realize.
