'AI SEO' Is Optimizing the Wrong Layer
Most GEO advice tunes the content that gets cited. The harder question is whether your page was ever in the room when the model chose.
Ask whether generative engine optimization is real and you get two camps talking past each other. One camp says it is snake oil, a rebrand of SEO with a new acronym. The other camp points at the Princeton GEO paper and says the effects are measured and reproducible, so of course it works. Both are describing the same study and reaching opposite verdicts, and the reason is that they are answering different questions. GEO produces real, measured lifts. It produces them at a stage of the pipeline that most advice never names, and the advice that gets sold as GEO mostly operates one layer too late.
Here is the distinction that resolves the argument. A retrieval-augmented answer is built in two separable phases. First the system decides which pages are even eligible: it rewrites your question into several sub-queries, pulls a candidate set of documents for each one from a search backend, and reranks that pool down to a handful of passages that fit the context window. Call that the retrieval layer. Then the model reads those passages and writes an answer, deciding which of the surviving sources to quote, paraphrase, or cite. Call that the generation layer. The famous GEO tactics all live in the second phase.
What the Princeton paper actually measured
Aggarwal and colleagues built GEO-Bench, roughly 10,000 queries across about 25 domains, and tested nine content interventions against a visibility metric. The winners are the ones everyone repeats: add citations, add statistics, add direct quotations, improve fluency. The best of them moved their visibility metric by up to around 40 percent, and lower-ranked sources gained the most, some by over 100 percent. Keyword stuffing did nothing or slightly hurt. The effects were domain-dependent, with citations helping most on factual queries and an authoritative tone helping on debate and history.
Read the method and one detail settles everything. Every source in that benchmark had already been retrieved before the intervention was applied. The experiment holds the candidate set fixed and rewrites the passages inside it, then measures how the model apportions visibility among documents it was always going to see. That is a clean and useful result. It tells you how to win the competition among retrieved sources. It says nothing about how to enter that competition, because entry was a precondition of the test rather than a variable in it.
The AI Visibility Funnel has four stages: Retrievability, Candidacy, Citation, Attribution. Retrievability is whether the crawler can reach and parse your page at all. Candidacy is whether your page makes the candidate set for a given sub-query. Citation is whether the model, having read the candidates, quotes or references you. Attribution is whether the visible answer names you as the source rather than folding your fact into unsourced prose. The Princeton interventions are Citation-stage moves. They assume Retrievability and Candidacy are already solved, and for the pages in a curated benchmark they were.
Why the content tweaks feel like the whole job
There is a reason the generation layer gets all the attention. It is the only layer you can see and touch. You can open your CMS, add an FAQ block, drop in a statistic with a source, tighten three paragraphs, and feel the satisfying click of work completed. The page looks better afterward. You can point to the diff. None of that tells you whether the page was one of the six passages the reranker kept for the query you care about, and most GEO audits never check, because checking requires observing the retrieval step, which does not happen inside your CMS.
Compare it to a hiring process. You can spend a week polishing a candidate's interview answers, coaching the handshake, rehearsing the closing pitch. All of it matters, and all of it is worthless if the resume never cleared the screen and the candidate is sitting at home during the interview. The Princeton study is a rigorous measurement of interview performance among people already in the building. The question practitioners actually need answered is how many rooms your name is being called into, and the polish work cannot answer it because it happens after the call.
This is where Share of Retrieval earns its place as the leading indicator. Citation counts tell you the outcome of contests you were already entered in. Share of Retrieval measures the fraction of relevant candidate sets your content appears in at all, upstream of any citation decision. A page can have flawless statistics and clean quotable passages and still hold a Share of Retrieval near zero for its target prompts, in which case every Citation-stage tactic is being applied to an audience the model never assembles. The lift is real. It is being spent on rounds you rarely reach.
The Retrieval Gap is where the losses actually happen
For any query there is a large set of pages on the open web that could correctly answer it, and a much smaller set the pipeline actually pulls into a candidate pool. The distance between those two sets is the Retrieval Gap, and for most brands it is enormous. Your product page might be one of a thousand documents that genuinely answer "best options for X," while the reranker for that sub-query keeps twelve. If you are not among the twelve, the answer gets written from sources that are, and it does not matter what your page says because the model never reads it.
Two facts about how these systems behave tell you where the gap gets decided. The fan-out step means a single user question becomes several sub-queries, so you compete for presence across a spread of candidate sets rather than for a single one, and a page tuned for the exact head phrase can miss the reformulations entirely. And the candidate pool for each sub-query is drawn from a search backend that indexes the whole web, so the sources competing for those slots reach far beyond your own site, into every review, forum thread, comparison article, and documentation page that mentions your category. Retrieval is decided across that whole surface, not on your domain.
Practitioner datasets often show third-party pages contributing heavily to brand mentions, but the share changes with the category, prompt set, and engine. The practical conclusion is that many retrieval opportunities live on property you do not control. You can influence those pages without being able to edit them, and no amount of on-site FAQ schema reaches them.
The Citation Surface is the lever the advice skips
The full set of third-party pages where a model can find and cite you is your Citation Surface, and it is the part of GEO that actually moves the retrieval layer. When a reranker assembles candidates for "is tool X worth it," the documents in contention are the comparison roundups, the community answers, the integration docs, the analyst notes. Getting mentioned, accurately and in a self-contained passage, on the pages that already retrieve well for your prompts does something the FAQ block cannot: it puts you into candidate sets you were absent from. That is a Candidacy-stage intervention, and it is the one most GEO checklists leave out entirely.
The mechanics reward a specific shape of content. Self-contained passages that answer a specific question are easier for retrieval systems to use, because a chunk that stands without its surrounding page survives the trip into a context window intact. Cited sources can include both old, durable references and recently refreshed pages; age is entangled with authority, links, and topic stability. Entity recognition and demand may correlate with citation frequency in observational datasets, but neither establishes a causal ranking factor. None of these levers live in your CMS. All of them shape whether you make the candidate set.
The counterargument, and where it stops
The honest objection is that content quality clearly does matter once you are retrieved, and the Princeton numbers prove it. That is correct, and it is worth stating plainly rather than waving away. Among retrieved sources, the page with the clean statistic and the quotable line really does win citations at a meaningfully higher rate, and a brand that ignores the generation layer will lose contests it entered. Content tuning is doing genuine work.
The objection only fails when it is treated as the whole strategy, because it confuses a necessary condition for a sufficient one. Citation-stage polish is necessary: skip it and you lose winnable rounds. It is not sufficient, because it cannot enter you into rounds you are absent from, and for most brands the larger loss is absence rather than defeat. The right sequence follows the funnel. Fix Retrievability first, so your pages and key third-party mentions are crawlable by the search-specific agents and fetchers used by the engines you measure. Then work Candidacy, by building and correcting your Citation Surface across the third-party pages that already retrieve for your prompts. Only then does Citation-stage content work pay its full return, because now it is being applied to contests you are reliably in.
What to do instead of another FAQ block
Start by mapping the real space of questions customers ask, not a flat keyword list. Prompt-Space Coverage means enumerating the actual prompts, including the messy conversational reformulations the fan-out step generates, and measuring your visibility across that space rather than for a handful of head terms. A keyword you rank for on Google can correspond to five distinct sub-queries in a RAG pipeline, and you may be present for one of them.
Then measure Share of Retrieval before you touch a single passage, so you know whether your problem is Candidacy or Citation. If you are absent from the candidate sets, no content tweak will help and the work is on your Citation Surface. If you are present but not cited, the Princeton playbook is exactly right and you should run it. Watching this properly means sampling prompts repeatedly across the four engines and recording where you enter candidate sets versus where you merely could have, which is the measurement shareof.ai automates so the retrieval layer stops being the invisible one. The point is to stop grading the interview and start counting the rooms.
A note on measurement
The retrieval layer is genuinely hard to observe from outside, and every figure here that comes from practitioner analysis rather than the Princeton benchmark is correlational and category-dependent. The third-party share, the passage-length effect, the age skew, the branded-search correlation: treat each as a hypothesis to test against your own prompt set, not a constant. The two-layer model holds regardless of the exact numbers, because it follows from how RAG pipelines are built rather than from any single measurement. Retrieval happens before generation, attribution is not the same as retrieval, and any optimization that only touches the words on your page is working on the last twenty percent of the problem while the first eighty decides whether that page is ever read.
Sources and further reading: Aggarwal et al., "GEO: Generative Engine Optimization," arXiv:2311.09735 (2023), for the benchmark and the Citation-stage interventions. For the pipeline mechanics, any current description of retrieval-augmented generation covering query fan-out, candidate retrieval, cross-encoder reranking, and citation attachment. Verify the crawler user-agents against each provider's current documentation before editing robots.txt.
Primary sources and evidence
- GEO: Generative Engine Optimization (Aggarwal et al.) — the original paper and benchmark.
- Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks — the foundational RAG paper.
Continue through the system
- The AI Visibility Funnel: A Diagnostic for Where You're Losing Citations
- Share of Voice vs Share of Model vs Share of Retrieval
- Schema Markup Won't Save You: What Structured Data Does and Doesn't Do for LLMs
Run the practical layer with the free AI visibility scan, explore the AI search library, or see the AI visibility benchmarks.