shareof.ai

What the Princeton GEO Paper Really Proved (and Where Practitioners Get It Wrong)

Everyone quotes the "+40%" number. Almost nobody read the method that produced it, so the advice built on top of it keeps misfiring.

The paper that launched a thousand LinkedIn posts is Aggarwal et al., 2023, "GEO: Generative Engine Optimization," out of Princeton and IIT Delhi. It gets cited as proof that you can boost your visibility in AI answers by 40 percent if you add statistics and citations to your pages. That summary is roughly true and badly compressed, to the point of being misleading, and the compression drops the two things a practitioner actually needs: the metric that 40 percent describes, and the fact that the winning move changes depending on the query.

Read the actual method and a more useful picture emerges. The paper measured one specific stage of how AI answers get made, held everything upstream of that stage constant, and found that the best edit for a legal query is not the best edit for a factual lookup. Most of the field took the headline and threw away the conditions. This piece puts them back.

What GEO-Bench actually is

The authors built a benchmark called GEO-Bench: roughly 10,000 queries drawn from existing datasets and spread across about 25 domains, from law and government to history, science, and opinion. For each query they assembled a set of source documents, the kind of pages a generative engine would retrieve and read before answering. Then they ran the query through a generative engine, edited the source content one method at a time, and measured how the edit changed that source's presence in the generated answer.

Two details in that setup decide everything downstream. The first is that the source set was fixed. The engine was not sent out to re-crawl the web; it was handed a candidate set of documents and asked to answer from them. Every result in the paper is about what happens once your page is already in the room. The second is how they measured visibility, because "visibility" in an AI answer means a share of the generated text and where in that text you appear, not a rank on a list.

They used two metrics. One counts words: how much of the final answer is attributable to your source, weighted so that being cited early and prominently counts for more than a passing mention near the end. The other is a subjective impression score, an LLM judging how relevant, influential, and useful your source appeared in the answer. Both are share-of-answer measures. Neither is a position on a list. When you see "+40%," it means one of these two shares rose by that fraction against the unedited baseline, on average, for the best-performing methods.

The nine methods, and which ones moved the number

The paper tested nine content edits. Naming them matters, because the popular retelling collapses nine into two.

Adding citations to reputable sources. Adding relevant statistics. Adding direct quotations. Improving fluency. Making the text easier to understand. Adding an authoritative tone. Inserting unique or technical vocabulary. Optimizing for specific keywords. Padding with keyword-stuffed repetition of the query terms.

Four of these consistently raised visibility across the benchmark: citations, statistics, quotations, and fluency. The top methods reached up to roughly a 40 percent relative gain on the visibility metric. Adding statistics stood out on one engine in particular, lifting the subjective-impression score on Perplexity by around a third. The mechanism is not mysterious. A generative engine assembling an answer favors passages that read as substantiated and self-contained, and a sentence carrying a figure or a cited claim survives summarization better than a vague one.

Keyword stuffing did nothing useful, and in places it hurt. This is the finding the industry keeps failing to internalize. The single most durable habit from a decade of search optimization, repeating the target phrase until the page reeks of it, is the one method the paper shows to be inert or counterproductive in a generative engine. The model does no term-frequency counting at all. It reads for something it can lift into an answer and stand behind.

The part the "+40%" retelling erases: it depends on the domain

Here is the finding that almost never survives the trip to a slide deck. The methods did not rank the same way across domains. The best edit was a function of the query type.

Citations helped most on factual queries, where the answer turns on whether a claim is backed. An authoritative tone did its best work on debate and historical topics, where the engine is weighing framing and stance rather than looking up a fact. Statistics paid off most on law, government, and opinion domains, where a number lends weight to an argument that is otherwise contestable. The averages that produce the headline figure are averaged over all of this variation, and the average conceals the rule that would actually help you: match the method to the question.

Query typeMethod that helped most
Factual lookupsCitations to reputable sources
Debate and historicalAuthoritative tone
Law, government, opinionRelevant statistics

Treat that table as the paper's real contribution and the "add statistics to everything" advice falls apart. Statistics were close to the strongest lever in a legal or policy context and a weaker one elsewhere. A recipe page or a product comparison does not become more citable because you bolted a percentage onto it. The generic prescription is a domain-specific result stripped of the domain, which is roughly the least useful thing you can do with a study whose whole point was that context decides.

Who gained the most, and why that reframes the ceiling

One more result deserves more attention than it gets. The sources that gained the most from these edits were the ones ranked lowest to begin with. Pages sitting near the bottom of the candidate set saw the largest relative jumps, some more than doubling their visibility, while pages already dominant had less room to move.

Read that carefully and it is genuinely encouraging for a challenger brand. If you are the incumbent everyone already cites, content edits give you modest headroom. If you are the fifth-best source on a topic and you are getting crumbs of the answer, a well-chosen edit can move you a long way in relative terms, because you started with so little to lose. The paper is describing a lever that works hardest exactly where a smaller player needs it to.

There is a quieter caveat inside the same result. A large relative gain from a small base is still a small absolute presence. Doubling your share of an answer when your share was five percent gets you to ten, which is real and worth having, and is not the same as owning the answer. The honest version of the finding is that these edits redistribute presence at the margin, and the margin is where most brands actually compete.

The boundary of the study, stated plainly

The single most important sentence about this paper is one the paper says clearly and the summaries omit. Retrieval was held fixed. The source documents were given to the engine. The experiment never tested whether a page gets pulled into the candidate set in the first place, only what happens to a page once it is already there.

In the terms of the AI Visibility Funnel, four stages run from Retrievability to Candidacy to Citation to Attribution. The GEO paper lives almost entirely in the Citation stage. It answers the question, "given that my page reached the model's context, how do I get more of the answer built from it?" It is silent on the two stages before that, whether the engine can crawl and index you at all, and whether you enter the candidate set for the sub-queries a real prompt fans out into. Those upstream stages are governed by Share of Retrieval, and no amount of adding statistics changes a candidate set you were never part of.

This is why "I followed the GEO paper and nothing happened" is such a common complaint, and why it is usually not the paper's fault. If your page is not being retrieved, the study's methods have no surface to act on. You optimized the shot without checking whether you were on the field. The paper is a strong guide to converting candidacy into citation and says nothing about earning candidacy, and conflating the two is the most expensive misreading in circulation.

What to actually do with it, per query type

The usable version of this research reads less like a checklist you apply everywhere and more like a small decision tree keyed to the kind of query you are trying to win.

For factual and how-to queries, invest in citations and self-contained, verifiable claims. The engine is looking for something it can attribute, so be the page that states the fact cleanly and backs it. For contested, editorial, or historical queries, a clear authoritative stance does more than a pile of numbers, because the model is arbitrating framing rather than fact-checking. For legal, regulatory, financial, and policy queries, relevant statistics carry unusual weight, so lead with the figure that anchors the argument.

Across all of them, stop stuffing keywords, and treat fluency as a real ranking input rather than a nicety, since the clearest passage is the one most likely to be lifted whole. And before any of this, confirm you are actually being retrieved, because the entire study assumes a gate you may not have cleared. Tools like shareof.ai exist to measure that upstream stage, your Share of Retrieval across the prompt space your customers really use, which is the number the GEO paper deliberately took off the table so it could study the one after it.

A note on measurement

The GEO paper is good research, and the way it is cited is a small case study in how findings decay. A conditional result ("statistics help most on legal and policy queries, when retrieval is already solved") hardens into an unconditional slogan ("add statistics for +40%") within about two forwards. If you are going to build a program on this study, build it on the conditions rather than the headline. The metric is share of the generated answer. Effects change with the domain. The biggest gains accrue to underdogs, and the whole result assumes you were retrieved in the first place. Verify each of these on your own category before you spend a quarter acting on any of them, because the one thing the paper proves beyond doubt is that the right move is not the same move everywhere.

Sources and further reading: Aggarwal, Murahari, et al., "GEO: Generative Engine Optimization" (arXiv:2311.09735, 2023), and the accompanying GEO-Bench dataset.

Primary sources and evidence

Continue through the system

Run the practical layer with the free AI visibility scan, explore the AI search library, or see the AI visibility benchmarks.