The AI Visibility Funnel: A Diagnostic for Where You're Losing Citations
Four stages sit between your page and a named mention in an AI answer. Find the one that is failing before you spend a week fixing the wrong layer.
Most brands that ask why they are invisible in ChatGPT or Perplexity treat it as one problem with one answer. Publish more content, add some schema, wait. The trouble is that "not showing up" describes at least four different failures, and the fix for each is unrelated to the fix for the others. A page that the crawler never fetched needs nothing you would do to a page that gets retrieved but never quoted, and both of those are separate again from a page that gets quoted without your name attached. Guess wrong about which stage is broken and every hour of work lands one layer away from the actual leak.
The way through is a funnel. A retrieval-augmented answer is assembled in stages, and your content can drop out at any of them. The AI Visibility Funnel names the four: Retrievability, then Candidacy, then Citation, then Attribution. Content flows left to right, and each stage passes only a fraction of what entered it. The value of the model is that each stage has its own symptom you can observe from outside, its own cause, and its own concrete fix. Work the funnel in order and you stop treating a Candidacy problem with a Citation-stage tactic.
The master diagnostic: retrieval versus naming
Before the four stages, one split settles most cases in about a minute. Ask a real customer prompt in the engine you care about and look at the sources panel, the list of citations the answer hangs off. Two questions decide everything downstream.
Do pages from your category appear in that source list at all? And when they do, are you the brand being named, or is it a competitor and a couple of roundup sites?
If your kind of page does not surface in the sources, the failure is early: Retrievability or Candidacy. The model never had your content in front of it, so nothing you write in the content itself can help until you are back in the candidate set. If your pages, or third-party pages that describe you accurately, do surface in the sources but the written answer names someone else, the failure is late: Citation or Attribution. You were in the room and the model chose to quote around you. Retrieval is the first half of the funnel and naming is the second, and almost every misdiagnosis comes from confusing the two. The rest of this piece is the detailed version of that one look at the sources panel.
Stage one: Retrievability
Retrievability is the plainest stage. It asks whether an AI crawler can reach your page, fetch it, and read the content that matters. If the crawler gets a login wall, a timeout, or an empty shell of HTML, your page does not exist as far as the index is concerned, and the three stages after this one never get a chance to run.
Symptom. You are absent from source panels across many different prompts, not just competitive ones. Even navigational or branded queries, where the answer should obviously involve you, pull from third-party pages about you rather than from your own site. Your server logs show few or no hits from the AI user-agents.
Cause. Usually one of two things. Either robots.txt or a firewall rule is blocking the crawlers, or your key content renders client-side and the fetched HTML arrives nearly empty. Many AI crawlers execute little or no JavaScript, so a page that paints its content after load looks blank to them. The real AI user-agents are worth checking against by name: OpenAI ships GPTBot, OAI-SearchBot, and ChatGPT-User; Anthropic ships ClaudeBot and Claude-User; Perplexity ships PerplexityBot and Perplexity-User. Google-Extended and CCBot are training controls, not direct evidence of search inclusion. A blanket "block all bots" rule can still catch search-specific crawlers unintentionally.
Fix. Server-render the content that carries meaning, so the substance is present in the initial HTML response rather than assembled in the browser. Audit robots.txt and your WAF against the user-agent list above and allow the crawlers you want. Then confirm from the other side: grep your access logs for those agent strings and check that the important URLs are actually being fetched with a 200. Retrievability is binary in a way the later stages are not. You are either fetchable or you are a hole in the index, and no amount of clever passage writing reaches a page the crawler skipped.
Stage two: Candidacy
Candidacy is where most brands actually lose, and it is the stage almost no advice names. A page can be perfectly crawlable and still never enter the small pool of documents the pipeline pulls for a given prompt. The model rewrites the user's question into several sub-queries, retrieves a candidate set for each from a search backend, and a reranker trims that pool to the handful of passages that fit the context window. Candidacy is whether you make that pool. Miss it and the model synthesizes its answer from documents that are not yours, and your content, however good, is simply not in the room.
The measure for this stage is Share of Retrieval: the fraction of relevant candidate sets your content appears in, upstream of any citation decision. It is the leading indicator, because it tells you whether you are even competing before the model decides who to quote. A page can hold flawless statistics and a clean quotable line and still sit at a Share of Retrieval near zero for its target prompts, in which case every downstream tactic is being applied to an audience the model never assembles.
Symptom. Your own pages get fetched, your logs prove the crawler visits, and yet you are missing from the sources on the exact commercial prompts you care about. Meanwhile the same handful of third-party pages, the comparison roundups and community threads, show up again and again. You appear for the literal name of your product and vanish the moment the query is phrased as a problem or a category.
Cause. Two mechanisms drive this. The first is the Retrieval Gap: for any query there is a large set of pages that could correctly answer it and a much smaller set the pipeline actually pulls, and you are on the wrong side of that distance. Often the reason is a prompt-space coverage gap, because the fan-out step turns one question into several sub-queries and a page tuned for a single head phrase misses the reformulations entirely, leaving you absent across much of the real space of prompts customers ask. The second mechanism is where retrieval physically happens. Practitioner datasets often show third-party pages contributing heavily to brand mentions, but the observed share varies with the category, prompt set, and engine. If that holds even roughly for your market, most of your candidacy is being decided on pages you do not own.
Fix. This is Citation Surface work, and it is unrelated to anything you do on your own domain. Your Citation Surface is the full set of third-party pages where AI can find and cite you: the roundups, the analyst notes, the integration docs, the forum answers that already retrieve well for your prompts. Map the prompt space first, enumerating the real sub-queries including the messy conversational ones, and find the sub-queries where you are absent from the candidate set. Then get accurately mentioned on the third-party pages that already win those sub-queries. Getting named, correctly and in a self-contained passage, on a page that already retrieves for "best tools for X" puts you into a candidate set you were locked out of. No FAQ block on your own site does that, because your own site is only a fraction of the surface where candidacy gets decided.
Stage three: Citation
Now the funnel crosses from retrieval into naming. Citation is the first stage where the content on the specific page genuinely matters, because by this point the model has your passage in its context window and is deciding whether to lean on it. You cleared retrieval. You are one of the few passages the reranker kept. The question is whether your chunk is the one the model chooses to build the sentence from.
Symptom. Your page does appear in the sources list, sometimes near the top, but the written answer draws its facts and phrasing from another source in the same list. You are cited as a footnote rather than used as the backbone of the answer. This is the retrieval-passed, naming-failed case from the master diagnostic, and it feels different from absence: you can see yourself in the panel, you are just not the one being quoted.
Cause. The passage the model retrieved from your page does not stand on its own, or a competing passage answers the sub-query more directly. The Princeton GEO team tested this at scale across roughly ten thousand queries and found that the content moves which do work are specific. Adding citations, adding statistics, adding direct quotations, and improving fluency each raised visibility, with the strongest interventions worth up to about a 40 percent lift on their metric and lower-ranked sources gaining the most. Keyword stuffing did nothing or slightly hurt. The effects were domain-dependent, with citations helping most on factual queries and an authoritative tone winning on debate and history. If your retrieved passage is vague where a competitor's is precise, the model reaches for the precise one.
Fix. Rewrite the passages that carry your key claims so they are specific and self-contained. A self-contained chunk that answers the question completely, without depending on distant context, is easier to use intact and gives the model something clean to quote. Back claims with a concrete figure and a source. Match the intervention to the domain: lead with data and citations on factual and commercial prompts, lean on clear authoritative explanation where the query is a matter of judgment. This is the layer the standard GEO checklist optimizes, and it is real work that pays, but only once retrieval is already solved.
Stage four: Attribution
Attribution is the last and subtlest stage. Here the model has used your content, the fact in the answer came from your page, and yet your name is not the one the reader sees. The answer states your claim as unsourced prose, or it collapses your fact into a general statement and hangs the visible citation on a different page that said something similar. Attribution is whether the visible answer credits you as the source rather than folding your contribution into the background.
Symptom. The answer contains a fact, a number, or a framing that is distinctly yours, phrased close to how you phrased it, but the citation chip points elsewhere or is missing. You provided the substance and a competitor or an aggregator got the visible credit. Over many prompts you notice your ideas propagating through AI answers while your brand name does not travel with them.
Cause. Attribution is not the same as retrieval, and it is not the same as citation either. The model synthesizes from several passages and then attaches citations by its own logic, and that logic favors the source that stated the claim most quotably and most self-containedly. When two pages carry the same fact, the one whose passage reads as a standalone, attributable unit tends to win the citation, and the other becomes uncredited background. Weak, hedged, or fragmentary phrasing on your side is easy for the model to absorb without a nametag.
Fix. Write the sentence you want quoted, as a sentence that can be lifted whole. Put the claim, the number, and your identifying context in the same self-contained passage, so the smallest chunk the model can extract still carries who said it. Where the fact is originally yours, publish it in a form specific enough that paraphrase loses information, which pushes the model toward quoting and crediting rather than absorbing. Strengthen entity clarity generally, while treating branded demand as an observational signal rather than a causal ranking factor. Attribution rewards the source that is hardest to restate without naming.
Running the diagnostic
Put the stages together and the flow is short. Look at the sources panel for a real prompt. If your category does not appear at all, you are losing at Retrievability or Candidacy, so check crawlability and logs first, and if those are clean the problem is Candidacy and the work is on your Citation Surface. If your category does appear but you are not the one appearing, check whether you show up anywhere in the panel: absent means Candidacy, present-but-not-quoted means Citation, and quoted-in-substance-but-not-credited means Attribution. Each verdict points at one fix and rules out the other three. The most expensive mistake in this whole field is running the Citation-stage playbook, the statistics and the quotable passages, on a brand that is losing at Candidacy and is never in the candidate set to begin with.
Two habits keep the diagnosis honest. Sample prompts repeatedly rather than once, because retrieval varies run to run and a single check is noise. And sample across the four engines separately, ChatGPT, Claude, Gemini, and Perplexity, because each runs a different backend and you can pass the funnel in one and fail it in another. This repeated cross-engine sampling is the measurement shareof.ai automates, turning the sources-panel check into a tracked Share of Retrieval number rather than a manual spot-check.
A note on measurement
Every figure here that comes from practitioner analysis rather than the Princeton benchmark is correlational and category-dependent. The third-party share, the passage-length range, the branded-search correlation: treat each as a hypothesis to test against your own prompt set, not a constant. What does not depend on the exact numbers is the shape of the funnel, because it follows from how retrieval-augmented systems are built. A prompt is fanned out, candidates are retrieved and reranked, passages are synthesized into an answer, and citations are attached last. Four stages, four distinct failures, four unrelated fixes. Diagnose the stage first and the fix is obvious. Skip the diagnosis and you will keep polishing passages the model was never going to read.
Sources and further reading: Aggarwal et al., "GEO: Generative Engine Optimization," arXiv:2311.09735 (2023), for the content-intervention results at Citation stage. For the pipeline mechanics behind Retrievability and Candidacy, any current description of retrieval-augmented generation covering query fan-out, candidate retrieval, cross-encoder reranking, and citation attachment. Verify the crawler user-agents against each provider's current documentation before editing robots.txt.
Primary sources and evidence
- GEO: Generative Engine Optimization (Aggarwal et al.) — the original paper and benchmark.
- Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks — the foundational RAG paper.
- Web search in the OpenAI API — current search-tool behavior and citations.
- OpenAI web crawlers — official user-agent controls for search and training.
- OpenAI web crawlers — GPTBot, OAI-SearchBot, and ChatGPT-User controls.
- Perplexity crawler documentation — PerplexityBot and Perplexity-User guidance.
- Google common crawlers — Googlebot and Google-Extended distinctions.
Continue through the system
- Share of Voice vs Share of Model vs Share of Retrieval
- Map Your Citation Surface: Every Page Where AI Can Find You
- The 15-Minute AI Visibility Audit You Can Run Right Now
Run the practical layer with the free AI visibility scan, explore the AI search library, or see the AI visibility benchmarks.