Your "AI Rank" Is a Vanity Metric
A single number pulled from one run of a stochastic system is a screenshot, not a measurement. Here is what to track instead.
A new category of tool will now sell you your "AI visibility score." You paste in a brand, it runs some prompts through ChatGPT and Perplexity, and it hands back a number. Maybe it is a rank, position 4 out of the brands mentioned. Maybe it is a 0-to-100 index blended across models. Either way you get one figure, a trend arrow, and the strong implication that watching that figure go up is the job. The figure feels like the old keyword rank did, familiar and clean, and that familiarity is exactly the problem, because the thing it claims to measure does not behave anything like a keyword rank.
Ask the same question of the same model twice and you can get two different answers, with two different sets of brands named in two different orders. No bug in the tool or glitch in the model explains this. The variation is the designed behavior of the system, and any measurement that ignores it is reporting noise as if it were signal. A point estimate drawn from a variable process tells you where the needle happened to land on one pull of the lever. Report it as "your rank" and you have dressed up a coin flip as a scoreboard.
The answer is a sample, not a lookup
When a classic search engine returns results, it is reading from a ranked index. Run the query again a second later and, barring a re-index, you get the same ten links in the same order. The output is a lookup. It is stable by construction, and that stability is what made a keyword rank a sensible thing to record: you were writing down a value that would still be true an hour from now.
A generative answer is produced differently. The model does not retrieve a stored answer, it generates one token by token, and at each step it samples from a probability distribution over possible next tokens. That sampling is governed by a temperature setting and related parameters the model host controls, and above temperature zero the process is deliberately random. Two runs of the same prompt walk two different paths through the distribution. The retrieval layer underneath adds its own variance on top: the prompt gets rewritten into several sub-queries by a fan-out step, each sub-query hits a search backend that may have re-crawled or reordered since the last run, and a reranker trims the pool to a few passages before the model sees any of them. Change the retrieved passages even slightly and the set of brands the model can mention changes with them.
So a single answer is one draw from a distribution of possible answers. Your brand appearing in it tells you that appearing was possible on that draw. Your brand missing from it tells you missing was possible too. Neither one, on its own, tells you the thing you actually want to know, which is how often you appear across the whole distribution of runs a real population of users will generate.
What a screenshot hides
Picture two brands competing on the same question. Brand A appears in 9 of every 10 runs. Brand B appears in 3 of every 10. On any given single run there is a real chance the screenshot catches Brand B present and Brand A absent, and a tool that samples once will, that day, report Brand B as winning. Run it again tomorrow and the ranking may flip. The underlying reality never moved. Only the sample did.
This is why a rank pulled from one run is not a measurement in any useful sense. A measurement comes with a claim about repeatability. If I tell you a table is 80 centimeters tall, I am implying you would get roughly 80 again if you measured it yourself. An "AI rank" carries no such implication. Measure again and you may well get a different answer, and the tool that gave you the first number never told you how wide that spread was. It sold you the mean of a distribution it never showed you, and often not even the mean, just a single sample standing in for one.
The honest unit here is a rate, not a position. Not "you rank third" but "you appeared in 62% of runs on this prompt against this model over the last N samples." A rate admits it came from counting. It invites the follow-up question that a rank suppresses: out of how many runs, and how much did it move.
The blended score hides the only reality that acts
The second trick in the vanity metric is the cross-model average. Because there are four assistants worth watching, ChatGPT, Claude, Gemini, and Perplexity, a tool that wants to give you one clean number has to blend them. So it runs each, scores your presence on each, and reports the average as your "AI visibility." The number looks comprehensive. It is the least informative form the data can take.
Retrieval and grounding differ so much between these systems that a blended score averages across incompatible worlds. ChatGPT uses provider-managed web retrieval. Perplexity operates its own crawler and a provider-managed retrieval stack. Gemini sits close to Google's own stack. Each runs its own crawler on its own freshness cycle, feeds its own reranker, and carries a different appetite for citing brand-owned versus third-party pages. A brand can appear in 70% of Perplexity runs and 10% of ChatGPT runs for reasons that trace to a single crawler being blocked in robots.txt. Average those and you get 40%, a number describing no system that exists and pointing at no action you can take.
The per-model split does not belong under the headline figure as a nice-to-have breakdown. It is the headline itself. A 40% blend tells you to "improve." A 70/10 split tells you a specific engine cannot see specific content, and it tells you which engine and roughly why. One of those is a diagnosis. The other is a mood.
Measure the distribution
The fix is to stop asking "where do I rank" and start asking "how often do I appear, and how much does it vary." That means treating each prompt-and-model pair as something you sample repeatedly, the way you would measure any noisy quantity, and reporting the shape of what comes back.
Concretely, for a given prompt and a given model, run it N times. Ten runs is a floor for a rough read, and more is better where the signal is close. Count the fraction of runs in which your brand appears. That fraction is your appearance rate, and it is the core number the whole approach is built on. Do the same for the competitors named alongside you, and you can rank by appearance rate rather than by whoever happened to surface on one pull.
Three numbers, reported per model, describe the situation honestly.
Share is the appearance rate itself, the fraction of runs you show up in. It answers the plain question of how present you are.
Variance is how much that rate wobbles, the spread across your N runs. A brand at 90% appearance with tight variance is genuinely locked in. A brand at 50% with wide swings is one re-crawl away from vanishing, and it needs a very different response than the raw share suggests. Variance is the number the single-run tools destroy by design, and it is often the most decision-relevant of the three.
Coverage is how much of your real prompt space you have this data for at all. A 70% appearance rate on the eight prompts you happened to test says little if your customers ask two hundred distinct questions. Coverage keeps the other two numbers honest by reminding you what slice of the actual demand they describe.
None of these collapses cleanly into a single scalar, and the resistance to being collapsed is the point. A trustworthy read of AI visibility is a small table, not a gauge. Down one side, your prompts grouped by intent. Across the top, the four models. In each cell, an appearance rate with its spread, backed by a run count you can see. The moment someone flattens that table to one number, ask what got averaged away, because the answer is usually everything you could have acted on.
Share of retrieval and share of voice, the honest pair
Even a well-sampled appearance rate measures only the visible surface, the finished answer. Two things happen before that answer exists, and separating them turns a symptom into a cause.
The final answer is where share of voice lives, the fraction of runs, or of the mention-space within them, that go to your brand. It is the outcome you ultimately care about. By itself it cannot tell you why a cell is dark, because there are two very different ways to be absent from an answer. The model retrieved your page and chose not to cite it, or the model never retrieved your page at all. Those have opposite fixes, and share of voice alone cannot tell them apart.
The upstream number is Share of Retrieval, the fraction of candidate sets your content lands in before the reranker and the model ever get to choose. This is the leading indicator. If you are absent from the candidate sets, no amount of prompt-side or answer-side tuning will help, because the model cannot cite what it never retrieved. If you are present in the candidate sets but absent from the answers, the problem sits at candidacy and citation, and that is a content and framing problem rather than a crawl problem. Watching both numbers together locates the failure on the AI Visibility Funnel, the four stages of Retrievability, Candidacy, Citation, and Attribution. Watching only the final answer leaves you tuning the last stage while the leak sits three stages up.
The per-model story sharpens here too. One crawler allowed and another blocked can, on its own, put a brand near zero on one engine. So the same brand holds strong Share of Retrieval on the engine that can reach it and sits invisible on the engine that cannot, a split a blended answer-level score erases completely. Track retrieval per model and it surfaces at once, and the fix might be one line in a robots.txt file rather than a quarter of content work.
What a trustworthy dashboard looks like
Strip away the styling and a defensible AI-visibility dashboard commits to a few things a vanity tool refuses to.
It shows a run count on every figure, because a rate without a denominator is a rumor. It never blends models into a single headline, keeping ChatGPT, Claude, Gemini, and Perplexity on separate columns. It reports variance next to every appearance rate, so you can tell a stable 60% from a coin-flip 60%. It distinguishes retrieval from citation, so a dark cell comes with a cause and not just a color. And it states its coverage plainly, telling you how much of your prompt space these numbers actually describe rather than implying the eight prompts it tested are the whole world.
Measuring this well is genuinely hard, and it is worth being honest about why. Running N samples across four models over a real prompt inventory is a lot of runs, and the cost is real. Retrieval visibility is partial, since not every system exposes its sub-queries or its candidate sets, so Share of Retrieval is sometimes inferred rather than read directly. The prompt space itself drifts as your category moves, so coverage is never finished. These are the reasons the honest version is harder to sell than a single glowing number. They are not reasons to prefer the number. A hard measurement that reflects reality beats an easy one that does not, and the easy one here is actively misleading, because it converts a wide, per-model distribution into a confident little scalar and hands it to you as fact.
This is the measurement shareof.ai is built to run, sampling each prompt many times per model and reporting share, variance, and coverage rather than one blended rank. The framework holds without any tool, though. You can do a rough version yourself this week: take five prompts your customers actually ask, run each one ten times through two models, and write down how often your brand appears and how much that number jumps around. The spread you see on that tiny sample is the thing every single-number "AI rank" was quietly hiding from you.
A note on measurement
Treat any AI-visibility figure the way you would treat a poll. Ask for the sample size, ask about the margin, and ask who got averaged together to produce the topline. A rank with no run count behind it is a screenshot of one moment in a system that will look different the next time you check. The useful quantity is a distribution: how often you appear, how much that rate moves, and across how much of the real question space you have looked, reported for each model on its own terms. It is more work to produce and less satisfying to glance at, and it happens to be true.
Primary sources and evidence
- Web search in the OpenAI API — current search-tool behavior and citations.
- OpenAI web crawlers — official user-agent controls for search and training.
Continue through the system
- Stop Tracking Keywords. Start Tracking Prompts.
- A/B Testing Content for AI Citation: A Methodology That Survives the Noise
- Share of Voice vs Share of Model vs Share of Retrieval
Run the practical layer with the free AI visibility scan, explore the AI search library, or see the AI visibility benchmarks.