shareof.ai

What 240 Buying Answers Reveal About Cross-Model AI Visibility

A measured four-model study of 60 commercial buyer questions, 12 companies, 45 visibility gaps, and the cost of treating one assistant as the market.

The most common mistake in AI-visibility reporting is to treat one assistant as the market. We ran the same 60 commercial buying queries across four answer engines—ChatGPT, Claude, Gemini, and Perplexity—to measure how often a relevant company survived all the way into the visible recommendation.

The cohort covered 12 active YC-backed companies across voice AI, developer infrastructure, legal AI, industrial software, data infrastructure, and adjacent software categories. It produced 240 live answers on August 22, 2026. This paper reports the aggregate findings; the company-level study remains available in the YC AI Visibility Ledger.

The average company appeared in 52.1% of the measured answer opportunities, but the average query contained that company in only 2.08 of four models. Cross-model agreement was the exception, not the default.

What we measured

For each company we wrote five buyer questions spanning category discovery, constrained use cases, alternatives, integrations, and budget or compliance needs. Every query was run once against ChatGPT / GPT-4.1 mini, Perplexity / Sonar, Claude / Haiku 4.5, and Gemini / 2.5 Flash.

An observation counted as present when the normalized answer named the company or an unambiguous product/domain alias. We did not infer presence from a citation alone. The resulting unit of analysis is a company-query-model observation.

MeasureResult
Companies12
Buyer queries60
Models per query4
Live answer observations240
Mean company answer share52.1%
Mean models containing the company per query2.08
Open visibility gaps45
Queries absent from all four models16

The 45 open gaps are the 60 company-query records minus 15 cases where all four models contained the company. Sixteen queries returned zero presence across all four models; another 29 appeared in only one, two, or three.

Finding one: a blended score hides model disagreement

A 52.1% average can sound like a stable mid-market position. It is not. It is the average of many split decisions.

Only 15 of 60 query records contained the relevant company in all four answers. Sixteen contained it in none. The remaining 29 sat between those poles: seven appeared in one model, eight in two, and fourteen in three.

That distribution matters operationally. A team looking only at a blended share can miss a clean model-specific failure. The correct reporting shape is a matrix—query by model—followed by a weighted aggregate. The matrix explains the aggregate; the aggregate cannot explain the matrix.

This is why a single AI rank is a vanity metric. The useful question is not “What is our rank?” but “Across which buyer jobs and engines are we consistently eligible to be named?”

Finding two: category strength does not transfer automatically

The company shares ranged from 0% to 85% in this cohort. Mintlify and FlutterFlow reached 85%; Resend and Whatnot reached 80%. At the other end, one company recorded 0%, while several established companies sat between 10% and 35%.

This is not a company-quality ranking. It is evidence that public buyer-language coverage varies sharply even among well-known companies.

Models answer the question they are given. A company can have strong entity recognition and still fail a compliance, workflow, or “best for a growing team” query if the public evidence does not connect the company to that constraint. The missing bridge can be an owned decision page, a third-party comparison, a marketplace listing, documentation, or customer proof.

Use the AI Visibility Funnel to separate access, candidacy, citation, and attribution before deciding which fix applies.

Finding three: zero-of-four gaps are build briefs

The 16 zero-of-four records deserve different treatment from one-off misses. When every measured engine omits a relevant company for the same buying job, the result is less likely to be a sampling accident.

That still does not prove causation. It does identify a reproducible evidence gap worth investigating:

  1. Does a dedicated page answer the buyer question directly?
  2. Can crawlers render and reach that page?
  3. Do independent sources corroborate the claim?
  4. Does the company appear on the category and comparison pages already cited for that query?
  5. Do product names, domains, and aliases resolve to one entity?

A zero-of-four result is therefore not a prompt to publish generic volume. It is a prompt to map the citation surface around a specific decision.

Why identical prompts still do not create identical systems

The four engines do not expose one uniform retrieval layer. OpenAI's Responses API can attach web-search tools and URL citations. Google describes grounded Gemini responses with search queries, grounding chunks, and supports. Perplexity returns structured search results for search-enabled answers. Anthropic's citation system represents source references as structured citation blocks.

Those interfaces make visible-source measurement possible, but they do not make the underlying candidate sets identical. Different indexes, query rewrites, ranking stages, freshness policies, and response-generation systems can produce different shortlists from the same buyer question.

That is why this study treats each model response as a separate observation rather than four votes from interchangeable replicas.

How to use the result

Build a prompt set from real buyer jobs, not head keywords. Run it across the engines your customers use. Preserve full answers and visible sources. Report at least three layers:

  • Coverage: how many prompt-model cells contain the brand.
  • Agreement: how many models agree for each prompt.
  • Evidence: which owned and third-party pages appear beside the decision.

Then repeat the same cells over time. The next paper in this series shows why: even within one model, the exact answer is not stable.

Limitations

This is a purposive cohort of 12 YC-backed companies and 60 commercial prompts, not a random sample of all brands or user questions. Each cell was observed once on August 22, so the four-model results measure cross-sectional disagreement, not temporal variance. Presence detection used normalized names and aliases, but edge cases can remain. Visible citations reveal presented sources, not every document retrieved internally.

The correct reading is directional and specific: for this cohort and these buyer questions, cross-model agreement was limited enough that one-engine monitoring would have materially misrepresented visibility.

Reproducibility and data

The aggregate release is available as JSON. It contains sample sizes, model labels, result totals, and limitations without customer or prompt-level identifiers.

Primary sources and evidence

Continue through the system

Run the practical layer with the free AI visibility scan.