shareof.ai

The Third-Party Citation Tax: 96.5% of Visible Sources Were Off-Site

A 13,593-citation production study showing where AI answer evidence lives, how ownership was classified, and why one connector was excluded.

When an AI answer cites evidence about a brand, the source is usually not the brand's own website. In Shareof.ai's production observation set, 13,116 of 13,593 usable visible citations—96.5%—pointed to a third-party domain.

That does not mean owned content is irrelevant. It means the measurable evidence surface is larger than the site a marketing team controls.

Across 1,189 cited answers and 2,449 distinct source domains, only 477 visible citations resolved to the tracked brand's own domain or a subdomain.

What counted

The study joins four production records: completed model observations, visible citations attached to those observations, normalized source domains, and the tracked project domain.

A citation was owned when its normalized hostname equaled the project domain or ended in that domain as a subdomain. Everything else was third-party.

Gemini was excluded from the ownership calculation because its connector returned 1,879 citations as the same Google proxy hostname, vertexaisearch.cloud.google.com, rather than exposing a consistently resolved publisher domain. Treating those proxy URLs as publishers would produce a precise but false result.

ModelVisible citationsOwnedThird-partyThird-party share
Perplexity6,7112826,42995.8%
Google AI2,835752,76097.4%
ChatGPT2,430852,34596.5%
Claude9551993698.0%
Copilot6621664697.6%
Combined13,59347713,11696.5%

The result is a surface map, not a causal claim

A visible citation tells us that the answer interface presented a source. It does not prove that the source caused the brand mention, nor does it reveal every document retrieved or used internally.

The 96.5% figure therefore answers a narrower and still useful question: where does visible evidence live when monitored answer systems show their work?

In this cohort, it overwhelmingly lived outside the tracked brand's domain.

That is the empirical basis for treating a brand's citation surface as a network: review sites, editorial comparisons, documentation ecosystems, marketplaces, community discussions, news releases, partner pages, and customer evidence.

Why owned pages still matter

Owned content performs at least four jobs even when it is not the visible citation:

  • it establishes canonical product facts and terminology;
  • it supplies documentation that third parties can verify;
  • it gives crawlers a stable entity and page structure;
  • it creates the original claims, datasets, and definitions other sites can cite.

The wrong conclusion is “stop investing in your site.” The right conclusion is “do not expect your site to be the whole distribution system.”

A first-party page can also be retrieved while a third-party source earns the visible attribution. Citation and retrieval are not the same stage. The AI Visibility Funnel keeps that distinction explicit.

The practical meaning of the tax

We call it a tax because access to a buying answer often depends on evidence controlled by someone else.

That dependency creates three kinds of work:

  1. Representation: ensure influential third-party pages describe the brand accurately.
  2. Eligibility: appear in the comparisons, directories, ecosystems, and discussions that answer the actual buyer question.
  3. Originality: publish evidence worth citing so independent pages have a reason to mention the brand.

The tax is not paid by buying links or mass-producing guest posts. It is paid by earning accurate placement on pages that already participate in the answer graph.

Which third-party sources appeared

Source concentration differed by model and project. Frequently observed domains included Reddit, YouTube, LinkedIn, TechRadar, G2, Capterra, Trustpilot, marketplace pages, vendor documentation, and category-specific publishers.

That list should not become a universal outreach checklist. A source matters only when it repeatedly appears for a relevant prompt family. The correct unit is source × query × model, not domain popularity in isolation.

This is also why reverse-engineering co-citation is more useful than collecting generic “high-authority” targets. A page that repeatedly appears beside the competitor in your exact category has direct observed relevance.

A defensible workflow

Start with the prompts that matter commercially. For each prompt-model cell:

  1. preserve the complete visible source list;
  2. normalize redirects, tracking parameters, and subdomains;
  3. classify sources as owned, competitor-owned, independent editorial, community, marketplace, documentation, or other;
  4. count recurring page-level and domain-level appearances;
  5. connect each source to the brands named in the same answer;
  6. prioritize corrections or contributions where the evidence is incomplete.

Do not collapse the analysis to a single percentage. The 96.5% aggregate establishes the scale of third-party dependence; the page-level graph tells you what to do.

Connector quality is part of research quality

The Gemini exclusion is not a footnote to hide. It illustrates a recurring measurement problem: citation interfaces can expose publisher URLs, redirectors, proxy URLs, or presentation-only links.

Google's own grounding documentation shows that a grounding chunk can contain a vertexaisearch.cloud.google.com URI while the title names the underlying publisher. Domain studies must resolve that indirection or exclude the record. Otherwise one proxy domain can dominate the result.

Perplexity has likewise moved from a simple citations field to structured search results with URLs and dates. Schema changes like these can break historical comparisons unless raw responses and parsing versions are retained.

Limitations

The dataset covers 10 projects, 69 or fewer distinct prompts per model, and an August 1–23 observation window. Project and model coverage is uneven. Domain matching can misclassify brands that use unrelated documentation or commerce domains. The result measures visible citations, not hidden retrieval. Citation counts are not unique-answer counts; one answer can show multiple sources.

The combined estimate excludes Gemini proxy URLs and should not be generalized to every industry or interface. Its value is as an observed baseline and a method other teams can reproduce.

Reproducibility and data

Download the aggregate counts, exclusions, and limits in the study JSON.

Primary sources and evidence

Continue through the system

See your own source graph with the free AI visibility scan.