Someone asked an assistant which company to trust in your category this morning. It named three. If yours was not among them, no report landed on your desk, no impression was logged, and nothing in your analytics recorded the loss.
That is the measurement problem at the centre of this channel. Brand mentions in AI answers are consequential and largely invisible, which is why the published research on citation behaviour has become the closest thing to a scoreboard most teams have.
The numbers below come from studies covering hundreds of millions of citations and prompts. They are useful, and they are also noisier than the headlines suggest, because samples, dates, query types, and platform versions differ from study to study.
This guide covers what the aggregate citation data actually shows, how often brands get mentioned as opposed to cited, why the platforms disagree so sharply, which source types dominate the pool, and how to read any of these figures without drawing the wrong conclusion.
What The Aggregate Citation Data Says About Where Brand Answers Come From
The source hierarchy behind AI answers is more consistent than the individual percentages are, and it is not the hierarchy most marketing plans assume. The pattern holds across independent measurement approaches, which is the part worth acting on.
1. Earned media accounts for roughly 84 percent of citations
Muck Rack’s May 2026 edition of its Generative Pulse study analysed more than 25 million links cited by ChatGPT, Claude and Gemini across 17 industries, and found earned media accounting for approximately 84 percent of all citations.
Read the definition before quoting the figure. Muck Rack’s “earned media” is a superset covering journalism plus academic, government, encyclopedic and other third-party sources, not press coverage alone.
That distinction matters because the headline is often repeated as though it means media relations specifically. The narrower journalism component is a separate number, covered further down.
2. Paid and advertorial content sits near 0.3 percent
The same study places paid and advertorial content at roughly 0.3 percent of citations, which is a useful figure to have on hand the next time sponsored content is proposed as an AI visibility tactic.
The gap between 84 and 0.3 is large enough that the direction of the finding survives most methodological objections. Whether the precise figure is 0.3 or 1 percent changes nothing about the decision it should inform.
Treat it as a statement about relative leverage rather than an argument against advertising, which serves purposes this channel does not measure.
3. Journalism alone supplies about 27 percent of cited sources
Journalism accounts for approximately 27 percent of cited sources in the May 2026 edition. That is a smaller share than the earned media headline implies, and a much larger one than most content budgets reflect.
Across the three editions of that study going back to July 2025, journalism citations reportedly stayed within a band of roughly 25 to 27 percent.
The stability is the interesting part. A share that holds across model updates is more likely to reflect how these systems weight credibility than a temporary quirk of one release.
4. Journalism citations span more than 20,000 distinct outlets
The same research reports journalism citations spanning more than 20,000 separate outlets, which means credible trade and regional coverage counts rather than only national titles.
That reframes the pitch strategy for most mid-market brands. A well-read industry publication is a realistic target in a way that a national business desk usually is not.
One outlet stood out in that data, and it is covered in the platform section below, because it illustrates how concentrated a single engine’s preferences can be.
5. The earned media share has held between 82 and 89 percent
Across all three editions of the study, the earned media share reportedly ranged from about 82 to 89 percent. Consistency across time is more valuable than precision at any single point.
A separate synthesis of independent studies put the range at roughly 82 to 95 percent depending on methodology, with Fullintel and Golin figures cited at the higher end. I have not verified those two underlying studies directly, so treat the upper bound as indicative rather than settled.
What survives is the ordering: third-party editorial content dominates, owned content follows, paid content is negligible.
6. Owned content remains a minority input across every study
No published study I found places brand-owned content anywhere near the top of the citation pool. Controlled testing referenced in earlier analysis put the baseline citation rate for content on a brand’s own domain at around 8 percent, rising when the same material appeared on third-party outlets.
I could not locate the primary source behind that 8 percent figure, so it belongs in the “directionally useful, verify before presenting” category. The broader point does not depend on it.
Your site still matters as the authoritative record for pricing, specifications, and policy, which is the argument developed in this guide to structuring content for citation.
How Often Brands Actually Get Mentioned Versus Cited
Mention and citation are different events, and conflating them is the most common reporting error on this surface. A brand can be named in an answer without its domain appearing anywhere in the sources.
7. Mention and citation overlap can fall to 30 percent
Semrush’s 2026 AI Visibility Index analysed 126 million United States AI search prompts between January and April 2026, and reports that on Gemini the overlap between mentioned brands and cited domains can be as low as 30 percent.
That means a brand can be recommended by name while the evidence behind the recommendation sits entirely on other people’s pages. Both problems are real, and they need separate fixes.
Being mentioned without being cited is an authority gap on your own domain. Being cited without being mentioned usually means your brand name is missing from the page the model chose to quote.
8. ChatGPT reportedly mentions brands more often than it cites them
BrightEdge tracking cited in secondary analysis puts ChatGPT’s brand mentions at roughly 3.2 times its brand citations, at approximately 2.37 mentions per response against 0.73 citations.
I am reporting this one through an intermediary rather than from BrightEdge’s own publication, so the exact ratio should be verified before it goes into a board deck. The pattern it describes is consistent with the Semrush overlap finding above.
The practical reading is that mention-only visibility is common, and it converts differently from cited visibility because the reader never gets a link to follow.
9. Only 36 brands held visibility across all four platforms
The Semrush index found that only 36 global brands maintained top-100 visibility across ChatGPT, Gemini, Google AI Mode and AI Overviews in every month of the study period, a group including YouTube, Google, Reddit, Amazon and Apple.
That is a striking number, and it should reset expectations about cross-platform consistency. If 36 brands in the world manage it, a single visibility score for your business is unlikely to move all four engines together.
The corollary is that platform-level reporting is not a refinement. It is the minimum unit of useful measurement.
10. Category concentration ranges from 41 to 83 percent
The same study measured how concentrated visibility is by industry. In News and Media, the three most visible brands accounted for approximately 82.9 percent of category visibility, and in Consumer Electronics approximately 76.9 percent.
Finance and Industrial were far more distributed, with the top three brands at roughly 41.4 and 42.2 percent respectively. Less concentrated categories leave more room for a challenger to earn share.
Before committing budget, find out which kind of category you are in, because the same effort produces very different returns on either side of that line.
11. Around 15 percent of retrieved pages get cited
Analysis reported by AirOps and referenced in per-engine research examined 548,534 pages ChatGPT retrieved during answer generation and found only about 15 percent appeared as citations in final responses.
I have not read the AirOps study directly, so treat the specific figure with caution. The implication is worth carrying regardless: retrieval is a qualifying round, not the outcome.
It also explains why “we rank well and still are not cited” is such a common complaint. Being read and being quoted are separate selections.
12. Median time to first citation was under seven days
Profound reportedly tracked roughly 900 newly published marketing pages and found a median of approximately 6.81 days to first citation by ChatGPT or Claude, with around 90 percent of eventually cited pages earning their first citation within 37 days.
If that holds, a page still uncited after five or six weeks is unlikely to become citable through patience alone. It needs different structure, different distribution, or both.
This is the most operationally useful timing figure available, and also one I would want to confirm at source before treating 37 days as a hard diagnostic threshold.
Why Platform Differences Make A Single Visibility Score Meaningless
If four assistants describe your company four different ways, that is the expected result rather than an anomaly. The engines are reading substantially different internets.
13. Only about 11 percent of cited domains overlap
An analysis of 680 million citations found that roughly 11 percent of domains cited by ChatGPT are also cited by Perplexity, with the overlap falling below 1 percent on individual queries.
Attribution for this dataset is inconsistent across the coverage I read, with different write-ups crediting Profound and Averi for the same 680 million figure. I cannot resolve which is correct, so I am reporting the finding without asserting the source.
Independent studies of smaller samples have reported similar overlap figures, which is the reason to take the number seriously despite the attribution confusion.
14. Citation rates range from roughly 55 to 96 percent
Muck Rack reports ChatGPT citing sources in approximately 96 percent of responses, Gemini in 82 percent, and Claude in 55 percent. Claude is the most selective about whether to cite at all.
That variance alone breaks any single AI visibility metric. A brand invisible on the engine that cites least is losing a different kind of opportunity from one invisible where citations are near-universal.
Selectivity also implies scarcity. Fewer citing responses means fewer slots to compete for on that platform.
15. Sources per response range from three to fifteen
Muck Rack puts ChatGPT at around five citations per response, Gemini at eight, and Claude at thirteen when it cites. Semrush’s larger dataset puts ChatGPT at approximately 15 sources per response and Gemini at three.
Those two findings point in opposite directions on both platforms, which is a genuine conflict rather than a rounding difference. Different sampling windows, query mixes, and product surfaces are the likely explanation.
I am reporting both rather than choosing, because picking the more convenient figure is how misleading decks get built. The safe conclusion is that answers draw on a small number of sources and the count varies substantially by platform and period.
16. Each platform has a different single top domain
Muck Rack found the most-cited domain differing by platform: Wikipedia for ChatGPT, PubMed Central for Claude, and Reddit for Gemini.
If you operate in health, finance, or law, the Claude finding is the one to note, because institutional sourcing narrows the pool considerably in regulated categories.
That platform-by-platform divergence is the practical reason work on ChatGPT visibility needs separate diagnosis from work on the others.
17. One outlet dominated ChatGPT across 13 industries
Muck Rack reports Axios appearing in ChatGPT’s top three cited domains in 13 of the 17 industries studied, the only journalism outlet to reach a top-three position across any provider.
A single outlet with that reach across unrelated categories tells you something about how narrow editorial preference can get inside one engine.
It is not a pitching instruction. It is evidence that concentration exists and that identifying your category’s equivalent is a worthwhile audit step.
18. Reddit’s share moved roughly 50 points in six weeks
Semrush tracked more than 230,000 prompts over 13 weeks and found ChatGPT citing Reddit in close to 60 percent of prompt responses in early August 2025, collapsing to around 10 percent by mid-September.
Any single citation percentage you read is therefore a dated snapshot rather than a constant. The underlying share can move by an order of magnitude inside a quarter.
That volatility is the strongest argument for continuous measurement over one-off audits, a point the broader GEO statistics make repeatedly.
Which Source Types Dominate The Citation Pool
Underneath the platform differences, a handful of domain types absorb a disproportionate share of citations. Knowing which ones tells you where your name needs to appear.
19. Wikipedia supplies nearly half of ChatGPT’s top-ten share
The 680 million citation analysis puts Wikipedia at approximately 47.9 percent of ChatGPT’s top-ten source share, with Reddit at approximately 46.7 percent of Perplexity’s.
Note the metric carefully. This is share within each engine’s top ten domains, not share of all citations, and the two are frequently confused in coverage of this dataset.
Encyclopedic sourcing is also the least gameable asset available, since attempting an entry without meeting notability standards tends to fail publicly.
20. Reddit reached roughly 40 percent of LLM references
A Semrush study of more than 150,000 AI citations across 5,000 keywords found approximately 40.1 percent of references pointing to Reddit, ahead of Wikipedia at 26.3 percent and YouTube at 23.5 percent.
Read that alongside the collapse figure above rather than instead of it. Reddit’s weight is both large and unstable, partly because platform access disputes have repeatedly changed the terms.
Those citations point at specific threads rather than company profiles, which means participation is the asset and promotional posting is a liability.
What The Measurement Data Says About Reporting Maturity
The final group of figures is about the observers rather than the engines, and it explains why so many teams cannot answer basic questions about their own AI visibility.
- Approximately 45 percent of marketing leaders surveyed by Semrush cannot accurately measure brand visibility inside AI-generated answers, and only about 9 percent report having tools that track all relevant metrics across platforms.
- Among organisations that fully integrate SEO and AI visibility into one workflow, approximately 81 percent reported increased traffic or leads from AI platforms, against approximately 36 percent of those managing the two separately.
- Adobe data cited in the same release reports AI traffic to United States retail sites up approximately 1,324 percent between October 2024 and May 2026, with travel up approximately 2,215 percent.
- Cross-platform analysis reported by Superlines documented citation volume variance of up to 615 times for the same brand between platforms, though I have not verified that study directly and would treat the multiple as illustrative.
The percentage-change figures deserve particular caution. Growth of that magnitude usually starts from a very small base, and none of these numbers tell you what share of total traffic AI referrals now represent.
How To Read Any Of These Numbers Without Being Misled
Every figure above carries a date, a sample, and a definition, and most of the bad decisions made with this data come from dropping one of the three.
Sample size is not the same as sample quality. A study of 680 million citations built from synthetic prompts can be less representative of your buyers than a study of 20,000 real ones in your category.
Vendor incentives are worth naming plainly. Most of this research is published by companies selling visibility tooling or media relations software, which does not make it wrong but does make independent corroboration more valuable than usual.
Definitions move the answer more than methodology does. “Earned media” at 84 percent and “journalism” at 27 percent come from the same dataset and describe the same reality.
Recency matters more here than in almost any other marketing statistic. Given the Reddit collapse and the divergence between the Muck Rack and Semrush citation counts, I would not present any figure in this article as current without checking the source publication first.
Finally, none of these percentages tell you what an assistant says about your business specifically. They tell you where to look. The only figure that describes your position is one you generate by running a fixed prompt set repeatedly and logging whether you were named, cited, and described accurately as three separate columns.
What These Statistics Change About A Brand Programme
Taken together, the data supports a fairly short list of conclusions, and most of them are unglamorous.
Third-party credibility is the primary input, so coverage, community presence and accurate inclusion in the databases already cited for your category do more than publishing volume on your own domain.
Platform-level measurement is mandatory rather than advanced, because the overlap figures and the “Universal 36” finding both indicate that one score cannot represent four engines.
Structured identity data underpins everything, because a system has to know what your business is before it can describe it accurately, which is the argument set out in this explanation of what AEO means for reputation work.
Continuous operation beats campaign work, since an answer is rebuilt from a fresh source pool on demand and one cleanup does not stay clean. Categories with long research cycles need it running permanently, which is the case made in this piece on SaaS reputation management.
Wider commercial context for that investment sits in these ORM statistics and in this breakdown of how AI chatbots source information about your brand.
How Professional Reputation Management Applies These Numbers
Running monthly prompt logs across four platforms, auditing crawler access, correcting entity records, and pursuing credible coverage while also operating a business is more than most teams can absorb.
The work itself is not exotic. It is accurate structured data, first-party content written to be quoted, credible third-party coverage, and a documented process for correcting a source when a description drifts.
Nadernejad Media Inc. treats visibility and reputation as one connected programme, pairing search work with the monitoring that catches an inaccurate AI description early, an approach detailed further in this case for a professional reputation management partner.
Handled that way, the statistics stop being industry trivia and become a specification for what your information environment needs to contain.
Frequently Asked Questions
1. How often do AI chatbots mention brands by name?
There is no single reliable figure, and I would be sceptical of anyone offering one. What the data does show is that mention frequency varies enormously by category, with the top three brands holding approximately 83 percent of visibility in News and Media against approximately 41 percent in Finance, per Semrush’s 2026 index. Your own mention rate has to be measured directly.
2. Is being mentioned the same as being cited?
No, and treating them as one metric is the most common reporting error here. Semrush found the overlap between mentioned brands and cited domains falling as low as 30 percent on Gemini, meaning a brand can be recommended while none of the supporting sources are its own pages.
3. Which single statistic matters most?
The earned media share, at roughly 84 percent of citations against approximately 0.3 percent for paid content, because it has held across three editions of the same study and points to a clear allocation decision. That said, it is one vendor’s dataset with a broad definition of earned media, so I would treat it as a strong signal rather than a proven law.
4. Why do different studies report different citation counts per response?
Because they sampled different prompts at different times on different product surfaces. Muck Rack puts ChatGPT at around five citations per response while Semrush puts it at approximately 15, and I cannot reconcile those from the published summaries. The honest answer is that the count is small and unstable.
5. Where should I check these numbers myself?
Muck Rack publishes its full report and methodology through Generative Pulse, and Semrush publishes its AI Visibility Index directly. For the figures I flagged as reaching me through secondary aggregators, including the Yext, AirOps, Surfer, BrightEdge, and Profound findings, I would go to each company’s own publication before citing the specific decimals.











