August 20, 2026

How AI Chatbots Source Information About Your Brand

How AI Chatbots Source Information About Your Brand

Somebody asked ChatGPT what your company does this morning. It answered in about eighty words, with confidence, and possibly without reading a single page you wrote.

That is the part most teams get wrong. They assume a chatbot describing their business has visited their website, when in practice the answer was assembled from journalism, encyclopedia entries, forum threads, and review pages, with the brand’s own site contributing a fraction of the material.

Understanding where the material actually comes from is the whole job. You cannot edit an AI answer, but you can change almost every input feeding it, and the inputs are more identifiable than they look.

This guide covers the two mechanisms by which a chatbot knows anything about you, which source types dominate in practice, why each platform reaches a different conclusion, which crawlers control your eligibility, and how to audit and correct what is currently being said.

The Two Completely Different Ways A Chatbot Learns Anything About Your Business

Every AI answer about your brand comes from one of two places, and confusing them leads to wasted effort.

The first is training. During training, the model absorbs an enormous body of text and ends up with statistical associations about your company: what category it belongs to, who competes with it, what reputation it carries. Nothing you publish today changes that until a future model is trained, and you cannot edit it retroactively.

The second is retrieval. When a chatbot searches the web mid-conversation, it fetches live pages and builds the answer from what it just read. This is the mechanism you can influence this quarter, and it is now the default behaviour on brand and product questions.

Google describes the underlying technique plainly in its own AI optimization guide, explaining that its generative features use retrieval-augmented generation over the same Search index that powers ordinary results.

The practical consequence is that there is no separate AI index to petition for inclusion. If your content is not retrievable for the relevant intent, it cannot be cited, no matter how well written it is.

Two symptoms tell you which mechanism produced a bad answer. If the chatbot cites sources and still describes you incorrectly, that is a retrieval problem, and the sources are the fix. If it confidently describes you with no citations at all, using details that are years out of date, you are looking at training memory, and the only remedy is flooding the current web with better, more consistent material.

Most brands need to work on both, but only one of them responds inside a quarter.

Which Sources AI Chatbots Actually Draw On When Describing Your Brand

The source hierarchy is remarkably consistent across studies, and it is not the hierarchy most marketing plans assume.

How AI Chatbots Source Information About Your Brand

1. Why Earned Media And Journalism Dominate AI Citations About Brands

Earned coverage is the single largest input. Muck Rack’s Generative Pulse analysis of more than 25 million links cited by ChatGPT, Claude, and Gemini found earned media accounting for roughly 84 percent of all AI citations, with journalism alone at about 27 percent.

The same research reports paid and advertorial content sitting near 0.3 percent of citations, which is worth remembering the next time sponsored content is proposed as an AI visibility tactic.

Note the definition before quoting the headline. Muck Rack’s “earned media” is a superset covering journalism plus academic, government, encyclopedic, and third-party sources, not press coverage alone.

Journalism citations in that study spanned more than 20,000 distinct outlets, which means credible trade coverage counts rather than only national titles.

2. How Wikipedia And Wikidata Anchor What A Chatbot Believes You Are

Encyclopedic sources do disproportionate work on identity questions. Muck Rack found Wikipedia to be the top cited domain for ChatGPT specifically, and earlier platform analysis placed it as ChatGPT’s leading source at around 7.8 percent of citations.

Wikidata matters alongside it, because structured properties there propagate into knowledge graphs and entity databases that several systems consult.

This is also the least gameable asset available. Attempting an entry without meeting notability standards tends to fail publicly and creates a problem of its own, so pursue it only when independent coverage genuinely justifies one.

3. Why Reddit Threads Carry So Much Weight In Chatbot Answers

Community discussion is heavily represented, and unevenly so between platforms. Muck Rack identified Reddit as Gemini’s leading cited domain, while separate synthesis placed Reddit at 20 to 24 percent of Perplexity’s citations.

Those citations point at specific threads rather than profiles, which means participation is the asset and a polished company page is not.

Reddit’s share has also swung sharply with platform access disputes, so treat any single percentage as a dated snapshot rather than a constant.

Promotional posting is a liability here rather than a tactic. A hostile thread naming your business becomes durable material across both search and AI answers.

4. How Video Transcripts Feed Descriptions Of Your Products

Video is a text source as far as these systems are concerned. One analysis of 36 million Google AI Overviews placed YouTube at roughly 23.3 percent of citations.

That reframes transcript quality as search infrastructure rather than an accessibility checkbox. Chaptered structure, accurate captions, and descriptive titles give retrieval something usable.

It also means an unanswered critical video about your product is feeding descriptions of you, which is why the mechanics of ranking videos belong in a sourcing audit.

5. Why Review Platforms Decide How Your Quality Gets Described

When a chatbot characterises your service quality, it is compressing review distributions into a clause. Recency and vividness carry more weight in that compression than volume does.

The result is that one detailed complaint can end up representing a pattern your own data does not support, simply because it was the most quotable text available.

Sustained review generation and consistent review responses are the only durable inputs. You cannot edit the sentence, only the corpus it summarises.

6. How Professional Profiles And Databases Establish Your Basic Facts

Structured business records answer the boring questions: size, sector, location, leadership, funding. LinkedIn company pages, professional databases, and industry registries are frequently the source for those details.

They are also where stale information persists longest. An outdated employee count, a departed executive still listed as current, or a legacy company name will be repeated back to buyers with total confidence.

Audit these annually at minimum, and immediately after any rebrand, acquisition, or leadership change.

7. Why Your Own Website Contributes Less Than You Expect

Owned media is a minority input across every study of this question. Controlled testing reported by one analysis found a baseline citation rate of about 8 percent for content sitting on a brand’s own site, rising substantially when the same content appeared on third-party outlets.

That does not make your site optional. It is the primary source for pricing, specifications, policies, and anything factual that only you can state authoritatively.

It does mean the site should be treated as the reference document rather than the whole strategy, written to be quoted rather than to persuade.

8. How Institutional And Academic Sources Shape Regulated Category Answers

Institutional sources carry outsized weight in sensitive categories. Muck Rack found Claude leaning most heavily on PubMed Central, and Perplexity has been reported to carry the highest share of academic domains among major engines.

If you operate in health, finance, law, or anything where accuracy has consequences, expect the answer pool to be narrower and more institutional than in consumer categories.

Original research, methodology pages, and genuine expert commentary are therefore stronger investments in those sectors than volume publishing.

Why Each Major Chatbot Reaches A Different Answer About The Same Brand

How AI Chatbots Source Information About Your Brand

If you have ever asked four assistants the same question about your company and received four different characterisations, that is expected rather than anomalous.

The platforms are reading different internets. An analysis of 680 million citations found only around 11 percent of domains cited by both ChatGPT and Perplexity, a result echoed by independent studies of tens of thousands of responses.

Citation behaviour differs too. Muck Rack reported Gemini citing sources in around 82 percent of responses, averaging eight links, while Claude cited in about 55 percent of responses but averaged thirteen sources when it did.

Editorial taste differs as well. The same research characterises ChatGPT as favouring large mainstream publications while Claude draws more heavily on niche and trade outlets whose most-cited titles carry substantially smaller audiences.

Recency windows differ again. Both ChatGPT and Claude were found to prefer recent material, with a meaningful share of citations coming from content published within the previous week, though the models front-load that recency differently.

The strategic conclusion is that “AI visibility” is not one channel. A pitch strategy tuned for one model will not transfer cleanly, which is why work on ChatGPT visibility needs separate diagnosis from work on the others.

The shared foundation is still worth building first, because clean structure, accurate entity data, and credible coverage improve your position on all of them simultaneously.

Which Crawlers Decide Whether Your Site Is Even Available To Be Quoted

Before any of the above matters, a crawler has to be able to reach you, and most vendors run separate bots for training and for retrieval.

The distinction is the most common self-inflicted wound in this channel. OpenAI’s GPTBot collects training material while OAI-SearchBot builds the index behind ChatGPT’s search answers, and blocking the second removes you from those answers even though the first is what most publishers meant to block.

Anthropic operates a comparable split, with ClaudeBot handling training, Claude-SearchBot handling search indexing, and Claude-User fetching pages at a user’s direct request. Each token requires its own directive, so disallowing ClaudeBot does nothing to the other two.

Perplexity documents two agents in its own crawler documentation: PerplexityBot for indexing and Perplexity-User for live fetches. Its help centre adds that a disallowed page may still surface as a domain and headline with a brief factual summary, meaning a blocked site can be described without ever being quoted.

Google’s arrangement is different again, since Google-Extended is a training control token rather than a crawler. Blocking it opts you out of Gemini training without affecting Googlebot or your Search rankings.

The workable pattern for most brands is to allow the retrieval agents and decide separately about the training ones. Allowing search bots while disallowing training bots is an explicitly supported configuration on both OpenAI and Anthropic.

Two caveats matter more than the syntax. Robots.txt is a request rather than access control, so non-compliant crawlers will ignore it, and a firewall is the only real defence. And permission in robots.txt means nothing if a CDN or bot-management rule refuses the request before it reaches your server.

Check server logs for these user agents over thirty days. Zero hits from a retrieval bot on a site that ranks well in search is a blocking problem, not a content problem, and no amount of rewriting will fix it.

How To Audit What AI Chatbots Currently Say About Your Business

How AI Chatbots Source Information About Your Brand

You cannot manage a description you have not recorded. This audit takes an afternoon to build and roughly two hours a month to maintain.

1. How To Build A Fixed Prompt Set That Reflects Real Buyer Questions

Write down twenty questions a real buyer would ask before choosing in your category, phrased the way they would type them rather than the way your marketing team writes headlines.

Include identity questions, comparison questions, trust questions, and pricing questions, since each pulls from a different part of the source hierarchy.

Then freeze the wording. Changing the prompt changes the answer, and you lose the ability to say whether anything improved.

2. Why You Must Run Every Prompt Several Times Per Check

Generation is non-deterministic. Research comparing repeated runs of identical prompts found meaningful variation in which sources appeared, even on the most reproducible engines.

Run each prompt three to five times in one sitting and record the proportion of runs that named you. That proportion is the measurement.

Always start a fresh conversation, because follow-up questions behave differently from opening ones and testing a follow-up measures something a buyer never experiences.

3. How To Log Whether You Were Named, Cited And Described Accurately

Record four separate outcomes per run: whether an answer appeared, whether your brand was named in the text, whether your domain was cited, and whether the description was factually correct.

These are genuinely different problems. Being named without being cited is an authority gap, while being cited without being named means your brand name is missing from the page the model chose to quote.

Conflating those four columns is the most common reporting error on this surface, and it leads teams to fix the wrong thing.

4. Which Server Log Checks Confirm Crawlers Can Actually Reach You

Filter thirty days of logs for OAI-SearchBot, Claude-SearchBot, PerplexityBot, and their user-triggered counterparts, then compare against your robots.txt and CDN rules.

Look for outright absence and for error responses, since a bot receiving a challenge page or a 403 is functionally blocked even when robots.txt permits it.

Also view the raw HTML of your pricing and services pages. Facts that only appear after JavaScript runs are facts these systems frequently never see.

5. Why Tracking Cited Domains And Competitors Matters As Much As Your Own Score

Record which domains get cited alongside you, not just whether you appeared. That column is what lets you distinguish a platform-level reweighting from a change your own team caused.

Log which competitors were named instead of you as well, since that is usually the line that gets budget approved.

Over three months, the pattern tells you exactly where to spend: if the same third-party pages keep appearing, those pages are your roadmap.

How To Correct An Inaccurate Answer At Its Source Rather Than Its Output

Once the audit shows a wrong description, the instinct is to look for a correction form. There isn’t a meaningful one, and that is the point.

Work in order of leverage instead. Start with your own site, since it is the only place you control absolutely: state the correct fact plainly, in server-rendered text, near the top of the relevant page, using the phrasing buyers actually use.

Then correct the structured record. Fix your Knowledge Panel, Wikidata properties, business profiles, and professional database entries, since those feed the identity layer that several systems consult before anything else.

Then address the specific source doing the damage. If one outdated article, review, or thread is supplying the wrong claim, that source is the project rather than the answer, and correction, updated coverage, or displacement all work on the input.

Then publish the counter-material. Because these systems favour recent content and earned coverage, a steady stream of accurate third-party mentions does more to overwrite a stale claim than a single definitive page on your own domain.

Finally, expect latency and variance. Retrieval-based answers update within days of the underlying sources changing, while training-derived claims persist until a future model absorbs the corrected web. The wider framing for that work sits in this explanation of what AEO means for reputation management.

What A Realistic Ninety Day Plan For AI Brand Sourcing Looks Like

How AI Chatbots Source Information About Your Brand

Weeks one and two are diagnostic. Build the prompt set, run every question several times across the major assistants, and log named, cited, and accurate separately for each.

Weeks two and three are technical. Audit crawler access in server logs, correct robots.txt for the training and retrieval split, confirm no CDN rule is refusing retrieval bots, and move key facts into server-rendered HTML.

Weeks three through six are entity work. Correct the Knowledge Panel, structured data, business profiles, and database records so the identity layer agrees with itself across every source.

Weeks six through ten are off-site. Given how heavily earned media dominates citations, pursue credible coverage and accurate inclusion in the third-party sources already cited for your category, and address any single hostile source doing the describing.

Weeks ten through thirteen are re-measurement. Re-run the full prompt set, compare against baseline, and separate your own progress from platform change by watching which other domains moved with you.

Then repeat monthly, because every answer is rebuilt from a fresh source pool on demand. One cleanup does not stay clean; a point the broader GEO statistics make repeatedly.

How Professional Reputation Management Keeps Your Brand Described Correctly By Chatbots

How AI Chatbots Source Information About Your Brand

Running monthly prompt logs, crawler audits, entity corrections, earned coverage, and source-level fixes while also operating a business is more than most teams can absorb.

The work itself is not exotic. It is accurate structured data, first-party answers written to be quoted, credible third-party coverage, and a documented process for correcting a source when a description drifts.

Categories with long research cycles need it running continuously rather than in campaigns, which is the case made in this piece on reputation management for SaaS companies. Wider context sits in these ORM statistics for 2026 and these AEO statistics on how answer engines source information.

Nadernejad Media Inc. treats visibility and reputation as one connected programme, pairing search work with the monitoring that catches an inaccurate AI description early. The same approach runs through its AI solutions and its business services.

Handled that way, what a chatbot says about your brand stops being a rumour you overhear and becomes an answer you supplied the material for.

Frequently Asked Questions

1. Do AI chatbots read my website when someone asks about my company?

Sometimes, but usually as one input among many. Studies consistently find brand-owned content making up a minority of citations, with earned media, encyclopedic entries, community threads, and review platforms carrying more weight. Your site remains the authoritative source for pricing, specifications, and policies, so it needs to be readable, but it is rarely the whole answer.

2. Why does ChatGPT describe my brand differently from Claude or Gemini?

Because they read different sources and cite at different rates. Research found only about 11 percent domain overlap between ChatGPT and Perplexity, and Muck Rack’s analysis found each platform favouring a different top domain, with ChatGPT leaning on Wikipedia, Claude on PubMed Central, and Gemini on Reddit. Different source pools produce different conclusions about the same company.

3. Can I stop AI chatbots from talking about my business?

Not realistically. Blocking crawlers removes your pages from citation without removing your brand from the conversation, since the model can still describe you from training memory and from third-party sources you do not control. Perplexity states explicitly that a blocked page may still surface as a domain and headline with a summary, which is usually the worst outcome available.

4. Does blocking GPTBot remove my brand from ChatGPT answers?

No. GPTBot is the training crawler, and OAI-SearchBot is what powers appearance in ChatGPT’s search answers, and the two are controlled independently. Anthropic operates the same split between ClaudeBot and Claude-SearchBot, so a business can decline model training while remaining citable in live answers.

5. How long does it take to change what an AI chatbot says about my brand?

Retrieval-based answers can change within days of the underlying sources changing, which is why technical and source-level fixes show up fastest. Claims coming from training memory persist until a future model is trained, so the only remedy there is consistent, accurate, recent material across the web. Expect fast movement on cited answers and slow compounding on uncited ones.

6. What is the single highest-return action for improving how chatbots describe my brand?

Earning credible third-party coverage, based on the consistency of the finding across independent studies that earned media accounts for the large majority of AI citations while paid content accounts for almost none. Before that, however, confirm that retrieval crawlers can actually reach your site, because no amount of coverage compensates for being technically unreadable.

Facebook
Twitter
LinkedIn
Pinterest