Somebody has probably told you that adding JSON-LD to your site will get you cited by ChatGPT. That claim is sold constantly, and the controlled evidence does not support it.
What the evidence does support is narrower and, for reputation purposes, more useful. Structured data does not appear to buy citations. It buys accuracy, and accuracy is what most brands are actually losing when an AI system describes them wrongly.
Those two things get conflated because they look similar from the outside. A brand that is described incorrectly and a brand that is not described at all both have a visibility problem, but they need different fixes, and only one of them is a schema problem.
This guide separates them. It covers what structured data is doing when a machine reads your page, what the published studies actually found about schema and AI citation, why the reputation case survives that finding intact, which types earn their place, and which common implementations create liability rather than value.
What Structured Data Is Doing When A Machine Reads Your Page
Structured data is a set of machine-readable assertions about a page, usually implemented as JSON-LD in the document head, using vocabulary defined by schema.org. It states in explicit terms what the visible content only implies.
A page can say “we have served the UK grocery sector since 2011” in prose. Organization markup says the same thing as typed fields a parser cannot misread: a legal name, a founding date, a URL, a set of identity links.
That distinction matters because natural language is ambiguous and typed fields are not. The company name in your header could be a brand, a product, a person or a location as far as a disambiguation system is concerned.
Two separate mechanisms consume that markup, and most confusion in this area comes from treating them as one. The first is rich result eligibility in traditional search, which is documented, testable and increasingly narrow. The second is entity understanding, which is largely undocumented and where the reputation value sits.
There is also a technical limit worth knowing before you build a strategy on markup alone. A searchVIU experiment reportedly hid product data in eight schema formats and asked five AI systems to read the pages, finding that the assistants could not use JSON-LD data that was not also visible on the page during a direct fetch, because they read the rendered page as text.
I am reporting that experiment second-hand rather than from its original publication, so treat the specifics with caution. The implication matches Google’s own guidance closely enough to act on: markup that contradicts or exceeds the visible page is not doing the work you think it is.
What The Evidence Actually Says About Schema And AI Citation
This is the part most guides skip, and skipping it is how businesses end up spending a quarter on markup and reporting no change. The published record is more consistent than the marketing around it.
a. Google states plainly that no special markup is required
Google’s documentation on AI features says there is no special schema.org structured data you need to add to appear in AI Overviews or AI Mode, and that no new machine-readable files or markup are required.
Its AI optimization guide goes further, listing overfocusing on structured data as a mistake, while recommending you continue using it as part of an overall SEO strategy for rich result eligibility.
Read that carefully rather than as a dismissal. Google is saying structured data is not an AI retrieval lever, not that it is worthless, and the distinction is exactly the one this article is built on.
b. The correlation between schema and citation is real and large
Ahrefs analysis of approximately six million URLs reportedly found that pages cited in AI answers were close to three times as likely to carry JSON-LD as uncited pages.
The authors’ own reading of that result was co-occurrence rather than causation, on the reasonable grounds that well-maintained authoritative sites tend to have both schema and citations.
I have not read the Ahrefs publication directly, so the multiple should be verified before it goes into a proposal. The interpretive caution attached to it is what matters more than the figure.
c. A controlled test found essentially no citation lift
The most useful evidence is experimental rather than correlational. A test tracking 1,885 pages that added schema reportedly found AI Overview citations moving approximately negative 4.6 percent, AI Mode approximately positive 2.4 percent, and ChatGPT approximately positive 2.2 percent.
Two of those three are statistically indistinguishable from zero, and the one significant movement went in the unhelpful direction. The researchers’ own summary was that AI citations barely moved.
They also flagged a limitation worth repeating, which is that the test pages were already heavily cited. It does not establish what happens to a new, undiscovered page, so the finding is narrower than “schema never helps”.
d. Schema prevalence is indistinguishable between cited and uncited pages
A cross-platform study published on SSRN in February 2026 ran a within-Google diagnostic and found schema prevalence among AI-cited and non-cited pages at approximately 43.1 percent against 44.8 percent, which collapsed the apparent effect entirely.
Its initial pooled analysis had produced a significant negative association between schema and citation, which the author attributed to a methodological artefact rather than a real penalty.
This is the closest thing to a peer-review-track paper in the area, and it is a single-author study, so I would treat it as strong evidence rather than settled fact.
e. Rank position predicts citation far better than markup does
The same study reports position-one pages being cited in approximately 43 percent of queries where they appeared, declining to around 5 percent at position seven, implying each position reduces citation odds substantially.
If that gradient holds, the retrieval backend is doing most of the work that schema is credited with. A page that cannot be found is not going to be cited whatever its markup says.
That reframes the priority order. Accessibility and ranking come first, extractable content second, markup third, and reversing that order is the most expensive mistake available here.
f. One exception involves concrete attribute fields
The same research found a genuine exception. Pages implementing Product or Review schema with populated concrete attributes, including pricing, aggregate rating and specifications, were cited at approximately 61.7 percent against 41.6 percent for pages using generic types such as Article or Organization.
The plausible mechanism is not that the markup is privileged. It is that pages carrying specific, verifiable, extractable facts make better sources than pages that do not, and the schema is a proxy for that.
The instruction that follows is about content rather than markup. Publish the specific facts, then mark them up, in that order.
Why Structured Data Still Matters For Reputation Specifically
Nothing above weakens the reputation case, because the reputation case was never about citation frequency. It is about what a system says once it has decided to describe you.
1. Entity disambiguation is the underlying problem
A system has to know which thing your business is before it can describe it accurately. Where that identity is fragmented, descriptions drift, merge with other companies, or fall back on whatever fragment happens to be most retrievable.
Inconsistent naming is the most common version of this, and it is usually invisible from the inside. One legal name on the website, a slightly different one on LinkedIn, and a trading name on a directory can genuinely read as separate entities to a disambiguation system.
That mechanism is developed further in this explanation of entity SEO and its relationship to reputation work.
2. The sameAs property declares your identity graph
The sameAs property is where structured data does its clearest reputation work. It lets a page assert that this organisation is the same entity as a named Wikidata item, a LinkedIn company page, a Crunchbase record and any other canonical profile.
That is a declaration rather than an inference. Instead of leaving a system to guess whether three similar records describe one company, you state that they do, from the domain you control.
Wikidata deserves particular attention here because, unlike Wikipedia, it carries no notability threshold, which makes it accessible to a business of any size. The wider role of these signals is covered in this piece on brand entities.
3. Knowledge Panel accuracy depends on record consistency
A Knowledge Panel is not something you switch on. It appears when enough consistent, corroborated information exists across sources for a system to present facts about you confidently.
Structured data on your own domain is one input to that, and the only one you control outright. It is also the input most likely to be treated as authoritative on questions only you can answer.
Where the panel is wrong, the fix is upstream of the panel. Correct the markup, the profiles and the third-party records so the identity layer agrees with itself, then wait for recrawl.
4. Stale markup becomes a reputation liability
Structured data ages badly and silently. A departed executive still listed in Person markup, a superseded legal name, an employee count from three funding rounds ago, or a service you no longer offer will be repeated back to buyers with total confidence.
This is the failure mode I would check first in any audit, because it is common and because nothing on the visible page necessarily reveals it. Templated markup in particular tends to outlive the facts it asserts.
Audit annually at minimum, and immediately after any rebrand, acquisition, leadership change or pricing change, since those are exactly the facts an assistant will be asked about.
5. Being mentioned and being cited are different failures
Semrush’s 2026 AI Visibility Index, built on 126 million United States prompts, reports that on Gemini the overlap between mentioned brands and cited domains can fall as low as 30 percent.
A brand can therefore be recommended by name while every supporting source belongs to somebody else. Structured data does not fix that gap, because the gap is about third-party authority.
What it does address is the accuracy of the description attached to the mention, which is the part reputation programmes are usually being asked about. The sourcing side of the problem is covered in this breakdown of how AI chatbots source information about your brand.
6. Markup constrains what a system can get wrong
The most defensible framing is negative rather than positive. Structured data does not reliably increase how often you are described. It reduces the range of ways you can be described incorrectly.
For a business where an inaccurate AI answer costs deals, that is worth doing on its own terms, independently of any citation-rate promise.
It also compounds with everything else, since earned coverage, review distributions and encyclopedic records all become easier to reconcile when your own assertions are unambiguous, a point developed in these AEO statistics.
Which Schema Types Earn Their Place In A Reputation Programme
Most sites need fewer types than vendors recommend, implemented properly, rather than many implemented shallowly.
- Organization on the homepage or About page, with legal name, URL, logo, founding date, contact points and a complete set of sameAs links. This is the foundation, and without it the rest has nothing to attach to.
- Person for named executives and authors, linked to the organisation and to their own canonical profiles, since leadership queries are among the most common identity questions asked about a business.
- Article with a real author reference and an accurate dateModified, which is how freshness and accountability get stated rather than implied.
- Product or Service with concrete populated attributes, given that this is the one type combination the SSRN study associated with a materially higher citation rate.
- LocalBusiness where physical locations exist, with hours and addresses that match every external profile exactly rather than approximately.
- WebSite with a canonical identifier, which is a small piece of housekeeping that helps tie the rest of the graph to one domain.
Every one of these should describe something a visitor can see on the page. Markup that asserts facts the page does not display is the single most common implementation error, and it is the one Google’s guidance is most consistently explicit about.
What Not To Do With Schema In A Reputation Context
Several widely recommended tactics are either dead or actively risky, and reputation work is exactly the context where a manual action hurts most.
- Do not mark up your own reviews with Organization or LocalBusiness types. Google’s review snippet documentation states that where the entity being reviewed controls the reviews about itself, pages using LocalBusiness or any Organization type are ineligible for the star review feature. This has been policy since a 2019 change and remains among the most violated rules in the field.
- Do not aggregate third-party ratings into your markup. The same documentation instructs sites not to aggregate reviews or ratings from other websites, and embedding a widget counts as controlling the process.
- Do not add FAQPage markup for search visibility. Google added a deprecation notice in May 2026, and FAQ rich results no longer appear in Search, with reporting and API support withdrawn through June and August of that year. FAQPage remains a valid schema.org type, and Google has said unused structured data does not cause problems, so there is no need to rip it out in a panic.
- Do not treat markup as a substitute for visible answers. Given the searchVIU finding on direct fetches, an answer that exists only in JSON-LD may not be readable by the assistants you are trying to influence at all.
- Do not add types you cannot maintain. Every additional type is a future accuracy liability, and stale markup damages the trust it was implemented to build.
How To Audit Structured Data As A Reputation Asset
The audit is short, and it looks different from a rich results check, because you are testing accuracy rather than eligibility.
Start with what a machine currently believes. Pull the JSON-LD from your homepage, About page, leadership page and main service pages, and read the assertions as statements of fact rather than as code.
Then check every fact against reality. Legal name, founding date, leadership, locations, service list and contact details are the ones that go stale, and the ones buyers ask about.
Then check the identity links. Confirm that your sameAs set is complete, that each target resolves, and that the name on each destination matches the name in your markup rather than approximating it.
Then confirm the markup matches the visible page. Fields asserting things a reader cannot see are the ones most likely to be ignored, and in the review case the ones most likely to attract a problem.
Then view the raw HTML rather than the rendered page. Facts that only appear after JavaScript executes are facts several retrieval systems will not see, which makes them absent regardless of how well marked up they are.
Finally, log what the assistants actually say. Run a fixed set of identity questions across the major platforms, several times each, and record whether the description was factually correct as a separate column from whether you appeared at all. That distinction is the whole point of the exercise, and the surrounding method is set out in these GEO statistics.
How Structured Data Fits Into A Wider Reputation Programme
Schema is infrastructure, and infrastructure alone does not produce visibility. It determines what happens to the material other people publish about you.
The largest single input remains third-party credibility. Muck Rack’s May 2026 analysis of more than 25 million cited links found earned media accounting for roughly 84 percent of AI citations, against approximately 0.3 percent for paid and advertorial content.
Structured data does not compete with that finding. It makes the coverage easier to attach to the right entity, which is the difference between credible third-party material describing your company and credible third-party material describing something adjacent to it.
Review distributions, encyclopedic records and professional databases work the same way. Each is more useful when your own assertions are unambiguous, and each is a separate workstream from markup.
Categories with long research cycles need all of this running continuously rather than in campaigns, which is the case made in this piece on SaaS reputation management, with wider commercial context in these ORM statistics.
How Professional Reputation Management Handles The Technical Layer
Maintaining an accurate entity graph across a website, a Knowledge Panel, Wikidata, professional databases and a dozen profiles is unglamorous, ongoing work that most teams deprioritise until something breaks.
The work itself is not exotic. It is accurate structured data, first-party facts stated plainly in server-rendered text, credible third-party coverage, and a documented process for correcting a record when a description drifts.
Nadernejad Media Inc. treats visibility and reputation as one connected programme, pairing entity and structured data implementation with the monitoring that catches an inaccurate AI description early, an approach set out further in this case for a professional reputation management partner and in this explanation of what AEO means for reputation work.
Handled that way, structured data stops being a citation gamble and becomes what it should have been described as all along, which is the record a machine checks before it speaks about your business.
Frequently Asked Questions
1. Does schema markup get my brand cited by AI chatbots?
The best available evidence says no, not directly. Google states there is no special schema.org markup required for its AI features, and a controlled test of 1,885 pages found citation changes that were statistically indistinguishable from zero. The correlation between schema and citation is real, but the researchers who found it attributed it to co-occurrence rather than cause.
2. So is structured data a waste of time for reputation work?
No, and this is where the two questions get conflated. Citation frequency and description accuracy are different outcomes. Structured data addresses the second by making your identity unambiguous, which is what most brands are actually losing when an assistant describes them incorrectly.
3. Which single schema type should a business implement first?
Organization markup on the homepage or About page, with a complete and accurate sameAs set. Nothing else in the graph has anything to attach to without it, and identity questions are the most common thing an assistant is asked about a business.
4. Should I remove my FAQPage markup?
There is no urgency. FAQ rich results stopped appearing in Google Search in May 2026 and the supporting reporting was withdrawn afterwards, but FAQPage remains a valid schema.org type and Google has said unused structured data does not cause problems for Search. Decide based on maintenance cost, and stop adding it in pursuit of search visibility.
5. Can I use review schema to show my star rating?
Not with Organization or LocalBusiness types where you control the reviews, which is the situation almost every business is in on its own site. Google’s documentation is explicit that those pages are ineligible for the star review feature, and embedding a third-party widget does not change that. Product pages are treated differently.











