You monitor whether AI talks about you. That is the right worry, in the wrong order.
The entire GEO industry is built around one question: how to get cited. Tools count mentions. Reports display share of voice. Vendors sell visibility. The assumption underneath all of it is that a citation is a win, and that the job is therefore to collect more of them.
The problem is not being absent from the answers. The problem is being present, and wrong.
This is not a theoretical risk. It is a measured property of the system, documented at scales that leave very little room for interpretation, and what follows gives you the rates, the failure modes, and the four moves that actually shrink your exposure.
The widest measurement available
Work published on May 27, 2026 under the acronym CITETRACE is the broadest measurement of the phenomenon that currently exists.
Its scale: 11,200 real queries drawn from twenty eight online communities, 112,000 answers produced by ten models from five different providers, and 761,495 evaluable citation pairs, each scored against a three-dimension grid validated by domain experts.
Provenance marking: arXiv preprint, not peer reviewed, but with a published protocol and a scale nothing in the commercial corpus comes close to.
Three findings.
30.6 percent of citations misrepresent their source. The link works, the source exists, it is often even relevant, but what the answer makes it say is not what it says.
27.1 percent come from sources unsuited to the domain. A medical question answered by citing a discussion thread, a legal question answered by citing a company blog.
And up to 96 percent of users meet at least one structurally misleading citation in a given answer.
The fourth finding is the most useful strategically, and it went almost unnoticed: between 88 and 96 percent of the quality variance is explained by the provider, not by the model. Citation reliability is a platform architecture decision, not a property of the underlying model. You cannot change it. You can find out which platform is failing you worst.
None of this is new, and it was measured before GEO existed
The first serious work on the subject predates the founding GEO paper. Published in 2023 and presented at a peer-reviewed conference, which makes it the strongest source in this file, it uses human judgment to evaluate four generative engines of that era.
Its results: 51.5 percent of generated sentences are fully supported by their citations, and 74.5 percent of citations actually support the sentence they are attached to.
One sentence in two. In 2023, before anyone was selling GEO.
Take the measure of what that means for the market. An entire industry organized itself around the goal of being cited by systems whose statements, at birth, were unsupported by the source displayed beside them half of the time.
What the journalism audits find
Before we get to what this costs a brand, two studies run by press institutions add a dimension the technical work does not capture: the nature of the errors.
The Tow Center for Digital Journalism at Columbia tested eight tools in March 2025 across twenty publishers with contrasting access policies, ten articles each, for 200 queries verified by hand.
Headline result: more than 60 percent incorrect citations, with a considerable spread between tools, from 37 percent for the best to 94 percent for the worst. And one counterintuitive observation, which the report is careful to document: the paid versions are more confident in their errors than the free ones. They are wrong with fewer qualifiers.
Four findings in that report are underused, and every publisher should read them twice.
Several tools appear to bypass robots.txt instructions. They fabricate URLs that do not exist. They cite syndicated or copied versions rather than the original, which sends the credit somewhere other than to you. And above all, licensing agreements guarantee nothing about citation accuracy. Paying to be a partner source does not protect you from distortion.
The second study, run jointly by the European Broadcasting Union and the BBC in October 2025, covers 3,000 answers evaluated in fourteen languages across four assistants.
45 percent of the answers contain at least one significant problem. 81 percent contain a problem of some kind. Sourcing is the leading cause, present in 31 percent of answers in the form of missing, misleading or false attributions. One assistant sits well below the others, with 76 percent of its answers flagged.
What complicates the picture
An honest article also goes looking for the lower numbers. Here is one.
An audit published in April 2026 covers 11,943 claim-source pairs in a multimodal context, evaluated by three independent automatic judges that agree with each other 87.7 percent of the time, themselves validated against human annotations. Provenance marking: arXiv preprint, not peer reviewed, protocol published.
Its rate of unsupported claims runs from 3.7 percent to 18.7 percent depending on the domain. Well below everything above.
Its dominant failure mode is the instructive part. It is not head-on contradiction, it is unverifiable specificity. The system injects one precise detail pulled from its internal memory while citing a source that does not contain that detail. The error is invisible on the page. It becomes visible only if you open the source, which roughly one reader in a hundred does.
The rates, side by side
Five measurements, five scopes, five methods. Here we line them up.
| Study | Scale | What is measured | Rate |
|---|---|---|---|
| Stanford, 2023, peer reviewed | 4 engines, human judgment | Sentences fully supported | only 51.5 percent |
| Tow Center, March 2025 | 8 tools, 200 queries checked by hand | Incorrect citations | more than 60 percent |
| EBU / BBC, Oct. 2025 | 3,000 answers, 14 languages | Answers with a significant problem | 45 percent |
| CITETRACE, May 2026 | 761,495 citation pairs | Citations misrepresenting their source | 30.6 percent |
| Multimodal audit, April 2026 | 11,943 pairs | Unsupported claims | 3.7 percent to 18.7 percent |
The rates move with what is measured and how it is measured. Not one of them is anywhere near zero. The most favorable number, on the narrowest scope, still leaves one claim in five without support in the worst domain tested.
What this actually produces for a brand
The errors that concern you do not look like spectacular hallucinations. They are ordinary, plausible, and expensive.
An obsolete price, pulled from an archived page, presented as current. A product line you discontinued, still described as available. A customer review from 2019, lifted from a forum, cited as established fact about your service today. A technical specification attributed to your product when it belongs to a competitor, or the reverse. A syndicated copy of your own article cited in place of the original, which hands the credibility to somebody else.
None of these errors triggers an alert. None of them shows up in a share-of-voice dashboard, which counts mentions without reading them.
What you do tomorrow morning
You cannot fix a model. You can shrink the error surface, and four actions have a plausible effect.
Publish canonical fact pages, dated. Pricing, scope of the offer, specifications, each carrying a visible update date. Consistency and a recent timestamp are among the very few factors the citation studies report a consistent effect for.
Clean up the third-party sources before your own site. Directories, comparison pages, partner listings still carrying a version of your offer you retired two years ago. Those are what feed the distortion, and they weigh more than your pages do.
Check it yourself, several times over. Ask your ten category questions on every platform, vary the wording, and read what is said about you. Not the mention count: the content. A single measurement is worth nothing.
Identify the platform that serves you worst. Since 88 to 96 percent of quality depends on the provider, the gap between platforms will be wider than the gap between anything you do. That tells you where a correction request is worth filing.
The industry sells you a share of voice. What you should be buying is an accuracy check. A false mention costs more than an absence, and nobody bills it as a risk.
Sources
- Seo, Y., Jeong, W., Kim, E., Jang, H. & Lee, D. (2026). Verified Misguidance: Measuring Structural Citation Failures in Search-Augmented LLMs, arXiv:2605.28565
- Liu, N. F., Zhang, T. & Liang, P. (2023). Evaluating Verifiability in Generative Search Engines, arXiv:2304.09848
- Jaźwińska, K. & Chandrasekar, A. (2025). AI Search Has a Citation Problem, Tow Center for Digital Journalism, Columbia
- European Broadcasting Union & BBC (2025). AI assistants misrepresent news content 45% of the time
- Samieyan Sahneh, E. & Aiello, L. M. (2026). Auditing the Reliability of Multimodal Generative Search, arXiv:2604.00944
- Vishwakarma, R., Kumar, S. & Jamidar, R. (2026). What Gets Cited: Competitive GEO in AI Answer Engines, arXiv:2605.25517
<strong>LaFactory</strong> measures AI visibility with a published protocol: repeated measurements, paraphrases, control group. No guaranteed placement, ever. Contact us to scope an audit.