Sixteen percent of the sources AI cites are themselves written by AI

by | Aug 26, 2026 | GEO

You hesitate to publish generated content. You have read that Google penalizes it. That readers spot it. That it is filler. So you pay humans to write your pages, which costs more and takes longer, and you accept the delay because you believe it protects the quality of what carries your name.

Meanwhile, a measurable share of what generative engines cite as a source is itself generated content.

The problem is not moral, or not only. The problem is structural: the system mechanically rewards the production that degrades it, and a perverse incentive does not correct itself.

This article gives you the measurement, the loop that produces it, what limits the damage, and four checks to run on your own pages.

The measurement

An audit published on May 22, 2026 by two researchers at Northwestern University covers four consumer generative engines, queried with 712 real human queries across three domains: politics, health and the environment.

Provenance marking: arXiv preprint, protocol published, not peer reviewed. The use of real queries rather than synthetic ones is a genuine strength of the work, and it is rare enough to be worth saying.

The result: roughly 16 percent of cited sources are themselves generated by AI. On all four engines.

One source in six. In domains where accuracy matters most.

The loop, and why it closes

The mechanism is simple, and it assumes bad faith from nobody.

Generated content is fast and cheap to produce, so there is a great deal of it. It tends to be built to a regular pattern, with self-contained claims and extractable phrasing, which is exactly the shape retrieval systems pick up most easily. So it gets cited. The citation proves to whoever produced it that the method works, so they produce more of it. And the corpus the engines draw from holds a rising proportion of that kind of text.

The important point is that it works. The producer is not wrong: the page is cited. That is precisely the problem. An incentive that rewards collectively damaging behavior is not corrected by individual virtue.

A model published in August 2026 by a team at Carnegie Mellon describes the same phenomenon through game theory. Competition between optimizers converges on what the authors call citation wars, where successive rewrites degrade document quality and introduce unsupported claims, until the whole thing settles into an inert steady state. Provenance marking: arXiv preprint, theoretical model, no field data.

That is not a hypothesis about the future. The 16 percent is a present-tense measurement of it.

What it produces downstream

Three documented consequences, and they stack.

Distortion. A May 2026 study covering 761,495 citation pairs, ten models and five providers, measures that 30.6 percent of citations misrepresent their source. Provenance marking: arXiv preprint, not peer reviewed. If a growing share of those sources is itself unverified generated text, then the distortion is being applied to material that was already fragile.

Substitution of the original. The Tow Center for Digital Journalism at Columbia University, in a March 2025 audit covering eight tools and two hundred manually verified queries, found that engines frequently cite syndicated or copied versions rather than the original. In practice: your article is cited, but through a copy, and the credit lands somewhere else. The same report notes that engines fabricate URLs, and that licensing agreements guarantee no correct citation whatsoever. Provenance marking: audit by a university research center, manual verification, published outside peer review.

Convergence among readers. A study published in the journal PNAS Nexus in September 2025, built on seven randomized experiments with thousands of participants, measures that people using language models produce reasoning that is shorter, less factual and more similar to one another than the reasoning produced with classic web links. Provenance marking: peer-reviewed journal publication, which makes it the strongest level of evidence in this whole file.

A corpus that looks more and more alike, summarized for readers who increasingly think alike.

The loop in one table

Step What happens Available measurement
Production Generated content is abundant and shaped for extraction Not applicable
Retrieval That regularity matches what engines pick up Formatting does not help citation, but structure helps retrieval
Citation A share of cited sources is generated content ~16 percent, 4 engines, 712 real queries
Rendering The citation frequently distorts what it cites 30.6 percent across 761,495 pairs
Attribution The original is sometimes replaced by a copy Observed across 8 tools, 200 queries
Reception The reader almost never checks ~1 percent click rate on displayed links
Reinforcement The producer sees that the method works Not applicable

What stops this from being a catastrophe

Two things temper the picture, and they matter.

Sixteen percent is not one hundred percent. Five sources in six are not generated content. The measurement describes real contamination, not collapse, and the difference between those two words is the difference between a problem you manage and a problem you flee.

Platforms have a direct interest in filtering. An MIT testbed built on product listings establishes that under a simple defense, optimization gains reflect a real improvement in the content rather than manipulation. And a team at Fudan University measures that production engines neutralize more than 99.78 percent of traditional search manipulation attacks, with the retrieval stage acting as the filter. Provenance marking on both: arXiv preprints, offline testbeds, not peer reviewed. Nothing suggests that the filtering of low-quality generated content escapes that same defensive dynamic.

The open question is timing. Filters arrive after the behavior they correct, exactly as they did in 2011 for artificial links.

What you do tomorrow morning

Four actions. Three of them are elementary hygiene that almost nobody performs.

Check who is being cited in your place. Take your category queries, list the URLs cited, and open every one of them. If you find your own content republished on a third-party site, you are feeding a competitor with your own work. The canonical tag and copy monitoring become useful again, for a reason that has nothing to do with classic search rankings.

Date everything. A visible, consistent update date is one of the very few signals the citation studies report a consistent effect for, and it is also the thing that separates an original from a late copy. Two reasons, one action.

Make your claims verifiable. Name your sources inside the text, with links that actually resolve. This is not an optimization technique. It is what separates you from the material this article describes.

And if you do publish generated content, review it for what it asserts, not for how it reads. The risk is not that a reader detects the machine. The risk is that an invented claim travels under your name, inside a system where 30.6 percent of citations already distort what they cite.

The question about generated content is no longer whether engines penalize it. It is whether you agree to add your share to the material you depend on.

Sources


<strong>LaFactory</strong> measures AI visibility with a published protocol: repeated measurements, paraphrases, control group. No guaranteed placement, ever. Contact us to scope an audit.

Cart