You see the same sentence everywhere. GEO supposedly lifts a piece of content by 30 to 40 percent in AI-generated answers. The number turns up in sales pages, LinkedIn posts, commercial proposals. It is presented as the proof that the discipline works.
The number is true. The study exists, it is serious, it is published.
And it does not say what it is made to say.
The problem is not the statistic. The problem is the experimental condition attached to it, the one nobody copies across. This article restores that condition, shows what the number actually proves, and what it does not.
The study, and what it actually measured
The founding GEO paper dates from late 2023, indexed as arXiv:2311.09735, produced by a team led out of Princeton. It carries the name that christened the entire discipline: Generative Engine Optimization.
Its protocol: take content, apply different modifications to it, then measure the visibility obtained in the answers of a generative engine. The winning modifications are well known, and they are reasonable. Add quantified statistics. Insert attributed quotations. Cite your sources explicitly.
The measured gain: 30 to 40 percent additional visibility in the generated answer.
So far, nothing to object to. That is a clean result, obtained inside a documented experimental frame.
The condition nobody copies across
Here is the sentence that changes everything, and it comes from the critical survey published on 15 July 2026, which reviewed forty-five studies across the window from November 2023 to July 2026.
The gains in the founding paper are valid within their experimental frame, but they are conditional on the source already being present in a fixed context. They establish neither organic discoverability nor a durable effect on traffic.
Read that condition again. The content tested was already inside the model’s context. The experiment does not measure how to get in. It measures what happens once you are in.
The analogy is a shop window. You measure the effect of a new window display on people who have already walked into the shop, then conclude that the new display brings more people onto the street. That is not the same measurement. It is not even the same question.
And it is precisely the question a company buying GEO is asking. It is not asking how to show up better in an answer where it already appears. It is asking how to appear at all.
The pipeline, and the stage where the number actually acts
The July 2026 survey makes a useful contribution to understanding why the confusion is so easy: it decomposes the production of a generative answer into successive stages, instead of treating it as a single ranking task.
The chain, in order: search activation, crawling and indexing, document retrieval, reranking, allocation inside the context window, citation, prominence, factual absorption, faithfulness, user behavior.
The 30 to 40 percent gain acts on the late stages: citation and prominence. It touches neither activation, nor crawling, nor retrieval, nor reranking. The first four stages, the ones that determine whether your page stands any chance of being seen at all, sit outside the scope of the experiment.
That is also what explains a counter-intuitive result in the same survey: citation-oriented rewrites can degrade retrieval. You optimize stage six while damaging stage three. The page becomes more citable and less findable. The net result can be negative, and nothing in the 40 percent figure will tell you so.
What the number proves, what it does not
| Claim | Status |
|---|---|
| Adding statistics and quotations increases the citation of content already present in the context | Established, within the original experimental frame |
| These modifications improve organic discoverability | Not established. Outside the scope of the study |
| These modifications produce a durable effect on traffic | Not established. No longitudinal measurement |
| The effect transfers from one platform to another | Not established. The survey reports that generic heuristics transfer poorly |
| The effect holds when competitors do the same thing | Contradicted. The individual gain collapses under competition |
| The 40 percent figure is the right order of magnitude to promise a client | No. It is a laboratory result sold as a field result |
The survey’s overall verdict, across the forty-five studies examined, fits in one sentence: content that has already been retrieved can have its citation causally modified, but no technique examined demonstrates a stable, longitudinal, cross-platform causal effect on organic discoverability or on downstream behavior.
Three years of literature. Not one technique.
What later work has measured since
Two subsequent results sharpen the picture, and they point the same way.
The first comes from a May 2026 study covering 252,000 trials, six models, a factorial design testing eighteen content factors, with brand anonymization and counterbalancing of source order. Provenance worth flagging: the work comes from a team at Sprinklr, a marketing software vendor, which is a declarable conflict of interest even though the protocol is published.
Its two main determinants of citation: topical relevance and position in the list. And its most awkward result for the market: pure formatting edits have almost no effect. Structure, bullet lists, chunking, markup. Nothing.
The second is more fundamental still. A paper published in the journal TACL in 2024, under the title Lost in the Middle, establishes that a model’s performance on a piece of information follows a U-shaped curve depending on where that information sits in the context. Well exploited at the beginning and the end, poorly exploited in the middle. Provenance marking: this one is a peer-reviewed journal publication, the highest level of evidence in the whole corpus.
Direct consequence: position in the context is one of the two most reproducible levers, and you have no control over it whatsoever. The engine decides the order. Part of what is sold to you as GEO visibility is a placement artifact.
The honesty owed to the original study
You have to be fair to the founding paper, because it did nothing dishonest.
It defined a protocol, measured an effect, published its experimental conditions. The result is real inside its own scope. The authors never claimed to establish an effect on organic traffic, and the survey that criticizes them explicitly acknowledges the validity of the result within its frame.
The distortion is not in the study. It is in the commercial retelling, which kept the number and threw away the condition.
That mechanism is ordinary, and it does not only concern GEO. A laboratory result circulates, sheds its methodological caveats with every copy, and lands in a commercial proposal as a promise. The only protection is to go back to the source, systematically, including when the number suits you.
What you do tomorrow morning
Take the last GEO proposal you received and look for the 30 or 40 percent figure. If it is in there, put three questions to whoever wrote it.
In the original study, was the tested content already present in the model’s context? The answer is yes, and it is written in the paper.
Does the study measure an effect on organic traffic? No, and the survey that examined forty-five studies confirms it.
What measurement does your agency propose to establish that same effect for me? That is the only question that counts, and it is the one that sorts the field.
A number without its experimental condition is not proof. It is a sales argument wearing the appearance of proof, which is worse.
Sources
- Aggarwal, P. et al. (2023). GEO: Generative Engine Optimization, arXiv:2311.09735
- Martinez, O. (2026). Optimizing Visibility in Generative Engines: A Critical Survey of Generative Engine Optimization (2023-2026), arXiv:2607.14035
- Vishwakarma, R., Kumar, S. & Jamidar, R. (2026). What Gets Cited: Competitive GEO in AI Answer Engines, arXiv:2605.25517
- Liu, N. F. et al. (2024). Lost in the Middle: How Language Models Use Long Contexts, Transactions of the ACL, vol. 12
- Chu, X. & Hou, Y. (2026). Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems, arXiv:2606.17443
<strong>LaFactory</strong> measures AI visibility with a published protocol: repeated measurements, paraphrases, control group. No guaranteed placement, ever. Contact us to scope an audit.