You read that GEO works. So you started doing it. You added statistics to your pages, slipped in expert quotations, restructured your paragraphs so a model could lift them cleanly. You may even have seen something move.
Your competitor read the same article. He did the same thing. So did the fifteen other players in your category.
The problem is not whether GEO works. The problem is that it only works while you are the only one practicing it. That is a measurable property, and it has just been measured.
This article sets out what the GEO gain becomes under competition, what cancels out, and the only two levers that hold. Read it before you sign an agency quote.
What the study measures: three experiments, three commercial models
The reference work is signed Xi Chu and Yupeng Hou, published on 16 June 2026 under the title Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems. Provenance marking, straight away: this is an arXiv preprint, not a peer-reviewed article. Its methodology is published and reproducible, which is not true of the agency reports circulating on the subject, but it is not a journal publication.
The protocol comes down to three experiments run on skincare products, with a robustness test on a different category. The choice of skincare is not incidental. It is what economists call a credence good, a product whose buyer cannot verify the quality even after using it. Exactly the terrain where a third-party recommendation weighs the most.
Three commercial models were queried: GPT-4o-mini, Claude Sonnet and Gemini 3 Flash. Not a laboratory model, not a simulation. The same systems that answer your customers.
The conditional monopoly: at equal specifications, the known brand always wins
The first result is brutal, and it will surprise nobody except the people selling GEO to small brands.
At identical composition, identical price and rigorously identical specifications, the known brands are recommended one hundred percent of the time. The authors call this a conditional monopoly. It is not a tendency and it is not a marginal preference: it is an absolute ceiling for as long as nothing else separates the products.
That result converges with another measurement, taken at a larger scale. A June 2026 analysis covering more than one hundred thousand answers and more than one hundred brands establishes three tiers of awareness from the very first query: global brands appear in 73 percent of answers, mid-tier brands in 44 percent, niche brands in 11 percent. Provenance worth flagging: the data comes from Ranqo, a vendor of GEO tracking software. The author himself concedes in his conclusion that causality remains to be established.
Thirty points of gap per tier. A small player does not start at zero, he starts at eleven percent, and that eleven percent is structural.
The collapse: from +0.802 to +0.007
Here is the part the market does not tell you.
The researchers measured the gain a brand obtains by adopting a GEO strategy, then ran the measurement again assuming that every brand adopts the same strategy. That is the real scenario, the one any sector drifts toward once three agencies are selling the same method to the same companies.
The individual gain falls from +0.802 to +0.007.
It does not decline. It does not halve. It disappears, to within one hundredth.
The structure of the problem is a classic social dilemma. Not taking part means the certainty of never being recommended while the others are. Taking part means a gain of roughly nothing, but a null gain beats a handicap. So everyone takes part, everyone pays, and nobody pulls ahead. The only player who wins for certain in that configuration is the one selling the service.
Two details of the study deserve quoting, because they show where the real levers sit.
The first: an advantage of 0.1 star on the average rating is enough to break the brand monopoly. One tenth of a star. Not a product page rewrite, not markup, not a configuration file: a genuinely better customer rating.
The second is more troubling: authority language is worth +0.17 point on the score the model assigns, and the effect persists when the clinical evidence invoked is fabricated. The model does not distinguish a real study from an invented one. That says something about the reliability of these systems, and it also says that part of the GEO market is in the process of discovering a lever it would be dishonest to exploit.
Game theory reaches the same place: citation wars
A second body of work, independent of the first, arrives at the same conclusion by another route.
Chen Xu, Zitian Guo and Chenyan Xiong, at Carnegie Mellon, published on 11 August 2026 a model of GEO competition as a repeated Stackelberg game under partial monitoring, tested on three benchmark datasets. Again: preprint, published methodology, no peer review.
Their finding is that competition degenerates into what they call citation wars. Each player rewrites his content to maximize his probability of being cited, which pushes the next one to rewrite further, and so on. The system converges on an inert steady state where nobody advances any more.
The collateral damage is the point that matters to you: these rewrites degrade document quality and introduce unsupported claims. In other words, the race for citation mechanically produces less reliable content. You damage your own page for a gain that cancels out.
That mechanism lines up with a finding in the critical survey of July 2026, which reviewed forty-five studies across the window from November 2023 to July 2026: citation-oriented rewrites can degrade retrieval of the document. You optimize the last stage of the pipeline by breaking an earlier one. The page becomes more citable and less findable.
What cancels under competition, what survives
The table below summarizes the documented levers, ranked by how well they resist generalized adoption.
| Lever | Effect in isolation | Effect once everyone applies it | Source |
|---|---|---|---|
| Pure formatting, structure, markup, lists | Near zero from the start | Zero | Sprinklr, 252,000 trials, 18 factors, May 2026 |
| Adding statistics and quotations | Positive in a fixed context | Cancels out | Princeton 2023 and July 2026 survey |
| Full generic GEO strategy | +0.802 | +0.007 | Chu and Hou, June 2026 |
| Authority language | +0.17 rating point | Holds, but exploits a flaw | Chu and Hou, June 2026 |
| Customer rating higher by 0.1 star | Breaks the brand monopoly | Holds | Chu and Hou, June 2026 |
| Genuine topical relevance | Primary determinant | Holds | Sprinklr, May 2026 |
| Pre-existing brand awareness | 73 percent against 11 percent | Holds by construction | Ranqo, June 2026 |
The reading is simple. Everything that amounts to surface manipulation cancels out under competition. Everything that amounts to a real improvement in the product or in the relevance of the content survives, because those are not optimizations: they are differences.
There is an irony in that table. The only two levers that survive are precisely the ones no agency can sell you as a monthly retainer.
The counter-evidence: MIT does not say the same thing
An honest article cites what contradicts it. Here is the best available opposing argument, and it is solid.
A team at MIT, Puneet Bagga, Vivek Farias, Tamar Korkotashvili, Tianyi Peng and Yuhang Wu, published a testbed called E-GEO: 13,747 realistic product queries crossed with ten Amazon listings, five generative engines, seven automated rewriters, fifteen manual heuristics, plus a red-teaming phase.
Their conclusion runs against everything above. A meta-optimization of the prompt surfaces a stable, domain-agnostic pattern, which suggests that a genuinely effective GEO strategy exists. And under a simple defense, the observed gains reflect a real improvement in the content, not a manipulation.
What that means: GEO would not be noise, but a correctly posed optimization problem, with a solution.
What it does not say: the experiment is an offline testbed on Amazon listings. It measures no longitudinal effect in production, no competitive dynamic, and it does not test what happens when the ten sellers in the category apply the same meta-optimization. That is precisely the variable Chu and Hou measure, and it is where the gain collapses.
So the two results are not contradictory. They describe two moments: before generalized adoption, and after.
What you do tomorrow morning
Stop reasoning in terms of optimization and start reasoning in terms of gap.
Faced with any GEO action anyone proposes to you, ask a single question: does this action produce a gap my competitors cannot replicate in three weeks? If the answer is no, you are buying a ticket in a race whose finish line sits at +0.007.
Three practical consequences.
Work on the customer rating before the product page. One tenth of a star weighs more than a full rewrite, and nobody can copy it off you.
Invest in genuine topical relevance rather than in formatting. The 252,000 trials in the Sprinklr study leave no room: pure formatting edits have almost no effect, and yet they fill half the checklists on sale.
And be wary of any engagement whose promised outcome is a visibility percentage. The number exists, it is measurable, and it trends toward zero the moment the sector moves.
GEO is not a scam. It is a zero-sum game sold to you as a competitive advantage.
Sources
- Chu, X. & Hou, Y. (2026). Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems, arXiv:2606.17443
- Xu, C., Guo, Z. & Xiong, C. (2026). Mechanism Design for Generative Engines: From Exploitation toward Win-Win Outcomes, arXiv:2608.11390
- Vishwakarma, R., Kumar, S. & Jamidar, R. (2026). What Gets Cited: Competitive GEO in AI Answer Engines, arXiv:2605.25517
- Martinez, O. (2026). Optimizing Visibility in Generative Engines: A Critical Survey of Generative Engine Optimization (2023-2026), arXiv:2607.14035
- Bagga, P., Farias, V., Korkotashvili, T., Peng, T. & Wu, Y. (2025). E-GEO: A Testbed for Generative Engine Optimization in E-Commerce, arXiv:2511.20867
- Kumar, P. (2026). Generative Engine Optimization at Scale: Measuring Brand Visibility Across AI Search Engines, arXiv:2606.20065
- Aggarwal, P. et al. (2023). GEO: Generative Engine Optimization, arXiv:2311.09735
<strong>LaFactory</strong> measures AI visibility with a published protocol: repeated measurements, paraphrases, control group. No guaranteed placement, ever. Contact us to scope an audit.