You get sold checklists. Twelve points, fifteen points, sometimes twenty. Structure your answers. Ship an llms.txt file. Deploy schema markup. Write self-contained paragraphs. Every point is presented on the same plane, with the same authority, as though they were all worth the same.
They are not worth the same. Some are measured and reproducible. Others are plausible but never established. Others still have been tested at scale and produce nothing.
The problem is not that these lists are wrong. The problem is that they blend the three categories together without ever saying which is which.
This article does the sorting, leaning on the only piece of work that examined the entire literature rather than copying the neighbor.
The reference document, and what it examines
On 15 July 2026 a critical survey indexed as arXiv:2607.14035 was published, reviewing forty-five studies selected across a window running from November 2023 to July 2026, supplemented by adjacent work on retrieval-augmented generation and on evaluation. Eighteen pages, eight tables, a literature matrix and a search protocol supplied as annex files.
Provenance marking: this is a preprint, with no peer review. Its strength is not its editorial status, it is that it publishes its selection method and its matrix, which makes it possible to contradict. No agency blog post on GEO does that.
Its structuring thesis deserves to be understood before the results. GEO is not a single ranking task. It is a partially observable stochastic pipeline, in ten stages: search activation, crawling, indexing, retrieval, reranking, allocation in the context, citation, prominence, factual absorption, faithfulness, then user behavior.
That decomposition explains half the sector’s misunderstandings. A technique that acts at stage seven says nothing about stage three. A number obtained at stage seven and sold as a global result is a category error, not an exaggeration.
The overall verdict, in one sentence
Here is the survey’s conclusion across the whole corpus.
Content that has already been retrieved can have its citation or its use causally modified. But no technique examined demonstrates a stable, longitudinal, cross-platform causal effect on organic discoverability or on downstream behavior.
Three adjectives, and all three count. Stable: the effect holds from one measurement to the next. Longitudinal: it holds over time. Cross-platform: it holds from one engine to another. Across forty-five studies and nearly three years, no technique ticks all three.
That is not the same thing as saying GEO does not work. It says the literature does not allow you to assert that it works, which is a weaker statement and a far more awkward one for a market that bills monthly retainers.
The sort, column by column
| Lever | Status | Evidence |
|---|---|---|
| Topical relevance of the content | Proven | Primary determinant across 252,000 trials, 18-factor factorial design, 6 models |
| Position in the context window | Proven, but outside your control | U-shaped curve established in a peer-reviewed journal |
| Modifying the citation of already retrieved content | Proven, narrow scope | Founding 2023 paper, confirmed by the survey within its frame |
| Brand mentions off your own site | Plausible, not established | Correlation of 0.664 across 75,000 brands, but the vendor writes “correlation is not causation” himself |
| Explicit pricing and recent timestamps | Plausible, converging signals | Help consistently in the 252,000-trial study |
| Presence on authoritative third-party sources | Plausible, not established | Systematic bias toward earned media, measured across several verticals and languages |
| Formatting, structure, lists, chunking | Folklore | Near-zero effect across 252,000 trials |
| An llms.txt file | Folklore | No correlation across ~300,000 domains; 97 percent of files never read across 137,000 sites; Google states it does not support it |
| Schema markup to obtain AI citations | Folklore | 1,885 pages against ~4,000 controls: zero effect on ChatGPT and AI Mode |
| Rewriting your pages “for the LLMs” | Counter-productive | Citation-oriented rewrites can degrade retrieval |
| Generic GEO strategy in a competitive market | Cancels out | Individual gain from +0.802 to +0.007 once every brand adopts it |
Three observations on that table.
The proven column holds three lines, one of which you do not control and one whose scope is narrower than it looks. What is left, in practice, is topical relevance.
The folklore column contains exactly what the market bills most easily, because those are the only things deliverable inside a monthly package. A file to drop in, markup to generate, paragraphs to restructure.
And the plausible column is where everything is decided. It points massively toward actions off your own site: being mentioned elsewhere, existing on third-party sources. That is public relations work, not on-page optimization.
What the commercial audits add, and why it is awkward
The survey does not confine itself to academic work. It also examines audits published by commercial players, and it draws three converging findings from them.
Source overlap is low. Two tools querying the same engine on the same prompt do not report the same cited sources. If your two vendors give you different numbers, it is not necessarily because one of them is lying.
Run-to-run variability is substantial. The same prompt, asked twice of the same model, does not produce the same citations. An April 2026 paper puts the principle plainly: visibility in generative search is a distribution, not a point. A single measurement has no value.
Faithfulness gaps persist. What the engine asserts is not always supported by the source it cites, including when the source itself is correct.
Those three findings carry a direct and rarely stated consequence: a GEO report that shows you a before and an after, each measured once, demonstrates nothing. It shows two draws from a random variable.
The evidence hierarchy, and how to use it
The survey’s most useful contribution for a decision-maker is not its verdict. It is the evidence hierarchy it proposes, which lets you sort any claim anyone puts in front of you.
From strongest to weakest: repeated measurement with a control group and human validation, then a controlled experiment without longitudinal control, then observational correlation, then case study, then assertion without data.
The reproducible protocol the document recommends comes down to five requirements: repeated measurements, paraphrases of the same intent, a control group, human validation, and accounting for interference between several players optimizing at the same time.
Ask which of those five your vendor applies. You will get a highly informative answer, often through his silence.
What contradicts this table, and should be read too
Two serious pieces of work pull the other way. Ignoring them would be doing exactly what this article holds against the checklists.
The first comes from MIT. A testbed called E-GEO, built on 13,747 product queries crossed with ten Amazon listings, five generative engines and seven automated rewriters, surfaces through meta-optimization a stable, domain-agnostic pattern. In other words: a general and effective GEO strategy would indeed exist. And under a simple defense, the observed gains correspond to a real improvement in the content, not to manipulation. Caveat for the file: it is an offline testbed, with no longitudinal measurement and no competitive dynamic.
The second is an analysis covering 75,000 brands, which measures a correlation of 0.664 between brand mentions across the web and visibility in Google’s generated summaries, against 0.218 for backlinks. The top three factors all sit off the site. Provenance: the study is published by an SEO tool vendor, who takes the trouble to write himself that correlation is not causation, and the relationship is probably confounded by brand size. Side figure worth keeping: 26 percent of brands have no mention at all in those summaries.
Neither of those two overturns the survey’s verdict. They indicate where to look: toward genuinely better content and off-site presence, not toward technical retouching.
What you do tomorrow morning
Take the GEO checklist you have to hand and file every line into one of the three columns in the table above. The exercise takes ten minutes and it is rarely pleasant.
Whatever lands in folklore, stop paying for it. The llms.txt file takes ten minutes to ship and costs nothing: it has no business appearing on an invoice. Schema markup has real virtues for crawling and rich results, but it does not get bought as a citation lever.
Whatever lands in plausible is your real budget. Being mentioned elsewhere, existing on third-party sources, improving the product itself. It is slower, it is less deliverable as a monthly package, and it is the only thing your competitors will not be able to replicate in three weeks.
And in the face of any new claim, one question only: how many repeated measurements, against what control group?
Forty-five studies in three years, and a single line holds in the column of actionable certainties. That is not a scandal. It is simply the state of a young discipline, which the market sells as a mature science.
Sources
- Martinez, O. (2026). Optimizing Visibility in Generative Engines: A Critical Survey of Generative Engine Optimization (2023-2026), arXiv:2607.14035
- Vishwakarma, R., Kumar, S. & Jamidar, R. (2026). What Gets Cited: Competitive GEO in AI Answer Engines, arXiv:2605.25517
- Liu, N. F. et al. (2024). Lost in the Middle: How Language Models Use Long Contexts, Transactions of the ACL, vol. 12
- Schulte, J., Bleeker, M. & Kaufmann, P. (2026). Don’t Measure Once: Measuring Visibility in AI Search, arXiv:2604.07585
- Chen, M., Wang, X., Chen, K. & Koudas, N. (2025). Generative Engine Optimization: How to Dominate AI Search, arXiv:2509.08919
- Bagga, P. et al. (2025). E-GEO: A Testbed for Generative Engine Optimization in E-Commerce, arXiv:2511.20867
- Ahrefs (2025). An Analysis of AI Overview Brand Visibility Factors (75K Brands Studied)
- Ahrefs (2026). We Tracked 1,885 Pages Adding Schema. AI Citations Barely Moved.
- SE Ranking (2025). LLMs.txt: analysis across ~300,000 domains
- Chu, X. & Hou, Y. (2026). Incumbent Advantage, arXiv:2606.17443
<strong>LaFactory</strong> measures AI visibility with a published protocol: repeated measurements, paraphrases, control group. No guaranteed placement, ever. Contact us to scope an audit.