You read a critique of GEO and you have no way of knowing whether the person writing it ever looked at the evidence pointing the other way. You cannot check every citation. You cannot read forty-five papers. So you decide who to trust on tone, which is exactly how you got sold the discipline in the first place.
Our position, published with the numbers behind it, is that most of what the market calls GEO is not demonstrated. That llms.txt does nothing. That schema markup buys no additional citations. That the gain from a generic strategy collapses the moment competitors adopt it.
The problem is not that critics of GEO are too harsh. The problem is that a critic who reads only what confirms the critique becomes exactly what is being criticized, and nobody on the outside can tell the difference.
So here are the five strongest published works that go against our reading, presented at their full strength, with their samples and their methods, plus the conditions stated in advance that would make us drop the position.
One. MIT finds a stable GEO strategy, and it rests on real content improvement
The single most awkward result for our position comes from a team at the Massachusetts Institute of Technology, with a testbed named E-GEO.
Its scale: 13,747 realistic product queries, crossed with ten Amazon listings, tested across five generative engines, with seven automatic rewriters, fifteen manual heuristics, and a red teaming phase. Provenance marking: arXiv preprint, not peer reviewed, offline testbed.
Two results, and both of them contradict our article on the competition paradox and our article on sorting the levers head-on.
A meta-optimization of the instruction produces a stable, domain-agnostic pattern. In other words, a general, transferable, effective GEO strategy would in fact exist.
And under a simple defense, the observed gains reflect a genuine improvement in the content rather than manipulation. GEO would then not be a zero-sum game at all, but a properly posed optimization problem whose solution also happens to serve the reader.
What limits this counter-proof. It is an offline testbed on marketplace listings. It measures no longitudinal effect in production, and above all no competitive dynamic: it does not test what happens when all ten sellers in the category apply the same meta-optimization. That is precisely the variable where the gain collapses elsewhere.
Two. The effect of AI recommendations is real, and our measurement misses it
We wrote that referral traffic from generative platforms is marginal, around one percent. That is accurate, and it is a bad argument.
A June 2026 study uses a rare protocol: a panel joining the browsing data of consenting users to their actual conversations with three assistants. The same people on both sides. With an event study against prior trends, classification of the recommendation stance, conditioning on non-customers, an internal control within the same answer, and matched backward-dated placebos. Provenance marking: arXiv preprint, observational design, not peer reviewed.
When a brand is recommended to a user who did not previously know it, Google searches for its name rise by 4.3 points, in a confidence interval of 3.1 to 5.5. Visits to its site rise by 2.4 points. Product pages at resellers rise by 1.0 point.
That effect leaves no referrer trace. It travels through a later search, which your analytics tool will file under brand organic.
What this invalidates on our side. Any conclusion drawn from referral traffic alone, including ours. If the influence travels through an invisible channel, then measuring the visible channel and concluding there is no effect is a reasoning error, and no amount of methodological care elsewhere repairs it.
What limits this counter-proof. Observational design, preprint status, and no transaction observed anywhere in the data. A rise in brand search is not a sale.
Three. Something does work, and it is measurable
We filed a great many levers under folklore. An analysis covering 75,000 brands shows that at least one signal correlates strongly.
Brand mentions across the web correlate at 0.664 with visibility in Google’s generated summaries, against 0.218 for backlinks. The top three factors identified are all off-site.
That does not say GEO works the way it is sold. It says a real, identifiable lever exists, and that a company can work on it. Against a position built on absence of evidence, that is not nothing.
What limits this counter-proof. The author of the study writes himself that correlation is not causation, and the relationship is plausibly confounded by brand size: large brands have more mentions and more visibility for the same underlying reason. Provenance marking: SEO tool vendor, correlational analysis, publishing about its own market.
Four. Our article on manipulation is too alarmist
We devoted a whole article to adversarial attacks, with spectacular results: fictional products doubling their presence, stealth injections, transfers to engines running in production.
A team at Fudan University tested ten genuinely deployed retrieval-augmented search products against a benchmark of one thousand real black-hat sites. Result: more than 99.78 percent of attacks neutralized, with the retrieval stage acting as the principal filter. Provenance marking: arXiv preprint, not peer reviewed, but the only work in the file testing production systems rather than laboratory ones.
We cited it in the article concerned, and that is exactly why it belongs here: it contradicts part of our own material. The adversarial papers describe isolated systems. Production holds.
What limits this counter-proof. The same work measures that attacks designed specifically for language models double the manipulation rate. The filter stops the old world, not the new one.
Five. Source concentration is weaker than the dominant narrative
We stressed the crushing weight of a handful of large platforms in generative answers.
A December 2025 study covering 55,936 queries, comparing six generative engines against two traditional engines, measures greater source diversity among the former, with 37 percent unique domains. More diverse than Google and Bing. Provenance marking: arXiv preprint, not peer reviewed.
The story about a lockout onto fifteen domains is therefore overstated, and we relayed part of it.
What limits this counter-proof. The same work establishes that generative engines do no better than classic engines on source credibility, political neutrality or safety. Greater diversity of mediocre sources remains a problem.
What would still hold if all five were fully established
| Our claim | Status if the five counter-proofs hold |
|---|---|
| The plus 40 percent figure is misquoted, its experimental condition dropped | Unchanged. That is a fact about how the paper is cited, not about its content |
| A single measurement is worthless, visibility is a distribution | Unchanged. No counter-proof addresses it |
| The llms.txt file has no measured effect | Unchanged. Three independent sources, including Google’s official position |
| Pure formatting edits have almost no effect | Unchanged. 252,000 trials, uncontradicted |
| AI use for news has stopped growing in France | Unchanged. University survey across 48 markets |
| The gain vanishes under general competitive adoption | Weakened by the MIT testbed, which does not test competition |
| AI referral traffic is marginal, therefore the effect is weak | Refuted. The effect travels through a channel invisible to attribution |
| Visibility is locked up by large brands | Qualified. Domain diversity is higher than Google’s |
Across eight central claims, five hold intact, one is weakened, one is qualified, and one is refuted.
That is an honest scorecard, and it is also an admission: a claim was published here that is now known to be poorly founded. It has been corrected in the article concerned.
What would change our mind
A position that cannot say what would refute it is not a position, it is a posture. Here are the conditions, stated in advance.
A longitudinal replication in production. A study running six to twelve months, on real sites, with a control group, showing a stable effect of identified techniques on organic discoverability. That is exactly what is missing from the forty-five studies examined by the reference survey.
An effect measured in a genuinely competitive market. If the gain measured by MIT survives when every actor in a category applies the same method, our article on the competition paradox falls.
A replication of the stable pattern outside marketplaces. The MIT result covers product listings. Transposed to owned sites and confirmed there, it would change our position on on-page optimization.
A renewed rise in adoption in the French market. If the annual Reuters Institute survey measures a clear increase in France, the usage-ceiling argument becomes void.
A platform commitment to a stable signal. If Google, OpenAI or Anthropic officially document a generative optimization lever, the debate over whether the discipline exists is closed. As of today, Google writes that it is still search engine optimization and points buyers of these services toward its own page on evaluating third-party advice.
What you do tomorrow morning
Apply to what you have just read the exact test we recommend applying to everything else. Open the sources, check the samples, look for what is missing.
You will find preprints that have not been peer reviewed, tool vendors publishing data about their own market, and a single peer-reviewed journal publication in the entire technical corpus. We have flagged it every time, and it remains a weakness of the whole file, not only of our side of it.
A three-year-old discipline cannot produce longitudinal evidence. That is normal. What is not normal is billing for it as though it already had.
Sources
- Bagga, P., Farias, V., Korkotashvili, T., Peng, T. & Wu, Y. (2025). E-GEO: A Testbed for Generative Engine Optimization in E-Commerce, arXiv:2511.20867
- Iannelli, M. & Ai, A. (2026). From Prompt to Purchase: How AI Brand Recommendations Move Consumers on the Open Web, arXiv:2606.10907
- Ahrefs (2025). An Analysis of AI Overview Brand Visibility Factors (75K Brands Studied)
- Chen, P., Hong, G., Wu, X. et al. (2026). Unveiling the Resilience of LLM-Enhanced Search Engines against Black-Hat SEO Manipulation, arXiv:2603.25500
- Zhang, P., Ye, Q., Peng, Z., Garimella, K. & Tyson, G. (2025). Source Coverage and Citation Bias in LLM-based vs. Traditional Search Engines, arXiv:2512.09483
- Martinez, O. (2026). Optimizing Visibility in Generative Engines: A Critical Survey of GEO (2023-2026), arXiv:2607.14035
- Reuters Institute (2026). Digital News Report 2026
- Google Search Central. Optimizing your website for generative AI features on Google Search
<strong>LaFactory</strong> measures AI visibility with a published protocol: repeated measurements, paraphrases, control group. No guaranteed placement, ever. Contact us to scope an audit.