You are told that if AI does not cite you, your content is not good enough. Not structured enough. Not clear enough. Not authoritative enough. Rewrite it, break it up, enrich it.
Sometimes that is true.
Often it is a question of placement. Your page was retrieved, it is sitting in the model’s context window, it is relevant, and it is seventh out of fifteen.
The problem is not the quality of your writing. The problem is its rank in a stack you do not control.
What follows lays out the mechanism, what it implies about the levers you actually hold, and the trap that half of all GEO engagements walk straight into.
The U-shaped curve
The reference work is titled Lost in the Middle: How Language Models Use Long Contexts, published in the journal Transactions of the Association for Computational Linguistics, volume 12, in 2024.
Provenance marking, and it counts here: journal publication with peer review, more than a thousand academic citations. That is the highest level of evidence in the entire GEO file, and it does not come from marketing.
The protocol is simple and clean. The model receives several documents, only one of which contains the answer. The position of that relevant document in the list is then shifted systematically, and performance is measured at each position.
The result: performance follows a U-shaped curve. The model makes good use of what sits at the start of the context. It makes good use of what sits at the end. It makes poor use of what sits in the middle.
Same document. Same content. Same relevance. Three positions, three different outcomes.
What that means for your page
A generative engine does not read your page on its own. It breaks your question into several internal queries, retrieves a set of documents for each one, stacks them in a context window, and composes an answer out of that stack.
Your page is one of those documents. Its place in the stack is decided by the retrieval and reranking stages, which is to say by the engine.
A study from May 2026 covering 252,000 trials, six models and a factorial design across eighteen factors confirms the phenomenon under real citation conditions: the two dominant determinants of the first citation are topical relevance and position in the list. Provenance marking: team at a marketing software vendor, protocol published, conflict of interest to declare.
Two determinants. You control one.
The table of what you control
| Pipeline stage | Who decides | Your room to act |
|---|---|---|
| Query fan-out into sub-queries | The engine | None. You do not even see the generated queries |
| Crawling and indexing | The engine, within your robots.txt | Allow or block |
| Document retrieval | The engine | Indirect, through topical relevance |
| Reranking | The engine | Indirect, same lever |
| Position in the context window | The engine | None |
| Citation and prominence | The engine | Weak, and the effect cancels out under competition |
The only column entry that is not some version of “none” is topical relevance, which acts upstream. Everything that happens after retrieval is out of your hands.
That is a complete reversal of classical SEO, where position was the outcome of a competition you could work on directly.
The consequence nobody draws: variance
This mechanism explains a large part of the instability every tracking tool suffers from.
If citation depends on position, and position varies from one run to the next, then citation varies from one run to the next. That is not a defect in the measurement tool. It is a property of the system propagating outward.
Work from April 2026, an arXiv preprint with a published method and no peer review, turns it into a practical rule: generative visibility is a distribution, not a point. A single measurement has no interpretation, precisely because the same content can land in second position on one run and eighth on the next.
When your monthly report shows a three-point drop, the first hypothesis to rule out is not the quality of your content. It is the draw.
The trap: optimizing one stage by breaking the previous one
Here we come to the most expensive mistake in the sector, and it follows directly from everything above.
The critical survey of July 2026, an arXiv preprint that examined forty five studies, notes that citation-oriented rewrites can degrade retrieval.
The reasoning is mechanical. You rewrite your page so it is easier to cite: self-contained sentences, decisive claims, extractable phrasing. In doing so, you strip out the semantic context that allowed the page to be retrieved across a spread of different formulations.
The outcome: the page is better cited when it is present, and it is present less often. Or retrieved lower, therefore closer to the middle, therefore inside the dead zone of the U-shaped curve.
The net balance can be negative. No dashboard will ever show you that, because it counts the citations gained and not the retrievals lost.
What complicates this
Two objections deserve examination before we close.
Context windows keep growing. True, they have grown enormously since 2024, and it is tempting to hope the problem dissolves. It does not dissolve. The original work shows that degradation is tied to relative position, and longer contexts mainly mean more documents entering the stack, which widens the middle rather than shrinking it.
The bias is not uniform across models. Some architectures are less sensitive to it than others. Work from May 2026, an arXiv preprint covering 761,495 citation pairs, establishes that 88 to 96 percent of the quality variance is explained by the provider rather than the model, which means platform architecture decisions dominate. Your exposure to position bias therefore depends mostly on which platform your audience uses. That is one more variable outside your control.
What you do tomorrow morning
Three practical consequences, and only one of them is comfortable.
Work on retrieval, not on citation. Topical relevance is the only documented lever that acts upstream, before position is fixed. In practice: actually cover a subject, in its real vocabulary, with the phrasings people really use, rather than rewriting for extractability.
Multiply your entry points. If your visibility rests on a single page, it rests on a single position in a single stack. Being present on authoritative third-party sources means offering several retrieval paths. A study run across several verticals and several languages, published as an arXiv preprint by a University of Toronto team, measures a systematic and pronounced bias toward exactly those third-party sources.
Distrust any engagement that promises to make your pages more citable. You are buying an optimization of stage six, paid for with a possible degradation of stage three, and sold without any mention of that trade.
Your content is not judged in a ring. It is stacked in a queue whose order is decided elsewhere, and a good share of what the market sells as editorial performance is a rank effect.
Sources
- Liu, N. F. et al. (2024). Lost in the Middle: How Language Models Use Long Contexts, Transactions of the ACL, vol. 12, DOI 10.1162/tacl_a_00638
- Vishwakarma, R., Kumar, S. & Jamidar, R. (2026). What Gets Cited: Competitive GEO in AI Answer Engines, arXiv:2605.25517
- Martinez, O. (2026). Optimizing Visibility in Generative Engines: A Critical Survey of GEO (2023-2026), arXiv:2607.14035
- Schulte, J., Bleeker, M. & Kaufmann, P. (2026). Don’t Measure Once: Measuring Visibility in AI Search, arXiv:2604.07585
- Seo, Y. et al. (2026). Verified Misguidance, arXiv:2605.28565
- Chen, M., Wang, X., Chen, K. & Koudas, N. (2025). Generative Engine Optimization: How to Dominate AI Search, arXiv:2509.08919
<strong>LaFactory</strong> measures AI visibility with a published protocol: repeated measurements, paraphrases, control group. No guaranteed placement, ever. Contact us to scope an audit.