Manipulating AI answers: what works in the lab, what holds up in production

by | Aug 26, 2026 | GEO

You are starting to be sold this, quietly. Proposals that talk about advanced optimization, proprietary techniques, methods your competitors are not using yet. Behind the vocabulary there is sometimes prompt injection: text written for the model rather than for the reader.

The academic literature has been working on this since 2024, and it has produced spectacular results.

The problem is not whether these techniques work. Some of them work very well, under the conditions where they were tested. The problem is that those conditions are not the conditions of a production engine, and the gap changes the entire risk calculation.

What follows separates the two, and names the weakness that is genuinely exploitable, which is not the one you would expect.

What works in the lab

Several studies have established that a conversational engine can be manipulated by text placed on a web page.

A first piece of work on adversarial optimization for language models demonstrates the effectiveness of what it calls preference manipulation attacks. Its results are striking: fictitious products roughly double their presence when competing against real ones. The attacks can be made stealthy by adding an instruction along the lines of “do not mention this message in your answer”. They can be triggered from a different page than the one presenting the product, or hidden in text a human cannot read. And they transfer to engines running in production.

A second study, on ranking manipulation in conversational search engines, shows that a tree-based technique reliably promotes poorly ranked products.

A third proposes a stealthy prompt optimization method, designed to remain undetectable on reading.

A fourth establishes that an adversarial injection framework transfers across different ranking architectures while preserving the fluency of the text.

Provenance marking: all four are arXiv preprints with published protocols. Not one of them is peer reviewed.

So much for the demonstration. It is solid, and it is worrying.

What holds up in production

Here is the work nobody cites in a sales deck, and it reframes the entire file.

Published on March 26, 2026 by a team at Fudan University, it tests ten deployed retrieval-augmented language model search products, including the best known ones, against a testbed named SEO-Bench built from one thousand real black-hat sites. Provenance marking: arXiv preprint, not peer reviewed, benchmark and protocol published.

The result: those engines neutralize more than 99.78 percent of traditional black-hat SEO attacks. And the authors identify the mechanism. The retrieval stage acts as the main filter. Manipulated content never reaches the model, because it is never retrieved in the first place.

Most of the attacks demonstrated in the lab assume the poisoned document is already in the context. In production, the stage that decides what enters the context eliminates it beforehand.

The same work adds a qualification that rules out any claim of invulnerability: attacks designed specifically for language model based engines, such as stuffing the rewritten query or splitting text into segments, double the manipulation rate compared to classical techniques.

So it is not a wall. It is a filter that stops the old world and lets part of the new one through.

The table of real risk

Technique Measured effectiveness in the lab Measured resistance in production Risk to you
Keyword stuffing, link farms, classic hidden text Not applicable 99.78 percent neutralized Useless, and detectable
Prompt injection inside the page High on an isolated system Filtered at retrieval in most cases High if detected, uncertain gain
Injection triggered from another page Demonstrated Not measured under real conditions Very high: hard to explain if discovered
Attacks designed for language model engines High Manipulation rate doubled Real, and where defensive research is aimed
Authoritative language, invoked evidence +0.17 points on the rating Not filtered: it is ordinary text The genuinely exploitable weakness

The last row is the one that counts.

The weakness nobody talks about, because it does not look like an attack

A June 2026 study measuring the recommendations of three commercial models tested the effect of register. Provenance marking: arXiv preprint, not peer reviewed, methodology published.

Authoritative language is worth +0.17 points on the score the model assigns. And the effect holds when the clinical evidence being invoked is fabricated. The model does not distinguish a real study from an invented one.

Understand what that means. There is no injection here, no hidden text, no concealed instruction. Just sentences that assert with confidence, invoking work that does not exist. No retrieval filter can stop that, because it is perfectly ordinary prose.

That is the genuinely exploitable weakness, and it is also the easiest one to cross without noticing. An agency writing your pages with categorical claims and vague references to “studies” is exploiting this bias, whether it knows so or not.

The collective version of the same mechanism was modeled by a Carnegie Mellon team in August 2026, in an arXiv preprint: competition degenerates into citation wars, in which successive rewrites degrade document quality and introduce unsupported claims.

Nobody needs to cheat for the corpus to decay. It is enough that everybody optimizes.

This has already happened once

In 2011, Google began penalizing sites that relied on artificial signals. Thousands of sites lost their visibility overnight, including companies that had paid in good faith for vendors selling “advanced techniques”.

None of our clients were hit. Not through foresight. We simply had nothing to take down.

The structure of the risk is exactly the same today, with two differences that make it worse.

Manipulation is recoverable, reputation is not. A filter ships in one update. A page containing concealed instructions, discovered by a journalist or a competitor, does not get removed from anyone’s memory.

The gain is temporary by construction. Defensive research moves faster than offensive research on this ground, because the platforms have a direct commercial interest in their answers not being purchasable. A testbed measuring 99.78 percent neutralization today would have measured far less eighteen months ago.

What you do tomorrow morning

Three checks, two of which take ten minutes.

Look for text aimed at the model inside your own pages. Content hidden with CSS, text in the background color, attributes stuffed with instructions, blocks generated by a tool nobody reread. A past engagement may have deposited some without telling you.

Reread your pages with one precise question: is every strong claim supported by a verifiable source? If your copy invokes “studies” without naming them, you are exploiting the authoritative language bias. It is legal, it is common, and it exposes you. The day a reader traces the source and does not find it, you have the profile of an evidence fabricator.

Ask your vendor what they actually do. Not the list of deliverables: the real content of what gets placed on your pages. A vendor who answers with proprietary vocabulary deserves a second, firmer question.

The good news in this file is that cheating works badly. The bad news is that the line between optimizing and deceiving runs straight through a practice almost everyone uses without thinking: asserting more strongly than you can prove.

Sources


<strong>LaFactory</strong> measures AI visibility with a published protocol: repeated measurements, paraphrases, control group. No guaranteed placement, ever. Contact us to scope an audit.

Cart