Test GEO on your own site in 30 days: our protocol

by | Aug 26, 2026 | GEO

You want to know whether GEO works. So you search for the answer. You find a case study, a vendor benchmark, an average lifted from somebody else’s market, and you try to apply it to yours.

Nobody can answer that question in the abstract. The literature establishes that no technique shows a stable, longitudinal, cross-platform effect on discoverability. It does not establish that nothing works anywhere. Your sector, your brand recognition, your queries and your audience are not the average.

The problem is not finding the right answer online. The problem is measuring it on your own site, with a method that survives scrutiny.

Here is the one we run. It fits in thirty days, it needs no paid tool, and you can reproduce it yourself.

The five non-negotiable requirements

The critical survey published in July 2026, which examined forty-five studies, proposes a reproducible protocol. Provenance marking: arXiv preprint, not peer reviewed, and the only systematic review of the field to date. Five requirements, and removing any one of them makes the result uninterpretable.

Repeated measurements. Generative visibility is a distribution, not a point. A single observation is noise.

Paraphrases. The answer changes with the phrasing, and your customers do not all phrase things the same way.

A control group. Comparable pages you do not touch, without which you will credit your own work with movements that belong to the platform.

Human validation. Read a sample of the answers, do not only count them.

Multi-actor interference. Your competitors are optimizing at the same time, and the gain from a generic strategy collapses when everybody adopts it.

Week 0: prepare, and change nothing

Pick ten intents, not ten keywords. An intent is a question a customer genuinely asks. “Which invoicing software for a small French business”, not “invoicing software”. Take five category intents, where your name has no particular reason to come up unprompted, and five brand or comparison intents.

Write three paraphrases per intent. One short phrasing, one long phrasing, one phrasing with a constraint attached. You now have thirty prompts.

Build two groups of pages. The test group holds five to ten pages you are going to modify. The control group holds a comparable number of pages of the same type, same age, same traffic, which you will not touch. This is the element nobody includes, and it is the one that makes the test valid.

Fix the list of platforms. Three is enough. Never aggregate them: between 88 and 96 percent of quality variance is explained by the provider, per the same July 2026 survey, so an average destroys the information.

Write down the date. Citations move in a matter of weeks. An undated reading is worth nothing.

Week 1: the baseline measurement

Run the thirty prompts, five times each, on every platform. That is 450 answers for three platforms. Budget three to four hours, or write a simple script.

For each answer, record four things and nothing else.

Is your brand named in the text, yes or no.

Is your site cited as a source, yes or no. That is not the same question, and the gap between the two is often the single most useful output of the whole test.

Which competitors are named.

Which sources are cited, with their URLs.

Then calculate, for each intent and each platform, your mention rate across the five runs, and the gap between the most favorable run and the least favorable one. That gap is your noise floor. Any later variation smaller than it means nothing at all.

It is the most important number in the protocol, and it is the one no commercial report will ever hand you.

Week 2: intervene only where the evidence exists

Apply to the test group, and only to the test group, the actions the literature reports an effect for. Nothing else, or you will not know what moved the result.

Real coverage of the topic, in its real vocabulary. Topical relevance is the primary determinant of citation in the largest available study, covering 252,000 trials. Provenance marking: arXiv preprint, largest sample in the field, not peer reviewed. In practice: fill in what the page does not currently address, using the words your customers actually use.

Canonical, dated facts. Pricing, scope of the offer, specifications, with a visible update date. Recent timestamping is one of the few factors with a consistent reported effect across studies.

Correction of third-party sources. Go through your baseline measurement and list the domains citing you with out-of-date information, then get it corrected. This is the highest-yield item and the slowest: start it in week 2 even though the effect will land well after day thirty.

Do not touch schema markup, do not deploy a configuration file, do not restructure your paragraphs for extraction. Those are the levers measured at zero, and including them would destroy your ability to attribute any result you do get.

Week 4: the final measurement, and how to read it

Repeat exactly what you did in week 1. Same prompts, same paraphrases, same number of runs, same platforms.

Then read the result in this order, in three questions.

Did the control group move? If it moved as much as the test group, your intervention explains nothing: the platform changed. A useful reminder here, from a three-month SEO tool vendor study: one major domain went from close to 60 percent of one engine’s answers to roughly 10 percent in six weeks, with no publisher doing anything at all.

Does the observed gap exceed your noise floor? If the dispersion you measured in week 1 was twelve points and you gained five, you gained nothing.

What did the answers actually say? Reread fifty answers at random. This is where you usually find the most actionable information in the test: an obsolete price, a discontinued offer, a feature attributed to a competitor. A May 2026 measurement across 761,495 citation pairs, published as an arXiv preprint, establishes that 30.6 percent of citations misrepresent their source.

The protocol in one table

Stage Action Mistake to avoid
Week 0 10 intents, 3 paraphrases, 2 page groups, 3 platforms Choosing keywords instead of intents
Week 1 5 runs per prompt, recorded in 4 columns A single run, or aggregating platforms
Week 1 Calculate the dispersion between runs Skipping this step: without it, nothing is interpretable
Week 2 Coverage, dated facts, correction of third-party sources Touching markup too, which makes attribution impossible
Weeks 2 to 4 Change nothing in the control group Editing it “while we are in there”
Week 4 Repeat identically, then read in order Comparing against the test group without looking at the control

What this test will tell you, and what it will not

It will tell you whether a documented intervention moves your mention rate beyond the noise, on your intents, on the platforms your audience actually uses. That is information nobody else can sell you, because it belongs to you alone.

It will not tell you whether the effect holds over time: thirty days is not a longitudinal measurement, and that is precisely what the entire literature is missing.

It will not tell you whether the effect survives your competitors. The gain from a generic strategy collapses from 0.802 to 0.007 under general adoption, per an arXiv preprint modeling competitive dynamics.

And it will not tell you what any of it earns. The link between a mention and revenue travels largely through a later brand search, invisible to last-click attribution.

An honest test states its limits. That is what separates it from a sales pitch.

What you do tomorrow morning

Open a spreadsheet. One row per prompt, one column per run, four columns of recorded observations. That is your tool.

Block three hours this week for the baseline measurement, and three hours a month from now for the final one. That is the real cost of the test, and it is considerably less than one month of retainer.

If the result is positive beyond the noise, you will know what to fund and on what evidence. If the result is flat, you will have saved a year of subscription fees and you will know exactly why.

Either way, you will have done what the market does not: measure before paying.

Sources


<strong>LaFactory</strong> measures AI visibility with a published protocol: repeated measurements, paraphrases, control group. No guaranteed placement, ever. Contact us to scope an audit.

Cart