Your dashboard displays a number. Twelve percent share of voice. Up three points since last month. A clean chart, a green arrow, a caption about momentum.
You feel informed. You are not.
The problem is not that the tool is lying. The problem is that it shows you a single draw from a random variable and calls it a measurement.
What follows explains why the same question asked twice does not produce the same answer, what that invalidates in the deliverables the industry sells, and the minimum protocol that makes a number defensible.
The principle: a distribution, not a point
Work published on April 8, 2026 states the problem without hedging. Visibility in generative search has to be treated as a distribution, not as a point. A single observation is not reliable.
Provenance marking: arXiv preprint, not peer reviewed, built on empirical measurements of the variance between successive runs.
The phrasing sounds technical. It is not. It says that a generative engine asked the same question twice in a row does not cite the same sources, and that the gap between the two answers is not an anomaly. It is how the system normally behaves.
A classical search engine is deterministic in the short term. You are in third position, you are in third position for everybody, and you will still be there in an hour. A generative engine does not work that way. It samples. It composes. Two identical runs produce two different answers, with different sources underneath them.
Measuring once, under those conditions, is rolling a die and recording the result as the value of the die.
What the audits reveal, and the reports do not show
The critical survey of July 2026 examined the audits published by the commercial players in this market alongside the academic work. Provenance marking: arXiv preprint, not peer reviewed, with the reviewed corpus listed. It draws three converging conclusions, and all three are awkward.
Source overlap is low. Two tools querying the same engine on the same question do not report the same cited sources. If your two vendors hand you different results, the most likely explanation is not dishonesty. It is that the system does not produce the same thing twice.
Run-to-run variability is substantial. This is not marginal background noise around a stable value. The spread can exceed the size of the movements your vendor presents to you as results.
Faithfulness gaps persist. What the answer asserts is not always supported by the source it cites, even when that source is correct and relevant.
The order of magnitude of the volatility
So much for the principle. Now we need the size of the effect, and an analysis published in November 2025 gives a concrete picture of it, at a scale that rules out any local accident.
Its protocol: more than 230,000 prompts, more than one hundred million citations recorded, across three engines, sampled weekly from July 14 to October 12, 2025. Provenance marking: published by an SEO tool vendor, with the methodology and the measurement window disclosed, and a commercial interest in the topic.
The striking result concerns a single domain, but it is spectacular. Reddit was cited in close to 60 percent of ChatGPT answers in early August. By mid-September, that rate had fallen to roughly 10 percent.
Six weeks. Fifty points.
No action by any publisher explains that. It is a platform change, invisible, unannounced, and it redistributed the visibility of millions of pages in the process. If your September report showed a decline, that was not your strategy.
One last element finishes off the measurement problem. Work from May 2026 covering 761,495 citation pairs, an arXiv preprint with a published protocol, establishes that 88 to 96 percent of the quality variance is explained by the provider, not by the model. Measuring one engine tells you almost nothing about what is happening on the others.
What your report shows, and what it should show
| What a typical GEO report shows | What it is worth | What it should be |
|---|---|---|
| A monthly share of voice | A single draw | A mean over repeated measurements, with dispersion |
| One query per intent | Sensitive to wording | Several paraphrases of the same intent |
| A month-over-month change | Confuses your action with platform changes | A control group of pages you leave alone |
| An automated mention count | Does not read what is said about you | Human validation on a sample |
| A single engine measured | Not transferable: variance sits between providers | One measurement per platform, not aggregated |
| No mention of active competitors | Ignores interference between optimizers | Accounting for the players optimizing at the same time |
Every line in the left column is standard in the deliverables sold today. Every line in the right column appears in the protocol the literature recommends.
The minimum protocol, in five requirements
The July 2026 survey proposes a reproducible protocol. It fits in five points, and none of them is optional if you intend to assert anything at all.
Repeated measurements. Ask the same question several times, and work on the distribution you obtain.
Paraphrases. Express the same intent several ways, because the answer changes with the wording, and because your customers do not all phrase things alike.
Control group. Comparable pages you deliberately do not touch, so you can separate your effect from platform movement. Without one, you will credit your own work with the September collapse described above.
Human validation. Read a sample of the answers instead of counting occurrences. A mention can be inaccurate, negative, or about a competitor the counter attributed to you.
Multi-actor interference. Account for the fact that your competitors are optimizing at the same time. That is what makes the individual gain from a GEO strategy collapse in a competitive market.
How many measurements, exactly
Here we owe you a caveat, and it illustrates the subject of this article better than any argument could.
Precise figures circulate. Somewhere between forty and one hundred and fifty runs would be needed for a tight confidence interval, or source overlap between two identical answers would sit at roughly 21 percent. Both numbers are plausible and consistent with the rest of the file.
Neither one can be traced to a primary source. They appear in tool vendor posts that never publish the protocol they came from. Provenance marking: self-reported, method undisclosed, commercial interest present.
The principle is solidly established by work whose method is public: a single measurement is worth nothing. The exact number of runs required is not established, and you get that caveat here instead of a round number dressed up as proof.
That is exactly the sorting you should be applying to everything presented to you on this subject.
What you do tomorrow morning
Open your latest GEO report and put five questions to whoever produced it.
How many times was each query run? If the answer is once, nothing else in the report has an interpretation.
What dispersion do you observe between runs? If no dispersion is reported, there were no repeated measurements.
Which pages serve as the control group? Without a control, you cannot separate your effect from a platform change.
What share of the answers was read by a human? A mention count says nothing about what was said.
And are the results aggregated across platforms? If they are, ask for the breakdown, because the variance sits mainly between providers and an average hides the only thing that matters.
A serious vendor answers those five questions without bristling, because they have already asked them internally. A vendor selling a monthly number changes the subject.
A number displayed to one decimal place is not a measurement. It is a number displayed to one decimal place.
Sources
- Schulte, J., Bleeker, M. & Kaufmann, P. (2026). Don’t Measure Once: Measuring Visibility in AI Search, arXiv:2604.07585
- Martinez, O. (2026). Optimizing Visibility in Generative Engines: A Critical Survey of GEO (2023-2026), arXiv:2607.14035
- Semrush (2025). The Most-Cited Domains in AI: A 3-Month Study
- Seo, Y. et al. (2026). Verified Misguidance: Measuring Structural Citation Failures in Search-Augmented LLMs, arXiv:2605.28565
- Vishwakarma, R., Kumar, S. & Jamidar, R. (2026). What Gets Cited: Competitive GEO in AI Answer Engines, arXiv:2605.25517
- Chu, X. & Hou, Y. (2026). Incumbent Advantage, arXiv:2606.17443
<strong>LaFactory</strong> measures AI visibility with a published protocol: repeated measurements, paraphrases, control group. No guaranteed placement, ever. Contact us to scope an audit.