Research · Measurement reliability Sources & citations

A single run is a point estimate, not a number: the published evidence

A report somewhere in your pipeline probably reads like this. "Your brand appears in the answers for 22% of tracked prompts." Ask one question before that number goes near a budget. How many times was each prompt actually run?

Is a single-run AI visibility number trustworthy?

No. One run gives you a single draw, not a settled number. A line like "your brand appears in the answers for 22% of tracked prompts" reads like a fact. Before it goes near a budget decision, ask how many times each prompt was run. AI answers vary between repeated runs, and the published evidence this article reviews is why repeated measurement with reported ranges is the standard.

The numbers behind that come in two parts. A major vendor's own experiment found ten repeated runs cut day-to-day citation-share noise by about 40%. Search Engine Land guidance recommends repeated measurement with confidence intervals. The procurement-grade format prints the range next to the rate. "Cited in 33% of runs (95% CI 12–65%, n=9)" is a different object from a bare "33%".

If the answer is "once," you are not looking at a measurement. You are looking at one draw from a spread of possible answers, which statisticians call a point estimate. That is not a criticism of any particular tool. It is a property of the thing being measured. By mid-2026 it is documented across academic preprints, search-industry guidance, and, most tellingly, a measurement vendor's own published experiment.

The review below collects that evidence in one place, with links and dates, so you can check each item yourself.

Does the same question return the same answer?

AI answer engines do not behave like a database. Ask the same question twice, under the same conditions, and the answer can differ. Different sources cited, different brands named, sometimes a different recommendation. The research linked below studies why. For a buyer of measurement the mechanism matters less than what follows from it. Any rate computed from one pass inherits that movement in full, whether it is a citation rate, a mention rate or a "visibility score."

How large is the effect? A measurement study posted to arXiv in January 2026 (arXiv 2601.21339, preprint) found that large-language-model outputs shift by 10–34% from sampling alone, before any change in the market, the model version, or your website. That movement is not an edge case. It is the floor you are standing on.

The published record

Four documents, from three very different kinds of authors, currently anchor this picture. None of them is ours.

1. An academic analysis calls single-run measurement "fundamentally unreliable."
A 2026 preprint analyzing AI-visibility measurement (arXiv 2603.08924) finds single-run measurement "fundamentally unreliable" due to nondeterminism. The same analysis estimates what statistical solidity would actually cost: a 95% confidence interval five percentage points wide on citation share requires roughly 40–150 repeated runs per platform. Keep that benchmark in mind. We return to it below, including how our own protocol measures against it.

2. A second study puts numbers on the variance.
The sampling study cited above (arXiv 2601.21339, 2026-01): 10–34% output variance from sampling alone. No prompt changes, no model updates, just asking the same thing again.

3. A vendor's own experiment, published on their own blog.
Profound, a global AI-visibility tool, explains their once-a-day reading methodology on their official blog ("Is once a day enough?", 2026-07-08). In that same post, their own experiment found that ten repeated runs cut day-to-day noise in citation share by about 40%. They deserve credit for publishing the experiment at all. Most public pages we checked in this category disclose no variance data. Look at what that implies. Say repeating a measurement ten times removes about forty percent of the day-to-day movement. Then a large share of what a single daily reading reports as change is, by the vendor's own numbers, noise.

4. Search-industry guidance now says the same thing.
Kevin Indig, writing in Search Engine Land ("Make prompt tracking more accurate", 2026-06-10) treats a single observation as a point estimate. He recommends repeated measurement, with confidence intervals reported alongside.

Academic preprints, a vendor's own published experiment, and trade-press methodology guidance. Three kinds of author with three different reasons to publish, and on this question the four documents above point the same way.

What does buying a point estimate risk?

Suppose the report on your desk is a single-run measurement. Three specific risks follow. All three are properties of the reporting format, not of any particular vendor.

The baseline risk. You commission a "before" measurement, spend a quarter on content and technical work, then commission an "after." Both are single runs. Given documented sampling variance of 10–34% (source: arXiv 2601.21339, 2026-01 preprint), the gap between them cannot be attributed. You cannot tell your program's effect from the measurement's own movement. The whole before-and-after story, the thing you bought the baseline for, cannot be read.

The accountability risk. A point estimate can never be wrong. If next month's number is different, that gets reported as "change." A number published with its confidence interval is something else — a claim you can prove wrong. If a re-measurement under the same protocol lands outside the stated interval more often than the confidence level allows, something is broken, and you can see it. Intervals are what make a measurement vendor accountable to you.

The allocation risk. AI-visibility reports are built on comparisons. Your brand versus competitors, one engine versus another, this month versus last. Budget then follows the differences. But if the gap between two figures is smaller than the noise band around each, moving spend on it is moving spend on static. Without a published range, you have no way to know which differences are real. The same discipline applies to before-and-after "lift" claims. We broke those down into the denominator questions that sort real metrics from vanity metrics.

What does a procurement-grade number look like?

None of this means AI-visibility measurement is not worth buying. It means the reporting format determines whether the number survives a finance review, a procurement questionnaire, or a skeptical CMO. A number built for that room discloses four things. Treat the four as a checklist and paste it straight into the questionnaire you already send.

Where do we stand against the benchmark we just cited?

We should apply the standard to ourselves before anyone else. CiteAngle's paid baseline runs every query in the panel seven times. It reports a Wilson 95% confidence interval next to every applicable rate, with the effective sample size beside it.

We know what one run costs because we ran a nine-brand panel that way and kept the output, a Korean-market clinic-sector run of ten questions across six answer surfaces on 2026-07-17. At one run per question the widest interval reached ±24.5 percentage points, and that width is why the panel flattened: all nine brands landed in the top band on all six surfaces, a leaderboard reading 6/6 for everyone with no rank inside it. The same design at seven runs stopped flattening them, and the panel spread across leadership counts instead. Brand names stay out of the public write-up; the band spread is the part that transfers.

We print the measured interval width next to every applicable rate, one by one, and where a width is wide the report says so. The 40–150 benchmark counts distinct questions, not repeats of one. Panorama asks fifty per platform, which sits inside that range, and the seven repeats stack on top as a separate axis. Each observation also carries its receipt. Measurement time, model and configuration hash, sealed so you can check the numbers yourself. And the external figures we cite publicly, including every figure in this article, are logged in our public claims registry with source and access date.

How many questions are enough?

The published academic estimate is roughly 40–150 queries per platform for a 95% confidence interval five points wide on citation share. The paper's own numbers run from about 40–50 on one engine to 150 or more on another. Panorama asks 50 distinct queries per platform, at the low end of that published range. It then repeats every one of them seven times, an axis the paper treats on its own. The report does not claim the sample is big enough. It prints the measured interval width next to every applicable rate, one by one, so you can see how much weight each number can bear before you lean on it. That width is the line most reports leave out, and it is the one that tells you whether a change is real.

The spread, and the runs it takes: figures from the published record

Output variance from sampling alone (scale 0 to 100 percent) — source: arXiv 2601.21339, 2026-01 preprint

Queries estimated for a ±5-point-wide 95% CI on citation share (academic estimate)

Sources: arXiv 2601.21339 (variance) · arXiv 2603.08924 (required-queries estimate) · links in the text above · evidence accessed July 16, 2026 · bar length proportional to each panel's maximum (zero baseline) · figures logged in our public claims registry.

The full protocol is written up on our methodology page: how many repeats, how the interval is built, what each status means, and what the receipt holds.

Start with the distribution, not the point

If you want to see what this looks like on your own brand, start with the free AI visibility check. It is a single-pass read, and the result carries that label on its face. What it hands you is the shape of your position. Which engines answer the question at all, which sources they pull from, and the format the report arrives in. Repeated runs, intervals and sealed receipts are what the paid baseline adds. When the number finally goes into a budget decision, it goes in with its uncertainty attached, the way a number you are about to spend against should.

Sources

What one run costs. Without repeats, a swing in the answers is indistinguishable from a result, and good work risks being cancelled on a number that was never measured twice. Every question here is asked 7 times and sealed with its conditions, which is why our deltas arrive as two ranges you can see separate.

Why the repeat count is the thing to ask about

Everything expensive about this failure happens quietly. A single-run figure goes into a budget request, the budget is approved, and the following quarter's single-run figure comes back lower — so the work gets judged a failure and stopped, when the honest reading is that nobody ever measured anything twice. Careers get shaped by that arithmetic. No error message is produced.

Our answer is a published number. Every question is asked 7 times, and the count is printed on the report instead of living in a methodology conversation. Rates ship with a Wilson 95% interval, so a delta is reported as two ranges and you can see whether they even separate before anyone calls it a result. Ask any vendor in this market how many times they ask. Whether the answer arrives as a number or as a sentence about rigour tells you most of what you need.

The next number you buy should carry its interval

See the format on your own brand first. A free AI visibility check, then a baseline that ships repeated runs, Wilson 95% intervals and sealed receipts.

Ready to turn US visibility evidence into a market plan? Tell us the US buyer, target state and decision date.

Request a written proposal