# A single run is a point estimate, not a number: the published evidence

원문: https://citeangle.com/en/research/single-run-point-estimate

Published July 17, 2026 · CiteAngle Research · Every external figure cited in
this article is logged in our [public claims registry](https://citeangle.com/en/claims) with source and
access date (sources accessed July 16, 2026) 
Updated July 17, 2026: range chart added
(same figures, drawn) · Updated July 22, 2026: question-led section headings, a run-count FAQ
and a lift-claims link added

## Is a single-run AI visibility number trustworthy?

No. One run gives you a single draw, not a settled number. A line like "your brand appears in the
answers for 22% of tracked prompts" reads like a fact. Before it goes near a budget decision, ask
how many times each prompt was run. AI answers vary between repeated runs, and the published
evidence this article reviews is why repeated measurement with reported ranges is the standard.

The numbers behind that come in two parts. A major vendor's own experiment found ten repeated runs cut day-to-day citation-share noise by about 40%. Search Engine Land guidance recommends repeated measurement with confidence intervals. The procurement-grade format prints the range next
to the rate. "Cited in 33% of runs (95% CI 12–65%, n=9)" is a different object from a bare
"33%".

Contents

1. [Same question, different answer](#sec-same-question)

2. [The published record](#sec-record)

3. [What a point estimate risks](#sec-risks)

4. [A procurement-grade number](#sec-procurement)

5. [Where we stand](#sec-benchmark)

6. [Start with the distribution](#sec-start)

7. [Sources](#sec-sources)

Questions this page answers: [Does the same question return the same answer?](#sec-same-question) [What does buying a point estimate risk?](#sec-risks) [What does a procurement-grade number look like?](#sec-procurement) [Where do we stand against that benchmark?](#sec-benchmark)

If the answer is "once," you are not looking at a measurement. You are looking at one draw from
a spread of possible answers, which statisticians call a point estimate. That is not a criticism
of any particular tool. It is a property of the thing being measured. By mid-2026 it is
documented across academic preprints, search-industry guidance, and, most tellingly, a
measurement vendor's own published experiment.

The review below collects that evidence in one place, with links and dates, so you can check each
item yourself.

## Does the same question return the same answer?

AI answer engines do not behave like a database. Ask the same question twice, under the same
conditions, and the answer can differ. Different sources cited, different brands named, sometimes
a different recommendation. The research linked below studies why. For a buyer of measurement the mechanism matters less than what follows from it. **Any rate computed from
one pass inherits that movement in full**, whether it is a citation rate, a mention rate or a
"visibility score."

How large is the effect? A measurement study posted to arXiv in January 2026
([arXiv 2601.21339](https://arxiv.org/abs/2601.21339),
preprint) found that large-language-model outputs shift by **10–34% from sampling alone**,
before any change in the market, the model version, or your website. That movement is not an edge
case. It is the floor you are standing on.

## The published record

Four documents, from three very different kinds of authors, currently anchor this picture. None
of them is ours.

**1. An academic analysis calls single-run measurement "fundamentally unreliable."** 

A 2026 preprint analyzing AI-visibility measurement
([arXiv 2603.08924](https://arxiv.org/html/2603.08924))
finds single-run measurement *"fundamentally unreliable"* due to nondeterminism. The same
analysis estimates what statistical solidity would actually cost: a 95% confidence interval five
percentage points wide on citation share requires roughly **40–150 repeated runs per
platform**. Keep that benchmark in mind. We return to it below, including how our own protocol
measures against it.

**2. A second study puts numbers on the variance.** 

The sampling study cited above
([arXiv 2601.21339](https://arxiv.org/abs/2601.21339),
2026-01): **10–34% output variance from sampling alone.** No prompt changes, no model updates,
just asking the same thing again.

**3. A vendor's own experiment, published on their own blog.** 

Profound, a global AI-visibility tool, explains their once-a-day reading methodology on their
official blog ([*"Is once a day enough?"*, 2026-07-08](https://tryprofound.com/blog/is-once-a-day-enough)). In that same post, their own
experiment found that **ten repeated runs cut day-to-day noise in citation share by about
40%.** They deserve credit for publishing the experiment at all. Most public pages we checked
in this category disclose no variance data. Look at what that implies. Say repeating a measurement ten times removes about forty percent of the day-to-day movement. Then a large share of what a single daily reading reports as *change* is, by the vendor's own numbers,
*noise*.

**4. Search-industry guidance now says the same thing.** 

Kevin Indig, writing in Search Engine Land
([*"Make prompt tracking more accurate"*, 2026-06-10](https://searchengineland.com/make-prompt-tracking-more-accurate-479708)) treats a single
observation as a point estimate. He recommends repeated measurement, with confidence intervals
reported alongside.

Academic preprints, a vendor's own published experiment, and trade-press methodology guidance.
Three kinds of author with three different reasons to publish, and on this question the four
documents above point the same way.

## What does buying a point estimate risk?

Suppose the report on your desk is a single-run measurement. Three specific risks follow. All
three are properties of the reporting format, not of any particular vendor.

**The baseline risk.** You commission a "before" measurement, spend a quarter on content and
technical work, then commission an "after." Both are single runs. Given documented sampling
variance of 10–34% (source: arXiv 2601.21339, 2026-01 preprint), the gap between them cannot
be attributed. You cannot tell your program's
effect from the measurement's own movement. The whole before-and-after story, the thing you
bought the baseline for, cannot be read.

**The accountability risk.** A point estimate can never be wrong. If next month's number is
different, that gets reported as "change." A number published *with* its confidence interval
is something else — a claim you can prove wrong. If a re-measurement under the same protocol lands outside the stated interval
more often than the confidence level allows, something is broken, and you can see it. Intervals
are what make a measurement vendor accountable to you.

**The allocation risk.** AI-visibility reports are built on comparisons. Your brand versus
competitors, one engine versus another, this month versus last. Budget then follows the
differences. But if the gap between two figures is smaller than the noise band around each,
moving spend on it is moving spend on static. Without a published range, you have no way
to know which differences are real. The same discipline applies to before-and-after "lift" claims.
We broke those down into [the
denominator questions that sort real metrics from vanity metrics](https://citeangle.com/en/research/ai-visibility-lift-claims-denominators).

## What does a procurement-grade number look like?

None of this means AI-visibility measurement is not worth buying. It means the *reporting
format* determines whether the number survives a finance review, a procurement
questionnaire, or a skeptical CMO. A number built for that room discloses four things. Treat
the four as a checklist and paste it straight into the questionnaire you already send.

- **The repetition count.** How many times was each prompt run? Stated, per protocol, not
 implied.

- **The interval, printed next to the rate.** "Cited in 33% of runs (95% CI 12–65%, n=9)"
 is a different object from "33%." The former tells you how much weight the number can bear.

- **The interval treated as information, not decoration.** A wide interval is a finding:
 *do not act on this number yet.* A narrow one is a different finding: *this is solid enough
 to build on.* Reports that publish widths honestly will sometimes tell you their own numbers
 are weak. That is the feature.

- **Enough of a paper trail to rebuild the number.** Measurement time, model and
 configuration, recorded so you can check the figure yourself instead of taking it on faith.

## Where do we stand against the benchmark we just cited?

We should apply the standard to ourselves before anyone else. CiteAngle's paid baseline runs
every query in the panel **[seven times](https://citeangle.com/en/methodology#why-7)**. It reports a
**Wilson 95% confidence interval next to every applicable rate**, with the effective sample
size beside it.

We know what one run costs because we ran a nine-brand panel that way and kept the output, a
Korean-market clinic-sector run of ten questions across six answer surfaces on 2026-07-17. At one run per
question the widest interval reached ±24.5 percentage points, and that width is why the panel flattened:
all nine brands landed in the top band on all six surfaces, a leaderboard reading 6/6 for everyone with no rank inside it.
The same design at seven runs stopped flattening them, and the panel spread across leadership
counts instead. Brand names stay out of the public write-up; the band spread is the part that
transfers.

**We print the measured interval width next to every applicable rate**, one by one, and
where a width is wide the report says so. The 40–150 benchmark counts distinct questions, not
repeats of one. Panorama asks fifty per platform, which sits inside that range, and the seven
repeats stack on top as a separate axis. Each observation also carries its receipt. Measurement time, model and configuration hash, sealed so you can check the numbers yourself. And the
external figures we cite publicly, including every figure in this article, are logged in our
[public claims registry](https://citeangle.com/en/claims) with source and access date.

### How many questions are enough?

The published academic estimate is roughly 40–150 *queries* per platform for a 95%
confidence interval five points wide on citation share. The paper's own numbers run from about
40–50 on one engine to 150 or more on another. Panorama asks 50 distinct queries per platform, at
the low end of that published range. It then repeats every one of them seven times, an axis the
paper treats on its own. The report does not claim the sample is big enough. It prints the measured interval width next to every applicable rate, one by one, so you can see how much
weight each number can bear before you lean on it. That width is the line most reports leave out,
and it is the one that tells you whether a change is real.

The spread, and the runs it takes: figures from the published record

Output variance from sampling alone (scale 0 to 100 percent) — source:
 arXiv 2601.21339, 2026-01 preprint

Output variance
 
 10–34%

Queries estimated for a ±5-point-wide 95% CI on citation share (academic estimate)

Academic estimate
 
 40–150 queries per platform

CiteAngle Panorama
 
 50 queries per platform, the low end of that range, as stated in the text

Sources: arXiv 2601.21339 (variance) · arXiv 2603.08924 (required-queries estimate) ·
 links in the text above · evidence accessed July 16, 2026 · bar length proportional to each
 panel's maximum (zero baseline) · figures logged in our [public claims
 registry](https://citeangle.com/en/claims).

The full protocol is written up on our [methodology page](https://citeangle.com/en/methodology): how
many repeats, how the interval is built, what each status means, and what the receipt holds.

## Start with the distribution, not the point

If you want to see what this looks like on your own brand, start with the
[free AI visibility check](https://citeangle.com/en/services). It is a single-pass read, and the result
carries that label on its face. What it hands you is the shape of your position. Which engines
answer the question at all, which sources they pull from, and the format the report arrives in.
Repeated runs, intervals and sealed receipts are what the paid baseline adds. When the number finally goes into a budget decision, it goes in with its uncertainty attached, the way a number you are about to spend
against should.

## Sources

- [arXiv
 2603.08924](https://arxiv.org/html/2603.08924) (2026, preprint): single-run AI-visibility measurement "fundamentally
 unreliable"; ~40–150 queries per platform estimated for a 5-point-wide 95% CI on citation
 share (§5.7; the unit is distinct questions, not repeats of one). Accessed 2026-07-16.

- [arXiv
 2601.21339](https://arxiv.org/abs/2601.21339) (2026-01, preprint): 10–34% LLM output variance from sampling alone. Accessed
 2026-07-16.

- [Profound official blog, "Is once a day enough?"](https://tryprofound.com/blog/is-once-a-day-enough) (2026-07-08): once-a-day
 reading methodology; own experiment, in which ten repeated runs cut day-to-day citation-share
 noise by ~40%. Accessed 2026-07-16.

- [Search Engine Land, Kevin Indig, "Make prompt tracking more
 accurate"](https://searchengineland.com/make-prompt-tracking-more-accurate-479708) (2026-06-10): single observation treated as a point estimate; repeated measurement
 with confidence intervals recommended. Accessed 2026-07-16.

**What one run costs.** Without repeats, a swing in the answers is indistinguishable from a
result, and good work risks being cancelled on a number that was never measured twice. Every
question here is asked 7 times and sealed with its conditions, which is why
our deltas arrive as two ranges you can see separate.

## Why the repeat count is the thing to ask about

Everything expensive about this failure happens quietly. A single-run figure goes into a budget
request, the budget is approved, and the following quarter's single-run figure comes back lower
— so the work gets judged a failure and stopped, when the honest reading is that nobody ever
measured anything twice. Careers get shaped by that arithmetic. No error message is produced.

Our answer is a published number. Every question is asked
7 times, and the count is printed on the report instead of living in a
methodology conversation. Rates ship with a Wilson 95% interval, so a delta is reported as two
ranges and you can see whether they even separate before anyone calls it a result. Ask any vendor
in this market how many times they ask. Whether the answer arrives as a number or as a sentence
about rigour tells you most of what you need.

The next number you buy should carry its interval

See the format on your own brand first. A free AI visibility check, then a baseline that ships repeated runs, Wilson 95% intervals and sealed receipts.

[Get your free AI visibility check](https://citeangle.com/en/#snapform)[See the measurement protocol](https://citeangle.com/en/methodology)
