# What actually gets cited by AI search: a sober review of the controlled evidence

원문: https://citeangle.com/en/research/what-gets-cited-evidence-review

An illustrative photograph for the argument above. It does not depict measured data.

Published July 15, 2026 · CiteAngle Research · Evidence base reviewed as of
July 14, 2026 · Controlled benchmarks, peer-reviewed studies, and platform docs only. Statistics
published by vendors were excluded from this review by design · Figures cited here are
logged in our public [claims registry](https://citeangle.com/en/claims) 
Updated July 17, 2026:
direct-answer summary added · Updated July 22, 2026: rank-vs-share-of-voice reconciliation
section added

**What counts as evidence here** — controlled benchmarks and peer-reviewed work, plus platform docs, reviewed as of July 14, 2026. Vendor-published statistics were left out by design.
Three controlled studies carry the weight: 252,000 pooled runs, 54 page edits of which 3 moved anything, and a peer-reviewed check of four commercial systems (EMNLP 2023) where only 51.5% of AI sentences held up against their citations.

**What actually gets cited by AI search:**
two things decide most of it. Whether the engine pulls your page in at all, and where your page
sits in the pile of text the model reads. Most rewrites of the wording move nothing you can
measure. And getting cited is a different event from shaping what the answer says. One study
looked at 602 prompts, 21,143 citations and 18,151 cited pages, and found that how often a page
was cited and how much it shaped the answer came apart. The EMNLP 2023 human evaluation measured
only 51.5% of generated sentences fully supported by their citations (74.5% of citations
supported their sentence). This review covers controlled benchmarks, peer-reviewed studies and
platform docs only. Statistics published by vendors were excluded by design, and every primary
source is linked.

Contents

1. [Fifty-four edits, three effects](#sec-cseo)

2. [252,000 controlled runs](#sec-252k)

3. [Rank vs share of voice](#sec-rank-vs-sov)

4. [Cited, absorbed, supported](#sec-events)

5. [Reading it like an operator](#sec-operator)

6. [Buy measurement for the chain](#sec-chain)

Questions this page answers: [What happened when 54 page edits were tested?](#sec-cseo) [What gets a page cited as a source?](#sec-252k) [Do the GEO studies contradict each other?](#sec-rank-vs-sov) [Are they the same event?](#sec-events) [How do you read this evidence like an operator?](#sec-operator)

## What happened when 54 page edits were tested? Three effects came out

The hardest controlled test of "conversational SEO" tactics we read is C-SEO Bench. It is a benchmark accepted at NeurIPS 2025 that evaluated 54 style and content edit conditions across
two tasks and six domains. Only 3 of the 54 conditions produced a significant positive
improvement in citation ranking. Official primary
[arxiv.org/abs/2506.11097](https://arxiv.org/abs/2506.11097)

The same benchmark found something the advice market rarely mentions. Moving a document toward
the front of the text the model reads produced a far larger effect than any of the
content edits tested. And when several competing documents adopted the same trick, the
gains shrank. A tactic that works because nobody else uses it is not a strategy. It is a
window, and it closes. Official primary
[arxiv.org/abs/2506.11097](https://arxiv.org/abs/2506.11097)

C-SEO Bench (NeurIPS 2025): 54 style and content edit conditions
tested; 3 showed a significant positive effect on citation ranking. Which squares are
highlighted is illustrative. The point is the ratio.

## What gets a page cited as a source? The same factors lead across 252,000 runs

A second controlled study, "What Gets Cited," ran 252,000 trials across 6 LLMs. It fed the
models two candidate documents at a time and varied 18 factors one at a time, with brand names
hidden and the order of documents balanced. Two factors led by a wide margin: how closely the
page matched the question, and where the page sat in the text the model read. Clear pricing
information and recent dates helped a little, and helped consistently. Changing only the
formatting did nothing worth measuring, and that is the cheapest, most heavily marketed lever
on the market. Official primary
[arxiv.org/abs/2605.25517](https://arxiv.org/abs/2605.25517)

Read those two results together and an order falls out. Getting pulled in as a candidate at
all, and landing early, decide most of it. Facts a reader can check, such as prices and current
dates, help at the edges. Cosmetic rewriting is noise. If your budget runs in the reverse order,
the controlled evidence says you are paying for the weakest lever first.

## Rank vs share of voice: do the GEO studies actually contradict each other?

Anyone reading the GEO literature side by side hits an apparent contradiction. The original
GEO paper, the KDD 2024 study that named the field, reported that adding quotations,
statistics and source citations improved a source's visibility in generated answers. C-SEO
Bench, reviewed above, then found that almost none of the tested content edits improved
citation ranking. Same tactics, opposite headlines. Official primary
[arxiv.org/abs/2311.09735](https://arxiv.org/abs/2311.09735)

The resolution is that the two studies score different events. The original GEO paper
measures share of voice: how much of the answer draws on your source, weighted by where in the
answer it lands. C-SEO Bench measures citation rank: whether your document moves ahead of
competing documents in the citation order. A content edit can stretch what the answer says with
your material without moving you past anyone. The reverse happens too. The C-SEO Bench
authors draw this distinction themselves when they discuss the earlier results. Read on
different rulers, "it works" and "it does nothing" stop being a contradiction.
Official primary
[arxiv.org/abs/2506.11097](https://arxiv.org/abs/2506.11097)

For a buyer, the lesson is the one this review keeps landing on. When two studies disagree,
or two vendors do, ask what each one is counting before you ask who is right. A number without
its definition is not evidence yet. It is a headline.

## Cited, absorbed, supported: are they the same event?

Even a citation is not the finish line. One study looked at 602 prompts, 21,143 citations and 18,151 cited pages. It found that how often a source gets cited and how much it shapes the answer come apart. Pages with real pull on the generated text
tended to be longer, well structured, and dense with definitions, numbers, comparisons, and
procedural evidence. Preprint study
[arxiv.org/abs/2604.25707](https://arxiv.org/abs/2604.25707)

And the citation link itself can overstate what the answer got right. Four commercial generative search systems went through a peer-reviewed human evaluation, published in Findings of EMNLP 2023. On average only 51.5% of generated sentences were fully supported by their
citations, and 74.5% of citations actually supported the sentence they were attached to. That
is a 2023 snapshot of earlier products, so it should not be quoted as a current error rate. It
did establish the structural point. Having a citation and having the sentence hold up are two
separate measurements. Official primary
[aclanthology.org/2023.findings-emnlp.467](https://aclanthology.org/2023.findings-emnlp.467/)

Sentences fully supported by their citations
 
 
 51.5%
 Citations that supported their linked sentence
 
 
 74.5%
 Human evaluation of four commercial generative search systems, Findings of EMNLP 2023.A 2023 product snapshot, not a current error rate.

Citation presence and statement support pull apart, which is why they
have to be measured as separate events instead of folded into one visibility score.

| Study | Scale | What it found | What it does not settle |

|---|---|---|---|

| C-SEO Bench (NeurIPS 2025) | 54 edit conditions, 2 tasks, 6 domains | 3 of the 54 lifted citation rank; context position mattered far more | Whether the same holds outside the benchmark's task set |

| "What Gets Cited" | 252,000 trials, 6 LLMs, 18 factors | Topical relevance and context position led; formatting-only edits were negligible | How much any one page can move its own retrieval |

| The original GEO paper (KDD 2024) | The study that named the field | Quotes, statistics and citations lifted share of voice | Nothing about citation rank, which is a different ruler |

Each row restates a source linked in the sections above, with its own scale and
date. The last column is the part most summaries drop.

## How do you read this evidence like an operator?

Every result above is causal only inside its test bed. C-SEO Bench used fixed candidate
documents and specific model snapshots; "What Gets Cited" measured first citations between two
injected documents, with no live retrieval or indexing in the loop. That is what a controlled
benchmark means. Cause and effect hold inside the test. The test makes no claim to be a
ranking law of any live commercial engine. Anyone quoting these numbers as guaranteed field
effects is stretching the evidence, and anyone waving them away is discarding the only clean
causal data the industry has. The same run-to-run caution applies to any number you are handed:
[a single run is a point estimate](https://citeangle.com/en/research/single-run-point-estimate), in
benchmarks and in dashboards alike.

**The takeaway you can act on.** The strongest controlled evidence puts retrieval
eligibility, topical relevance, and context position first; checkable facts like pricing and
current dates second; and generic style rewriting last. And it shows that **being cited, being
absorbed into the answer, and being supported by the answer are three different events** that
move independently. A single blended "AI visibility score" hides exactly the distinctions the
evidence says matter.

One more filter worth copying: this review used controlled benchmarks, peer-reviewed
studies and platform docs only. We set aside studies built on data gathered and published by GEO tool vendors. Not because they are wrong, but because they cannot separate
observation from sales motive, and effect claims deserve sources without a horse in the race.

That rule applies to us too, so the one first-party number here sits outside the review.
In our own run of 50 US category and brand questions across all 15 answer surfaces, measured
July 27, 2026, 8,662 citations resolved to 1,987 separate domains, and 1,088 of those domains,
54.8% of them, were cited exactly once. It describes the ground rather than testing what causes
a citation. Cause is what the controlled studies above are for; the shape of the ground is what
tells you how much of it a single page can realistically reach.

## Buy measurement for the chain, not tricks for one link

The practical consequence of this evidence is an ordering discipline. Before you spend on content rewrites, find where your brand drops out of the chain. Are your pages retrieved as candidates at all, are they cited, do the answers absorb your facts, and do
the statements about you hold up against your actual pages? Each stage has a different fix, a
different cost, and a different payoff, and the controlled studies above show the stages are
not interchangeable.

That chain is what CiteAngle measures. We watch real AI-search answers for the questions your buyers ask in your market. Appearance, brand mention and source citation stay apart as [distinct measured events](https://citeangle.com/en/methodology#why-7). You get the stage-by-stage picture the controlled evidence says you should be managing — so your next dollar goes to the link that is actually broken.

When a result from this kind of measurement is written up, the
[evidence write-up format](https://citeangle.com/en/case-study) keeps every figure tied to the grid it came from:
baseline, re-measurement, window and scope. A citation gain reads as evidence, not a testimonial.

See where your brand stands in the citation chain. Which questions surface
you, which answers cite you, and which competitors hold the source slots you want.
[See the delivered report format](https://citeangle.com/en/samples) or
[review the measurement programs](https://citeangle.com/en/services).

**What the trick economy costs.** Without measurement wide enough to see the chain, a year
of rewriting can be lost with nobody able to say whether it failed or was simply invisible from
where they looked. Our panel is sealed and repeated, competitors read side by side with you,
which is why a null result here is information you can act on.

## Why bring this to us

The practical cost of the trick economy is not the money spent on the tricks. It is the year
you lose. A team rewrites a hundred pages in the recommended style, nothing observable happens,
and there is no measurement in place capable of telling anyone whether the work failed or simply
was not visible from where they were looking. By the time that becomes clear the budget is gone,
and the argument for trying again has gone with it.

What we sell is the instrument that ends that loop. Questions, engines, run count and judging
rule are fixed before the first reading and sealed with it, so the second reading is genuinely
comparable, and not merely later. Rates arrive with the interval around them, which is what
stops a swing in the answers from being read as a result. And the same team that reads the
evidence goes on to do the work it points at, so a finding does not have to survive a handover to
become a change.

Stop guessing what gets cited. Measure it on your own queries.

The controlled evidence points at sources and eligibility, not writing tricks. A measured baseline shows which sources answer your buyers' questions today, engine by engine.

[Get your free AI visibility check](https://citeangle.com/en/#snapform)[See the measurement protocol](https://citeangle.com/en/methodology)
