Research · Evidence review Sources & citations

What actually gets cited by AI search: a sober review of the controlled evidence

The GEO advice market runs on writing tricks. Add statistics, add quotes, sound authoritative, get cited. One controlled study ran 54 page edits and found 3 that moved anything; another pooled 252,000 runs and put retrieval and position at the top. Being cited turns out not to be the same event as shaping the answer. Here is what the strongest public evidence supports, with every primary source linked.

Aerial view of a river delta at sunset, many channels converging toward the sea
An illustrative photograph for the argument above. It does not depict measured data.

What counts as evidence here — controlled benchmarks and peer-reviewed work, plus platform docs, reviewed as of July 14, 2026. Vendor-published statistics were left out by design. Three controlled studies carry the weight: 252,000 pooled runs, 54 page edits of which 3 moved anything, and a peer-reviewed check of four commercial systems (EMNLP 2023) where only 51.5% of AI sentences held up against their citations.

What actually gets cited by AI search: two things decide most of it. Whether the engine pulls your page in at all, and where your page sits in the pile of text the model reads. Most rewrites of the wording move nothing you can measure. And getting cited is a different event from shaping what the answer says. One study looked at 602 prompts, 21,143 citations and 18,151 cited pages, and found that how often a page was cited and how much it shaped the answer came apart. The EMNLP 2023 human evaluation measured only 51.5% of generated sentences fully supported by their citations (74.5% of citations supported their sentence). This review covers controlled benchmarks, peer-reviewed studies and platform docs only. Statistics published by vendors were excluded by design, and every primary source is linked.

What happened when 54 page edits were tested? Three effects came out

The hardest controlled test of "conversational SEO" tactics we read is C-SEO Bench. It is a benchmark accepted at NeurIPS 2025 that evaluated 54 style and content edit conditions across two tasks and six domains. Only 3 of the 54 conditions produced a significant positive improvement in citation ranking. Official primary arxiv.org/abs/2506.11097

The same benchmark found something the advice market rarely mentions. Moving a document toward the front of the text the model reads produced a far larger effect than any of the content edits tested. And when several competing documents adopted the same trick, the gains shrank. A tactic that works because nobody else uses it is not a strategy. It is a window, and it closes. Official primary arxiv.org/abs/2506.11097

C-SEO Bench: 3 of 54 edit conditions showed a significant positive effect
C-SEO Bench (NeurIPS 2025): 54 style and content edit conditions tested; 3 showed a significant positive effect on citation ranking. Which squares are highlighted is illustrative. The point is the ratio.

What gets a page cited as a source? The same factors lead across 252,000 runs

A second controlled study, "What Gets Cited," ran 252,000 trials across 6 LLMs. It fed the models two candidate documents at a time and varied 18 factors one at a time, with brand names hidden and the order of documents balanced. Two factors led by a wide margin: how closely the page matched the question, and where the page sat in the text the model read. Clear pricing information and recent dates helped a little, and helped consistently. Changing only the formatting did nothing worth measuring, and that is the cheapest, most heavily marketed lever on the market. Official primary arxiv.org/abs/2605.25517

Read those two results together and an order falls out. Getting pulled in as a candidate at all, and landing early, decide most of it. Facts a reader can check, such as prices and current dates, help at the edges. Cosmetic rewriting is noise. If your budget runs in the reverse order, the controlled evidence says you are paying for the weakest lever first.

Rank vs share of voice: do the GEO studies actually contradict each other?

Anyone reading the GEO literature side by side hits an apparent contradiction. The original GEO paper, the KDD 2024 study that named the field, reported that adding quotations, statistics and source citations improved a source's visibility in generated answers. C-SEO Bench, reviewed above, then found that almost none of the tested content edits improved citation ranking. Same tactics, opposite headlines. Official primary arxiv.org/abs/2311.09735

The resolution is that the two studies score different events. The original GEO paper measures share of voice: how much of the answer draws on your source, weighted by where in the answer it lands. C-SEO Bench measures citation rank: whether your document moves ahead of competing documents in the citation order. A content edit can stretch what the answer says with your material without moving you past anyone. The reverse happens too. The C-SEO Bench authors draw this distinction themselves when they discuss the earlier results. Read on different rulers, "it works" and "it does nothing" stop being a contradiction. Official primary arxiv.org/abs/2506.11097

For a buyer, the lesson is the one this review keeps landing on. When two studies disagree, or two vendors do, ask what each one is counting before you ask who is right. A number without its definition is not evidence yet. It is a headline.

Cited, absorbed, supported: are they the same event?

Even a citation is not the finish line. One study looked at 602 prompts, 21,143 citations and 18,151 cited pages. It found that how often a source gets cited and how much it shapes the answer come apart. Pages with real pull on the generated text tended to be longer, well structured, and dense with definitions, numbers, comparisons, and procedural evidence. Preprint study arxiv.org/abs/2604.25707

And the citation link itself can overstate what the answer got right. Four commercial generative search systems went through a peer-reviewed human evaluation, published in Findings of EMNLP 2023. On average only 51.5% of generated sentences were fully supported by their citations, and 74.5% of citations actually supported the sentence they were attached to. That is a 2023 snapshot of earlier products, so it should not be quoted as a current error rate. It did establish the structural point. Having a citation and having the sentence hold up are two separate measurements. Official primary aclanthology.org/2023.findings-emnlp.467

Evaluating Verifiability (Findings of EMNLP 2023): citation presence vs. statement support Sentences fully supported by their citations 51.5% Citations that supported their linked sentence 74.5% Human evaluation of four commercial generative search systems, Findings of EMNLP 2023.A 2023 product snapshot, not a current error rate.
Citation presence and statement support pull apart, which is why they have to be measured as separate events instead of folded into one visibility score.
StudyScaleWhat it foundWhat it does not settle
C-SEO Bench (NeurIPS 2025)54 edit conditions, 2 tasks, 6 domains3 of the 54 lifted citation rank; context position mattered far moreWhether the same holds outside the benchmark's task set
"What Gets Cited"252,000 trials, 6 LLMs, 18 factorsTopical relevance and context position led; formatting-only edits were negligibleHow much any one page can move its own retrieval
The original GEO paper (KDD 2024)The study that named the fieldQuotes, statistics and citations lifted share of voiceNothing about citation rank, which is a different ruler

Each row restates a source linked in the sections above, with its own scale and date. The last column is the part most summaries drop.

How do you read this evidence like an operator?

Every result above is causal only inside its test bed. C-SEO Bench used fixed candidate documents and specific model snapshots; "What Gets Cited" measured first citations between two injected documents, with no live retrieval or indexing in the loop. That is what a controlled benchmark means. Cause and effect hold inside the test. The test makes no claim to be a ranking law of any live commercial engine. Anyone quoting these numbers as guaranteed field effects is stretching the evidence, and anyone waving them away is discarding the only clean causal data the industry has. The same run-to-run caution applies to any number you are handed: a single run is a point estimate, in benchmarks and in dashboards alike.

The takeaway you can act on. The strongest controlled evidence puts retrieval eligibility, topical relevance, and context position first; checkable facts like pricing and current dates second; and generic style rewriting last. And it shows that being cited, being absorbed into the answer, and being supported by the answer are three different events that move independently. A single blended "AI visibility score" hides exactly the distinctions the evidence says matter.

One more filter worth copying: this review used controlled benchmarks, peer-reviewed studies and platform docs only. We set aside studies built on data gathered and published by GEO tool vendors. Not because they are wrong, but because they cannot separate observation from sales motive, and effect claims deserve sources without a horse in the race.

That rule applies to us too, so the one first-party number here sits outside the review. In our own run of 50 US category and brand questions across all 15 answer surfaces, measured July 27, 2026, 8,662 citations resolved to 1,987 separate domains, and 1,088 of those domains, 54.8% of them, were cited exactly once. It describes the ground rather than testing what causes a citation. Cause is what the controlled studies above are for; the shape of the ground is what tells you how much of it a single page can realistically reach.

Buy measurement for the chain, not tricks for one link

The practical consequence of this evidence is an ordering discipline. Before you spend on content rewrites, find where your brand drops out of the chain. Are your pages retrieved as candidates at all, are they cited, do the answers absorb your facts, and do the statements about you hold up against your actual pages? Each stage has a different fix, a different cost, and a different payoff, and the controlled studies above show the stages are not interchangeable.

That chain is what CiteAngle measures. We watch real AI-search answers for the questions your buyers ask in your market. Appearance, brand mention and source citation stay apart as distinct measured events. You get the stage-by-stage picture the controlled evidence says you should be managing — so your next dollar goes to the link that is actually broken.

When a result from this kind of measurement is written up, the evidence write-up format keeps every figure tied to the grid it came from: baseline, re-measurement, window and scope. A citation gain reads as evidence, not a testimonial.

See where your brand stands in the citation chain. Which questions surface you, which answers cite you, and which competitors hold the source slots you want. See the delivered report format or review the measurement programs.

What the trick economy costs. Without measurement wide enough to see the chain, a year of rewriting can be lost with nobody able to say whether it failed or was simply invisible from where they looked. Our panel is sealed and repeated, competitors read side by side with you, which is why a null result here is information you can act on.

Why bring this to us

The practical cost of the trick economy is not the money spent on the tricks. It is the year you lose. A team rewrites a hundred pages in the recommended style, nothing observable happens, and there is no measurement in place capable of telling anyone whether the work failed or simply was not visible from where they were looking. By the time that becomes clear the budget is gone, and the argument for trying again has gone with it.

What we sell is the instrument that ends that loop. Questions, engines, run count and judging rule are fixed before the first reading and sealed with it, so the second reading is genuinely comparable, and not merely later. Rates arrive with the interval around them, which is what stops a swing in the answers from being read as a result. And the same team that reads the evidence goes on to do the work it points at, so a finding does not have to survive a handover to become a change.

Stop guessing what gets cited. Measure it on your own queries.

The controlled evidence points at sources and eligibility, not writing tricks. A measured baseline shows which sources answer your buyers' questions today, engine by engine.

Ready to turn US visibility evidence into a market plan? Tell us the US buyer, target state and decision date.

Request a written proposal