US Measurement Methodology

Choose any state. Preserve every denominator.

A reproducible en-US protocol for national baselines and state-level targeting without pooling incompatible surfaces — every applicable rate ships with its confidence interval and a receipt you can verify.

ChatGPT · Perplexity · Claude · Grok · Google AI Mode

Numbers a board can act on, and the work that moves them

One standard covers what we read and what we change, and you can verify it.

Want to see this standard on your own brand first? Check one page free →

Why any of this has to be written down. Picture the report on the table. A citation rate sits on the front page and somebody asks the obvious thing: out of what? If the document cannot answer that in one line, the room moves to the next slide. Nothing gets flagged. The number quietly stops being used, and the work behind it goes with it.

Three gaps do that damage, and they compound. Without a written denominator, a rate that rose because fewer observations were read and a rate that actually rose look identical on the page — and you cannot decide a budget between them. Without a fixed run count, last month and this month are two different measurements, so the gap between them costs you the very decision it was collected to inform. And without the conditions stored beside the reading, the figure is lost the moment anyone asks where it came from, which is usually the moment it mattered most.

So why bring the question here. Every part above is written down: the denominator, the run count, the judging rule, and the conditions on each capture, sealed at the moment it was taken, because that seal is the reason the same setup can still be read a year from now. Observations we planned and could not read are printed as a count and left in the denominator, because taking them out is the cheapest way to make a rate look better, and that is a lever we would rather not own. Your brand is read side by side with your competitors, inside the same answer, on the same question, in the same run, which is what turns a score into a place on a map. And we put our own pages through this gate before we publish them — that is the reason we can print the limits, in our own words, in the table below.

The method is a discipline, stated plainly. Five promises make the numbers hold. Judgment kept apart from collection. Each rate read only against the asks we could judge. Seven runs per question. A receipt on each capture. And a public registry behind each figure we print. That is why a CiteAngle report can go straight into the room where the call gets made. The same rules then carry the work it orders: the pages and content we change, the brand facts engines repeat about you, the social and paid channels we run, and a second reading on the grid that found the gap.

17

places we read on each audit tier — six AI answer engines, Google AI Overviews, Bing Copilot and DuckDuckGo’s search assist, plus organic results on Google, Yahoo, Bing and DuckDuckGo, Google and Bing news, YouTube and short-form video

Each surface read as itself · the math counts 15 engine captures

7

runs per question, so we catch how much an answer moves instead of guessing

Standard on each paid audit

95%

Wilson score confidence intervals on the citation rates we can compute — a range that says where the true number most likely sits

Reported with the effective sample size

50

buyer questions in a full Panorama — each one put to 15 engine captures, 7 runs, and reported with its status

Compass asks 20, Radar 5 — same rules, three question widths

Judging kept apart from collecting
A citation is checked against the source. Mention and tone are judged on their own track, and we label how sure we are.
Where each number came from
Each run carries interface, model, location, capture time and source receipt — so the same setup can be read again.

The verdict is yes or no — a binary rule. There or not there — a citation, a mention — not a blended score out of ten. Outside research on judge reliability finds binary calls stay stable under stress where multi-point scales do not (RAND, arXiv 2603.05399).

A measurement premise — we read the baseline response, before personalization. Our runs are logged out and hold no state, so the numbers describe what each place shows before it knows anyone. One person's answer can differ with account state and history. That is exactly why we lean on repeated reads under matched setups, and on a series over time rather than a single capture. This API baseline is measured with no persona or memory in play; on real user screens, the swing between answers runs widest for brands in the middle of the pack for awareness.

Baseline visibility is the definition, not a disclaimer. The stateless setup we read sits in the same layer the platforms themselves call "before personalization". OpenAI documents Temporary Chat as a mode that neither creates nor uses memories (OpenAI Help Center), and Google describes Gemini's temporary chats the same way (Google blog). Each real answer starts here and then branches with account state and history. We have found no published work where personalization rescues a brand that is invisible at the baseline. Outside benchmarks pull the other way. In one, models follow a user's plainly stated wishes less than 10% of the time inside ten turns of chat (PrefEval), and the best model in a long-running chat test scores 52.8% (HorizonBench). The baseline is where answer presence is won first.

ChatGPT has since opened opt-in location sharing (OpenAI release notes, 2026-03-26), so an answer in the app can now turn on where the user is. What we read here is the stateless baseline. Reads that pin a location are sold as a grid option and kept on their own count.

The comparison sheet you were going to build

Ten rows, our column already filled in — including the parts that do not flatter us.

Buying in this category tends to end the same way. You open a spreadsheet, put vendors across the top and questions down the side, and then spend a month trying to get straight answers into the cells. So we keep that sheet here with our own column done. Lift the rows into your own file and ask everyone the same ten things, us included. Each row carries three parts: what we do, what the number is counted out of, and where it stops being useful. That third part is the one that rarely survives a sales call, which is why it sits in the table rather than in a footnote.

The row on your sheetOur answerCounted out of — and where it stops
Places readSeventeen places on every paid audit tier: six AI answer engines, Google AI Overviews, Bing Copilot and DuckDuckGo’s search assist, plus organic results on Google, Yahoo, Bing and DuckDuckGo, Google and Bing news, YouTube and short-form video.The grid arithmetic counts 15 engine captures, because one Google capture is read as two places. A door outside the roster for that market is not measured, and we write that down instead of scoring it as absent.
Runs per questionSeven, on every paid audit. The run count ships locked inside the read’s condition hash, so it travels with the evidence rather than with the invoice.Seven runs buy a measured range, not a precision claim. Widths print at whatever they come to, and the wide ones stay wide on the page.
Grid sizeA full Panorama asks 50 buyer questions, 7 times each, on 15 engine captures — and every ask is made. The narrower tiers run the same rules on shorter question panels.That is the number of asks, not the number of answers. Observations that failed, came back blocked or could not be judged keep their own status and stay out of the rate.
What counts as a hitA yes-or-no call, checked against the source. Citation and mention run on separate tracks, and collection is kept apart from judging.A blended score out of ten would hide which half moved. The cost of a binary call is that a near miss reads the same as a miss, so the report shows the answer text beside the verdict.
UncertaintyA Wilson 95% interval and the effective sample size beside every rate we can compute.An interval covers sampling noise. It says nothing about whether the question panel was the right one — that part is a judgment, and it is ours to defend in the report.
Provenance per observationInterface, model, location, language, device where it applies, cache policy, capture time, run number and a content-addressed receipt.A receipt proves what came back and when. It cannot prove why an engine picked what it picked, so the report separates what we saw from what we infer.
Checking it yourselfThe formula behind each published rate, its interval and the exact command to compute it again all sit on the public claims registry.What is open is the measurement standard and the sums behind the printed numbers. The delivery playbook — which queries to enter first, and in what order — ships inside the engagement.
Markets and languages3 locales — United States, Japan and Korea. Each one is written in its own language rather than translated, and a machine contract fixes what every locale’s pages have to carry.A checker compares the 3 against that contract on every run, so a page cannot go missing in one market without something turning red. It compares structure, not writing, so a person still reads each locale.
What stops the method drifting189 improvement rules live in one ledger, and every one of them carries the incident or the source that put it there. On each deploy a gate reads the ledger and checks 141 of them against the pages we actually render — that is the count you can check yourself, by opening the page. The remaining 48 look at our side instead: whether the instrument that measures a technique is in the repository at all. You can also turn the same kind of page check on an address of your own. Check one page free: paste a URL, skip the sign-in, and the scoring table opens under the result.The gate checks that a technique is present, not that it worked. It reads rendered HTML only, and it knows nothing that is not written in the ledger. Those three edges are ours to watch by hand.
Who checks usEvery figure printed on this site is filed at the same value in the claims registry, and runs that came out against us stay in.Filing is something we do, not something a third party did to us. The receipts are checkable by anyone; the discipline of filing them is a practice, and you should ask any vendor how theirs is enforced.

If a row on this sheet matters to you and the wording here is not precise enough to paste into a procurement document, write to support@citeangle.com and we will send the exact contract language for that row.

One run is a point estimate. Seven runs make a range.

AI answers vary run to run. A single run shows you one answer out of many. So we ask each question again and again, and we print the range next to each rate we can. You see how solid each number is before you act on it, and the same range is what tells us whether our own work moved it later.

Published by a vendor itself

Asking again cuts the noise, and it shows

Profound, a global AI-visibility tool, lays out its once-a-day reading method on its own blog — and their own experiment found that ten repeated runs cut day-to-day noise in citation share by about 40% (Profound blog, 2026-07-08).

Published in the research

Single runs are point estimates

One academic study finds single-run AI-visibility reads “fundamentally unreliable” (arXiv 2603.08924). A second finds LLM output shifting by 10–34% from sampling alone (arXiv 2601.21339). Trade guidance says to read more than once and print the range beside the number (Search Engine Land, 2026-06-10).

Where US buying already sits

The answer forms before the click

Pew Research Center measured US Google visits: people who met an AI summary clicked a normal result link on 8% of visits, against 15% without one — and clicked a link inside the summary on just 1% of such visits (Pew Research Center, 2025-07-22). And G2’s The Answer Economy report — from G2, which itself has a stake in AI-search visibility — finds 51% of surveyed B2B software buyers now start research with an AI chatbot more often than with Google (n=1,076, March 2026 survey; announcement).

Interval-width convergence — runs 1 → 7computed at a fixed report-format example rate of 33% · a schematic, not a live feed

One run can only say “somewhere in a very wide band.” Ask again and the band shrinks. That is why each rate we can report carries its measured Wilson 95% interval and effective sample size, printed at whatever width it comes out. Worked out with the Wilson score interval (z=1.96); the upright mark is the fixed example rate. Full rules in the method.

Our own rules — seven runs, Wilson 95% confidence intervals, widths printed as they come out — are written up in the method. Every outside figure above is listed with its source and date on the registry of published claims.

Why repeated runs — the published evidence

The case for asking again is public, and we hold ourselves to it in the open: we publish the confidence interval and effective sample size next to every applicable rate, and widths go out as they come out — wide ones stay wide.

The public proof, source by source — four separate lines
Single runs are point estimates
One academic study finds single-run AI-visibility reads “fundamentally unreliable”, and puts the cost of a five-point-wide 95% confidence interval on citation share at 40–150 queries per platform (arXiv 2603.08924, §5.7 — distinct questions, not repeats of one). That is one paper’s estimate, not a settled standard, and it sits on a different axis from how often we repeat a question. Our seven-run default sits where an outside convergence study saw single-run error (±0.37) first fall under the 0.10 stability line. We print measured widths with each rate instead of claiming that seven is enough.
The swing is measured, not guessed at
LLM outputs shift by 10–34% from sampling alone (arXiv 2601.21339).
Trade guidance points the same way
Search Engine Land treats one observation as a point estimate and tells you to read again, with the range printed beside the number (Kevin Indig, 2026-06-10).
Even a daily-reading tool measured it
Profound’s own experiment found that ten repeated runs cut day-to-day noise in citation share by about 40% (Profound blog, 2026-07-08).

Why seven runs — the precision ladder

Ask once and the range maths says almost nothing. Worked out with the same live Wilson function our reports use (95%, z=1.96), a query cited on its single run ([20.7%, 100%]) and one never cited ([0%, 79.3%]) give ranges that overlap. You cannot tell them apart. At seven runs they split: the 7/7 lower bound (64.6%) clears the 0/7 upper bound (35.4%), so always-cited and never-cited fall into different bands.

Our own rerun shows the same thing on live data. Read a market once and the bands stretch wide enough to overlap: every brand on the panel sits in the top band, on every surface, tied. That is not a ranking, it is a coin toss with a chart around it. Run the same panel with the repeats in place and the field separates — the top band thins out, and the leaders stop being the same names from one surface to the next.

Outside research puts a number on the same design point. A university convergence study of four AI search engines (Schulte et al. 2026, arXiv 2604.07585) measured the standard error of a single-run brand-detection read at ±0.37 — coin-flip ground. Stack runs and the error falls. The curve first crosses the stability line, a standard error under 0.10, at seven runs. That is exactly where our default sits. Each paid audit asks each query seven times per surface. The run count ships locked inside the read’s condition hash — a default, not an option. Published rates stand on the whole set that comes back — on Panorama that is 50 questions read on every surface, seven times each — each rate with its confidence interval.

Past seven, the returns bend. Forty runs cost 5.7× the work, and the range narrows by less than that either way you count it. In the always-cited case charted above it is 4.0×. Under the square-root law that holds in the middle of the range, about 2.4× (√5.7). So the next step up is not more runs. It is more waves on the same k7 protocol. A monthly re-read on the matched rules adds the time axis a single audit cannot have. A real shift shows up as ranges that pull apart, not as one number that moved. Every in-house figure in this section can be worked out again from the live aggregation function. Outside research figures carry their source inline. The sums and the commands to run them again sit on the claims registry.

Runs come standard — and we build precision past them

Three separate lines of 2026 research converge on the same design point, and we hold both halves of it: asking a query more than once is necessary, and asking again on its own is not sufficient.

The three research lines, and why asking again is not enough on its own
Reruns drift before a word changes
Run the same prompt again — not one word changed — and the lists that come back overlap at a Jaccard of only 0.50–0.61, in a live study of OpenAI and Anthropic models (arXiv 2605.27440). One capture is shaky before wording even enters the picture.
The pipeline is stochastic by nature
A 2026 critical survey recasts generative-engine optimization as a “stochastic, partially observable pipeline” rather than one ranking task (arXiv 2607.14035) — so the honest unit is a distribution, not a point.
Even temperature zero is not deterministic
With greedy decoding at temperature 0, one model (Qwen3-235B-A22B) still gave 80 distinct answers across a thousand runs, from batch-level nondeterminism alone (Thinking Machines, 2025-09) — a swing the user cannot switch off.
…and yet repetition is not sufficient
Because that noise stays, seven runs do not buy a precision claim — they let us report the measured range. The next step up comes from widening the surface and query axes and from more program waves, not from stacking runs.

The standard is public — formulas, ranges and the commands to check them

As of July 2026 the trade still has no agreed metric for AI-recommendation visibility. A recent academic map says plainly that “no standardised metrics exist” for it (arXiv 2606.23057). We treat that gap as a reason to open the box. The formula behind each published rate, its Wilson confidence interval, and the exact command to run it again all ship on the public claims registry. One reference guide argues the same design point we build on. Track presence per surface and per model rather than folding it all into one blended “share of model” score (graph.digital). The playbook for the work itself — which queries to enter first, and in what order — ships with the engagement, not on this page.

What SEO means here: crawl, indexing and rendering kept clean, gated on every deploy of our own site; plain search places read on the same grid, with the same runs, as AI answers; and the hand-off from search visibility to AI citations, measured in our own data. Each number carries a receipt you can check.

The US surface rails we measure

A few doors carry most of US search, and we weight our rails to match. As of June 2026, StatCounter share puts Google at 86.67% and Bing at 8.73%, with Yahoo Search at 2.55% and DuckDuckGo at 1.53%. So our US search rail leads with Google Search and Bing Search, and the smaller doors stay in as second-line checks under the en-US contract.

Plain search is read in the same grid, not bolted on. Google Search and Bing Search results run as their own surfaces beside the AI answers. Search-engine optimization is home ground for this method. So where you stand on the results pages buyers still use is read side by side with AI citation, query for query.

US search

Google Search and Bing Search lead; Yahoo Search and DuckDuckGo as checks

Google AI Overviews and plain results are read on Google Search. Bing Search carries its own AI answer surface. Yahoo Search and DuckDuckGo run on the same queries as door checks, so the smaller doors show up in your numbers next to the big ones.

US AI answers

Named AI-answer interfaces

ChatGPT, Perplexity, Claude and Gemini are logged by exact interface, target, receipt and evidence namespace. Location-capable direct API rails and approved regional observation are logged the same way — official API, consumer chat UI and clean-session tests each on their own rail.

Where it came from

The exact name of what we read

Each run records interface, model, location, language, device where it applies, cache policy, capture time, run number and source receipt — so the same setup can be read again.

Each observation keeps its exact interface label and a content-addressed receipt. Official API, consumer chat UI, clean-session tests and causal checks each use their own evidence rail and their own count.

The national baseline and state targets stay separate

United States national baseline

The national panel keeps its own query set, interfaces, runs and count — the one view of where the brand stands across the market.

50 states + Washington, DC

A 51-unit add-on maps state-by-state AI answers, mentions, direct citations and search openings on its own count, read next to the national headline.

State selection fixes the intended target. A verified state read then adds location-capable API or exact browser receipts, contracted coverage and acceptance gates before a state or DC score goes out. Each state score is reported on its own count.

Measurement lead

A measurement firm that hides its own numbers has nothing to sell. Each figure printed on this site is filed at the same value in the claims registry. The numbers behind it can be pulled from the public dataset. Runs that came out against us stay in. We widen the range and say where it widened.

— Jay Sim, measurement lead, CiteAngle

Jay Sim — CEO of ARTIFEX Co., Ltd. and CiteAngle's measurement lead. He owns the measurement system end to end: query design, repeated-run operation, judging rules and final report review. Before anything goes out he checks that each public number matches the claims registry. The author line on CiteAngle research articles points to that job — contact support@citeangle.com.

Everything published under that byline is collected on the author page.

Methodology questions, answered

Why do you ask the same query more than once?

Because one run only tells you where a single draw landed. We ask the same locked question again and again to catch the citation, the mention and the swing between answers. And we publish the confidence interval next to every applicable rate. The public case for asking again is gathered above. The sum that shows seven runs pulling always-cited apart from never-cited is public in Why seven runs, worked out with the same live function our reports use.

Does a logged-out read match what real users see?

What we read is the baseline, before any personalization — and holding that baseline still is exactly what makes the numbers line up run to run. One person's answer can differ with account state and history, which is built into the rules. That means matched no-state setups, repeated runs and a series over time instead of a single capture (see the measurement premise above). A held baseline is the surface where before-and-after change can be read with the same ruler.

Can an outsider check the numbers?

Yes — the product is built to be checked. Each statement behind our public copy sits on the public claims registry with its source type and its date. Each run we deliver carries its provenance — interface, model, capture time and settings — with the run count, 7 reads of the same question, locked inside the read's condition hash. The integrity layer checks by hash, the same way every time, whenever you ask. The method is public and every figure carries its receipt, so you can confirm every observation on your own.

There is a second layer, and it promises something different in kind. Asking an engine again today draws a fresh sample rather than replaying the old one. Repeat a run inside the reproduction window written into your report and the confidence ranges should overlap — that is the claim we stand behind on this layer. Run it after that window closes and we file the gap as drift and label it as drift, because by then the engine itself has moved. So the deterministic promise stays where it belongs, on the stored bytes and their hashes, and the live layer promises the method and the range instead.

Why can't rank trackers see AI visibility?

Because citation is decided at a different stage than ranking. In an outside US desktop study of buying queries, only 19% of Google AI Mode citations came from the same query's top-20 plain results (seoClarity, data 2025-09). Even a strong rank table leaves most citations decided off the table. That is why we read the answer surfaces head on, with a plain-search cross-check on the same queries, so you see where the chain breaks.

Where our own instruments were wrong

When the tool is wrong, every number built on it is wrong. So we keep a record of the times ours was, and of what we changed each time.

One sentence decides what goes on this page and what does not. If this number were removed, would another number we published lose its basis? If yes, it goes on the page. If no, it does not.

An accessibility checker died on our own pages, and the tally read that failure as zero violations.

How it surfaced. The same checker ran fine on competitors' pages. Ours failed all 123 times, yet the result table showed zero violations. The tell was that the error only ever leaned our way.

What changed. We injected the checker in a way our security settings allow, and all 246 checks ran. Counting a failure as a zero is now banned, and a result that flatters us gets doubted first.

2026-08-01

A pixel-based contrast check marked every site as failing, competitors included.

How it surfaced. Blurred pixels at the edge of each letter landed in the background sample. All 28 sites failed, not just ours, which meant the number could not decide anything.

What changed. Pass or fail now comes from a different check that reads each element. The pixel number is kept only for ranking sites against each other, with that limit written at the head of the table.

2026-08-01

We compared samples of different kinds — our 120 pages against 28 pages on the other side.

How it surfaced. Above-the-fold text came out at 79.31% against 0.96%. The gap was too wide to be real: ours was a whole site, theirs a handful of lead pages. Measured like against like, our pricing page ranked first of nineteen.

What changed. Every comparison now states what kind of sample the denominator is. A comparison that does not say whether it is whole-site or page-to-page is not used.

2026-08-01

Test-made observations sat in the pile of real ones, and the tally was quietly contaminated.

How it surfaced. A count came out at 1,400 where it had no business being. The shape matched the real records, so no tally caught it; plotting answer length is what finally showed a second population.

What changed. We did not delete the fakes, we tagged them twice. Deleting them would make the same mistake untestable. A separator now runs ahead of every tally, so nothing is counted while the two are mixed.

2026-08-01

We wrote a polarity test to prove that our suite-coverage checker notices when someone deletes a test runner. It had stopped testing anything. The line it targeted had moved into a helper earlier, so the injection matched nothing at all. Its own liveness guard missed that, because comparing the file before and after with `splitlines()` and a newline join drops the trailing newline — the two strings differ even when the filter removes nothing.

How it surfaced. The deploy gate failed on this single test, so we reproduced the injection by hand. It removed zero lines. The guard still reported a difference, and the checker returned the same three paths for the original file and for the supposedly damaged copy.

What changed. The guard counts removed lines now and requires exactly one. Reading keeps line endings so the trailing newline survives, and the target line sits in a constant, which gives a future refactor one place to update. We proved both directions: the web path appears for the real source and vanishes when we cut that line.

2026-08-11

Our graph search warned that its provenance evidence was older than the graph itself. The warning was on for every input. We measure evidence per root and then merge it, and the merge carries the base timestamp through, so the freshness comparison could never advance.

How it surfaced. The warning fired on all three evidence sets we had on disk, whose share carrying a source measured 99.31%, 98.28% and 72.52%. A warning that is always on cannot separate a healthy evidence set from a rotten one, so we retired the comparison. Retiring it stranded the two tests that read it. One could no longer fail. The other could no longer pass.

What changed. The pair measures coverage now, which is the question freshness stood in for: is anything rendering without a provenance grade? When every node carries a grade, the check must produce no warning. When one node in four has none, it must produce one. The threshold is configurable, so the tests lock the sentence rather than the number.

2026-08-11

In our state-drift gate, an undeclared anchor and a detected drift both exit with code 1. Both polarity tests expected that code and received it for the wrong reason. They passed while measuring nothing.

How it surfaced. The sibling tests started failing once the gate began requiring anchor declarations. Adding the declaration made the two quiet tests exercise the path their names promise. Where two failure routes return the same code, a test cannot report which one it saw.

What changed. Planting a synthetic fact goes through one helper now, and that helper also declares the anchor and disables the ratchets meant for the real list. All 18 tests in the file pass. We also found 40 registered facts against a known-risk list of 27, which left the denominator smaller than the numerator. We listed the 17 that were missing, each with the reason it drifts, and the list stands at 44.

2026-08-11

The US answer layer is forming now

An aperture opened wide over concentric measurement rings, its centre filled with gold — measurement posture, not guesswork.

What your buyers read before they click is no longer an experiment.

What put this shift on the record

At I/O 2026 Google said AI Mode reached 1 billion monthly users inside its first year, with query volume more than doubling each quarter (Google’s own announcement, 2026-05-19 — blog.google; platform self-reported figure). What your buyers read before they click is no longer an experiment. Where your brand sits in it can be read, sealed and read again on the same panel.

Bing's index reaches more than one door. It supplies Yahoo Search results under a long-standing partnership (Search Engine Land), sits under Microsoft Copilot, and — by multiple industry analyses — still shapes which sources ChatGPT search pulls up. Our US panel reads Bing and Bing Copilot head on, each on its own rail.

A US federal court has put this shift on the record. In the remedies opinion in United States v. Google (September 2, 2025), the court found that "GenAI chatbots grounded in general search perform an information-retrieval function that is similar to GSEs" — general search engines — and that chatbots "often include citations and links to websites when responding to information-seeking queries" (US v. Google, case 1:20-cv-03010, remedies opinion (PDF)). What this whole read stands on — AI answers citing and linking sources the way search does — is now written into a federal antitrust ruling, next to the industry theses. That ruling is under appeal, so we cite it as pending, not final.

What comes back from one read

You get names and addresses, not a grade out of a hundred.

Here is the shape of it. We put a buyer question to an engine. The engine writes an answer, and a few brands end up inside it along with the pages it leaned on. We record which brands got named, which pages got cited, and which engine did the naming, with the time it happened.

One answer proves nothing. Engines rewrite themselves through the day, so the same question goes back in, again and again. What you read is the spread across those runs rather than the single screenshot that happened to flatter someone.

By the end you are not holding an opinion about your visibility. You are holding a list: this question, that engine, the rival sitting in the slot, and the page of theirs the answer trusted.

Your own row is the part worth reading. Start it below.

The measurement detail behind the home page

The same sentences, in full, one screen deeper.

What Panorama asks50 buyer questions, each one asked again on 17 places
Observed Cited / mentioned Indeterminate Unmeasured A sketch — not a live feed.

01 · Read the rivals

Recommended, reviewed or only named

We name the role the answer gave your brand. Then we count how often each rival took that same slot in the same run.

Google Search
AI Overview
1
your brand
source2 your domainsource
Organic results
3
you

1 named · 2 cited as a source · 3 plain search rank — read apart.

That list is chosen, not fixed. A screen earns a place on it by passing four checks: people in this market actually read it, we can ask it the same way again next round, a name being said can be told apart from a link being cited, and we can reach a state that is clean of one person’s account history. Screens that miss a check go to the fix-and-publish side of the work instead. Because buyers in each country read different screens, every market gets its own list, and the list we agree on is written into your contract before the first run.

We hand back the pages they cited and the rivals they named, then fix those pages and run the campaigns.

Each run is sealed. When it ends we take a fingerprint of the setup and the raw data, so any later edit shows. Fourteen engines give fifteen places an answer can turn up, because one Google search returns both a results page and an AI Overview.

One run is a point estimate, not a number, so each rate ships with its range — the math is in the method. Every buyer query ends in one of four states, set out on the definitions page.

Layout sketch. Answer engines, their source rails and plain Google and Bing results each run on their own track. All fifteen places are named in the method.

A paid audit puts the same question to fourteen engines seven times over — five questions in Radar, twenty in Compass, fifty in Panorama, read across the fifteen places an answer can appear. Every one gets asked, and it names the competitors standing beside you. The report arrives one to three business days after measurement starts, and questions after that are answered in writing within two business days.

Because those four are linked, you check the numbers rather than our sentences.

Browse the public claims registry →

Fourteen engines, fifteen places — one Google search returns both a results page and an AI Overview.
Step three sets the order for the next round — each lap leaves a before and an after.
Same grid, three depths — only the number of questions changes.

One observation is one question, put to one place, one time. AI answers change from one day to the next, so the same question goes out 7 times and every reply is kept. A full Panorama plans 5,250 of these observations — and makes every single one of them.

snapform diagram — 1 · Pick your industry · 2 · Up to 3 questions · 3 · We run it for you

why diagram — 1 · Decide what to fix · 2 · Do the work · 3 · Measure it again

best-fit diagram — Free visibility check · same business day · Paid audit · the baseline you keep · Program · work done, then measured again

surface-map diagram — 1 buyer question · asked seven times · 14 · engines asked · 15 · places a name can appear · one Google = results page + AI Overview · named in the answer and cited as a source are logged apart

The free check runs one to three of your buyer questions across six places an AI answer shows up, and it starts the same business day you send it.

evidence chain diagram — Claims registry · statement + source + date · Measurement run · bound by run identifier · Sealed fingerprint · a later edit shows · Run again · same conditions

three depths diagram — 5 questions · 5 · Radar · 20 questions · 20 · Compass · 50 questions · bar length = buyer questions · each depth asks 15 engines, seven times per question · 50

Written by , measurement lead at CiteAngle. Every number on this page is checked against its source before it goes out. As of .

Ready to turn US visibility evidence into a market plan? Tell us the US buyer, target state and decision date.

Request a written proposal

The marker CTNG-CANARY-EN-20260725-7QF3 on this page is a unique string we publish so we can measure how long a published change takes to reach AI answers. It carries no other meaning.