What should you verify before buying AI visibility measurement? Line up what each vendor states on its own public pages, on the same axes. Published prices run from $20 a month to enterprise contracts with no price at all. Price is not the widest gap. The widest gap is how much each vendor says about how the number was made. That is what the grid below and the five questions cover (all vendor pages read 2026-07-16, vendor names as initials). Here is our own row on those axes. We run each prompt-surface pair 7 times. We print the measured Wilson 95% CI width in the report. We publish every planned observation in the grid with its status, the failures included.
How was this comparison built?
Methodology note. Every fact below comes from a vendor's own public pages: price pages, help centers, official docs. We read all of them on 2026-07-16. Where a page carries a date, we note it. Vendor names are shown as initials. CiteAngle wrote this page. We sell measurement and we sell the fixes. That is why a receipt rides along with every number we sell. "Not stated" means we did not find the item on the public pages we read. It is not a claim about what the tool can do.
How do AI visibility vendors compare on published facts? (public pages only, accessed 2026-07-16)
| Vendor | Published pricing | Engines measured | Cadence vs. same-prompt repeats & confidence intervals | Raw data access | Methodology & verification artifacts |
|---|---|---|---|---|---|
| Company S | AI add-on $99/mo/domain (stated on an annual-billing basis), 25 prompts; bundles $117–455 | Inconsistent across its own docs (help center lists 4; pricing page differs) | Daily/weekly refresh. No same-prompt repeats or CIs stated | Export/API not stated (per help docs) | Not stated |
| Company A | $199/mo per index (or $1,990/yr); $699/mo for all platforms ($6,990/yr, includes 2,500 custom prompt checks per month); check add-ons at $50, $100 and $250 Company A row re-checked 2026-08-02 | 7 engines indexed (one marked paused, one custom-prompts-only) | Chatbot data refreshed monthly (90-day window); some surfaces continuous. No repeats/CIs; self-described as "directional indicators" | Not checked in our re-verification scope | Not checked |
| Company B | Fully undisclosed; contact-sales only, with AI features bundled into all subscriptions | No single official list; 3 to 5 engines depending on the document (union: 6) | Collection cadence, sampling volume, repeats, CIs: none disclosed | Metrics APIs only; no evidence of raw response logs | A trademarked parser is named; how it works is not disclosed |
| Company C | Amounts undisclosed; three tiers with published quantity gates (credits, keywords) | 9 AI surfaces with a published per-surface credit price table | Daily/weekly/monthly per topic. No same-prompt repeats or CIs stated | Response viewing, XLSX, API; bulk raw export not specified | States it uses "official APIs"; no white paper or audit |
| Company P | $99 (1 engine) / $399 (3 engines) / custom enterprise | 11 engines on the features page vs. "up to 10" on enterprise, which is internally inconsistent | Daily, officially. Its own blog (2026-07-08) publishes a repeat-sampling experiment with error math, but no always-on CI display in the product is stated | Not checked | Not checked |
| Company O | $29/$189/$489/custom (billed by prompt count) | 7 engines = 4 base + 3 paid add-ons | Once daily, automatic. No repeats or CIs stated | Full responses + CSV/JSON + public API | Partial disclosure (measurement method stated for one engine); no standalone methodology page |
| Company E | Pro $800/mo (100k prompts analyzed); custom enterprise | 10 engines listed; tracks base-model APIs and consumer apps separately | States it samples each prompt 100 times per model, with a 30-day default refresh, and does publish margin-of-error figures (±9 pt at one reading, ±1 pt at 100) — though the interval method behind those numbers is not named | No reporting API (on the roadmap, per its docs) | Methodology described across scattered pages; no external audit stated |
| Company R | $20–780/custom (credit-metered) | 10 engines officially listed (a "17+" claim is never fully enumerated) | 8 monitoring schedules mentioned. No repeats or CIs stated | REST API from mid tier; raw response export unconfirmed | No methodology page (404) |
| CiteAngle (us) | Fixed, published one-time diagnostic pricing, with no retainer | 15 engine captures, read as 17 separate surfaces — the Google and DuckDuckGo captures each yield two (a result page and an AI answer) | 7 repeated runs per prompt-surface pair, with the measured Wilson 95% CI width printed in the report | Every planned observation published with 1 of 3 statuses: success, failure, or indeterminate | Sealed hash ledger plus a public /claims registry |
The grid rows rearrange what each vendor states about itself; they are facts, not rankings. Three more early-stage vendors were left out pending re-verification.
What do these services cost, and how often do they repeat a run? (the same grid, re-plotted)
Published price ranges, on one axis
USD, monthly subscription
Same-prompt repeats & confidence intervals, as stated on public pages
Stated Partial Not stated / not checked
Five questions to ask before you buy: which ones prove the number?
1. How many runs produced that number?
AI answers vary run to run. The movement is documented in academic work and acknowledged on at least one vendor's official blog. A single-run score is a point estimate. Ask for the repeat count and whether an error bound (a confidence interval) appears in the report itself. Across the pages we checked, some vendors state repeat sampling and some discuss error bounds in their methodology write-ups. What we did not find was a vendor that prints that width in every deliverable and pairs it with a sealed hash ledger and a public verification desk. The published record behind this question, meaning the variance studies and the run-count math, is collected in our evidence review of single-run measurement.
There is an arithmetic answer to this question, so we ran it through our own measurement function on 2026-07-17 instead of arguing about it. Repeats are the expensive axis. The work is questions × surfaces × runs, so moving from seven runs to forty multiplies the measuring work by 5.7. Sweep every observed rate and the interval it buys narrows by at most 4.0. Precision does not track spend here, which is why "as many runs as possible" is a worse answer than a fixed number written into the contract. Ask for the number, then ask what it was chosen against.
2. Do you get to see the failed measurements?
A report that only shows successful readings hides the denominator. Ask whether every attempted reading is disclosed with its status: success, failure, indeterminate. Denominators are also where lift claims go soft, so we turned that problem into denominator questions you can ask a vendor in writing.
3. Is the measurement sealed?
To check later that "this is what it said on that date," the deliverable needs to exist in a tamper-evident form: hash ledgers, a public claims registry. Two things get mixed up here. One is checking the file has not changed, which a hash does, every time. The other is getting the same number from a fresh run, which fades as the answers move. If a vendor promises "run it again, same number" with no end date, ask which of the two they mean.
4. Does it measure the surfaces your customers actually use?
Check the official engine roster against your market. Rosters differ more than marketing pages suggest — several vendors' own documents disagree internally about which engines are included.
5. Have you converted the billing unit into a total cost?
Prompts, credits, engine add-ons, seats. The metering units differ everywhere. Work out what "from $29/mo" becomes at your prompt volume and engine mix, and if pricing is fully undisclosed, ask why.
Where does CiteAngle stand on the same axes? Fixed one-time pricing
CiteAngle's diagnostic is a one-time, fixed-price product, not a retainer. It runs 50 prompts across 15 engine captures, each prompt-capture pair repeated 7 times, and reports the result across all 17 US surfaces. It prints the measured Wilson 95% CI width in the report, publishes the status of every planned observation including failures, and seals the results in a hash ledger backed by a public /claims registry. Repeat sampling itself is something other vendors state as well. Company E states 100 samples per prompt per model. What we offer is the combination: an auditable grid with sealed receipts, published CI widths, and a public claims registry, at a fixed published price. Everything we show as evidence is labeled for what it is: demonstrations we built ourselves.
How this document is maintained: one address, fixed axes
This URL is the canonical home of this comparison. A new version does not get a new address. The page updates in place, with each fix and its date shown here, so cite and link this URL. The axes are fixed too: public price, surfaces measured, repeats and confidence intervals, raw-data access, sealing and checking. Every version reads those same axes. And the scope, stated plainly: a comparison document answers what each product measures and how. Which sources AI answers actually pick is the next question, and a measurement grid answers that one. This format exists so a reader can re-check the facts later.
Side by Side: public-page comparisons
The article you are reading is Part 1 of the public-page comparison series. Next: Part 2, the US edition of the roster check (forthcoming). The full series lives in the series index. The grid in this article was accessed 2026-07-16.
What choosing on confidence costs. Without a public grid you risk buying the best deck instead of the widest measurement. We sit on this grid under the same rules, with every reading sealed and competitors read side by side with your brand, which is the reason we were willing to publish it.
Why we published a grid that includes us
The reason to build this comparison at all is that the alternative is choosing on confidence. Every vendor in this market sounds certain and the decks look alike. What separates them is how much each one will say about how its number was made. That stays invisible unless somebody lines the public pages up side by side. A buying team without that grid ends up rating presentations.
We sit on the grid under the same rules, and that is the point. Our panel is published as arithmetic: 50 buyer questions, 15 engines, 7 runs each, landing on 17 surfaces, with search results and AI answers held in one grid and competitors read inside the same answers. Conditions are sealed at measurement time, and every public figure is registered with its source and date. Ask any vendor here for those same four things in writing, us included. The ones who can answer are the shortlist.
Ask every vendor the same five questions, including us
Our answers are on the record: the protocol on the methodology page, every public figure on the claims registry, and the delivered screens on samples.
Want the same five answers for your own shortlist? Contact us and we will scope it against your category.