Where these numbers come from — public studies, linked in the text and verified July 14, 2026. Across 1.5 million real ChatGPT conversations, more than 70% of use now happens outside work, up from 53% two years earlier; G2's survey of 1,076 B2B buyers puts 51% of vendor research starting in an AI assistant.
Is one keyword still one contest in US AI search? No. One keyword is now dozens of questions. Google's own developer docs describe how its AI search features take a single query and issue multiple related searches in parallel, a technique Google calls query fan-out. Each of those questions is a separate contest your pages either enter or miss, and a one-row keyword report cannot see any of it.
The demand shift behind fan-out has been measured. A study of 1.5 million real ChatGPT conversations found more than 70% of usage now happens outside of work, up from 53% two years earlier. Pew put clicks on a plain result at 8% of visits with an AI summary, against 15% without. And in G2's Answer Economy survey (n=1,076), 51% of the B2B software buyers asked said they now start research with an AI chatbot more often than with Google.
How does query fan-out work? Google documents the mechanism itself
Query fan-out is not vendor talk. Google's guide for succeeding in AI search says its AI features sit on the Search index and core ranking systems. It also says they can fire off several related searches at once for one user query. That is the query fan-out technique. Official primary developers.google.com
The same docs set the entry rules. To be eligible for AI experiences, a page must be indexed and eligible to show as a search snippet. Official primary developers.google.com Put those two statements together and the new contest takes shape. One buyer query becomes a family of sub-questions. Candidacy is decided per sub-question, not per keyword. Where those contests physically happen, meaning AI Overviews, AI Mode, Copilot and the surfaces around them, is mapped in our US AI search surface map.
The research lineage runs years deep
Decompose-and-recombine is not a marketing buzzword. It is a published line of research. Work presented at ACL 2019 showed something useful. Break a hard question into sub-questions, then score the answers together, and reasoning across many documents gets better. Peer-reviewed aclanthology.org What was once a research trick is now part of how commercial AI search finds and grounds its answers, per Google's own docs. The path has been in plain sight for years. Measurement practice just hasn't caught up.
Do zero rows in your keyword table mean zero demand?
Here is the uncomfortable part for anyone running US demand planning off a query report. Google Search Console's own docs explain two things. Rare queries are anonymized and left out of performance results, to protect user privacy. And the report centers on top rows, so the sum of the rows you see can differ from the chart totals. Official primary support.google.com · developers.google.com
Long, specific buyer questions are exactly the kind that land in that rare, anonymized tail. So here is how to read the official docs. A query showing zero rows in your table is not proof of zero demand. The question family your buyers actually use can be invisible in the very tool most US teams treat as the census of demand. That is the tool's own stated design.
Three unit errors that follow from measuring keywords alone
1. Coverage error. The engine contests a family of sub-questions; you track one
head term.
2. Visibility error. The question-shaped tail is anonymized out of your query
report, so the demand you most need to see is the demand you least observe.
3. Inference error. Repeated checks of the same query add observations, not
independent evidence. The unit that supports a decision is the question family crossed
with the page that answers it, observed over repeated windows.
Where are the questions moving? Into conversations
This question-shaped demand is not a guess. An NBER working paper looked at 1.5 million real ChatGPT conversations. More than 70% of usage now happens outside of work, up from 53% two years earlier. These are everyday people asking everyday questions, including the ones that end in a purchase. NBER working paper nber.org
Buyer behavior points the same way (added July 17, 2026). Pew Research Center measured US Google visits and found users who encountered an AI summary clicked a traditional result link on 8% of visits versus 15% without a summary, and clicked a link inside the AI summary on just 1% of such visits. Pew Research Center pewresearch.org And in G2's The Answer Economy survey (March 2026, n=1,076), 51% of surveyed B2B software buyers said they now start research with an AI chatbot more often than with Google. G2 itself has a stake in AI-search visibility, so read it as a vendor-adjacent survey, stated at its own predicate level. Industry survey prnewswire.com
What should you measure instead?
The fix is a change of unit, and it pays off immediately in decision quality. In our own run of 50 US category questions read on all 15 answer surfaces in a single pass, measured July 27, 2026, citations per surface ran from 14 to 2,675 on the same questions on the same day. Pick the wrong surface and you have measured a different market. The full per-surface breakdown of that run is published. Before the unit changes, it helps to see the two side by side. Here is the same comparison table we use when a keyword report and an AI answer disagree.
| Question | A keyword table answers | A question family answers |
|---|---|---|
| What is the unit? | One typed string. | The family of questions the engine runs from it. |
| What does zero mean? | Nobody typed that string. | Nobody asked, or nobody asked it that way. |
| Who is the competitor? | Whoever ranks on that string. | Whoever gets cited anywhere in the family. |
| How often do you read it? | Once a month is normal. | Repeatedly, because the answer moves between runs. |
| What decision does it support? | Which page to optimise. | Which question to be present for at all. |
Every row restates a mechanism documented in the sections above. The table is the short version to send to whoever owns the keyword report.
Measure question families, not keywords. Map the buying questions in your category, meaning the comparisons, fit checks, implementation and switching questions your buyers actually pose, and treat each family as one demand object. That is the unit the engine is reasoning in.
Measure candidate entry, not just rankings. Google's documented entry conditions mean the first question for any page is simple — does it enter the contest at all, for each sub-question in the family? A page can rank respectably on a head term and still be absent from most of the family the engine actually runs.
Ask again and again, and read the unit honestly. One look at one query is a single point, not a picture. The published work on single-run spread is blunt about how far one pass can wander. Choices hold up when they rest on the question family and the page that answers it, watched over repeated windows. That is the bar CiteAngle builds its US work on.
Fan-out families are the demand your keyword tools never show you. The families you have not mapped. The sub-questions where you never even enter. The rivals who are already the default answer there. Every one of those can be fixed, once you have counted it.
What the undercount costs. If demand never reaches your keyword table, you lose a year to a category that looks flat and is not, and no report anywhere will flag it. We measure at the question level, 7 runs each and sealed with its conditions, which is why a rate here separates a real pattern from a single lucky reading.
The second cost is the argument you cannot win. Without question-level readings, a colleague who says "nobody searches for that" is not wrong about the keyword table — the demand is real and the evidence for it is invisible, so the budget goes elsewhere and the category is conceded by default.
Why measure this with us
The reason this gap persists is that nothing on a marketing team's desk is built to see it. A keyword tool reports the query a person typed. Fan-out happens after that, inside the engine, and the searches it issues are never typed by anyone — so they appear in no volume table, and a category quietly generating demand can read as flat for a year.
Measuring at the question level is the only way round it, and it is what our panel is for. We put 50 buyer questions to 15 engines, 7 times each, and read what came back. The typed query is not the unit any more. Because the same question is asked repeatedly, a name that appears in some replies and not others shows up as a rate with an interval instead of a coin flip you happened to observe once. Your competitors are read inside those same answers, so you learn who currently owns each question and not merely whether you were in it.
Your buyers ask dozens of questions. Measure which ones you win.
A free AI visibility check reads one to three of your buyer questions across six places where AI answers show up, same day. The paid grid widens to fifty queries, seven runs each, across all 17 US surfaces.