Most AI Mode answers show a source sidebar averaging seven domains — and only a small share of citations overlap with AI Overviews.
It is tempting to treat Google’s two AI surfaces as one thing wearing different clothes. This study establishes that they are not. AI Mode answers carry a source sidebar averaging around seven domains — considerably more than a typical Overview shows — and, more consequentially, only a small share of the sources cited in one appear in the other. Same company, same index underneath, two substantially different sets of chosen sources. That has direct consequences for how you measure and where you can win.
The research examines citations across both of Google’s generative surfaces at scale, recording which domains are surfaced in AI Mode responses and which appear in AI Overviews for comparable queries, then measuring how much the two sets have in common. Working across a very large query volume gives the overlap measurement some weight, since citation sets are noisy at small samples.
Two findings emerge. The first is a matter of breadth: AI Mode surfaces a wider set of sources per answer than an Overview typically does. The second, and the more important one, is a matter of composition: the sources chosen are substantially different, not merely more numerous. It is the second finding that changes how you should work.
The natural assumption is that a single company running a single index would draw on the same sources when answering the same question in two places. The data does not support that assumption. The retrieval behind a conversational, multi-turn experience appears to differ meaningfully from the retrieval behind a summary placed on a results page — different enough that being chosen for one implies remarkably little about being chosen for the other.
This is less strange than it first sounds. The two surfaces do different jobs. An Overview compresses an answer into a small space above a page of links, which favours a handful of highly corroborating sources. AI Mode supports exploration across turns, which favours breadth — more sources, covering more facets, capable of supporting follow-up questions. Different jobs produce different selection, even from the same underlying index.
The wider source set in AI Mode follows from its conversational design. A dialogue that will be extended needs enough grounding to support wherever it goes next, and a broader retrieval provides that headroom. There is also more interface room: a dedicated conversational surface can display a sidebar of sources in a way a compressed summary above organic results cannot.
For brands this breadth is straightforwardly good news. More sources per answer means more citation slots per question, which makes AI Mode a less winner-take-all surface than an Overview that names three. A source that is genuinely relevant but not the single most authoritative option has a materially better chance of appearing. The competition is real but the door is wider.
Breadth is useful; low overlap is structural. If the sources cited in AI Mode were largely the same as those in Overviews, you could reasonably optimise once and measure once. Because they are largely different, you cannot. Presence on one surface is weak evidence about the other, which means a brand can be confidently visible in Overviews and effectively absent from the conversational surface without ever knowing.
This is the practical core of the study. It establishes that Google’s AI answer layer is not one target but at least two, each of which must be measured on its own terms. Any measurement programme reporting a single figure for “Google AI visibility” is averaging across two systems that demonstrably disagree about who to cite.
Only a small share of citations overlap between AI Mode and AI Overviews. Presence on one tells you remarkably little about the other, which means they have to be measured separately.
If you are tracking AI visibility on Google by checking Overviews, you have measured one surface and inferred the other. That inference is not supported. The same applies in reverse, and it applies to any dashboard presenting a consolidated Google AI number without distinguishing where the citations came from.
The correction is to treat each surface as a separate tracked entity, with its own citation rate and its own share of voice, reported separately even if they roll into a summary figure. This makes the measurement more expensive, which is a real cost. It also makes it correct, and the alternative — a confident number that averages two disagreeing systems — is worse than a slightly more complex report.
What it cannot tell you: why the two surfaces select differently, whether the gap is narrowing as both mature, or how overlap varies for your specific vertical. It establishes that the difference exists and is substantial — not the mechanism producing it.
The instinctive reading of a second surface is a second problem. The better reading is that low overlap means the competitive landscape on each surface is genuinely different — and being absent from one that competitors have not colonised either is a position that can be taken rather than a deficit to be recovered.
A brand that has fought its way into Overview citations in a competitive category may find AI Mode surprisingly open, because the incumbents there are a partly different set. And the wider source count means the surface accommodates more winners. For a challenger, this is a more promising place to invest attention than a compressed Overview where three highly-corroborated sources take everything.
Presence in Overviews tells you little about AI Mode. DUNkē tracks citations across eight AI engines and surfaces — separately, per prompt, against competitors — so you can see where you are actually present.
The study does not explain the mechanism, but the shape of the difference is consistent with what is understood about conversational retrieval. An experience designed for multi-turn exploration decomposes topics more aggressively and retrieves across more sub-questions, which favours sources with comprehensive coverage of a topic rather than a single strong answer to the headline query.
That would predict exactly what is observed: a wider source set, drawn from a partly different pool, weighted toward sources that can support an unfolding conversation. It also suggests where to invest for this surface specifically — depth across a topic’s facets and answers to the natural follow-up questions, rather than a single well-optimised page.
The practical consequence for measurement design is that your tracked prompt set needs to run against both surfaces rather than one. The same questions, posed in each, recorded separately — which roughly doubles the observation cost for Google but produces the only data that reflects reality.
It also argues for including the conversational follow-ups in your prompt set, not merely the opening questions. If AI Mode is where multi-turn research happens, a prompt set consisting only of standalone queries under-samples the behaviour that surface actually exhibits. Capturing the arc — opening question plus the questions that typically follow — gives a truer picture of where you appear.
Both surfaces are actively changing, and a comparison of this kind is a snapshot of a moving relationship. The overlap figure could reasonably narrow as retrieval converges, or widen as the surfaces specialise further.
Treat the structural finding — that the surfaces select differently and must be measured separately — as the durable takeaway. Treat the specific overlap percentage as a reading taken at a particular moment, worth re-measuring rather than assumed to persist.
Low overlap does not mean the fundamentals differ. Both surfaces reward being retrievable, relevant, comprehensive, well-evidenced, and credible — the shared foundation covered elsewhere in this Academy. What differs is which sources clear that bar in each context and how many are shown, not the qualities being assessed.
Nor does it mean you need two content strategies. It means one strong foundation, tuned toward comprehensive topical depth for the conversational surface, and measured separately in both places. Concluding that AI Mode requires an entirely separate programme overreads a finding about citation composition into a claim about optimisation mechanics that the data does not support.
The sidebar format is worth considering on its own terms, because it differs from how Overview citations are presented. A dedicated panel listing the sources behind an answer gives each cited domain persistent visibility while the user reads, rather than a compressed set of links attached to a summary. The citation is more prominent and more durable within the interaction.
For brands that changes what a citation is worth. Being one of seven sources in a visible panel delivers real brand exposure even when no click follows, and it invites the verification behaviour that produces clicks — a user who wants to check a claim has an obvious route. This is one of the surfaces where the zero-click framing is least absolute, because the sources are presented as part of the answer rather than beneath it.
A surface citing around seven sources per answer has a materially different competitive structure from one citing three. In a three-slot answer the sources are likely to be the most corroborated, most authoritative options — which tends to mean incumbents. Widening to seven admits sources that are genuinely relevant without being the single strongest option.
For a challenger brand this is the more promising surface to contest. The marginal cost of being the fifth-best source rather than the second-best is much lower, and comprehensive topical coverage can earn a slot that domain authority alone would not. It is a rare instance where the newer surface is more open than the established one, and it argues for prioritising the conversational surface if you are competing from behind.
A single Overview is one opportunity to be cited. A conversation is several, because each turn triggers its own retrieval, and a source that supports the opening question may or may not support the follow-ups. Over a research session of five or six exchanges, the number of citation slots available is a multiple of what a single answer offers.
This changes what content earns presence. Depth across a topic’s facets matters more than a single excellent page, because the conversation will move to adjacent questions and whoever answers those gets cited there. A brand with one strong page appears at the start and disappears; a brand with genuine cluster coverage can be present throughout — which is a considerably more valuable outcome from the same research session.
It is reasonable to ask whether the two surfaces will converge as both mature, and there are grounds to think the gap is structural rather than transitional. The surfaces serve different jobs — compression versus exploration — and those jobs impose genuinely different selection criteria. A source ideal for a three-line summary is not necessarily ideal for supporting a multi-turn investigation.
If that reading is right, the divergence is a design consequence rather than an artefact of immaturity, and it will persist as long as the surfaces remain distinct experiences. That has planning implications: treating separate measurement as a temporary inconvenience until things converge would be a mistake. The safer assumption is that a search ecosystem containing several answer surfaces will keep requiring several measurements.
The practical build follows from the finding. Your prompt set should run against both surfaces, recorded separately, and it should include conversational arcs rather than only standalone questions — an opening question plus the follow-ups that typically succeed it, since that is the behaviour AI Mode exhibits and the behaviour that generates additional citation slots.
Doing this roughly doubles the observation cost for Google alone, which is a genuine constraint. The efficient compromise is a core set of high-value questions tracked on both surfaces with their follow-ups, and a broader set tracked on whichever surface matters more for your category. That keeps the primary comparison honest without making the programme unaffordable.
The diagnostic pattern this study enables is finding yourself well cited in Overviews and absent from AI Mode, or the reverse, and each points somewhere different. Strong in Overviews but absent from AI Mode typically indicates a depth problem — you have an authoritative answer to the headline question but not the coverage to support an exploration around it.
The reverse pattern — present in AI Mode but not Overviews — usually indicates a corroboration or authority gap: broad enough coverage to be retrieved for many sub-questions, but not the concentrated credibility that a compressed three-source summary selects for. Reading which side you are weak on tells you whether to invest in breadth or in authority, which is a genuinely useful diagnosis that a single blended metric would have hidden.
Separate measurement produces a reporting question: how to present two numbers without either confusing the audience or burying the difference. The workable approach is a headline figure covering the surfaces you care most about, with the split available immediately beneath, and a note explaining why they differ.
What fails is silently averaging them, which produces a number that describes neither surface and moves for reasons nobody can explain. It also fails to present them so separately that the audience cannot form an overall view. Understanding that this is a presentation problem rather than a measurement one is why the reporting design deserves attention alongside the tracking setup.
If conversational surfaces continue growing as a share of information-seeking, the properties this study identifies become more important rather than less: breadth of source selection, multi-turn citation opportunities, and a competitive field that is partly distinct from the one on the results page.
That argues for treating investment in conversational-surface visibility as forward-looking rather than incremental. The comprehensive topical depth it rewards is the same investment that serves topical authority generally, so it is not a speculative bet on one surface — it is work that pays across the answer layer while positioning particularly well for the direction the layer appears to be heading.
It is worth quantifying what the single-surface shortcut actually costs, because the argument for doubling measurement effort has to survive a budget conversation. If overlap between the surfaces is low, then measuring one and inferring the other means your reported visibility is roughly uncorrelated with your visibility on the unmeasured surface.
That is not a small imprecision; it is a number that could be badly wrong in either direction. A brand reporting healthy Overview citation could be entirely absent from the conversational surface and would have no indication of it. Given that the remedy is additional observation rather than additional strategy, the cost of measuring both is modest against the risk of confidently reporting half a picture.
If the surfaces did begin to converge, the signals would be visible in your own tracking: overlap rising over successive measurement periods, source counts moving toward each other, and the same competitors appearing in both. Watching for that is worthwhile, because convergence would simplify measurement considerably.
The reason not to assume it is that the surfaces serve genuinely different jobs, and design differences that follow from purpose tend to persist. But this is an empirical question rather than a settled one, and a tracking programme that records both surfaces separately will answer it for your category without any additional work — which is a further argument for building it that way from the start.
For teams that genuinely cannot resource both equally, the prioritisation should follow audience behaviour rather than industry attention. If your buyers research through extended conversational sessions — common in considered, technical, or high-value purchases — AI Mode is where their process actually happens. If they ask short factual questions in the course of ordinary search, Overviews matter more.
The wider source count is a secondary consideration favouring the conversational surface for challengers, and the concentrated authority requirement favours Overviews for established brands defending a position. Neither is a general answer. What is general is that the choice should be made deliberately on evidence about your buyers, not defaulted to whichever surface is discussed most.
Although this study examines two Google surfaces, the underlying lesson extends across the answer layer: different surfaces retrieve differently, and presence on one is weak evidence about another. The same has proven true across engines, and there is no reason to expect the pattern to stop at Google’s internal boundaries.
That makes this study a useful proof of a general principle rather than a narrow comparison. Any measurement programme reporting a single blended AI visibility figure — across surfaces, across engines, or both — is averaging systems that demonstrably disagree about who to cite. The finding here is simply the cleanest available demonstration, because it holds the company and the index constant and still finds substantial divergence.
Read alongside the rest of this hub, this study performs a specific job: it establishes that the answer layer is plural. The click-impact research tells you what displacement costs, the prevalence research tells you how widely it applies, and this one tells you that the thing you are being displaced by is not a single system with a single set of preferences.
That has a compounding effect on measurement design. If surfaces disagree about who to cite, and coverage varies by query set, and click impact varies by intent, then any single headline number about AI visibility is an average of averages across dimensions that all move independently. The practical conclusion running through all three studies is the same: the only figures worth planning with are the ones measured on your own queries, on the surfaces your buyers actually use.
Reduced to its simplest form, this study says: measure both, because one does not predict the other. Everything else — the wider sidebar, the multi-turn opportunities, the challenger advantage, the diagnostic asymmetry — follows from that single structural fact about how the two surfaces select sources.
It is a modest-sounding conclusion for a large piece of research, and it is the kind that saves the most effort. A measurement programme built on the assumption of a single Google AI surface will produce numbers that look authoritative and describe something that does not exist. Building it correctly costs additional observation and nothing else, which makes this one of the cheaper corrections available in the whole discipline.
AI Mode answers surface a wider set of sources than AI Overviews — around seven domains in the sidebar — and only a small share of citations overlap between the two surfaces. Same company, same underlying index, substantially different chosen sources, because a conversational experience built for exploration retrieves differently from a compressed summary above a page of links.
The practical consequences are direct. Measure the surfaces separately, because presence on one is weak evidence about the other and any consolidated Google AI figure averages two systems that disagree. Extend your tracked prompt set to include conversational follow-ups, since that is the behaviour AI Mode actually exhibits. And treat the wider source count as an opening rather than a second front: more citation slots, a partly different competitive field, and a surface where comprehensive topical depth is what earns a place.
“Two surfaces from the same company, drawing on the same index, largely disagree about who to cite. Which means one number for “Google AI visibility” is an average of two systems that don’t.” The Age’X Research Team
Get a free GEO audit — the same analysis behind every article here.