What is actually measured, what is merely asserted, and where the migration from links to answers really stands.
Two years into the answer layer, the debate has moved past whether AI search matters and into a harder question: what, specifically, has changed, and by how much. The honest answer is that the evidence is uneven — some effects are now well measured, others are asserted far more confidently than any dataset supports. This is an attempt to separate the two, and to describe where the migration of visibility from links to answers actually stands as of 2026.
Four findings have enough independent measurement behind them to be treated as settled. Click-through on queries carrying an AI Overview falls sharply — position-one click-through more than halves, and longitudinal work shows an initial collapse followed by partial recovery to a persistent gap rather than a return to baseline. That shape matters: the disruption is smaller than the first measurements suggested and considerably larger than nothing.
The second is that publisher referrals have fallen materially without substitution. The third is that different answer surfaces select different sources, including two surfaces operated by the same company on the same index. The fourth is that answer coverage is actively managed by the platform rather than fixed, having expanded aggressively and then been pulled back. Each of these is supported by work with a stated sample and method.
The changes above are not independent events; they are stages of a single migration, and naming the stages makes the current position legible. We use a four-phase curve: displacement, divergence, disintermediation, delegation. Most categories are somewhere between the second and third phase, and the fourth has begun in narrow domains.
The value of the sequencing is predictive rather than descriptive. A category early in displacement can see what divergence looks like by observing categories ahead of it, and the responses that work at each phase differ — optimisation still pays in displacement, measurement becomes the constraint in divergence, and business-model adjustment becomes unavoidable in disintermediation.
Four phases in the movement of visibility from ranked links to generated answers. Categories progress at different speeds; the correct response differs by phase.
Displacement
Answers appear above results and absorb clicks. Rankings still deliver, at reduced return. The work is still recognisably SEO.
Marker: click-through falls on covered queries while rankings hold.
Divergence
Surfaces begin selecting different sources from each other and from the ranked list. Rank stops predicting presence.
Marker: you rank well and are absent from the answer above you.
Disintermediation
The visit stops being necessary for a growing share of demand. Referral-based measurement misdescribes the business.
Marker: influence rising while sessions fall; substitution does not arrive.
Delegation
Assistants act rather than describe — shortlisting, comparing, transacting. Being named becomes commercially decisive.
Marker: buyers arrive pre-shortlisted, having never visited.
Category progression along the migration curve
Horizontal timeline 2024–2026 with category bands (news/publishing, informational B2B, D2C retail, local services, regulated industries) plotted at their current phase. Shows publishing furthest along and regulated categories least affected. Mark as directional positioning, not measured.
Divergence is the phase causing the most organisational confusion, because it breaks the instrument everyone relies on. A rank tracker keeps reporting accurately while the thing it measures decouples from the outcome it used to predict. Teams see stable rankings, falling traffic, and no explanation in their reporting, which produces the characteristic failure of this phase: months spent optimising metadata against displacement that metadata cannot address.
The exit from phase two is measurement rather than optimisation. Until a brand can see whether it is cited, separately per surface, it cannot tell whether a traffic change is its own doing or the environment’s. This is the single most common gap we find, and it is why the state of the field in 2026 is better described as a measurement crisis than a visibility crisis.
The layer is plural and getting more so. Google operates at least two distinct generative surfaces with meaningfully different citation behaviour. ChatGPT retrieves live and cites, drawing heavily on community and reference sources. Perplexity is citation-first and rewards passage-level quality over domain authority. Copilot draws on a different index entirely. Each has a distinct character and a partly distinct set of favoured sources.
The practical consequence for 2026 is that “AI visibility” as a single metric has stopped being meaningful. A brand can be strong on one surface and absent from another with no contradiction, and any consolidated figure averages systems that demonstrably disagree. The reporting standard that survives scrutiny is per-surface, per-prompt, benchmarked against named competitors.
Displacement is measured. Substitution has not arrived. Surfaces disagree. What almost nobody has is a reliable, per-surface view of whether their own brand is being named — which is the prerequisite for every other decision.
Several claims circulate with a confidence the data does not carry. That AI search has ended organic traffic: a substantial share of clicks persists even on covered queries, and many queries carry no answer feature at all. That a specific percentage of searches now trigger an Overview: prevalence is a property of the query set measured, and published figures differ by a wide margin while all being defensible.
Also unsupported: that particular on-page techniques reliably produce citations. The correlational work identifies signals that accompany AI visibility — branded mentions, video presence — without establishing that manufacturing those signals produces the outcome. The gap between “cited brands look like this” and “do this to be cited” is where most of the field’s bad advice currently lives.
Confidence levels across commonly cited claims
Horizontal bar chart. Rows: CTR displacement, referral decline, cross-surface divergence, coverage prevalence, on-page causation, correlational remedies. Bars indicate evidence strength (multiple independent studies / single study / correlational / assertion only). Colour-code the bottom two categories in orange as caution.
The most consequential weakness in the field is that almost nothing establishes causation. We have good measurement of what answers displaced and poor measurement of what makes a brand appear in them. The available work is correlational, drawn from broad panels, and cannot distinguish signals that cause citation from signals that accompany the brand prominence that causes it.
This matters practically because it means most optimisation advice is inference. Closing it requires controlled work: holding a brand’s content constant while varying a single signal, or comparing matched cohorts where one receives an intervention. That is expensive and slow, which is why nobody has published it — and why the honest position in 2026 is that we know a great deal about the effect and comparatively little about the mechanism.
Every reliable finding in this piece comes with the same caveat: measure it on your own queries. DUNkē tracks citations across eight AI engines — per prompt, per surface, against competitors.
The underexamined half of the story is the buyer. Research that once involved comparing several sources now frequently involves one conversation, and the shortlist arrives assembled. That compresses the consideration stage and moves the decisive moment earlier — before any brand-owned property is visited, and often before the buyer could name where the shortlist came from.
For go-to-market this is more consequential than the traffic effect. A brand absent from the assembled shortlist is not competing later on a website or a sales call; it has been excluded before the process it was built to win begins. Categories where buyers research conversationally are seeing this first, and it is the mechanism that turns an AI visibility question into a pipeline question.
The disciplines that survive this period unchanged are the ones that were always about being genuinely the best available source: comprehensive coverage, real expertise, structural clarity, earned credibility. What changes is which of them bind. Technical access and extractability matter more because a retrieval system is now the reader. Entity clarity matters more because naming requires resolution. Earned corroboration matters more because assertion is discounted.
What matters less is the page-level tuning that dominated the previous decade. Not because it stopped working — it still helps the ranked list — but because its marginal return has fallen relative to the work that decides answer presence. The reallocation implied by the 2026 evidence is away from on-page optimisation and toward entity, evidence, and measurement.
The platform behaviour underneath these effects moved considerably. Coverage of generated answers expanded aggressively and was then deliberately pulled back, which established that prevalence is a managed policy rather than a fixed property. Conversational surfaces matured from experiment into a distinct product with its own retrieval. And crawler controls proliferated, giving publishers finer per-bot decisions than the blunt allow-or-block choice of the previous period.
Each of these has a planning consequence that outlasts the specific change. Managed coverage means any strategy predicated on a category staying uncovered is fragile. A distinct conversational surface means measurement must be per-surface. Granular crawler controls mean access decisions are now deliberate choices with visible consequences rather than defaults nobody examined. The direction across all three is toward more variables, not fewer.
If phase two is a measurement problem, it is worth stating what an adequate stack contains. At minimum: a fixed prompt set drawn from real buyer questions rather than a keyword export, run across the surfaces your audience uses, sampled repeatedly because answers vary, benchmarked against a competitor set defined by who actually appears, and archived so trend exists.
Almost no organisation has this. The common state is rank tracking plus Search Console, which together describe the ranked list accurately and the answer layer not at all. That gap is the single most consequential omission in 2026 marketing measurement, because it means the fastest-growing surface in search is the one nobody is instrumented for. Closing it is cheap relative to what it informs.
The strongest available signals associated with AI visibility are brand-level rather than page-level: how often a brand is mentioned across the web, and how present it is in formats like video. That is a genuinely useful finding and it is repeatedly overread. Correlation of this kind cannot distinguish a signal that causes citation from one that accompanies the overall prominence that causes it.
The honest use is as a description of what cited brands look like, which points investment in a defensible direction without licensing a mechanical programme. A brand that becomes genuinely more discussed and more present is doing something worth doing regardless; a brand manufacturing mention volume without the underlying substance is reproducing a proxy and should not expect the outcome.
It is worth acknowledging that practitioners are acting on patterns the literature has not confirmed, and that this is not necessarily wrong. Audit work surfaces regularities — entity resolution failures, description inconsistency, extractability problems — that are diagnostically reliable long before anyone publishes a controlled study establishing them.
The discipline that keeps this honest is labelling. A pattern from practice is useful and should be described as a pattern from practice. The failure mode is dressing it as research, which the field does routinely by attaching invented percentages to genuine observations. Our own frameworks in this publication are labelled as ours, and where we have no measurement we say so rather than estimating.
Evidence status of common claims
Rows: displacement magnitude, referral decline, surface divergence, coverage prevalence, entity effects, extractability effects, mention causation. Columns: claim, evidence type (multi-study / single study / correlational / practice observation / assertion), confidence. Makes the article’s central argument scannable.
Three developments would materially alter the 2026 assessment. Regulatory or commercial arrangements that require answer surfaces to drive traffic to sources would reduce the disintermediation effect, and licensing deals between AI companies and publishers point weakly in that direction. A reversal of coverage expansion would restore click volume on affected queries.
The third and most consequential would be agentic behaviour becoming mainstream — assistants transacting rather than describing. That would convert being named from a visibility question into a distribution question, and would make the evidence base that determines naming considerably more valuable than it currently is. None of these is predictable, which argues for building the durable capabilities rather than betting on a scenario.
The executive translation is narrow. Establish per-surface citation measurement, because you are almost certainly not instrumented for the surface that is growing. Segment your query portfolio by answer coverage so you know your exposed share. Reallocate a portion of content budget toward entity clarity and earned corroboration, which are the pillars the 2026 evidence supports and the ones most organisations underfund.
And change what you ask for in reporting. A single AI visibility figure should be treated as a warning sign rather than a deliverable, because it averages surfaces that disagree. Per-surface citation share against named competitors, trended, is the reporting standard that survives the current evidence. Anything less specific is describing a landscape that does not exist.
Every question about this period eventually becomes a request for a forecast, and the honest answer is that the variables are controlled by platforms making product decisions rather than by trends with momentum. Coverage has already moved in both directions. Surfaces have appeared and been restructured. Extrapolating any of it is unsound.
What can be forecast is the durability of the underlying requirements. Whatever the surfaces look like in two years, they will still need to determine what a source is, whether independent evidence supports it, and whether a usable passage can be extracted. Building those capabilities is a bet on the mechanism rather than the interface, which is the only kind of bet the current evidence supports.
A minor but persistent obstacle in 2026 is that the field cannot agree what to call itself. GEO, AEO, LLMO and AI SEO are used interchangeably, inconsistently, and often to describe the same work. This matters less than the arguments about it suggest, but it does create real friction when briefing, hiring, and buying.
The pragmatic position is that the terms describe emphases within one practice rather than distinct disciplines, and that the underlying requirements — be reachable, resolvable, corroborated, extractable — are the same whichever label is used. When evaluating a vendor or a hire, ignore the terminology entirely and ask which of those four they can actually affect.
Most organisations are funding this out of existing search budget, which produces a predictable constraint: the work that most needs funding — earned corroboration — sits outside what a search budget conventionally covers, while the work a search budget covers naturally is the least binding.
The organisations making progress have generally done one of two things: moved a portion of PR or brand budget under the same objective, or accepted that their search function now commissions earned coverage. Neither is a large financial change. Both are organisational ones, which is why they are harder than they look and why they are the most common bottleneck we encounter.
The capability the answer layer demands does not map cleanly onto existing roles. It requires someone comfortable with technical diagnosis, entity and structured data work, earned media relationships, and measurement design — a combination that sits across three conventional job descriptions and belongs entirely to none.
In practice, teams solve it either by pairing a technical SEO with a communications lead under a shared objective, or by developing measurement capability internally and outsourcing the earned work. What does not work is assigning it to whoever currently owns content, which reliably produces a content-shaped response to a problem that is usually not content-shaped.
If the field could commission one study, the highest-value target would be causal rather than descriptive: matched brand cohorts where a single signal is varied — entity declaration, description consistency, extractability — with citation outcomes tracked over quarters. That would settle questions currently answered by inference.
The second would be a stratified prevalence and citation study across categories, published with its sampling frame, so that practitioners could locate their own category rather than importing an average built from someone else’s query mix. Both are expensive and neither is glamorous, which is why the field has neither — and why 2026 remains a period of confident claims resting on thin evidence.
A quieter shift sits underneath the measurement story. Buyers who research conversationally arrive having already been told, by a system they broadly trust, what the options are and roughly how they compare. That changes the first conversation a brand has with them: it starts from a position the brand did not set and may not know about.
Sales teams are encountering this before marketing teams are measuring it — prospects citing comparisons nobody in the company wrote, or ruling out capabilities the product has. The practical instruction is to find out what the engines say about you before your buyers do, because that description is now part of your positioning whether you authored it or not.
Reduced to actions rather than analysis: build a fixed prompt set from real buyer questions and start measuring citation per surface, because you almost certainly cannot see the fastest-growing surface in search. Audit how third-party sources describe you and correct the contradictions, which is mechanical work with unusually good returns.
And segment your query portfolio by whether an answer feature appears, so you know what share of your demand is exposed. None of the three requires new budget or new headcount, all three produce information you currently do not have, and together they convert the general anxiety about AI search into a specific, bounded picture of your own position.
If this entire assessment reduced to a sentence for someone with no time: the answer layer has taken a measurable, durable share of the visibility that links used to carry, nobody has established what causes a brand to appear in it, and almost no organisation can currently see whether they do.
That combination — a real effect, an unsettled mechanism, and an unmeasured position — is the actual state of AI search in 2026. It argues for scepticism toward anyone selling certainty about the mechanism, and urgency about closing the measurement gap, because every decision that follows depends on being able to see your own position rather than importing someone else’s average.
Is it too late to start on AI visibility?
No, but the slow components compound, so starting later means arriving later. Access and extractability work immediately; corroboration and category association accumulate over quarters, which is an argument for beginning those now rather than when the traffic impact becomes acute.
Which engine should we prioritise in 2026?
Whichever your buyers actually use, which varies by market and ecosystem far more than industry commentary suggests. The shared fundamentals make you eligible across all of them, so engine choice affects emphasis rather than requiring separate strategies.
Should we expect the displacement to worsen?
Coverage is actively managed and has moved in both directions, so extrapolating the current trend is unsafe. The defensible planning position is to build for the environment as measured while monitoring your own coverage, since the platform has demonstrated it will change it.
What single metric should a board see?
Citation share on a fixed set of commercially important prompts, benchmarked against named competitors and trended. It is the only measure that reflects the thing you control and moves before the downstream demand effects appear.
As of 2026, the displacement of clicks by answers is measured and stable, substitution has not materialised, the answer layer is plural with surfaces that disagree about sources, and coverage is a managed variable rather than a fixed property of search. Those four findings are supported. Most of what else circulates — particularly claims about what causes a brand to be cited — is inference presented as evidence.
The practical position that follows is unglamorous: locate your category on the migration curve, fix the measurement gap that phase two exposes, report per-surface rather than in aggregate, and reallocate effort from page-level tuning toward entity clarity, earned corroboration, and the ability to see whether you are actually being named. The field will produce better causal evidence eventually. Until it does, measuring your own position is worth more than any published average.
Methodology note: this piece synthesises published third-party research and states no original measurement. Where we describe distributions across phases of the migration curve, those are qualitative positioning from audit work rather than measured data, and are labelled as such. No figure in this article is ours unless attributed to our own practice.
“We know a great deal about what the answer layer took and remarkably little about what it rewards. Treating the second as settled is the most common error in the field right now.” The Age’X Research Team
Get a free GEO audit — the same analysis behind every article here.