The string is degrading as a measurement. The demand it stood for is entirely intact — which changes the planning unit, not the discipline.
Every few years someone declares keywords dead, and every time the claim is both directionally right and practically useless. Keywords are not dying. What is dying is the assumption that a keyword is the unit of demand — that a string of words maps cleanly to a need, and that winning the string wins the need. Answer engines broke that mapping, and the replacement is not "no keywords" but a different unit entirely.
A keyword was never the thing anyone wanted. It was a compressed, observable artefact of a need — someone with a problem typed the shortest string they thought might work, and search systems matched strings to documents. The whole discipline was built on that compression being stable enough to plan against.
It worked remarkably well for two decades because the compression was consistent. Similar needs produced similar strings, tools could aggregate them into volume figures, and a content plan built on those figures addressed real demand. None of that was wrong. It was a good proxy in an environment where the proxy held.
Two things, and only one of them is AI. The first is that engines stopped matching strings and started matching meaning, which happened years before answer engines and already weakened the one-keyword-one-page logic. Semantic understanding meant a page could rank for terms it never contained, and the keyword became a hint rather than a target.
The second is conversational input. When someone asks an assistant a full question — with their situation, their constraint, and their actual decision in it — the compression that made keywords tractable is gone. The same need now produces a hundred different phrasings, none of which any tool aggregates into a volume figure you can plan against.
This is the practically consequential part. Keyword tools work by aggregating similar strings into countable terms. Conversational phrasing defeats that: a question expressed a hundred ways fragments across a hundred low-volume variants, most of which report as zero, none of which sums to the real demand behind them.
So the tools are not becoming wrong, they are becoming incomplete in a specific direction — they systematically undercount the fastest-growing form of information-seeking. A plan built purely on reported volume is therefore aiming at the shrinking half of demand while reporting confidence, which is a worse failure than aiming at nothing.
What replaces the keyword as the planning unit, and what each layer is good for. Keywords survive as an input to the top layer rather than as the plan itself.
Keyword
A compressed string. Countable, aggregatable, and increasingly incomplete as phrasing fragments.
Still useful for: sizing traditional search demand.
Question
The full conversational form, with context and constraint attached. Not reliably countable.
Useful for: knowing what to actually answer.
Job
The underlying need several questions express. Stable across phrasings and interfaces.
Useful for: deciding what content should exist at all.
Entity & topic
The things and subjects the job concerns, which engines reason about directly.
Useful for: what you need to be known for.
The workable planning unit is the job — the underlying need that many differently-phrased questions express. Jobs are stable in a way strings are not: someone deciding between two categories of product has the same job whether they type three words or ask a paragraph, and whether they use a search box or an assistant.
Planning at the job level produces content that serves every phrasing of the same need rather than one variant of it. It also survives interface change, which strings do not. The practical exercise is to cluster the questions you can observe by what the asker is trying to accomplish, and treat those clusters as the plan.
Keywords were always a compressed proxy for demand. Conversational input broke the compression, which means the artefact is degrading while the thing it measured is unchanged.
More than the obituaries suggest. Volume data still sizes traditional search demand, which remains substantial and is where a large share of commercial queries still sit. Difficulty estimates still indicate competitive reality. And the vocabulary keywords surface — how customers actually name things — remains directly useful, because engines reason about entities and your customers name those entities in ways your marketing department frequently does not.
What should stop is treating the tool output as the plan. A keyword list is evidence about demand, not a content strategy, and the difference matters more now than it did when the proxy was tighter. Use it as one input among several rather than as the artefact everything else is derived from.
The planning artefact that works is a question map organised by job: the real questions your buyers ask, clustered by what they are trying to accomplish, prioritised by commercial value, with the phrasings recorded rather than compressed. It is built from support tickets, sales conversations, community discussion, People Also Ask data, and direct observation of what engines return.
It is less tidy than a keyword export and considerably more actionable, because each entry states what needs answering rather than what needs targeting. It also feeds citation measurement directly, since the same questions become the tracked prompt set — which a keyword list cannot do, because nobody asks an assistant a three-word string.
Questions are what buyers actually ask assistants, and they are what you can measure presence against. DUNkē tracks citations across eight AI engines — per prompt, against competitors.
A specific failure worth naming. Question-shaped queries frequently report zero or near-zero volume in keyword tools, because the phrasing is too specific to aggregate. Teams reading that as absence of demand systematically decline to answer the questions their buyers most often ask.
The corrective is to trust observed questions over reported volume when the two disagree. A question appearing repeatedly in support tickets and community discussion has demand regardless of what a tool reports, and the tool’s zero reflects a measurement limitation rather than a market fact. This is one of the clearest cases where the instrument is now misleading in a knowable direction.
The unit shift has a structural consequence beyond planning. If the demand unit is a question, the content unit should be an answer — a self-contained passage that resolves one question completely. Pages become collections of answered questions rather than treatments of a topic.
That is also, not coincidentally, the structure retrieval systems reward, since they extract passages rather than documents. The planning shift and the structural shift are the same shift viewed from two ends, which is a useful consistency: organising around questions improves what you write and how it is found simultaneously.
The practical arrangement is not to choose between keywords and questions but to use each where it works. Keywords remain the better instrument for sizing traditional search demand, identifying customer vocabulary, and assessing competitive difficulty on the queries that still behave like queries.
Questions are the better instrument for deciding what to answer, structuring the content, and building the prompt set you measure against. A planning process using keywords for sizing and questions for specification produces better output than either alone, and the transition most teams need is not abandoning one but adding the other.
The artefact is a table rather than a list. Each row records the question as actually asked, the underlying job it expresses, the decision stage, where it was observed, its commercial value, and which page currently answers it. That structure makes it a content plan, a prompt set, and a gap analysis simultaneously.
Sources are internal first: support tickets, sales call notes, chat logs, and site search queries, all of which record real phrasing from real buyers. Then external: People Also Ask, community discussion, and direct observation of what engines return for category questions. Two days of collection produces something considerably more useful than a keyword export.
The step that makes the map manageable is collapsing phrasings into jobs. Ten differently-worded questions frequently express one need, and treating them as ten targets produces redundant content that competes with itself — the same failure one-page-per-keyword produced, repeated in a new vocabulary.
The test is whether one well-written passage would satisfy all of them. Where it would, they are one job. Where it would not, the distinction is real and worth separate treatment. Applied honestly this usually reduces a list of two hundred observed questions to thirty or forty jobs, which is a plan rather than a backlog.
The second thing that survives the shift is entities, and they matter more than they did. Engines reason about things rather than strings, which means the products, categories, problems, and organisations your content concerns are what it is understood to be about.
That has a planning consequence: naming entities the way the wider world names them matters more than using the keyword variant with the highest volume. A brand describing its category in language only it uses has an entity problem dressed as a keyword decision, and the volume figure that justified the choice is measuring the wrong thing.
Keyword list versus question map
Side-by-side columns: what each records, what it is good for, what it misses, and what artefact it produces. Rows for sizing demand, specifying content, structuring pages, and measuring presence. Makes visible that only one of them doubles as a measurement instrument.
An honest observation about the current state: no tool does question research well. Keyword platforms are adding question data and it is largely derived from search-box phrasing rather than from conversational input, which means it captures the shorter, more compressed end of the distribution — exactly the part that was never the problem.
That gap will close, and until it does the work is manual and the advantage goes to teams willing to do it. Collecting real questions from first-party channels is unglamorous and produces better input than any current tool, which is an unusual and temporary situation worth exploiting while it lasts.
A calendar built from a keyword list is ordered by volume, which systematically front-loads informational content and defers commercial questions. A calendar built from a question map is ordered by job value and decision stage, which usually inverts it.
That inversion is the practical output of this whole argument. Same team, same capacity, different sequence — and the second sequence answers the questions buyers ask while deciding rather than the ones they ask while browsing. Most of the benefit here comes from reordering rather than from producing more.
It would be overstating the case to imply questions are always the better unit. Navigational demand is genuinely string-shaped — people type brand names and product names, and no conversational reframing improves that. Transactional queries are frequently short and compressed for the same reason: the user knows exactly what they want.
So the string remains the right unit at both ends of the funnel and loses its grip in the middle, where consideration happens and where phrasing expands. That middle is also where most content investment goes, which is why the shift matters despite the string surviving at the extremes.
If we overstated this, it would sound like: keyword research is obsolete, volume data is worthless, and everything should be rebuilt around conversational questions. That would be wrong on all three counts and would cost a reader real capability.
The defensible version is narrower. The string is degrading as a proxy in a specific segment of demand, tools undercount that segment in a knowable direction, and the planning unit that survives both interfaces is the job. Everything else — the tools, the volume data, the difficulty estimates — keeps working for what it was always good at.
A brand with hundreds of pages built on a keyword plan does not need to rebuild them. It needs to map what it has against the question map and identify three states: pages that answer a real job well, pages that target a string variant of a job another page already covers, and jobs with no page at all.
The first group needs structural work only. The second should be consolidated, since multiple pages targeting variants of one job compete with each other and dilute signals. The third is the actual content gap, and it is usually smaller than expected and weighted toward decision-stage questions the keyword plan never surfaced.
The underrated benefit of the question map is that it is simultaneously a measurement artefact. The same questions that specify what to write are the prompts you track citation against, which means content planning and visibility measurement stop being separate exercises with separate inputs.
A keyword list cannot do this, because nobody asks an assistant a three-word string. That single property — that the planning unit and the measurement unit are the same thing — is probably the strongest practical argument for the shift, independent of any claim about keyword decline.
Less changes than the framing suggests. The work is still finding out what people want, producing something that serves it, and being findable. What changes is the artefact in the middle: a compressed string list gives way to a question map, and the map is more faithful to the demand it represents.
That is an improvement in instrumentation rather than a revolution in practice. The industry has been planning against a lossy proxy for twenty years because it was the only thing countable. The proxy is now lossier and the alternative is finally observable, which makes this a good moment to switch rather than a crisis requiring one.
Keywords are not dead. The keyword as the unit of demand is finished, because conversational input broke the compression that made strings a workable stand-in for needs. Tools now undercount that demand in a direction we can name, which makes volume a partial input rather than a plan.
Plan at the job level, build a question map from real observed phrasing, use keyword tools for sizing and vocabulary, and structure content as answered questions. The map doubles as your citation prompt set, which is the property that makes the shift worth making regardless of what you believe about the pace of everything else.
From observed question to content brief
Linear flow: collect real phrasings from first-party channels → cluster into jobs → assign decision stage and commercial value → map to existing pages → identify gaps → brief with the question as the specification. Annotate the point where the same artefact becomes the citation prompt set.
The shift bites hardest where consideration is longest. B2B buyers researching a complex purchase ask conditional, multi-part questions with their situation embedded — exactly the phrasings that fragment beyond what any tool aggregates. A B2B keyword plan therefore misses a larger share of real demand than a consumer one does.
Consumer categories retain more compressed, string-shaped demand, particularly in transactional and navigational queries. So the practical urgency of this shift varies by business model, and B2B organisations with long consideration cycles should treat the question map as the primary artefact rather than as a supplement.
Pull the last fifty questions your sales or support team actually received, written as the customer wrote them. Cluster them by what the person was trying to accomplish. Check which of those jobs has a page that answers it directly, in a passage that would stand alone.
That exercise takes two hours, requires no tooling, and consistently surfaces two or three high-value jobs with no adequate page against them. It is also the first step of building the question map, which means the smallest useful version of this argument is available immediately.
This is a planning-layer argument. The technical pieces in this publication cover how retrieval works and why passages rather than pages are the competing unit; this covers what you should be planning to write in the first place, and the two connect directly — the question map specifies content, and passage structure delivers it.
Readers who want the research method for building the map should read the prompt research material in our Academy. Readers wanting the measurement side should read the benchmark and index reports, both of which use question-set construction as their foundation. The question is the unit throughout, which is the consistency worth noticing.
Stop deriving your plan from a keyword export and start deriving it from questions your buyers actually asked, clustered by what they were trying to accomplish. Keep the tools for sizing and vocabulary, where they remain the best available instrument.
That one change reorders the content calendar, improves the structure of what gets written, and produces the prompt set you need for measurement — three benefits from one artefact, which is unusual enough to be worth the two days it takes to build.
The strongest counter-argument is that this is a distinction without a difference — that good keyword research always meant understanding intent, that experienced practitioners have always planned around needs rather than strings, and that we are describing best practice as though it were a discovery.
That is fair for the best practitioners and describes a minority of the discipline. The observable reality is that most content plans are still keyword exports ordered by volume, most briefs still specify a target term, and most calendars still front-load informational topics because that is where the volume sits. Where the objection holds, this article is a restatement. Where it does not, which is most places, it is a correction.
There is something instructive in the fact that our planning artefact was determined for twenty years by what happened to be countable. The keyword became the unit not because it was the right one but because it was the one tools could aggregate, and the entire discipline organised around that constraint.
The constraint has changed and the artefact has not, which is the actual subject of this article. Worth asking, periodically, which of your other planning units exist because they are correct and which exist because something could count them.
There is a difference between a keyword being a bad target and a keyword being bad information. It remains excellent information — it tells you people search, roughly how many, and in what words. What it has stopped being is a specification for what to build.
Treating information as specification is the actual error, and it predates answer engines. A volume figure was never an instruction to write a page; it was evidence that demand existed in a particular vocabulary. Reading it as evidence rather than as a brief is most of the correction this article argues for.
Should we stop paying for keyword tools?
No. They remain the best available instrument for sizing traditional search demand and for surfacing how customers name things. The change is what you do with the output, not whether you gather it.
How do we size demand for questions we cannot get volume on?
By frequency of observation rather than by reported volume: how often does this appear in support, in sales calls, in community discussion. It is less precise and it is measuring the right thing, which is the better trade.
Does this mean one page per question?
No — one self-contained answer per question, and several can live on one page. That is the structure retrieval rewards and it avoids the thin-page proliferation that one-page-per-keyword produced.
Will keyword volume data get better at capturing this?
Possibly, though the fragmentation is structural rather than a coverage gap. A hundred phrasings of one need do not aggregate cleanly without knowing they share a need, which requires the semantic step tools are only beginning to apply.
Keywords are not dying. The assumption that a keyword is the unit of demand is, because conversational input broke the compression that made strings a workable proxy for needs. The same need now produces a hundred phrasings, none of which aggregates into a volume figure, which means keyword tools are becoming incomplete in a specific and knowable direction rather than becoming wrong.
What replaces the keyword as the planning unit is the job — the need behind many questions — expressed as a question map rather than a keyword list. That artefact is less tidy, more actionable, doubles as your citation measurement prompt set, and survives interface changes that strings do not. Keep the tools for sizing and vocabulary. Stop deriving the plan from them.
Methodology note: we state no figure for how far keyword volume undercounts conversational demand, because no dataset establishes it — the fragmentation is precisely what makes it unmeasurable with current instruments. The claim that question-shaped queries frequently report zero volume is observational. A study would compare observed question frequency in first-party channels against reported search volume for the same underlying job, which would quantify the gap for a given category.
“The keyword was always a compressed proxy for a need. Conversational input broke the compression — which means the measurement is degrading while the demand it measured is entirely intact.” The Age’X Research Team
Get a free GEO audit — the same analysis behind every article here.