The proxy worked for twenty years. It now permits a number to rise while the client gets worse — and no dishonesty is required.
The ranking report is the most successful product our industry ever built and the reason it is now structurally exposed. It made an abstract service legible, gave clients something to hold, and let agencies demonstrate progress before revenue arrived. It also trained an entire market to buy a proxy — and the proxy has started decoupling from the outcome it stood for. Agencies still selling rankings are selling a measurement that is quietly ceasing to mean what everyone agreed it meant.
It is worth defending the practice before criticising it. Search results were invisible to clients, outcomes took months, and attribution was contested. The ranking report solved a real problem: it made an intangible service demonstrable, gave a client something to check between invoices, and provided a shared language for progress.
It also had genuine predictive validity. For most of the industry’s history, moving from page two to position three reliably produced more visits, and more visits reliably produced more revenue. Selling the proxy was defensible because the proxy worked. The criticism that follows is not that the industry was foolish to adopt it; it is that the conditions justifying it have changed and the practice has not.
Two things, both structural. Answer features now sit above the ranked list and satisfy a substantial share of queries before anyone reaches it, which severs position from click. And the sources those answers cite are frequently not the sources ranking beneath them, because citation is decided by different properties than ranking is.
The result is a measurable divergence: a brand can hold the top position, see its rankings reported as excellent, and be entirely absent from the answer occupying the screen above its listing. Nothing in the ranking report is wrong. It is measuring a real thing that has stopped standing for the thing it was purchased to stand for.
This is the part that deserves the discomfort. An agency compensated on ranking improvement is rewarded for a number that can rise while the client’s commercial outcome falls. That is not a hypothetical — it is the expected case on queries where answer coverage is expanding, and it will become more common rather than less.
No dishonesty is required for this to happen. The agency does good work by the terms of the contract, the report shows improvement, the client’s traffic declines for reasons the report does not contain, and the relationship ends in mutual confusion. The problem is the contract rather than the conduct, which is precisely why it will not be solved by better intentions.
Four questions for assessing whether a metric still deserves to be sold. Rankings passed all four for two decades and now fail two.
Does it predict the outcome?
Is movement in this metric reliably followed by movement in the commercial result?
Rankings: weakening. Position no longer predicts visits on covered queries.
Can it improve while the outcome falls?
Is there a plausible state where the metric rises and the client is worse off?
Rankings: yes, and increasingly the expected case.
Does it capture the whole surface?
Does it measure all the places the outcome is now decided?
Rankings: no. It cannot see inside the answer above the list.
Would the client choose it if they understood it?
Presented with what it does and does not measure, would they still want to be sold on it?
The uncomfortable one, and the one that settles it.
The standard defence is that rankings still drive traffic, which is true and insufficient. A substantial share of queries carry no answer feature, positions on those still deliver, and abandoning rank tracking entirely would be an overcorrection that discards a useful instrument.
But “still useful for part of the surface” is a different claim from “adequate as the headline deliverable”. A report leading with rankings implies rankings represent the outcome. Once they represent a shrinking portion of it, leading with them is a presentational choice that systematically flatters, and the fact that it flatters in the seller’s favour is not incidental.
An agency compensated on rankings is rewarded for a number that rises as answer coverage removes the clicks beneath it. No dishonesty required — the contract produces the outcome on its own.
The tempting fix is to substitute citation share for rank and carry on as before. That would repeat the error one layer up. Citation share is a better instrument for the current surface and it is still a proxy, and selling it as the deliverable recreates the same structure with a newer number.
What should replace the metric is a different contract. The deliverable becomes a diagnosis: here is what is currently preventing the outcome, here is the single constraint we will address, here is what should change and by when, and here is how you will know if we were wrong. The measurement supports that claim rather than being sold as the product.
Concretely, an agency should be willing to state: this brand is failing at independent category association; we will address it; within two quarters, description specificity should improve and presence on segment questions should stabilise; if neither moves, our diagnosis was wrong and we will say so.
That is a considerably more exposed position than a ranking report, which is why it is rare. It is also the only version of this service that a client can genuinely evaluate. An engagement that cannot be wrong within a defined window is asking for trust it has not earned, and the industry has been asking for exactly that for a long time.
Falsifiable engagements need measurement the client can see. DUNkē tracks citations across eight AI engines — per prompt, against named competitors — which is what makes a prediction checkable.
This cuts both ways, and buyers hold more responsibility than they usually accept. Four things are reasonable to require of any search engagement: a measurement baseline established before work begins; a single named constraint rather than a prioritised list of forty findings; a falsifiable prediction with a timeframe; and a stated separation between diagnostic and remediation incentives.
Most agencies cannot supply all four, which is precisely what makes the requirement useful as a filter. It is also a fair test of us. Any reader applying it to The Age’X should expect the same answers, and should treat an inability to produce them as disqualifying regardless of who is being assessed.
The resistance is structural rather than cultural. Ranking reports scale: they are automated, comparable across clients, and produce a monthly artefact with no marginal cost. Diagnostic engagements do not scale in the same way — they require judgement, they produce narrow recommendations, and they occasionally conclude that the client should spend less.
An industry optimised for the first will not adopt the second because someone argued it should. It will adopt it when clients stop buying the first, which is beginning to happen for a reason that has nothing to do with ethics: the reports stopped predicting the outcomes, and buyers noticed. Commercial pressure will accomplish what argument will not.
Before dismantling it, it is worth crediting properly. The ranking report solved an information asymmetry problem: clients could not observe the work, could not evaluate the expertise, and could not wait months for revenue evidence. A weekly artefact showing movement made a trust relationship possible where none could otherwise have formed.
It also disciplined agencies. A visible metric that moves is harder to fake than a narrative, and it forced practitioners toward work that produced observable results. The industry professionalised partly because of it. Understanding what it solved is necessary before proposing a replacement, because the replacement has to solve the same underlying problem: how does a client evaluate work they cannot see.
The replacement addresses the asymmetry differently. Rather than a frequent artefact showing movement, it offers a specific prediction with a deadline: this constraint is binding, we will address it, these indicators should move by this date, and if they do not our diagnosis was wrong.
That gives a client something better than a dashboard — a test they can run on the supplier. It requires more expertise to offer and more courage to stand behind, which is precisely why it filters. An agency unable to make a falsifiable prediction about the work it is proposing has revealed something useful about how well it understands the problem.
This model tends to reduce contract value rather than increase it, which is worth stating since the obvious cynical reading is that consultancies invent frameworks to charge more. Diagnostic engagements are short. Their most common output is a narrow programme, and a meaningful minority conclude that the client should spend less or spend elsewhere.
That is commercially inconvenient and it is the mechanism that makes the diagnosis trustworthy. An auditor whose recommendation always happens to be the service they sell has an incentive problem no amount of methodology fixes. Separating diagnosis from remediation, and being willing to hand the work elsewhere, is what resolves it — at the cost of revenue that was never legitimately yours.
Two contract structures
Columns for ranking-based and diagnostic engagements. Rows: deliverable, reporting cadence, what the client can verify, failure mode, incentive alignment, typical scope trajectory. Makes visible that one produces artefacts and the other produces decisions.
The buyer side carries responsibility that is rarely acknowledged. Clients ask for rankings because rankings are what they were taught to ask for, then judge on traffic, then evaluate on revenue — three different measures at three stages, none agreed at the outset. Agencies optimise the one in the contract and get criticised on the one in the boardroom.
The fix is available to buyers immediately: agree the measure before the engagement, require it to be the same measure throughout, and require it to be something the agency can plausibly influence within the contract period. Most disputes in this industry come from that alignment never having been done, and it costs nothing but a conversation.
An agency wanting to change this faces a genuine commercial obstacle: clients are buying rankings from competitors, and refusing to sell them looks like refusing to sell. The practical route is not refusal but reordering — report rankings as context, lead with the diagnosis and the citation measurement, and let the comparison make the argument.
Within a quarter or two the reordered report explains movements the ranking report cannot, and the client stops asking for the old order. That is a slower path than declaring a new philosophy and it works, because it demonstrates the deficiency rather than asserting it.
The strongest counter is that rankings remain a good proxy for most brands in most categories, and that the divergence is concentrated in a subset of query types large enough to matter to specialists and small enough to be a rounding error for everyone else. If that is right, this argument is premature.
Our reading of the published evidence is that the affected share is substantial and growing, but we cannot demonstrate the threshold at which the proxy becomes untenable for a given brand — that depends on its specific query mix. The honest position is that every brand should measure its own divergence rather than accept either our argument or the industry default, which is a conclusion that does not favour us particularly.
Concretely, a report that survives the decay test leads with the outcome measures the client actually cares about — qualified traffic, conversions, revenue where attributable — followed by competitive position across both surfaces, followed by the trend, followed by a plain-language account of what changed and what is being done.
Rankings appear, as context, alongside citation presence for the same query set. Neither is presented as the outcome. The narrative section is the deliverable and the dashboard supports it, which is the inversion of how most reports are constructed and the whole substance of the change being proposed.
Any argument about incentives eventually reaches how people are paid, and this one has an uncomfortable answer: performance compensation tied to any single proxy recreates the problem regardless of which proxy is chosen. Citation-share-based compensation would produce the same distortion one layer up.
The structures that hold up tie compensation either to genuine business outcomes, accepting the attribution difficulty, or to the delivery of diagnostic work whose quality can be assessed independently of the metric. Neither is clean. Both are better than paying for movement in a number that can rise while the client’s position deteriorates.
The argument applies internally with a sharper edge, because in-house teams set their own targets. A search function measured on rankings will optimise rankings, report progress accurately, and be blamed for a revenue outcome its objectives never pointed at.
The fix is available without any supplier negotiation: change what the function is measured on. Qualified traffic and commercial outcomes as the objective, rankings and citation presence as diagnostic instruments. That is a conversation with a CMO rather than a procurement exercise, and it removes the misalignment that produces most of the frustration in the discipline.
Practically, replacing a ranking promise with a diagnostic one changes what an agency can say in a sales conversation. Rather than projecting position improvements, the offer becomes: we will establish where you actually stand, identify what is preventing the outcome, predict what should change and by when, and be wrong in public if we are.
That loses deals to competitors promising numbers, which is a genuine commercial cost and not one we can argue away. It also selects for clients who will be good to work with, since a buyer who prefers a falsifiable diagnosis to a projected ranking has already demonstrated how they will evaluate the work.
Rankings were a good proxy for two decades because position predicted visits. Answer features severed that relationship and citation is decided by different properties, so the metric can now improve while the client’s outcome degrades — and an agency compensated on it is structurally rewarded for that.
The replacement is not a better metric but a different contract: a baseline, one named constraint, a falsifiable prediction with a deadline, and separated diagnostic and remediation incentives. Clients should require all four and should apply the requirement to us as readily as to anyone else. An engagement that cannot be wrong is asking for trust it has not earned.
It would be inconsistent to argue this and not accept the standard. Applied to us: we should be able to produce a measurement baseline before work begins, name a single constraint rather than a list, state a falsifiable prediction with a timeframe, and separate diagnostic from remediation incentives.
Readers should require that, and should treat an inability to supply any of the four as disqualifying regardless of who is being assessed. An argument about industry standards that exempts its author is marketing wearing the clothes of criticism, and the only way to avoid that is to state the test in a form that applies to us first.
Rankings were an honest proxy that has stopped tracking the outcome it stood for, and continuing to sell them as the deliverable creates an incentive to be rewarded while a client gets worse. That is a structural problem in the contract rather than a moral failing in the practitioners, which is why it persists among people acting in good faith.
The replacement is a falsifiable diagnosis with separated incentives — harder to sell, easier to evaluate, and frequently smaller in scope than what the client expected to buy. The industry will not adopt it because it is argued for. It will adopt it when buyers start asking the four questions, which costs them nothing and is the only lever that has ever moved this business.
Should this metric be your headline deliverable?
Four sequential gates from the Proxy Decay Test: does it predict the outcome, can it rise while the client falls, does it cover the whole surface, would an informed client choose it. Terminal nodes: headline deliverable, supporting context, or retire. Run rankings and citation share through it as worked examples.
If you take one thing from this as a client rather than as a practitioner, it is four questions to ask any prospective supplier. Will you establish a measurement baseline before starting? Will you name one constraint rather than a list? What specifically should change, and by when? And are your diagnostic and remediation incentives separated?
Four questions, no expertise required to ask them, and they filter more effectively than any credential check. Most suppliers cannot answer all four, which tells you something useful before you have spent anything. Ask them of us as readily as of anyone else — an argument like this earns nothing if its author is exempt from it.
An agency publishing an argument against the industry’s standard deliverable invites an obvious question about motive. The answer is that we sell the alternative, and readers should weight the argument accordingly rather than treating it as disinterested.
What makes it worth publishing anyway is that the underlying claim is checkable without trusting us. Rank-citation divergence is documented in published third-party research. The incentive problem follows logically from it. And the four questions we propose apply to us identically. An argument that survives its author’s conflict of interest being disclosed is one worth making.
The empirical basis for this argument is the rankings-versus-recommendations analysis elsewhere in this publication, which sets out the four populations produced by crossing the two axes and demonstrates the divergence this piece treats as given.
The alternative deliverable is described in the audit breakdown and the framework articles: a baseline, a named constraint, a falsifiable prediction. This piece makes the case for why that contract should replace the ranking report; those describe what it actually looks like in practice.
A metric that can improve while the client’s outcome degrades should not be the deliverable, and rankings have become one. The replacement is not a newer number but a different contract: measure first, name one constraint, predict what should change and by when, and separate the incentive to diagnose from the incentive to remediate.
Buyers hold the lever here, not the industry. Four questions, asked of every supplier including us, would change this business faster than any amount of argument about what it should be.
Should we stop rank tracking entirely?
No. It measures a real surface that still delivers traffic. The change is presentational and contractual: report it as context alongside citation presence, and stop selling it as the outcome.
How do we price a diagnostic engagement?
Usually as a short fixed-scope diagnostic followed by a separate decision about remediation. Separating them removes the incentive to find the constraint you happen to sell against, which is the point.
What if the diagnosis says the client needs PR, not SEO?
Then say so. An agency willing to hand work to another function is demonstrating that the diagnosis is real, and in our experience it is the single strongest thing you can do for a client relationship.
Is this not just repositioning consultancy as a premium product?
It is a narrower deliverable that frequently reduces scope. If it were a pricing strategy it would be an unusually poor one, since its most common output is a recommendation to do less than the client was expecting to buy.
Rankings were a defensible proxy for two decades because position reliably predicted visits. Answer features severed that relationship, and citation is decided by properties ranking does not measure — so a brand can hold the top position and be absent from the answer above it. The report is not wrong; it has stopped standing for what it was bought to stand for.
The structural problem is the incentive: an agency compensated on rankings is rewarded for a number that can rise while the client’s outcome falls, with no dishonesty required. The replacement is not a newer metric but a different contract — a baseline, a single named constraint, a falsifiable prediction with a timeframe, and separated diagnostic and remediation incentives. Clients should demand all four, including of us, and should treat an inability to supply them as disqualifying.
Methodology note: this is an argument about industry practice, not a study of it. We have no measurement of how many agencies lead with ranking reports or how client outcomes correlate with contract structure, and we state no figures. The empirical basis is limited to the published research on rank-citation divergence, which is cited; everything built on it is interpretation.
“An agency paid on rankings is paid for a number that can rise while the client gets worse. No dishonesty is required — the contract does it on its own.” The Age’X Research Team
Get a free GEO audit — the same analysis behind every article here.