Is your citation rate good? Nobody has published the distributions that would answer it. This sets out how we are building them.
Every brand measuring AI visibility eventually asks the same question: is this number good? Citation on a fifth of your priority prompts could be dominant or negligible depending entirely on your category, and nobody has published the distributions that would answer it. This report exists to build them — and this edition sets out the methodology, because a benchmark whose sampling frame is undisclosed is worse than no benchmark at all.
A brand cited for twenty per cent of its priority prompts has learned almost nothing from that figure alone. In a category dominated by reference sources and community discussion, twenty per cent may be exceptional. In a category where commercial brands routinely appear, it may be poor. The number without a distribution is uninterpretable.
This is the gap the report addresses. Every brand doing this work has its own measurement and no way to situate it, which means internal debates about whether performance is adequate are conducted on intuition. A distribution converts that into a position — not a target, but a location relative to what is actually achieved in comparable conditions.
The single most important methodological decision is that everything is reported by category rather than in aggregate. Citation behaviour varies enough between categories — because source mix, coverage, and competitive structure all differ — that a cross-category average describes nothing anyone operates in.
A blended figure would be arithmetically correct and strategically useless, in exactly the way a market-wide prevalence figure is. So the report reports distributions within categories, and where a category has too few brands sampled to support a distribution, it says so rather than reporting a thin one.
The sampling and reporting design. Published in full because a benchmark is only as trustworthy as the frame that produced it.
Stratified by category
Brands sampled within defined categories, with the category definitions published.
Cross-category averages are never reported.
Stratified by size
Within category, brands sampled across size bands, because size correlates with corroboration.
Prevents large-brand skew from setting the benchmark.
Standard prompt construction
Prompts built to a fixed template per category so measurement is comparable between brands.
Not brand-supplied prompt sets, which self-select.
Percentile reporting
Distributions rather than averages, since the shape matters more than the centre.
Median, quartiles, and range published.
Minimum sample rule
Categories below a threshold sample size are reported as insufficient rather than estimated.
No thin distributions presented as benchmarks.
The most serious threat to any benchmark in this field is that the easiest brands to measure are the ones already doing the work. A benchmark built from clients, or from brands who volunteered, describes an unusually competent population and would systematically make everyone else feel worse than the market warrants.
The frame therefore samples brands within a category independently of any relationship with us, using a standard prompt construction rather than brand-supplied question sets. That is more work and it is the only version that produces a benchmark rather than a customer survey. Where we cannot achieve an unbiased sample in a category, the correct output is to report the category as insufficiently sampled.
Per category: the distribution of citation rate across the sampled brands, reported as median and quartiles rather than as an average; the distribution of description quality, since being cited badly differs from being cited well; the composition of the cited set, showing what proportion of citations go to commercial brands versus reference, community, and editorial sources; and the concentration — whether citations cluster among a few brands or spread.
That last measure is underrated. A category where three brands take most commercial citations is a different competitive environment from one where twenty share them, and it changes whether entering is realistic. Concentration tells a brand whether the ceiling above it is a wall or a queue.
Citation rate on its own is uninterpretable. Category variance dwarfs brand variance, which means a cross-category average describes a market nobody operates in.
The underlying measures move slowly. Corroboration accrues over quarters, entity confidence changes over months, and category competitive structure shifts more slowly still. A monthly benchmark would report sampling variance as change and invite exactly the over-reaction this publication argues against elsewhere.
Quarterly matches the rate at which the thing being measured actually moves, which is the correct basis for reporting cadence. It also allows a sample large enough to support distributions, which monthly production would not.
Three cautions. A benchmark is a location, not a target — being at the median is not automatically adequate if your commercial position requires more. Categories are approximations, and a brand at the boundary of two should read both. And a distribution describes what is currently achieved, not what is achievable, which are different things in a young discipline where most participants are early.
Used properly it answers whether your number is unusual and in which direction. Used improperly it becomes a target that anchors ambition to the current average, which in a formative period is a low bar. The right reading is diagnostic: where am I, and is the gap explicable.
Comparison requires consistent measurement. DUNkē tracks citations across eight AI engines — per prompt, per surface, against competitors — on a method compatible with how this benchmark is built.
It cannot tell you whether your citation rate is commercially sufficient, because that depends on your margins, your category’s purchase cycle, and what proportion of your buyers use these interfaces — none of which a benchmark observes. A brand at the seventy-fifth percentile in a category where AI-influenced purchasing is minimal may still be over-investing.
It also cannot establish causation. Brands at the top of a distribution differ from those at the bottom in many ways, and the benchmark records position rather than explaining it. Where we describe apparent patterns among high performers, that is observation offered as a hypothesis rather than as a finding.
The second stratification is by brand size, and omitting it would produce a benchmark that quietly sets large-brand performance as the standard. Size correlates with corroboration because large brands are written about more, which means an unbanded distribution would place most brands below a median driven by companies with structural advantages.
Banding makes the comparison useful: a mid-market brand compares against mid-market brands, where the achievable range is genuinely comparable. It also reveals something interesting in itself — the size-performance relationship is weaker than expected, which is consistent with what audit comparisons show and worth reporting as a finding rather than assuming.
Citation rate alone treats all citations as equivalent, and they are not. A brand named as one option among five is in a different position from one named as the recommended choice, and a brand described accurately differs from one described generically or with stale facts.
The benchmark therefore reports description quality as a separate distribution: specificity, accuracy, and hedging assessed on a consistent scale. A brand at the median on citation rate and the top quartile on description quality is in a materially stronger position than the reverse, and a single blended score would hide that entirely.
The concentration measure — how citation share distributes across sources within a category — is the most strategically loaded number in the report. High concentration means a few sources hold most of the ground and entry requires displacement. Low concentration means the field is open and joining is achievable.
It also predicts how much benchmark position matters. In a concentrated category, being at the median is being nowhere, because the meaningful share sits with the top few. In a dispersed one, median position represents real presence. That interaction between concentration and percentile is worth reading before drawing any conclusion from a position.
Citation rate against brand size, by category
Points plotted by size band on the x-axis and citation rate on the y-axis, coloured by category, with the category median as a horizontal line. The expected finding — wide vertical spread within each size band — is the evidence that size is not the determining variable.
For a benchmark to be usable, a reader has to be able to measure themselves the same way. That means publishing the prompt construction template, the sampling protocol, and the definition of citation used — because a brand computing its own figure on a different basis is comparing incomparable numbers and will reach a wrong conclusion confidently.
This is the practical argument for publishing methodology in full rather than summarising it. A benchmark whose method is described but not specified produces a number readers cannot replicate, which makes the comparison decorative. We would rather publish something checkable than something impressive.
Coverage expands as category samples reach the minimum threshold, which means early editions will report a small number of categories properly rather than many categories thinly. Categories will be added as their samples qualify, and the covered list is stated in each edition.
The alternative — launching with broad coverage built on thin samples — would produce distributions that later corrections would have to undo, and benchmark figures are unusually sticky once published. A number that gets quoted for two years and then revised has done more harm than a delayed release.
Three failure modes worth pre-empting. Treating the median as a target, which anchors ambition to the current average in a discipline where most participants are early. Comparing across categories, which the report structure discourages and readers do anyway. And using a benchmark position to justify investment level without reference to commercial context.
That third is the most consequential. A brand at the twenty-fifth percentile in a category where its buyers barely use these interfaces may be correctly invested. Benchmarks measure relative position, not commercial adequacy, and conflating the two produces both over-investment and complacency depending on which side of the median a brand lands.
The benchmark answers whether your number is unusual. The index answers who holds the category. The trends report answers what changed environmentally. The movement report answers who gained and lost. They are deliberately separate because each requires a different sampling design and combining them would compromise all four.
A brand using them together would read the benchmark to situate its performance, the index to understand the competitive structure, the trends report to check for environmental explanation, and the movement report to see whether anyone is gaining. Four instruments, four questions, one reason not to collapse them into an index number.
The most likely failure is sampling drift: categories where the initial sample was representative becoming unrepresentative as the market changes, without the frame being revisited. Benchmarks are unusually vulnerable to this because they are refreshed rather than rebuilt, and drift accumulates invisibly.
The protection is periodic frame review published alongside the data, and stating when a category’s sample was last reconstructed. A benchmark whose sampling frame has not been revisited in two years is reporting a distribution from a market that may no longer exist, and readers should ask that question of any benchmark they rely on.
The practical value of knowing your percentile is narrower than it sounds. It settles internal arguments about whether performance is adequate, which are otherwise conducted on intuition. It calibrates ambition, by showing what is actually achieved rather than what sounds impressive. And it identifies when a result is anomalous enough to warrant investigation in either direction.
What it does not do is tell you what to change. That comes from your own diagnostic, and a brand that knows it is at the twentieth percentile and does not know why has learned only that a problem exists. Benchmarks locate; they do not diagnose.
Because a benchmark’s credibility rests entirely on its sampling, and publishing numbers first invites readers to evaluate the findings before they can evaluate the method. Once a figure is in circulation it gets quoted regardless of its frame, and correcting it later is nearly impossible.
Publishing the frame first also invites criticism while it can still change the design, which is the point of publishing methodology at all. A benchmark whose method arrives after its numbers has already made the decisions that matter and is describing them rather than submitting them.
Citation rates are uninterpretable without distributions, and none exist publicly. This report builds them by category and size band on a sample drawn independently of client relationships, using a standard prompt construction, reported as percentiles rather than averages because the shape matters more than the centre and category variance dwarfs brand variance.
Read your own category, check concentration before drawing conclusions from rank, treat the median as a location rather than a target in a discipline where most participants are early, and interrogate the sampling frame of any benchmark you rely on. The frame is not a footnote — it determines whether the numbers describe anyone at all.
How the sampling frame is constructed
Nested diagram: category definitions at the outer level, size bands within each, sampled brands within each band, and the standard prompt construction applied uniformly. Mark where relationship-based sampling would enter and is deliberately excluded. Annotate the minimum-sample threshold below which a category is not reported.
A properly-sampled benchmark will show most brands performing worse than they believe, including brands doing genuinely good work, because the discipline is young and almost everyone is early. That is not a pleasant message and it is the accurate one.
It also means the benchmark will occasionally contradict a vendor’s claims about their clients’ results, including ours. We have built the frame to exclude our own clients partly for that reason: a benchmark we could influence would be worthless, and a benchmark that can embarrass us is the only kind worth publishing.
If unbiased sampling at sufficient scale proves impractical across enough categories, the honest response is to stop rather than to publish thin distributions. A benchmark built on a sample too small or too skewed to represent its category is worse than none, because it gets quoted.
We would rather publish two categories properly than twelve inadequately, and if it turns out we can only do two, that is what will appear. Stating the abandonment condition in advance is a form of precommitment, and it is the only thing preventing the gradual slide from rigorous to convenient that afflicts most industry benchmarks.
Second of four. The trends report covers the environment; this covers distributions — whether your number is unusual. The index covers who holds the category and the movement report covers who is changing. Reading all four in sequence answers most of what a brand needs to know about its position.
This one is the slowest to produce because unbiased category sampling at scale is genuinely difficult, and it is the one most likely to be released incrementally. Categories appear as their samples qualify rather than on a schedule, which is the correct behaviour for a benchmark and an awkward one for a publication calendar.
Before concluding that your citation rate is good or bad, find out what the distribution looks like in your category and at your size band. Absolute numbers in this discipline are uninterpretable, and most internal arguments about whether performance is adequate are conducted without the one input that would settle them.
Where no distribution exists yet for your category, the honest position is that you do not know — which is more useful than a confident judgement built on an intuition about what a good number looks like.
If we could publish only one figure per category, it would be the median citation rate for mid-size brands — because that is the number most readers are implicitly asking for when they ask whether their performance is adequate, and because it is the least distorted by the structural advantages of large incumbents.
We publish distributions rather than that single number because the spread within a band is wide enough that a median alone would mislead in both directions. But if a reader takes one thing from an edition, that is the figure to find, with the quartile range beside it for context.
Benchmarks change behaviour, which is a responsibility rather than a feature. A published median becomes a target, a target becomes a ceiling, and a discipline anchors its ambition to the average of an early cohort. That has happened in every marketing measurement that became standard.
The mitigation is to report distributions rather than centres, to state repeatedly that a median is a location rather than a standard, and to publish the concentration data that shows when a median position represents nothing. Whether that is sufficient is uncertain, and it is the reason we approached this report last rather than first.
The narrow purpose is settling one internal argument: is our performance adequate. That argument is currently conducted on intuition in almost every organisation doing this work, and intuition in a young discipline is calibrated against nothing.
A distribution replaces intuition with a location. It does not replace judgement about whether that location is commercially sufficient, which depends on factors no benchmark observes. But it removes the part of the disagreement that was never about judgement in the first place, which is usually most of it.
Citation rates mean nothing without a distribution to place them in, and no public distributions exist. This builds them by category and size band, on a sample drawn independently of any client relationship, reported as percentiles because the spread matters more than the centre.
It answers one question: is our number unusual, and in which direction. It deliberately does not answer whether that number is commercially sufficient, because that depends on margins, purchase cycles, and buyer behaviour no benchmark can observe. Read your own category, check concentration, and interrogate the sampling frame of anything claiming to benchmark you.
Of the four reports in this series, the benchmark is the one we have been slowest to publish, and the reason is sampling. Producing distributions that describe a category rather than a client list requires identifying and measuring brands with no relationship to us, at sufficient scale, per category.
That is genuinely laborious and it is the entire value of the report. A benchmark that took a week to produce would have been built from data we already held, which is exactly the population that would make it useless. The delay is the methodology working, and stating that seems better than implying the timeline was arbitrary.
There is a reasonable position that says nobody should publish benchmarks in a discipline this young, because the numbers get quoted, become targets, and anchor a field’s expectations to an early and unrepresentative cohort. We have some sympathy with it.
What tips the balance is that brands are already making investment decisions against imagined benchmarks — a sense of what a good number looks like, formed from vendor claims and conference anecdotes. Replacing an unexamined intuition with a measured distribution and its stated limits is an improvement even if the distribution is imperfect, provided the limits travel with it. Whether they will is the risk we are accepting.
Which categories are covered?
Those where an unbiased sample of sufficient size is achievable. Categories below the threshold are reported as insufficiently sampled rather than estimated, and the covered list is published with each edition rather than assumed to be stable.
Can we get our brand included?
Inclusion is by sampling frame rather than by request, because brands opting in would bias the distribution toward those already invested. The benchmark is more useful to you if you are not in it than if participation were self-selected.
How do you handle brands in multiple categories?
They are sampled within the category the prompt construction targets, which means a brand can appear in more than one distribution with different results. That is informative rather than contradictory — it usually indicates uneven category association.
Why publish the method before the data?
Because the method determines whether the data means anything, and readers should be able to assess it independently of whether the findings flatter or challenge them. A benchmark whose frame arrives after its numbers has the order backwards.
A citation rate is uninterpretable without a distribution, and no public distributions exist. This report builds them by category and size band, using a standard prompt construction on a sample drawn independently of any client relationship, reported as percentiles rather than averages because category variance dwarfs brand variance and the shape matters more than the centre.
It answers one question — is our number unusual, and in which direction — and deliberately does not answer whether that number is commercially sufficient, which depends on factors no benchmark observes. Read your own category, check the concentration measure, treat the median as a location rather than a target, and interrogate the sampling frame of any benchmark you rely on, this one included.
Methodology note: this edition publishes the sampling frame and reporting design rather than a completed data run. Building an unbiased category sample at sufficient scale takes longer than publishing a client-based benchmark would, and we consider the delay preferable to publishing a distribution that describes the wrong population. Categories will be released as their samples reach the minimum threshold, and the covered list will be stated in each edition.
“A benchmark built from an agency’s own clients describes an unusually competent population and makes everyone else feel worse than the market warrants. The sampling frame is not a footnote — it is the entire methodology.” The Age’X Research Team
Get a free GEO audit — the same analysis behind every article here.