Four traits recurred across brands the models name — and neither size nor publishing volume was one of them.
Run enough brand audits and the same profile keeps appearing on the winning side. Not the largest brands, not the ones with the biggest content operations — a specific and initially counterintuitive combination of traits that recurs whether the category is D2C skincare or enterprise software. This piece sets out that profile, the study design behind it, and the four traits that separate brands the models name from brands they do not.
The first thing an audit programme of this kind destroys is the assumption that size predicts presence. It does not, reliably. We repeatedly found category leaders — brands with dominant organic rankings, large content libraries, and substantial marketing budgets — absent from the recommendation set in their own category, while smaller competitors with a fraction of the traffic were named consistently.
The second casualty is content volume. Brands publishing at high frequency were no more likely to be cited than brands publishing rarely, and in several cases less so, because volume without differentiation produced a large library of pages that resembled everything else written on the same subjects. Neither size nor output is the variable, which is what makes the actual profile worth setting out.
The design is straightforward and reproducible, which matters more than its scale. For each brand: a fixed set of category-level prompts posed across multiple engines, recording whether the brand is named and how it is described; an entity resolution check asking each engine to identify the brand directly; an independent description audit collecting how the brand is described across third-party sources; and a structural assessment of whether key pages are extractable.
The output per brand is a profile rather than a score — which of the four traits are present, which are absent. Aggregating those profiles across categories is what surfaces the pattern. We are explicit that this is a qualitative comparative design, not a controlled study: it establishes what winners have in common, not that those traits caused the outcome.
The four-part audit protocol
Vertical flow showing the four assessments per brand (prompt battery, entity resolution, description audit, structural assessment) feeding into a four-trait profile. Include the fixed prompt structure so readers can replicate it.
Every consistently cited brand could be identified accurately by every engine tested, without hedging and without confusion with any similarly-named organisation. That sounds like a low bar. It is not one that most brands clear — particularly those with common-word names, recent rebrands, or a presence in more than one category.
Among losing brands, resolution failure was the most frequently overlooked problem, precisely because nobody thinks to test it. Teams assume the model knows who they are because their own site says so clearly. The test takes ninety seconds and disproportionately often returns a description that is vague, outdated, or about a different company entirely.
Winners were connected to their category by sources they did not control — press coverage, community discussion, review platforms, reference entries, analyst writing. Losers asserted the same connection extensively on their own properties and almost nowhere else.
This was the single most common structural difference between the two groups, and the least addressable by anything a marketing team can publish. It is also the clearest illustration of why content programmes aimed at AI visibility so often produce nothing: they add to the self-published side of the evidence while the deciding evidence sits on the independent side.
The differentiator that surprised us most was consistency rather than volume. Brands with moderate coverage that described them the same way across sources outperformed brands with substantially more coverage that described them inconsistently — different categories, different positioning, different claims about who they serve.
The mechanism is intuitive once observed: a model assembling a picture from conflicting evidence has no confident description to work with and hedges. A brand described identically in twelve places is more legible than one described in forty different ways. This finding has an unusually cheap remedy, which is why it is the first thing we now correct in any engagement.
Brands with moderate but coherent third-party description outperformed brands with far more coverage that contradicted itself. Conflicting evidence produces hedging; hedging produces omission.
The fourth trait is structural and the only one fully within the brand’s control. Winning brands had key pages organised into self-contained, clearly-labelled passages that answered specific questions directly — not necessarily shorter or simpler content, but content a retrieval system could take a bounded piece from without it losing meaning.
Losing brands frequently had excellent information distributed across long, unsegmented prose that made complete sense to a human reading sequentially and offered a machine nothing clean to extract. This is the trait most commonly mistaken for a content quality problem when it is a content structure problem, and it is the fastest of the four to fix.
Four traits that recurred across brands the models name. The first three are earned or corrected; the fourth is built. All four appeared together in consistently cited brands.
Resolvable
Every engine identifies the brand accurately and without hedging. No confusion with similarly-named organisations.
Test: ask each engine who the brand is.
Associated
Independent sources connect the brand to its category. The claim is made by parties other than the brand.
Test: does the brand appear when the category is discussed without it?
Consistent
Third-party descriptions agree on category, audience, and positioning. Conflicting evidence is absent.
Test: collect ten independent descriptions and compare.
Extractable
Key pages break into self-contained passages that answer specific questions and survive being lifted.
Test: read a passage in isolation. Does it still make sense?
Winner profile against brand size
Scatter plot: x-axis brand size proxy (organic traffic or revenue band), y-axis citation rate across the prompt battery. Overlay marker shape for number of the four traits present. Expected pattern: trait count predicts vertical position far better than size does.
The losing profile is more uniform than the winning one and worth stating separately, because it is not what teams expect. Losers were usually technically sound, frequently well designed, often publishing more than the winners, and almost always confident that their category leadership was self-evident. The failure was rarely visible on their own property.
The recurring shape was a brand that had invested heavily in everything it controlled and almost nothing in what it did not. Excellent site, extensive content, no independent evidence base. From a model’s position that is a brand making a strong claim with no corroboration — which is a claim to be cautious about, not one to repeat to a user.
The protocol in this article is reproducible on your own brand. DUNkē handles the measurement half — citation tracking across eight engines, per prompt, against competitors — so the profile rests on data rather than impression.
The four traits held across categories, but their difficulty did not. D2C brands cleared resolution easily and struggled with independent category association, because consumer coverage tends to discuss products rather than positioning. B2B software cleared association readily — the ecosystem is well documented — and struggled with consistency, because analyst, review, and press descriptions drift apart over time.
Local and service brands struggled most with resolution, where name collisions are common and the independent record is thin. Multi-category businesses struggled with association for a structural reason: being connected to many categories reads as being definitive in none. Knowing your category’s characteristic weakness narrows the diagnosis before you run it.
A comparative design of this kind is only useful if the obvious confounds are addressed, so it is worth being explicit. We compared brands within categories rather than across them, because citation behaviour differs enough by category that cross-category comparison would be meaningless. We recorded brand size proxies so that any size effect would be visible rather than hidden.
We used a fixed prompt structure per category so that the questions were comparable, and sampled repeatedly because generated answers vary between runs. What we could not control for is the direction of causation, which is the design’s central limitation and the reason we describe traits that co-occur with citation rather than traits that produce it.
Of the four, description consistency has by far the best effort-to-effect ratio, and it is worth separating from the others for that reason. It requires no earned coverage, no relationships, and no content production. It requires collecting how you are described across third-party sources, cataloguing the contradictions, and correcting them through each source’s legitimate channel.
The work is tedious and largely mechanical, which is exactly why it goes undone — nobody is excited by updating an outdated directory entry. But a brand described coherently across a dozen sources gives a model a confident picture to work from, and one described in a dozen conflicting ways gives it a reason to hedge. It is the cheapest of the four traits to acquire.
Independent category association is the hardest, because it cannot be manufactured and cannot be bought convincingly. It requires that sources with no commercial relationship to you state, in the course of writing about something else, that you are one of the options in your category. That happens when you are genuinely one of the options people discuss.
The routes that work are the slow ones: original data others need to cite, expertise visible enough to be quoted, products discussed because they are used, and coverage earned by being newsworthy rather than by pitching. The routes that do not work are the ones that produce the appearance of association without the substance, which trusted sources filter and models discount.
Trait acquisition difficulty and sequence
Four nodes for the traits, positioned by effort (x-axis) and typical time to acquire (y-axis), with dependency arrows. Extractability bottom-left (cheap, fast), consistency next, resolution moderate, association top-right (expensive, slow). Annotate dependencies showing association is only worth pursuing once resolution clears.
The most consistent surprise was confidence. Losing brands were not uncertain about their category position — they were emphatic about it, and could point to substantial evidence: revenue, customer count, rankings, awards. The evidence simply was not the kind a model can retrieve and corroborate independently.
This produces a difficult conversation, because the diagnosis sounds like a claim that the brand is not a category leader when it is a claim that no independent source has documented that it is. Those are different statements and only the second is being made. Brands that grasp the distinction move quickly; brands that hear it as a challenge to their market position tend to argue rather than act.
One pattern worth extracting for smaller brands: association does not require being described as the leader. It requires being described as an option. A model assembling a set of candidates needs sources that place you in the consideration set, and “a specialist in X for Y” performs that function as well as “the leading X” does.
This is why narrow, specific positioning outperforms broad claims in this contest. “The leading platform” is a claim nobody independent will make about a challenger. “Widely used by mid-market manufacturers for compliance reporting” is a description a trade publication will write, a review platform will reflect, and a model can rely on. The second gets you into the set; the first gets you nothing.
The protocol needs no tooling. Write ten category-level prompts a real buyer would ask — not your brand name, the category question. Run them across the engines your buyers use, three times each, recording which brands are named. Then ask each engine directly who you are. Then collect ten third-party descriptions of your brand and compare them. Then read three of your key pages’ passages in isolation.
By the end you will have your own four-trait profile and a list of the competitors being named instead of you. That comparison is usually the most useful artefact, because it converts an abstract problem into a specific one: these sources are being cited for our category, and here is what they have that we do not.
The obvious commercial move with audit data is to publish striking percentages. We have deliberately not, and the reason is that we cannot support them to the standard this publication holds. A percentage implies a representative sample and a measured distribution; what we have is a set of engagements with brands who came to us, which is a biased sample by construction.
What we can support is the protocol and the traits, both of which are reproducible by anyone. That is a less impressive claim and a more useful one, because a reader can verify it rather than trusting it. If we later run a properly stratified study, the figures will be published with the sampling frame attached. Until then, the method is the contribution.
The four traits held across the categories we audited, and it is worth marking where we would not expect them to. Brands whose category is barely queried conversationally can satisfy all four and remain effectively invisible, because eligibility for an empty question set is worth nothing. That is a demand problem rather than a visibility one.
Highly regulated categories behave differently too, since engines are more conservative about naming specific providers where the stakes are high, and coverage of those queries is thinner. In both cases the traits are still necessary and considerably less sufficient, which is worth establishing before committing to a programme built on them.
Having run the profile many times, the remediation order that produces results fastest is not the order the traits are listed. We start with extractability, because it is fully controlled and returns within weeks where the other traits are already present. Then resolution, because it is cheap and blocks everything above it.
Then consistency, which is mechanical and underrated. Association comes last despite mattering most, because it takes quarters and benefits from the other three being in place first — corroboration attaching to a resolved, consistently described entity is worth considerably more than the same coverage attached to an ambiguous one.
The brands whose citation presence moved after intervention share a pattern worth reporting, with the caveat that this is observation rather than controlled measurement. The first visible change is usually in how they are described rather than whether they appear — hedged, generic descriptions become specific ones.
Appearance follows, typically first on narrower questions and later on the general category question. That progression is consistent enough to use as a leading indicator: if the descriptions engines return about you are sharpening, the underlying evidence is accumulating even before citation rate moves. Teams that know to watch for it stay funded through the lag.
This study establishes co-occurrence and cannot establish causation, and we would rather state that plainly than dress the finding up. Brands with the four traits are cited more; whether the traits produce the citation or both follow from something upstream — genuine market prominence, most plausibly — is not settled by this design.
What makes it useful despite that limit is that the traits are individually defensible on other grounds. Being resolvable, consistently described, independently associated with your category, and structurally extractable are good outcomes whether or not they cause citation. The design is weak on causation and strong on identifying a profile worth having.
The negative findings are as useful as the positive ones. Winners had no shared publishing cadence, no common content management platform, no consistent site architecture, and no shared approach to schema beyond entity declaration. Several had visibly dated websites. Several published rarely.
This matters because each of those is routinely sold as an AI visibility requirement. If a factor genuinely determined citation we would expect it to appear across winners in multiple categories, and these did not. The absence is not proof they are irrelevant, but it is evidence against treating them as the lever — which is how they are frequently marketed.
Stripped of framework, the study answers one question: if you had to predict which of two brands a model will name, what would you look at? Not size, not traffic, not output. You would check whether each can be identified cleanly, whether someone independent has placed them in the category, whether those sources agree, and whether their pages yield a liftable passage.
That is a genuinely different prediction rule from the one most marketing organisations operate on, and it is the study’s contribution. Whether the four traits cause citation remains open. That they predict it better than the intuitive alternatives is what the comparison establishes, and it is enough to change where a reasonable team would spend.
Four traits, four tests, one afternoon. Can each engine identify you accurately and without hedging. Does any independent source place you in your category. Do those sources describe you consistently. Do your key passages survive being read in isolation. That is the entire diagnostic, and it needs no tooling.
What makes it worth running is that the answers are usually surprising in a specific direction: brands assume their problem is content and discover it is identity or association. The four tests are cheap enough that there is no defensible reason not to have run them, and the result reliably redirects the next quarter of spend toward something that can actually move.
Brands that commission an audit are not a random sample of brands. They are disproportionately organisations that already suspect a problem, have budget to investigate it, and operate in categories where AI visibility is commercially salient. That biases the population toward brands with something wrong, which is worth holding when reading any pattern drawn from it.
It probably does not bias the trait comparison itself, since winners and losers within that population are compared against each other rather than against the wider market. But it almost certainly means the failures we see are more severe and more frequent than they would be across all brands. We would expect a stratified market sample to show the same traits and a healthier distribution.
Can we run this protocol ourselves?
Yes — that is why it is set out here. The prompt battery, resolution check, and structural assessment need no tooling. The description audit is manual but tractable. The part that benefits from tooling is running it repeatedly over time so you can see movement.
How many of the four traits do we need?
Consistently cited brands had all four. Brands with three were cited intermittently, and the missing trait was almost always association or consistency. We would not describe any single trait as sufficient.
Does fixing extractability alone help?
It helps where extractability is the binding constraint — typically brands already resolvable and associated but not being drawn from. Where the failure is upstream, structural work on your pages changes nothing, because the brand is not a candidate yet.
How long before the profile changes after intervention?
Extractability and resolution can move within weeks. Association and consistency move over quarters because they depend on third parties publishing and on old descriptions ageing out. Judge the programme on the leading indicator, not on quarterly traffic.
Across audits, the brands models name share four traits: they can be resolved unambiguously, they are connected to their category by sources they do not control, those sources describe them consistently, and their key pages break into passages that survive extraction. Neither size nor publishing volume separated winners from losers, and most losing brands were technically sound and confident in a category leadership that no independent source had documented.
The most actionable finding is that consistency of third-party description beat volume of coverage, because it is both the cleanest differentiator and among the cheapest to correct. The protocol is reproducible in an afternoon, and running it on your own brand produces a profile rather than a score — which of the four you have, which you lack, and therefore which piece of work is the only one currently worth doing.
Methodology note: this is a qualitative comparative design, not a controlled study. It establishes which traits co-occur with citation; it does not establish causation, and we make no claim about the proportion of brands failing each trait. Figures describing distribution are deliberately absent rather than estimated. A causal version would require matched cohorts with a single trait varied and outcomes tracked over time — a design we are building toward. Readers should treat the profile as a reproducible diagnostic, and the trait ordering as observed tendency rather than measured effect size.
“The brands models name are not the biggest in their category. They are the ones whose category membership somebody else has written down, consistently, in more than one place.” The Age’X Research Team
Get a free GEO audit — the same analysis behind every article here.