The sequence we run, why measurement comes first, and why the output is one instruction rather than forty findings.
Most audits produce a document. The useful ones produce a single instruction. This is the process we run before recommending anything — what we look at, in what order, and why the sequence is the part that matters. The output is deliberately narrow: one binding constraint, one owner, one timeline. Everything else we find gets recorded and explicitly deferred, because a report that lists forty problems has told a client nothing about where to start.
An AI visibility audit exists to answer one question: what is currently preventing this brand from being named, and what would it take to change it. Everything else — the inventory of issues, the technical findings, the content observations — is supporting material. If the audit does not end with a single binding constraint and a realistic timeline, it has produced a document rather than a decision.
This is a sharper objective than most audits set, and it changes the method. A general site audit surveys everything and reports what it finds. A diagnostic audit tests a specific sequence and stops at the first failure, because in a conditional system the findings above the failure cannot be acted on productively. The discipline is in stopping.
Before any assessment we build a baseline: a fixed set of commercially meaningful prompts, run across the engines the client’s buyers actually use, recording whether the brand is named, how it is described, and which competitors appear instead. Sampled repeatedly, because generated answers vary between runs and a single observation is not evidence.
This comes first for a reason that is easy to underrate. Without a baseline, every subsequent claim about improvement is unfalsifiable, and the client cannot distinguish work that succeeded from an engine that changed its behaviour. It also frequently changes the brief: a brand convinced it is absent sometimes turns out to be cited for the wrong things, which is a different problem with a different remedy.
The audit sequence
Vertical flow: measurement baseline → Access tests → Identity tests → Evidence tests → Expression tests, with an exit arrow at each stage marked "stop here if failing". Show the deferred-findings register running alongside as a parallel track that collects everything above the exit point.
We verify that the engines can physically reach and read the content: crawler permissions checked per bot rather than in aggregate, server logs examined to confirm the crawlers are actually arriving, and the rendered output inspected to confirm the content exists without client-side JavaScript execution. We also check that resources required for rendering are not themselves blocked.
This takes about an hour and is almost never the binding constraint for an established brand — which is exactly why it is worth doing first. The failures we do find here are usually inherited from a security policy or a default configuration nobody chose, and the cost of discovering that after six months of content investment is considerable.
We ask each engine directly who the brand is, and read the answer for three things: accuracy, specificity, and hedging. Then we check for collisions — other organisations sharing the name or operating adjacently — and examine how the brand is declared in structured data, how consistently it is named across its own properties, and whether it is connected to authoritative external profiles.
The failures here are more common than clients expect and almost always invisible internally, because everyone inside the company knows who they are. A description that is outdated, generic, or about a different organisation is a hard blocker: corroboration cannot attach to an entity the model cannot resolve, so nothing above this pillar can be built until it clears.
This is the largest part of the work and the part clients are least prepared for. We collect how the brand is described across independent sources — press, reference entries, review platforms, community discussion, analyst writing — and assess two things separately: whether those sources connect the brand to its category at all, and whether their descriptions agree with each other.
The two failures are distinct. Absent association means the model has no independent basis for raising the brand when the category is discussed. Inconsistent description means it has conflicting evidence and hedges. The first requires earned coverage over quarters; the second is frequently correctable in weeks by fixing profiles and requesting updates, which makes separating them worth the effort.
Access and Identity are cheap and commonly clear. Evidence binds most often — and it is the pillar that cannot be fixed on the client’s own property, on the client’s own timeline, by the client’s own team.
We assess whether the brand’s key pages can be extracted from: whether headings describe their sections, whether opening sentences state the answer, whether passages survive being read in isolation, whether comparative content sits in tables and procedural content in steps, and whether anything important is trapped in images.
This is genuinely important and it is fourth, because excellent extractability on a brand the model cannot resolve or corroborate changes nothing observable. When Expression is the binding constraint — which happens with brands that are already well established and well described — it is the most satisfying engagement we run, because the remedy is fast and the effect is visible within weeks.
What we check per pillar
Four columns, one per pillar. Rows: what is tested, how it is tested, who owns the remedy, typical timeline, frequency as binding constraint. The row to make prominent is the last, showing Evidence dominating.
Everything above is also run against a named competitor set, because absolute findings are hard to interpret. A brand cited for a fifth of its priority prompts might be leading or trailing depending entirely on what the alternatives achieve, and a description inconsistency matters more if competitors are described coherently.
The competitor set is defined by who actually appears in the answers rather than by who the client considers a rival, which routinely surfaces sources the client had not considered competitors at all — publishers, communities, specialists. Those are frequently the most instructive part of the audit, because they demonstrate what the engines in that category are actually rewarding.
Every audit we run begins with measurement, because without it no finding is falsifiable. DUNkē tracks citations across eight AI engines — per prompt, against named competitors — which is the baseline the rest depends on.
We do not produce a composite score. A number averages the pillars and conceals the binding one, which inverts the entire purpose. We do not deliver a prioritised list of forty findings, because prioritisation implies the items are independent and they are not. And we do not recommend content production before the diagnosis, however much the client expects it as the natural output of an audit.
The most common friction in this process is the last one. A client commissioning an audit frequently expects a content plan, and receiving instead a fortnight of entity work and a two-quarter evidence programme feels like a smaller answer than they paid for. It is a considerably more useful one, and the alternative is a content programme aimed before anything told it where to point.
The output is three things: the measurement baseline, the pillar profile with the binding constraint identified, and a deferred register listing everything else found, explicitly marked as not-yet-actionable with the condition that would make it actionable. The plan addresses only the binding constraint, with a named owner and a realistic timeline.
The deferred register matters more than it sounds. It records that the other issues were seen and assessed, which prevents them being rediscovered as failures of the audit later, and it converts into the next phase of work automatically once the constraint clears. It is also what allows the plan to be genuinely narrow without appearing to have missed anything.
An audit is a reading, and the conditions move: engines change retrieval, competitors publish, descriptions drift, deployments regress the technical pillars. We re-run the full sequence periodically and the measurement continuously, because the leading indicator moves long before any downstream business metric does.
The re-audit also validates the previous diagnosis, which is the only honest way to know whether the framework is working for a given client. If the binding constraint was correctly identified and remedied, citation presence on the affected prompts should move within the expected window. If it does not, the diagnosis was wrong, and that is worth knowing early rather than defending.
The inputs determine the quality of the diagnosis. We ask for the query and prompt set the business actually cares about — not a keyword export, but the questions buyers ask before deciding — along with access to Search Console, the list of properties and profiles the brand controls, and a named competitor set as the client understands it.
That last item is deliberately collected before we define competitors ourselves, because the gap between who a client considers a rival and who actually appears in their category’s answers is frequently the most revealing finding of the engagement. When those two lists barely overlap, the client learns something about their market that has nothing to do with search.
Generated answers vary between runs. The same prompt can name different brands on consecutive observations, which means a single check is a sample rather than a measurement. We run each prompt multiple times across each surface and record a rate, because the difference between being named occasionally and being named reliably is the difference between a boundary position and a secure one.
This is also the most common flaw in internal audits we review. A team checks twenty prompts once, finds themselves absent, and concludes they have no presence — or checks once, finds themselves present, and concludes the opposite. Both conclusions are drawn from insufficient sampling, and both lead to misallocated effort.
The most laborious part of the process deserves explaining, since it is the part clients most often want to skip. We collect how the brand is described across every third-party source we can find — press, directories, review platforms, reference entries, partner sites, community discussion — and record the category descriptor, the audience described, and the positioning claimed in each.
Then we compare them. What emerges is usually a spread rather than a consensus: three different category descriptors, two incompatible audience descriptions, and several entries describing a business that no longer exists. Presented as a table, this is frequently the single most persuasive artefact in the audit, because it makes an abstract problem immediately concrete.
The description spread
One row per third-party source, columns for category descriptor, audience described, positioning claimed, and last-updated date. Highlight contradictions in orange. Include the brand’s own description as the top row for comparison. This is the artefact that most often changes a client’s mind.
We build it from observation rather than from the client’s list. Running the prompt battery and recording every brand and domain named across all runs produces the set of sources actually competing for the category’s answers. That list routinely includes publishers, community threads, reference entries, and specialists the client had never considered.
This is more useful than a commercial competitor list for a specific reason: those sources are demonstrating what the engines reward in that category. A community thread being cited repeatedly is evidence that candid comparative discussion is what the category’s questions call for, which is actionable in a way that knowing your commercial rival’s position is not.
Every audit surfaces issues that are real, fixable, and not the constraint — a slow template, a thin content section, inconsistent internal linking, missing alt text. These go into the deferred register with the condition that would make them worth acting on, and they do not appear in the recommendation.
Resisting the urge to lead with them is a discipline rather than an oversight. A list of twenty legitimate findings dilutes a single binding instruction, and clients reliably act on the easiest items rather than the most important one. Narrowness is the deliverable. The register exists so that nothing is lost, not so that everything is scheduled.
The honest framing of the engagement is that the analysis is days and the remedy is months, and the client’s attention is required disproportionately at the start and the end. At the start, defining the real buyer questions and the properties they control. At the end, accepting a narrow recommendation that may sit outside the function that commissioned the work.
That last point is where engagements succeed or fail. An audit commissioned by a content team that concludes the binding constraint is entity resolution or earned coverage requires the finding to travel to another function. Building that expectation in at the start — that the answer may not belong to whoever asked the question — substantially improves the odds of it being acted on.
Occasionally the correct output is that the brand should not invest in AI visibility yet. That happens when the prompt set reveals that the category is barely queried conversationally, when answer coverage on the relevant queries is minimal, or when a more fundamental commercial problem makes search visibility the wrong constraint to work on.
Saying so costs us the follow-on engagement and it is the right output. A programme built on a category where the demand does not exist will produce measurements that go nowhere and will eventually be cancelled, taking the credibility of the whole discipline with it inside that organisation. Declining is rare, and it is worth being willing to do.
Four documents, deliberately few. The measurement baseline, which is the client’s to keep running. The pillar profile, one page, with the binding constraint named. The deferred register, listing everything found above the constraint with its activation condition. And the remediation plan, which addresses only the constraint and names an owner and a timeline.
What we do not hand over is a hundred-page audit document, because those are read once and filed. The constraint on usefulness is not how much we found; it is how much gets acted on, and narrow documents get acted on. If a client wants the exhaustive version, the deferred register is it — complete, and explicitly marked as not-yet-actionable.
Clients disagree with the diagnosis reasonably often, usually when it points away from the function that commissioned the audit. The productive response is not argument but a testable prediction: if the diagnosis is right, remedying the named constraint should move citation on the affected prompts within a stated window.
That converts a disagreement about interpretation into an experiment with a deadline. It also protects the client from us — if we are wrong, the measurement shows it and the engagement adjusts. Any audit method that cannot be falsified within a quarter is asking for trust it has not earned, which is a poor basis for spending.
Second audits have a recognisable pattern. The original constraint has usually cleared, because it was specific and someone owned it. What has frequently regressed is Access or Identity, through deployments and organisational change, which is why we re-run the full sequence rather than only checking the pillar we worked on.
And the deferred register has usually aged: some items resolved themselves, some became more serious, and new ones appeared. That evolution is the argument for re-auditing on a cadence rather than treating the first engagement as definitive. A diagnosis is a reading of conditions, and the conditions move.
The structure we have converged on is a short diagnostic engagement producing the four artefacts, followed by a decision point, followed by remediation work that may or may not involve us. Separating diagnosis from remediation matters, because an auditor who profits from a particular remedy has an obvious incentive to find that constraint.
It also produces better outcomes when the binding pillar is Evidence, since the remediation belongs with communications rather than with search. Being willing to hand the work to someone else is what keeps the diagnosis honest, and it is the clearest signal that the method is a method rather than a sales process.
The method has moved in three ways worth recording. Measurement was originally the last step and is now the first, because too many engagements produced recommendations nobody could validate. Competitor definition moved from the client’s list to observation, because the two rarely matched and the observed set was consistently more instructive.
And the deferred register was added after we noticed clients acting on the easy findings rather than the binding one. Each change came from a failure rather than a theory, which is generally the case with process. We expect further revisions, and the most likely is separating the two halves of the Evidence assessment formally.
Whether or not the auditor is us, four things are reasonable to require. A measurement baseline, so the recommendations can be validated. A single named constraint rather than a ranked list, because a ranked list is an abdication of the diagnostic responsibility. A falsifiable prediction with a timeframe. And a stated separation between diagnosis and remediation incentives.
An audit that supplies none of these is a document rather than a diagnosis, and the industry produces a great many of them. The four requirements are not onerous and they filter effectively, which makes them a useful specification to put in front of any prospective supplier — including anyone reading this considering hiring us.
An audit is worth running only if it changes what happens next. Everything in this process — measuring before diagnosing, running bottom-up, stopping at the first failure, deferring the rest, predicting a falsifiable outcome — exists to protect that single property against the natural tendency of audits to become comprehensive documents nobody acts on.
The test we apply to our own work is simple: three months later, did the client do one specific thing because of the audit, and did the measurement move as predicted. Where both are true the process worked. Where the client did five things and nothing moved, we produced a document, which is the failure mode this entire method is built to avoid.
An audit is premature for brands whose category is barely queried conversationally, whose product is not yet in market, or whose commercial constraint is demand rather than visibility. In each case the audit will produce accurate findings that do not matter, and acting on them consumes capacity that belongs elsewhere.
The cheap pre-check is to run ten category-level prompts and see whether the engines produce substantive, cited answers at all. Where they do not, the category has not yet moved into the answer layer and the correct decision is to defer. We would rather say that at the outset than deliver a competent diagnosis of a problem the client does not yet have.
The five stages of the audit, run in order, stopping at the first material failure. Everything above the stop point is recorded and deferred rather than recommended.
Baseline
Fixed prompt set from real buyer questions, run repeatedly across the engines the client’s buyers use. Competitors defined by who actually appears.
Runs first. Without it, nothing that follows is falsifiable.
Access
Crawler permissions verified per bot, arrival confirmed in server logs, content confirmed present without client-side rendering.
An hour of work. Rarely the constraint, always worth eliminating.
Identity
Direct identification tested on each engine. Read for accuracy, specificity, and hedging. Name collisions catalogued.
Ninety seconds. Fails more often than clients expect.
Evidence
Independent descriptions collected and compared. Category association and description consistency assessed separately.
The largest block. Where the audit most often stops.
Expression
Key pages assessed for extractability: heading descriptiveness, passage self-containment, format-intent match.
Assessed last, and genuinely fast to remedy when it is the constraint.
The ordering encodes the conditionality that makes the audit useful. A brand failing identity cannot be helped by expression work, because corroboration and extraction both attach to an entity the engine must first resolve. Running the stages out of order produces findings; running them in order produces an instruction.
The baseline sits at stage zero rather than stage five for the same reason. An audit that measures after diagnosing has no way to validate its own recommendation, which turns the engagement into an assertion the client has to take on trust. Measuring first is what makes the rest checkable.
How long does an audit take?
The assessment itself is days once the measurement baseline exists, and building a reliable baseline takes a couple of weeks because answers must be sampled repeatedly to distinguish signal from run-to-run variation.
Can we run it internally?
The Access, Identity and Expression assessments are entirely reproducible internally. The Evidence assessment is laborious but tractable. The part that benefits most from external work is competitor definition, because internal assumptions about who competes are usually wrong.
What if several pillars are failing?
Only the lowest matters. Multiple failures are common and do not change the instruction: fix the lowest, re-measure, then reassess. Attempting several at once dilutes effort and makes attribution impossible.
How do we know the diagnosis was right?
Citation presence on the affected prompts should move within the pillar’s expected window. If it does not, the diagnosis was wrong. Building that validation into the engagement is what keeps the method honest.
The audit runs a fixed sequence — measurement, then Access, Identity, Evidence, Expression — and stops at the first failure. Its output is a baseline, a pillar profile naming the single binding constraint, and a deferred register recording everything else with the condition that would make it actionable. It is deliberately narrower than most audits, because a report listing forty findings has told the client nothing about where to begin.
The two things that most surprise clients are that measurement precedes diagnosis, and that roughly half the assessment happens off their own property. Both follow from the same fact: AI visibility is decided by whether engines can resolve a brand and find independent agreement about it, and neither of those is visible in a site crawl. Diagnose the constraint, fix only that, re-measure, and let the deferred register become the next phase.
Methodology note: this describes our process rather than reporting a study. Where we characterise how often a pillar binds, that is qualitative observation from audit work and is stated without a figure deliberately. We do not claim a measured distribution across a brand population.
“A report listing forty problems has told you nothing about where to start. The audit is finished when it produces one instruction, one owner, and one timeline.” The Age’X Research Team
Get a free GEO audit — the same analysis behind every article here.