The complete inventory, grouped into five conditional stages — and roughly a third of it never touches your website.
This is the complete inventory we work through before recommending a single change. It is deliberately exhaustive and deliberately ordered — grouped into the four pillars of the AI Visibility Framework™ plus the measurement layer that has to exist before any of it means anything. Most of these checks take minutes. A few take days. The value is not in any individual item but in running them in sequence and stopping at the first pillar that fails.
The temptation with any long checklist is to work it end to end. That is the wrong use here, because the groups are conditional rather than independent: items in Evidence cannot produce a result while Identity is failing, and Expression work on an unresolvable brand changes nothing observable. The list is a reference inventory, and the diagnostic sequence is what tells you which section you are currently in.
The productive method is to run the Measurement group first to establish a baseline, then work upward from Access, stopping at the first group where a material failure appears. Record everything you find above that point in a deferred register. Return to it when the constraint clears. Used that way the list is a genuinely complete map; used as a to-do list it is mostly wasted motion.
Whether the engines can physically reach and read the content. Binary, cheap to verify, and total in its consequences when it fails. Almost never the binding constraint for an established brand, which is precisely why it is worth eliminating first rather than assuming.
Whether the engines can resolve what the brand is and distinguish it from everything similarly named. This group fails far more often than teams expect and is invisible from inside the organisation, because everyone internally already knows who they are.
The 127 checks as a conditional sequence
Five stacked bands, one per group, sized by item count. Arrows show conditional progression upward with an exit at each band labelled "stop and remedy". Annotate the Evidence band to show that roughly a third of all items in the list sit off the brand’s own property.
What independent sources say about the brand: whether they connect it to its category at all, and whether their descriptions agree. This is the largest off-property group, the slowest to remedy, and the one that most frequently binds. It is also the group internal audits most commonly skip, because none of it appears in a site crawl.
Evidence and parts of Identity are assessed entirely on third-party sources. Any audit built from a site crawler covers at most two of the five groups — and misses the one that usually decides the outcome.
Whether the content answers real questions in passages that survive extraction. The largest group by item count and the fastest to remedy, which is why it dominates most content programmes — and why it produces so little when the pillars beneath it have not cleared.
The Measurement group is the precondition for the other 111 meaning anything. DUNkē handles the citation half — per prompt, per engine, against named competitors, trended over time.
Whether you can observe the outcome at all. This group runs first, not last, because without a baseline no subsequent finding is falsifiable and no remedy can be distinguished from an environmental change. It is the shortest group and the one whose absence invalidates everything else.
Two things are worth noticing about how the 127 distribute. Expression carries the most items and is the easiest to complete, which is why audits that lack a diagnostic sequence gravitate there — it produces the longest list of completed work. Evidence carries nearly as many, almost all requiring third parties, and is where the majority of established brands are actually stopped.
The second observation is that a conventional technical SEO audit covers most of Access and a portion of Expression — roughly forty per cent of this list — and none of Identity’s external checks or Evidence at all. That gap is not a criticism of technical auditing, which does its job well. It is the reason a technically clean brand can be entirely invisible to the answer layer.
If the list has to be compressed, a handful of items catch a disproportionate share of real failures. The per-bot crawler verification catches silent access blocks that invalidate everything else. The direct identity test catches resolution failures that are invisible internally. The description agreement comparison catches the evidence conflict that produces hedging.
Add the passage-in-isolation test for extractability and the repeat-sampled citation baseline, and you have five checks that between them identify the binding constraint for most brands. The remaining 122 refine the diagnosis and specify the remedy, which matters, but the five are what tell you which pillar you are working in.
It is worth being precise about tooling, since most teams assume their existing stack covers more than it does. A conventional site crawler handles most of the Access group and a meaningful portion of Expression — heading structure, alt text, thin pages, internal linking, canonical and directive checks. That is genuinely useful and it is roughly forty per cent of the inventory.
It cannot test whether an engine can resolve your entity, because that requires querying the engine. It cannot assess whether third-party descriptions agree, because that requires collecting and comparing sources. And it cannot tell you whether you are cited, because that requires running prompts. The three things that most often determine the outcome are all outside what a crawler can see.
The inventory should be adapted rather than adopted. A local business needs the location and listing items expanded and most of the product-data items removed. An ecommerce catalogue inverts that. A B2B software brand needs the analyst and review-platform items weighted heavily and the community items less so.
What should not change is the grouping and the ordering, because those encode the conditionality that makes the list work. Add and remove items within groups freely; do not reorder the groups. A version tailored to your category with the sequence intact is considerably more useful than this list applied literally.
Coverage by tool type
Rows: site crawler, rank tracker, Search Console, analytics, citation tracker, manual review. Columns: which of the five groups each covers, and what proportion of the 127 items. Makes visible that the common stack covers Access and part of Expression and nothing else.
The groups have different half-lives. Access should be re-verified after significant deployments, because it regresses silently and the cost of a delay is total. Identity warrants quarterly checks and immediate re-verification after any rebrand, acquisition, or category change. Expression is assessed at publication for new content and audited periodically for the existing library.
Evidence needs quarterly review because third-party descriptions drift and new sources appear. Measurement runs continuously by definition. Setting these cadences explicitly is what turns the inventory from a one-time exercise into a maintained state, and the cadence differences are large enough that treating them uniformly wastes effort on the stable groups while under-serving the volatile ones.
Without claiming a measured distribution, the items that fail most often in our experience cluster tightly. In Access: blocked render-critical resources and content that only exists after client-side rendering. In Identity: hedged or inaccurate direct identification, and inconsistent category descriptors across the brand’s own properties.
In Evidence: contradictory third-party descriptions and outdated entries nobody has corrected. In Expression: opening passages that depend on preceding text, and comparative content written as prose where a table would serve. In Measurement: the absence of a stable prompt set entirely. Checking those nine first is a defensible shortcut when time is short.
If a scorecard is required for reporting, build it per group rather than as a composite, and record pass, fail, or not-yet-assessed for each item. The group-level view then shows which group is failing without averaging the failure away, which is the property that matters.
What to avoid is weighting items to produce a single percentage. It invites the reassuring outcome where a brand failing every Identity check still scores well because Expression is thorough. The inventory exists to locate a binding constraint; any presentation that obscures which group is failing has inverted its purpose regardless of how satisfying the number looks.
The count is what the inventory contains rather than a target we wrote toward, and it changes. Items get added as engines expose new controls or as failure modes appear; items get retired when a practice becomes obsolete or a check proves consistently uninformative. The number is a snapshot of a working document.
We mention this because round-number checklists tend to be constructed backwards from the number, which produces padding at the margins. The useful test of any inventory is whether every item has ever caught something real. Ours have, which is the only justification for their presence — and the reason we would rather publish an awkward number than a tidy one.
The inventory grows as the landscape does, and the recent additions cluster around two developments. Per-bot crawler granularity produced several new Access items, since access is now a set of individual decisions rather than one. And the divergence between answer surfaces produced new Measurement items, because a single tracked surface is no longer adequate.
The additions worth noting for anyone maintaining their own version: conversational follow-ups in the prompt set, per-surface separation for Google properties, and recording how a brand is described rather than only whether it appears. All three exist because the assumption they replaced — that AI visibility is one thing, measured once — stopped holding.
Equally instructive is what has come off. Checks built around keyword density, meta keyword tags, and exact-match placement were retired because they measure a mechanism that no longer operates. Several schema checks were consolidated once it became clear that structured data supports understanding rather than acting as a lever.
Retiring items is harder than adding them, because a check that has been in the process for years feels like diligence. The test we apply is whether the item has ever changed a recommendation. Where it has not, it is documentation rather than diagnosis, and keeping it dilutes the sequence that makes the inventory work.
For agencies or groups running this across many brands, the inventory becomes more useful aggregated. Recording which group binds for each brand produces a portfolio view that identifies systemic issues — a shared platform causing Access failures, a positioning ambiguity affecting every sub-brand, a category where Evidence binds universally.
That aggregate view frequently justifies a single intervention across many properties, which is a considerably better return than remediating brand by brand. It also builds an internal evidence base about which remedies actually move citation in which categories, which is the closest thing to causal knowledge available without a controlled study.
A list of checks encodes what is currently known to matter, which means it lags the landscape by however long it takes for a new failure mode to become visible. Items exist because something failed and someone noticed; a failure nobody has yet encountered has no corresponding check.
The protection is to treat unexplained results as inventory gaps rather than as noise. A brand clearing all 127 and remaining uncited is telling you the list is incomplete for that case, and investigating why is how items get added. Every check here originated that way, which is both the inventory’s justification and its ceiling.
For a team that will realistically never run 127, the defensible compression is five: verify each AI crawler by name and confirm arrival in logs; ask each engine to identify you and read for hedging; collect ten third-party descriptions and check whether they agree; read three passages from key pages in isolation; and establish a repeat-sampled citation baseline.
Those five span all four pillars plus measurement, take under a day, and between them locate the binding constraint for most brands. The other 122 are how you specify the remedy precisely once you know where you are working. Running five properly beats running 127 without a sequence, which is the practical concession this article should end on.
Any checklist accumulates items faster than it retires them, because adding feels like rigour and removing feels like sloppiness. The discipline that keeps it useful is an annual review asking of each item whether it has changed a recommendation in the past year. Items that have not are documentation.
We apply that test to this list and it shrinks it every time before the additions grow it again. The count is therefore stable in magnitude and unstable in composition, which is the correct behaviour for an inventory tracking a landscape that keeps moving. A checklist that has not changed in two years is describing a search environment that no longer exists.
The inventory is long and the idea behind it is short: a brand becomes citable by being reachable, identifiable, independently corroborated, and structurally extractable, and each of those can fail in a specific, checkable way. The 127 items are simply the ways we have found those four things to break.
Held that way the list stops being intimidating. You are not running 127 tests; you are answering four questions, and the items are the evidence you gather to answer them. That framing also tells you what to do when you encounter a failure the list does not cover — work out which of the four it belongs to, and add it.
No inventory of this kind is complete, and claiming otherwise would be the first thing to distrust about it. These are the checks that have caught real failures in our work. A brand operating in a category we have not audited, on a platform we have not seen, will encounter failure modes absent from this list.
That is an argument for treating it as a starting structure rather than a specification. The four-pillar grouping is the durable part; the items within each are a working set that should grow with your own experience. An inventory that never changes is not being used.
The inventory produces findings, and findings are not a plan. Convert them by identifying the lowest group with a material failure, selecting the specific items within it that are failing, and assigning each an owner and a date. Everything above goes to the deferred register with the condition that would activate it.
That conversion step is where most checklist exercises collapse, because a completed inventory feels like a finished piece of work. It is not — it is the input. The output is one group, a handful of items, and a named person, which is a considerably smaller document than the 127 and the only version anyone will act on.
How the 127 checks are organised. Groups are conditional rather than independent — items in a higher group cannot produce a measurable result while a lower group is failing.
Measurement
Prompt set, engine coverage, repeat sampling, competitor definition, baselines for citation and share of voice.
Runs first. The other 111 are unfalsifiable without it.
Access
Per-bot crawler permissions, log-confirmed arrivals, rendering, blocked resources, directives, sitemaps, response health.
Binary and cheap. Total consequences when it fails.
Identity
Direct identification testing, name collisions, schema declaration, linking properties, description consistency across owned properties.
Fails invisibly. Nobody internally thinks to test it.
Evidence
Independent category association, description consistency across third parties, coverage inventory, review and community presence, mention profile.
Almost entirely off-property. Where most brands stop.
Expression
Answer-first structure, heading descriptiveness, passage self-containment, format-intent match, product data completeness, coverage depth.
Largest group, fastest to fix, and fourth for a reason.
The distribution reflects how many discrete, checkable things each group contains rather than how much each matters. Expression carries the most items because content has the most inspectable properties; Measurement carries the fewest because it is a small number of design decisions.
Reading item count as importance inverts the actual weighting. Evidence has fewer items than Expression and stops considerably more brands, and Measurement has fewest of all while being the precondition for every other check meaning anything. The counts describe the inventory; the ordering describes the priority.
Do we need to run all 127?
Not in one pass. Run Measurement, then work upward and stop at the first material failure. The remaining checks are recorded as deferred and become the next phase once the constraint clears.
How long does the full inventory take?
Access and Identity are hours. Expression is days for a substantial site. Evidence is the longest because it involves collecting and comparing third-party descriptions manually, and Measurement takes a couple of weeks to produce a stable baseline.
Which checks give the fastest visible result?
Expression items, but only where the pillars beneath have cleared. That conditionality is the reason the list is ordered rather than prioritised by speed.
Can this be automated?
Most of Access, parts of Expression, and much of Measurement automate well. Identity resolution testing and the Evidence group require judgement about whether descriptions genuinely agree, which does not reduce cleanly to rules.
The inventory is 127 checks across five groups: Access, Identity, Evidence, Expression, and the Measurement layer that must exist before any of the others can be validated. The counts are uneven by design — Expression carries the most items because it contains the most discrete, checkable things, while Evidence carries nearly as many and stops far more brands.
What makes the list useful is the ordering rather than the completeness. Run Measurement, work upward, stop at the first group that materially fails, and record the rest in a deferred register. Roughly a third of these checks happen off your own property, which is the portion a site crawler cannot see and the reason technically clean brands are so often absent from the answers their buyers are reading.
Methodology note: the grouping and ordering reflect our audit practice, not a measured hierarchy. Item counts describe the inventory as we run it and are not a claim that these are the only 127 things worth checking — the list evolves. We make no claim about what proportion of brands fail each group.
“An unordered checklist produces findings. An ordered one produces an instruction — which is the only output an audit is actually for.” The Age’X Research Team
Get a free GEO audit — the same analysis behind every article here.