DUNkē tracking 12,847 prompts globally·+34% AI mentions for Mysthelle this week·WeaverStory now cited in 4/5 engines·Banana Club ranking #2 on Perplexity·Linen Trail · 11x backlink growth · Q2·DUNkē tracking 12,847 prompts globally·+34% AI mentions for Mysthelle this week·WeaverStory now cited in 4/5 engines·Banana Club ranking #2 on Perplexity·Linen Trail · 11x backlink growth · Q2·
Blog
Reference

127 things we audit before optimising a brand for AI search

The complete inventory, grouped into five conditional stages — and roughly a third of it never touches your website.

MMohabbat Khan
16 min read

This is the complete inventory we work through before recommending a single change. It is deliberately exhaustive and deliberately ordered — grouped into the four pillars of the AI Visibility Framework™ plus the measurement layer that has to exist before any of it means anything. Most of these checks take minutes. A few take days. The value is not in any individual item but in running them in sequence and stopping at the first pillar that fails.

Executive summary
  • 127 checks across five groups: Access (24), Identity (22), Evidence (31), Expression (34), Measurement (16).
  • The grouping is the method. Items are only actionable once the pillars beneath them have cleared.
  • Evidence and Expression carry the most items — and Evidence is where audits most often stop.
  • Roughly a third of the list runs off your own property, which is the part most internal audits omit entirely.
  • Use it as a reference, not a to-do list. Working it top to bottom without diagnosing the binding pillar wastes most of the effort.

How to use this list

The temptation with any long checklist is to work it end to end. That is the wrong use here, because the groups are conditional rather than independent: items in Evidence cannot produce a result while Identity is failing, and Expression work on an unresolvable brand changes nothing observable. The list is a reference inventory, and the diagnostic sequence is what tells you which section you are currently in.

The productive method is to run the Measurement group first to establish a baseline, then work upward from Access, stopping at the first group where a material failure appears. Record everything you find above that point in a deferred register. Return to it when the constraint clears. Used that way the list is a genuinely complete map; used as a to-do list it is mostly wasted motion.

Group one: Access — 24 checks

Whether the engines can physically reach and read the content. Binary, cheap to verify, and total in its consequences when it fails. Almost never the binding constraint for an established brand, which is precisely why it is worth eliminating first rather than assuming.

  • Robots directives checked per AI crawler user-agent, not in aggregate
  • GPTBot access verified
  • ClaudeBot access verified
  • PerplexityBot access verified
  • Google-Extended setting confirmed against intent
  • Bingbot access verified for Copilot eligibility
  • Server logs examined to confirm AI crawlers are actually arriving
  • Crawl frequency per bot assessed
  • Blocked render-critical JavaScript resources identified
  • Blocked CSS resources identified
  • Content presence in raw HTML confirmed without JS execution
  • Server-side or static rendering verified on key templates
  • Client-side-only content inventoried
  • Hash-based routing identified where present
  • Lazy-loaded content tested for crawler visibility
  • Interaction-gated content identified
  • Soft 404s and error responses on key paths
  • Redirect chains on priority URLs
  • Canonical directives verified on key templates
  • Stray noindex directives across templates
  • XML sitemap accuracy and currency
  • Sitemap submitted and referenced in robots.txt
  • Response times on priority templates
  • CDN or firewall rules blocking legitimate crawlers

Group two: Identity — 22 checks

Whether the engines can resolve what the brand is and distinguish it from everything similarly named. This group fails far more often than teams expect and is invisible from inside the organisation, because everyone internally already knows who they are.

  • Direct identification test on each target engine
  • Accuracy of returned description
  • Specificity versus generic category language
  • Presence of hedging in the description
  • Name collisions with other organisations identified
  • Adjacent-market confusions identified
  • Legal versus trading name consistency
  • Brand name consistency across own properties
  • Category descriptor consistency across own properties
  • Organization schema present and valid
  • Entity identifiers and reference properties declared
  • Links to authoritative external profiles declared
  • Logo and brand asset declaration
  • Person schema for named authors and executives
  • Author-to-organisation relationship declared
  • Knowledge panel presence and accuracy
  • Reference source entry presence where warranted
  • Reference source factual accuracy
  • Historical names and rebrand traces still in circulation
  • Acquired or parent entity confusion
  • Subsidiary and sub-brand disambiguation
  • Location and market declaration where relevant
Recommended visual — Flowchart

The 127 checks as a conditional sequence

Five stacked bands, one per group, sized by item count. Arrows show conditional progression upward with an exit at each band labelled "stop and remedy". Annotate the Evidence band to show that roughly a third of all items in the list sit off the brand’s own property.

Group three: Evidence — 31 checks

What independent sources say about the brand: whether they connect it to its category at all, and whether their descriptions agree. This is the largest off-property group, the slowest to remedy, and the one that most frequently binds. It is also the group internal audits most commonly skip, because none of it appears in a site crawl.

  • Independent sources connecting brand to category identified
  • Volume of independent category association assessed
  • Consistency of category descriptor across third-party sources
  • Consistency of audience or segment description
  • Consistency of positioning claims
  • Contradictory descriptions catalogued
  • Outdated third-party descriptions catalogued
  • Press coverage inventory and recency
  • Quality and authority of covering publications
  • Trade and analyst coverage presence
  • Review platform presence in relevant categories
  • Review volume and recency
  • Review sentiment and recurring themes
  • Review response practice
  • Community platform presence and sentiment
  • Authenticity of community presence assessed
  • Signs of manufactured or astroturfed presence
  • Video platform presence on category topics
  • Third-party video mentioning the brand
  • Reference source coverage of the category itself
  • Unlinked brand mention inventory
  • Linked mention inventory and referring domain quality
  • Topical relevance of referring domains
  • Anchor text naturalness across the profile
  • Original data or research assets published
  • Citation of the brand as a data source
  • Executive or expert visibility in the field
  • Partnership and integration listings
  • Directory and association listings accuracy
  • Competitor corroboration profile for comparison
  • Share of independent category discussion versus competitors
Where the list weights
A third of the checks are not about your website

Evidence and parts of Identity are assessed entirely on third-party sources. Any audit built from a site crawler covers at most two of the five groups — and misses the one that usually decides the outcome.

Group four: Expression — 34 checks

Whether the content answers real questions in passages that survive extraction. The largest group by item count and the fastest to remedy, which is why it dominates most content programmes — and why it produces so little when the pillars beneath it have not cleared.

  • Answer-first structure on priority pages
  • Opening answer length and completeness
  • Self-containment of opening passages
  • Unresolved pronouns and back-references in key passages
  • Heading descriptiveness across templates
  • Question-shaped headings where intent warrants
  • Single H1 and logical heading hierarchy
  • Paragraph segmentation into single-idea units
  • Key fact positioned near its heading
  • Comparative content presented as tables where applicable
  • Table header clarity and unit consistency
  • Enumerable content presented as genuine lists
  • List item self-containment
  • Procedural content presented as ordered steps
  • Genuine FAQ sections built from real questions
  • Information trapped in images without text equivalents
  • Alt text presence and descriptiveness
  • Format matched to query intent per priority page
  • Topic coverage breadth against the question set
  • Coverage of natural follow-up questions
  • Depth per subtopic assessed
  • Thin or duplicative pages identified
  • Cannibalising pages competing for one intent
  • Internal linking to priority pages
  • Anchor text descriptiveness on internal links
  • Orphan pages among priority content
  • Cluster structure and pillar presence
  • Content currency on time-sensitive pages
  • Update cadence and maintenance evidence
  • Product data completeness where commercial
  • Product schema accuracy and currency
  • Specification and attribute completeness
  • Evidence density: data, figures, citations
  • Authorship and credential presence on content
The 16 that come first

Measurement before anything else

The Measurement group is the precondition for the other 111 meaning anything. DUNkē handles the citation half — per prompt, per engine, against named competitors, trended over time.

Explore DUNkē →

Group five: Measurement — 16 checks

Whether you can observe the outcome at all. This group runs first, not last, because without a baseline no subsequent finding is falsifiable and no remedy can be distinguished from an environmental change. It is the shortest group and the one whose absence invalidates everything else.

  • Prompt set defined from real buyer questions
  • Prompt set coverage across funnel stages
  • Prompt set stability for trend comparison
  • Conversational follow-ups included in the set
  • Engines tracked matched to actual audience use
  • Per-surface separation for Google properties
  • Repeat sampling to handle answer variability
  • Named competitor set defined by who appears
  • Citation rate baseline established
  • Share of voice baseline established
  • Description quality recorded, not just presence
  • Coverage prevalence measured on own queries
  • Branded search trend captured
  • Direct traffic trend captured
  • Self-reported attribution question in place
  • Historical archive retained for trend analysis

What the distribution tells you

Two things are worth noticing about how the 127 distribute. Expression carries the most items and is the easiest to complete, which is why audits that lack a diagnostic sequence gravitate there — it produces the longest list of completed work. Evidence carries nearly as many, almost all requiring third parties, and is where the majority of established brands are actually stopped.

The second observation is that a conventional technical SEO audit covers most of Access and a portion of Expression — roughly forty per cent of this list — and none of Identity’s external checks or Evidence at all. That gap is not a criticism of technical auditing, which does its job well. It is the reason a technically clean brand can be entirely invisible to the answer layer.

The checks that catch the most problems

If the list has to be compressed, a handful of items catch a disproportionate share of real failures. The per-bot crawler verification catches silent access blocks that invalidate everything else. The direct identity test catches resolution failures that are invisible internally. The description agreement comparison catches the evidence conflict that produces hedging.

Add the passage-in-isolation test for extractability and the repeat-sampled citation baseline, and you have five checks that between them identify the binding constraint for most brands. The remaining 122 refine the diagnosis and specify the remedy, which matters, but the five are what tell you which pillar you are working in.

Which checks a crawler can actually do

It is worth being precise about tooling, since most teams assume their existing stack covers more than it does. A conventional site crawler handles most of the Access group and a meaningful portion of Expression — heading structure, alt text, thin pages, internal linking, canonical and directive checks. That is genuinely useful and it is roughly forty per cent of the inventory.

It cannot test whether an engine can resolve your entity, because that requires querying the engine. It cannot assess whether third-party descriptions agree, because that requires collecting and comparing sources. And it cannot tell you whether you are cited, because that requires running prompts. The three things that most often determine the outcome are all outside what a crawler can see.

Building your own version of this list

The inventory should be adapted rather than adopted. A local business needs the location and listing items expanded and most of the product-data items removed. An ecommerce catalogue inverts that. A B2B software brand needs the analyst and review-platform items weighted heavily and the community items less so.

What should not change is the grouping and the ordering, because those encode the conditionality that makes the list work. Add and remove items within groups freely; do not reorder the groups. A version tailored to your category with the sequence intact is considerably more useful than this list applied literally.

Recommended visual — Comparison table

Coverage by tool type

Rows: site crawler, rank tracker, Search Console, analytics, citation tracker, manual review. Columns: which of the five groups each covers, and what proportion of the 127 items. Makes visible that the common stack covers Access and part of Expression and nothing else.

How often to re-run each group

The groups have different half-lives. Access should be re-verified after significant deployments, because it regresses silently and the cost of a delay is total. Identity warrants quarterly checks and immediate re-verification after any rebrand, acquisition, or category change. Expression is assessed at publication for new content and audited periodically for the existing library.

Evidence needs quarterly review because third-party descriptions drift and new sources appear. Measurement runs continuously by definition. Setting these cadences explicitly is what turns the inventory from a one-time exercise into a maintained state, and the cadence differences are large enough that treating them uniformly wastes effort on the stable groups while under-serving the volatile ones.

The items most often failed

Without claiming a measured distribution, the items that fail most often in our experience cluster tightly. In Access: blocked render-critical resources and content that only exists after client-side rendering. In Identity: hedged or inaccurate direct identification, and inconsistent category descriptors across the brand’s own properties.

In Evidence: contradictory third-party descriptions and outdated entries nobody has corrected. In Expression: opening passages that depend on preceding text, and comparative content written as prose where a table would serve. In Measurement: the absence of a stable prompt set entirely. Checking those nine first is a defensible shortcut when time is short.

Turning the inventory into a scorecard

If a scorecard is required for reporting, build it per group rather than as a composite, and record pass, fail, or not-yet-assessed for each item. The group-level view then shows which group is failing without averaging the failure away, which is the property that matters.

What to avoid is weighting items to produce a single percentage. It invites the reassuring outcome where a brand failing every Identity check still scores well because Expression is thorough. The inventory exists to locate a binding constraint; any presentation that obscures which group is failing has inverted its purpose regardless of how satisfying the number looks.

Why 127 and not a round number

The count is what the inventory contains rather than a target we wrote toward, and it changes. Items get added as engines expose new controls or as failure modes appear; items get retired when a practice becomes obsolete or a check proves consistently uninformative. The number is a snapshot of a working document.

We mention this because round-number checklists tend to be constructed backwards from the number, which produces padding at the margins. The useful test of any inventory is whether every item has ever caught something real. Ours have, which is the only justification for their presence — and the reason we would rather publish an awkward number than a tidy one.

The items we added most recently

The inventory grows as the landscape does, and the recent additions cluster around two developments. Per-bot crawler granularity produced several new Access items, since access is now a set of individual decisions rather than one. And the divergence between answer surfaces produced new Measurement items, because a single tracked surface is no longer adequate.

The additions worth noting for anyone maintaining their own version: conversational follow-ups in the prompt set, per-surface separation for Google properties, and recording how a brand is described rather than only whether it appears. All three exist because the assumption they replaced — that AI visibility is one thing, measured once — stopped holding.

The items we retired

Equally instructive is what has come off. Checks built around keyword density, meta keyword tags, and exact-match placement were retired because they measure a mechanism that no longer operates. Several schema checks were consolidated once it became clear that structured data supports understanding rather than acting as a lever.

Retiring items is harder than adding them, because a check that has been in the process for years feels like diligence. The test we apply is whether the item has ever changed a recommendation. Where it has not, it is documentation rather than diagnosis, and keeping it dilutes the sequence that makes the inventory work.

Using the inventory across a portfolio

For agencies or groups running this across many brands, the inventory becomes more useful aggregated. Recording which group binds for each brand produces a portfolio view that identifies systemic issues — a shared platform causing Access failures, a positioning ambiguity affecting every sub-brand, a category where Evidence binds universally.

That aggregate view frequently justifies a single intervention across many properties, which is a considerably better return than remediating brand by brand. It also builds an internal evidence base about which remedies actually move citation in which categories, which is the closest thing to causal knowledge available without a controlled study.

The honest limitation of any inventory

A list of checks encodes what is currently known to matter, which means it lags the landscape by however long it takes for a new failure mode to become visible. Items exist because something failed and someone noticed; a failure nobody has yet encountered has no corresponding check.

The protection is to treat unexplained results as inventory gaps rather than as noise. A brand clearing all 127 and remaining uncited is telling you the list is incomplete for that case, and investigating why is how items get added. Every check here originated that way, which is both the inventory’s justification and its ceiling.

The five checks that would catch most of it

For a team that will realistically never run 127, the defensible compression is five: verify each AI crawler by name and confirm arrival in logs; ask each engine to identify you and read for hedging; collect ten third-party descriptions and check whether they agree; read three passages from key pages in isolation; and establish a repeat-sampled citation baseline.

Those five span all four pillars plus measurement, take under a day, and between them locate the binding constraint for most brands. The other 122 are how you specify the remedy precisely once you know where you are working. Running five properly beats running 127 without a sequence, which is the practical concession this article should end on.

Keeping the inventory honest

Any checklist accumulates items faster than it retires them, because adding feels like rigour and removing feels like sloppiness. The discipline that keeps it useful is an annual review asking of each item whether it has changed a recommendation in the past year. Items that have not are documentation.

We apply that test to this list and it shrinks it every time before the additions grow it again. The count is therefore stable in magnitude and unstable in composition, which is the correct behaviour for an inventory tracking a landscape that keeps moving. A checklist that has not changed in two years is describing a search environment that no longer exists.

How to hold all of this

The inventory is long and the idea behind it is short: a brand becomes citable by being reachable, identifiable, independently corroborated, and structurally extractable, and each of those can fail in a specific, checkable way. The 127 items are simply the ways we have found those four things to break.

Held that way the list stops being intimidating. You are not running 127 tests; you are answering four questions, and the items are the evidence you gather to answer them. That framing also tells you what to do when you encounter a failure the list does not cover — work out which of the four it belongs to, and add it.

A note on completeness

No inventory of this kind is complete, and claiming otherwise would be the first thing to distrust about it. These are the checks that have caught real failures in our work. A brand operating in a category we have not audited, on a platform we have not seen, will encounter failure modes absent from this list.

That is an argument for treating it as a starting structure rather than a specification. The four-pillar grouping is the durable part; the items within each are a working set that should grow with your own experience. An inventory that never changes is not being used.

What to do with the results

The inventory produces findings, and findings are not a plan. Convert them by identifying the lowest group with a material failure, selecting the specific items within it that are failing, and assigning each an owner and a date. Everything above goes to the deferred register with the condition that would activate it.

That conversion step is where most checklist exercises collapse, because a completed inventory feels like a finished piece of work. It is not — it is the input. The output is one group, a handful of items, and a named person, which is a considerably smaller document than the 127 and the only version anyone will act on.

Proprietary framework

The Five Conditional Groups™

How the 127 checks are organised. Groups are conditional rather than independent — items in a higher group cannot produce a measurable result while a lower group is failing.

16

Measurement

Prompt set, engine coverage, repeat sampling, competitor definition, baselines for citation and share of voice.

Runs first. The other 111 are unfalsifiable without it.

24

Access

Per-bot crawler permissions, log-confirmed arrivals, rendering, blocked resources, directives, sitemaps, response health.

Binary and cheap. Total consequences when it fails.

22

Identity

Direct identification testing, name collisions, schema declaration, linking properties, description consistency across owned properties.

Fails invisibly. Nobody internally thinks to test it.

31

Evidence

Independent category association, description consistency across third parties, coverage inventory, review and community presence, mention profile.

Almost entirely off-property. Where most brands stop.

34

Expression

Answer-first structure, heading descriptiveness, passage self-containment, format-intent match, product data completeness, coverage depth.

Largest group, fastest to fix, and fourth for a reason.

Why the counts are uneven

The distribution reflects how many discrete, checkable things each group contains rather than how much each matters. Expression carries the most items because content has the most inspectable properties; Measurement carries the fewest because it is a small number of design decisions.

Reading item count as importance inverts the actual weighting. Evidence has fewer items than Expression and stops considerably more brands, and Measurement has fewest of all while being the precondition for every other check meaning anything. The counts describe the inventory; the ordering describes the priority.

Common misconceptions

MisconceptionA longer checklist means a better audit.
What the mechanics sayOnly if it is ordered. An unordered 127-point list produces 127 findings and no instruction. The sequence is what makes it useful.
MisconceptionWe can complete this with a site crawler.
What the mechanics sayA crawler covers Access and part of Expression. Identity’s external checks and the entire Evidence group require assessing sources you do not control.
MisconceptionExpression has the most items so it matters most.
What the mechanics sayItem count reflects how many discrete things can be checked, not how often the group binds. Evidence has fewer items and stops more brands.
MisconceptionWe should fix everything we find.
What the mechanics sayFindings above the binding pillar are not actionable yet. Record them, defer them, and return when the constraint clears.
Key takeaways
  1. Run Measurement first — the other 111 checks are unfalsifiable without a baseline.
  2. Work upward and stop at the first failing group. The sequence is the method.
  3. Expect a third of the work to be off your own property. That is the portion internal audits usually omit.
  4. Do not confuse item count with importance. Expression is longest; Evidence binds more often.
  5. Keep a deferred register so nothing is lost while the plan stays narrow.

Frequently asked questions

Do we need to run all 127?

Not in one pass. Run Measurement, then work upward and stop at the first material failure. The remaining checks are recorded as deferred and become the next phase once the constraint clears.

How long does the full inventory take?

Access and Identity are hours. Expression is days for a substantial site. Evidence is the longest because it involves collecting and comparing third-party descriptions manually, and Measurement takes a couple of weeks to produce a stable baseline.

Which checks give the fastest visible result?

Expression items, but only where the pillars beneath have cleared. That conditionality is the reason the list is ordered rather than prioritised by speed.

Can this be automated?

Most of Access, parts of Expression, and much of Measurement automate well. Identity resolution testing and the Evidence group require judgement about whether descriptions genuinely agree, which does not reduce cleanly to rules.

The bottom line

The inventory is 127 checks across five groups: Access, Identity, Evidence, Expression, and the Measurement layer that must exist before any of the others can be validated. The counts are uneven by design — Expression carries the most items because it contains the most discrete, checkable things, while Evidence carries nearly as many and stops far more brands.

What makes the list useful is the ordering rather than the completeness. Run Measurement, work upward, stop at the first group that materially fails, and record the rest in a deferred register. Roughly a third of these checks happen off your own property, which is the portion a site crawler cannot see and the reason technically clean brands are so often absent from the answers their buyers are reading.

References & sources
  1. The Age’X audit practice — this inventory and the AI Visibility Framework™ grouping it follows are ours.
  2. Ahrefs — brand-level correlates of AI visibility, informing the weight of the Evidence group.
  3. Wix — cross-engine citation format analysis, informing several Expression checks.
  4. Semrush — cross-surface citation divergence, informing the per-surface Measurement checks.

Methodology note: the grouping and ordering reflect our audit practice, not a measured hierarchy. Item counts describe the inventory as we run it and are not a claim that these are the only 127 things worth checking — the list evolves. We make no claim about what proportion of brands fail each group.

“An unordered checklist produces findings. An ordered one produces an instruction — which is the only output an audit is actually for.” The Age’X Research Team

Key takeaways

  • 127 checks across Access (24), Identity (22), Evidence (31), Expression (34), Measurement (16).
  • The groups are conditional, not independent — ordering is the method.
  • A site crawler covers roughly 40% of the list and none of Evidence.
  • Expression has the most items; Evidence binds more often.
  • Run Measurement first or the other 111 checks are unfalsifiable.
Sources
  1. 1The Age’X audit practice
  2. 2Ahrefs
  3. 3Wix
  4. 4Semrush
M
Mohabbat Khan
The Age’X builds AI search visibility infrastructure. We track the answer engines every week so your brand stays cited.

See how your brand shows up in AI answers.

Get a free GEO audit — the same analysis behind every article here.