DUNkē tracking 12,847 prompts globally·+34% AI mentions for Mysthelle this week·WeaverStory now cited in 4/5 engines·Banana Club ranking #2 on Perplexity·Linen Trail · 11x backlink growth · Q2·DUNkē tracking 12,847 prompts globally·+34% AI mentions for Mysthelle this week·WeaverStory now cited in 4/5 engines·Banana Club ranking #2 on Perplexity·Linen Trail · 11x backlink growth · Q2·
Blog
Research series

The Age’X AI search experiment series

Nineteen times in this publication we have said the cause is not established. These are the pre-registered experiments meant to close six of them.

MMohabbat Khan
15 min read

Across this publication we have repeatedly reached the same sentence: we can observe the pattern and we cannot establish the cause. It appears nineteen times, attached to nineteen questions the field answers confidently and nobody has tested. This series is the attempt to close some of them — pre-registered experiments with hypotheses stated before the data is collected, published whether or not the results support what we currently believe.

Executive summary
  • Almost everything the field asserts about causation is inference. Correlational studies and audit patterns dressed as mechanism.
  • Pre-registration is the fix. Hypothesis, method, and success criteria published before data collection, so a null result cannot be quietly discarded.
  • Six experiments open the series, each isolating one variable the rest of this publication treats as important.
  • Null results will be published, which is the only commitment that makes the rest credible.
  • The point is not to be right. It is to make the field’s claims testable, including ours.

Why the field needs this

The evidence base for AI visibility currently has two layers. There is measurement of effects — how click-through changed, how referrals declined, how surfaces differ — which is genuinely good and cited throughout this publication. And there is everything about causes, which is correlational at best and practitioner inference at worst.

That gap would matter less if the industry were honest about it. Instead the inference layer is presented with the confidence of the measurement layer, and brands are making budget decisions on the basis of claims nobody has tested. We have contributed to this — every framework in this publication is a pattern from audit work, and we have labelled them as such, which is a lower bar than testing them.

What pre-registration actually fixes

The specific problem it solves is that outcomes get reinterpreted after the fact. An experiment that fails becomes an exploratory study. A hypothesis that was not supported gets replaced by one the data does support. Success criteria drift toward whatever moved. None of this requires dishonesty; it is what happens when the analysis follows the data.

Publishing the hypothesis, the method, the sample, the analysis plan, and the success criteria before collecting anything removes the flexibility. The result is then interpretable: either the prediction held or it did not, and both outcomes are publishable. That is a low bar in most empirical fields and an unusual one in marketing, which is precisely why it is worth doing.

Proprietary framework

The Evidence Ladder™

Four levels of evidence for a causal claim. Almost all AI visibility advice sits at the bottom two while being presented as though it sat at the top.

01

Practitioner inference

A pattern observed across engagements. Useful, unfalsifiable, and the level most frameworks occupy — including ours.

Claim strength: this is worth trying.

02

Correlational

Signals that co-occur with the outcome across a population. Directionally informative, cannot distinguish cause from proxy.

Claim strength: cited brands look like this.

03

Natural experiment

An external change affecting some subjects and not others, observed rather than induced.

Claim strength: this appears to matter.

04

Controlled intervention

One variable varied deliberately, others held constant, outcome measured against a pre-registered prediction.

Claim strength: this causes that, within stated limits.

The six experiments

The series opens with six, chosen because each isolates a variable this publication treats as consequential and because each is tractable at a scale we can actually run. They are deliberately unglamorous — single-variable tests of things practitioners assert daily.

Each will be published as a pre-registration first, with hypothesis and method, and as results second, whatever those results are. The gap between the two publications is the commitment: once a prediction is public, a null result cannot be reframed as an interesting exploratory finding.

Experiment one: does entity resolution gate corroboration?

This publication’s strongest claim is that corroboration attaches poorly to an unresolved entity — that fixing entity confidence retroactively increases the value of coverage a brand already has. It is the load-bearing assertion in three articles and it is inferred, not measured.

The test: identify brands with comparable independent mention volume but materially different entity resolution quality, and compare their citation rates. If resolution gates corroboration, the well-resolved cohort should convert mentions into citations at a substantially higher rate. If the conversion rates are similar, the gating claim is wrong and entity work is a parallel investment rather than a precondition — which would require revising several articles here.

The commitment that makes this credible
Null results get published

An experiment programme that only reports confirmations is marketing. Publishing the pre-registration first, and the result second whatever it says, is the only structure that makes either worth reading.

Experiment two: does restructuring for extractability cause citation?

The most frequently given advice in this discipline is to restructure content into self-contained, answer-first passages. It is asserted everywhere, including here, and rests on a mechanistic argument about how retrieval works rather than on measurement.

The test: within a single site, take matched pages addressing comparable questions with comparable authority, restructure a randomised half, hold everything else constant, and track citation presence on the specific questions those pages address over two quarters. A within-site design controls for entity, domain authority, and corroboration simultaneously, which is what makes it cleaner than a cross-brand comparison.

Experiment three: does any schema type affect citation?

Our structured data article argues that schema declares rather than persuades, and that most types add little. That is a mechanistic argument with no measurement behind it, and it contradicts a good deal of industry advice, which makes it exactly the sort of claim that should be tested rather than asserted.

The test: matched pages differing only in the presence of a specific markup type, tracked for citation over time. Run separately per type, starting with Organization linking properties — which we claim is the highest-value declaration — and with FAQPage, which we claim adds little. If the results invert our ordering, the article is wrong and will be revised publicly.

Recommended visual — Comparison table

The six experiments at a glance

Rows for each experiment. Columns: the claim being tested, the design, the pre-registered prediction, the sample requirement, the timeframe, and which articles in this publication would need revision if the result is null. That last column is the accountability mechanism.

Experiment four: does description consistency beat coverage volume?

Our audit work suggests that brands described coherently across a dozen sources outperform brands described inconsistently across forty. It is one of the more useful findings we report and it comes from qualitative comparison rather than measurement.

The test: score a cohort of brands on independent mention volume and on description consistency separately, then model citation rate against both. If consistency predicts citation more strongly than volume, the claim holds. If volume dominates once consistency is controlled for, we have been over-weighting a variable because it made a satisfying story.

Experiment five: does product attribute completeness change recommendation?

The ecommerce argument in this publication rests on product data determining whether an item can be represented in a recommendation. It is mechanically plausible and untested, and it asks brands to fund an operations project on the strength of that plausibility.

The test: within a single catalogue, complete the constraint attributes on a randomised subset of comparable products, leave a matched subset unchanged, and track product-level citation on conditional questions over two quarters. A within-catalogue design controls for brand entity, domain, and corroboration entirely, which makes it the cleanest experiment in the series.

Experiment six: which passage property decides the tie-break?

Our citation article claims directness usually decides selection among eligible sources — that the passage answering the exact question beats the better treatment of the topic. That is the strongest assertion in that piece and it rests on comparative observation.

The test: construct matched passage pairs differing on exactly one property — directness, specificity, evidence, currency — publish them on comparable pages, and observe which is selected for the same question. This is the hardest experiment to run cleanly because it requires publishing content designed to be tested, and it is the one whose result would be most immediately actionable.

Follow the results

Pre-registrations first, results second

Each experiment publishes its hypothesis before its data. DUNkē supplies the citation measurement these designs depend on — across eight AI engines, per prompt, over time.

Explore DUNkē →

What we expect to find, stated in advance

Pre-registration means predicting, so: we expect entity resolution to show a substantial gating effect, extractability restructuring to produce a measurable but smaller effect than practitioners claim, and most schema types to show no detectable effect with the linking properties as a possible exception.

We expect consistency to outperform volume, product attribute completeness to show a clear effect on conditional questions and none on general ones, and directness to dominate the passage properties. Writing these down is uncomfortable, which is the point — if three of the six come back null, that is a finding about how much of this discipline rests on plausible reasoning rather than evidence, and it would include our own.

The limits of what these can establish

Every one of these is a within-domain or matched-cohort design with a modest sample, which means results will be suggestive rather than definitive. They cannot rule out confounds we have not thought of, they will not generalise cleanly across categories, and a single null result does not disprove a mechanism that might operate at a scale we cannot reach.

What they can do is move specific claims from level one on the evidence ladder to level three or four, which is a meaningful improvement over the current position. We would rather establish six things partially than continue asserting twenty confidently, and we would rather be contradicted by better work than remain untested.

Why we are publishing the programme rather than just the results

Publishing the pre-registrations makes them checkable and reproducible. It also invites the criticism that improves a design while it can still change, which is the entire function of methodological transparency and is worth more to us than the head start of running quietly.

There is a commercial argument against this: the results might undermine claims we make in engagements, and a competitor could run the same designs. Both are true. The first is the point — a claim that cannot survive testing should not be sold. The second matters less than a field where several parties test openly, which reaches reliable knowledge faster than one where everyone holds private data and asserts from it.

What a pre-registration will contain

So the commitment is checkable, here is the template each will follow. The claim being tested, stated as a falsifiable proposition rather than a topic. The design, including how subjects are matched and what is held constant. The sample requirement and how it will be assembled. The measurement instrument and sampling cadence.

Then the analysis plan, specified before data exists, including how confounds will be handled. The success criteria, stated as what result would support the hypothesis and what result would contradict it. And the revision commitment, naming which published articles a null result would affect. That last field is unusual and is the one that makes the rest binding.

Why within-domain designs come first

Four of the six experiments vary something within a single site or catalogue rather than comparing brands, and that choice is doing most of the methodological work. A within-domain design holds entity resolution, domain authority, corroboration, and technical configuration constant automatically — because they are properties of the domain, not of the page.

Cross-brand comparisons cannot control those without matching on variables that are themselves hard to measure. So where a question can be posed within a domain, it should be, and the cross-brand designs are reserved for questions about brand-level properties that a within-domain design structurally cannot address.

The confound we are most worried about

Time. Answer engines change during any experiment long enough to be meaningful, which means an observed difference between a treated and untreated group could reflect an engine change interacting with page characteristics rather than the treatment itself.

The mitigation is concurrent controls rather than before-and-after comparison: treated and untreated pages observed in the same periods, so an engine change affects both. That is why every design here is a comparison between groups rather than a comparison against a baseline, and readers should treat any AI visibility experiment using before-and-after without a concurrent control as uninterpretable.

Sample size, honestly

These will be small. Assembling matched brand cohorts is laborious, within-domain experiments are limited by how many comparable pages a site has, and citation is a noisy outcome requiring repeated sampling. We expect to be reporting suggestive differences rather than tight confidence intervals.

Stating that in advance prevents the usual manoeuvre where an underpowered study reports a striking result and the sample size appears in a footnote. Where an experiment is too small to distinguish a moderate effect from noise, we will say so and report it as inconclusive rather than as a finding — which is a third possible outcome alongside support and contradiction.

Recommended visual — Flowchart

The pre-registration to publication pipeline

Linear flow: claim identified from published article → pre-registration published with prediction and revision commitment → data collection → analysis per stated plan → result published (supported / contradicted / inconclusive) → affected articles revised. Mark the point after which the hypothesis can no longer change.

Questions we are not testing yet

Several of the nineteen remain out of reach. Whether the visibility pillars are strictly ordered would require varying three while holding one constant across a large matched cohort. Whether the recommendation ladder’s rungs are genuinely sequential needs a similar design. Whether branded mention volume causes citation or proxies for prominence needs a sample we cannot assemble.

And the population-level questions — how brands distribute across failure modes, what proportion fail each pillar — need a stratified sample of the market rather than of brands we can reach. Those stay at level one on the ladder and will keep being labelled as inference until someone with a larger sample runs them.

What would count as this series failing

Not null results, which are the point. It fails if pre-registrations stop appearing before data, if an experiment is quietly dropped after an unfavourable interim look, if success criteria are restated after the fact, or if a null result is published without the corresponding article being revised.

Each of those is observable from outside, which is deliberate. A reader can check whether the pre-registration predates the result, whether the prediction matches what was later reported, and whether the named article changed. Publishing the failure conditions is the closest thing to a guarantee available when the party making the commitment is also the one running the tests.

Why this belongs at the end of the publication

The twenty-nine articles preceding this one set out frameworks, protocols, and arguments, and every one of them is honest about resting on observation rather than measurement. Read as a body of work, the publication makes a strong case and repeatedly declines to claim more than it can support.

This article is what that repeated caveat obliges. Having said nineteen times that the mechanism is unestablished, the choice is to keep saying it or to try to establish some of it. The second is slower, riskier, and more likely to prove parts of the publication wrong, which is a reasonable description of what doing this properly looks like.

An invitation to disagree with the designs

The most useful response to this article is a criticism of a design before it runs. A confound we have not considered, a matching variable we should control for, an outcome measure that would be more informative — all of those are worth more now than after the data is collected and the analysis is locked.

That is the practical reason for publishing pre-registrations rather than only results. A design improved by criticism produces better evidence than one defended after the fact, and we would considerably rather revise a method than a conclusion.

How a reader should use this before any results exist

The immediately usable part is the ladder rather than the experiments. Applied to any claim about AI visibility — ours, a vendor’s, a conference talk’s — ask which level it sits at. Is this someone’s experience, a correlation across a population, an observed natural experiment, or a controlled test with a stated prediction?

Almost everything currently in circulation is level one or two, and almost all of it is delivered with the confidence of level four. Simply noticing that changes how a reader weights advice, and it costs nothing to apply. It is also the reason we put the ladder near the top of this article rather than the experiment list — the classification is more useful than any single result will be.

What this does not license

Two misreadings to head off. This is not an argument for inaction until evidence exists — the frameworks in this publication are unproven and they are also the best available structure for a decision that has to be made now. Waiting for causal evidence in a field this young means waiting years while competitors accrue.

Nor is it an argument that untested claims are worthless. Practitioner inference from many engagements is genuine information, and dismissing it because it sits at level one would discard most of what is currently known. The instruction is to weight claims by their level and to be honest about which level you are on, not to act only on level four.

The uncomfortable part for us

Publishing this creates an obligation we may come to regret. If experiment one returns null, three articles in this publication need revising, including the entity work we describe as the highest-leverage neglected discipline. If experiment two returns null, the extractability advice that runs through a dozen pieces is overstated.

We think those outcomes are unlikely and we cannot rule them out, which is the definition of a claim worth testing. Publishing the revision commitment in advance means those revisions will happen visibly rather than through quiet edits, and a reader can check. That is the whole mechanism, and it only works because it is inconvenient.

The last word

This publication has argued throughout that the field’s central problem is confident claims resting on thin evidence — unfalsifiable audits, ranking reports that stopped predicting, benchmarks with undisclosed frames, movement commentary built on variance. Each of those criticisms applies to us in some form, and the only honest response is to submit our own claims to the standard we have been demanding.

Six experiments will not settle a discipline. What they can do is establish that the claims here are the kind that could be wrong and that we will say so when they are. In a field where almost nothing is testable, being testable is itself a contribution — and it is the one we would rather be judged on than any framework in the twenty-nine articles before this.

Where this sits in the series

Last, and deliberately. The twenty-nine articles before it establish what we believe and label the evidence behind each claim; this one submits a portion of that to testing. Read in order, the publication moves from frameworks to protocols to arguments to the experiments that would validate the frameworks — which is the correct direction and an unusual one to travel publicly.

Readers arriving here first should read the state-of-the-field assessment and the visibility framework next, since those contain most of the claims these experiments target. Readers who have read everything else should treat this as the accountability document for the rest.

The single instruction

Before acting on any claim about what causes AI visibility — from this publication or anywhere else — ask which level of evidence it rests on. Practitioner inference, correlation, natural experiment, or controlled test with a pre-registered prediction.

Almost everything in circulation is one of the first two and is presented as though it were the last. That single classification habit will improve your decisions more than any framework here, and it costs nothing but the willingness to hold good advice and weak evidence in mind at the same time.

Common misconceptions

MisconceptionPractitioner experience is evidence.
What the mechanics sayIt is level one on the ladder: useful, unfalsifiable, and routinely presented as though it were level four. Ours included — every framework in this publication sits there until tested.
MisconceptionCorrelational studies establish what works.
What the mechanics sayThey identify signals that accompany the outcome. Distinguishing a cause from a proxy for overall brand prominence requires varying one thing deliberately, which correlational designs cannot do.
MisconceptionAn experiment that fails was a waste.
What the mechanics sayA published null result removes a claim from circulation, which is more valuable than another confirmation. The waste is running it and not publishing.
MisconceptionThese will settle the questions.
What the mechanics sayThey will move six specific claims up the evidence ladder with modest samples and stated limits. Settling requires replication by parties other than us.
Key takeaways
  1. Locate any claim on the evidence ladder before acting on it — including every framework in this publication.
  2. Demand pre-registration from anyone publishing causal claims about AI visibility.
  3. Treat null results as findings. A programme that only confirms is not a programme.
  4. Prefer within-domain designs, which control for entity and authority in ways cross-brand comparisons cannot.
  5. Watch which of our articles get revised. That is the accountability this series creates.

Frequently asked questions

What happens if the results contradict your frameworks?

We revise the articles publicly and say what changed and why. That commitment is stated here so it can be checked later, and the comparison table identifies in advance which articles each null result would affect.

Can others participate?

Yes, and replication by parties with different samples is the most valuable contribution available. The designs are published specifically so they can be run independently, and we would publish a contradicting replication alongside our own result.

Why these six first?

Because each isolates a variable this publication treats as load-bearing, and each is tractable at a scale we can reach. The questions we are not testing yet are mostly ones requiring samples or controls we cannot currently achieve.

How long until results?

The within-domain experiments report over two quarters; the cohort comparisons take longer because they depend on assembling matched samples. Pre-registrations publish first, which means the predictions are on record well before any data exists.

The bottom line

This publication has said nineteen times that a pattern is observable and its cause is not established. That is an honest position and a weak one to keep occupying, because the field is making budget decisions on inference presented as mechanism — and our own frameworks sit at the bottom of the evidence ladder alongside everyone else’s.

The series pre-registers six experiments, each isolating one variable we treat as consequential, with hypotheses and success criteria published before data collection and null results committed to in advance. Six partial answers is a better position than twenty confident assertions, and the accountability mechanism is explicit: the comparison table names which articles here would need revising if each result comes back null. Watch which ones do.

References & sources
  1. Seer Interactive, Ahrefs, Semrush, Conductor, Reuters Institute, Wix — the measurement layer this publication rests on, and the standard the causal layer does not yet meet.
  2. The Age’X practice — the Evidence Ladder™ is our framework, and sits at level one of itself until tested.

Methodology note: this article contains no results, by design. It is a programme statement and a set of commitments: pre-registration before data collection, publication of null results, public revision of affected articles, and open designs so others can replicate. Each experiment will publish its own pre-registration with full method, sample requirements, analysis plan, and success criteria before any data is gathered. Readers should judge this series on whether those commitments are kept rather than on the results themselves.

“We have written nineteen times in this publication that we can see the pattern and cannot establish the cause. That is an honest thing to say once and an embarrassing thing to keep saying.” The Age’X Research Team

Key takeaways

  • Almost every causal claim in this field is inference presented as mechanism.
  • Pre-registration removes the flexibility that lets null results be reinterpreted.
  • Six experiments, each isolating one variable this publication treats as load-bearing.
  • Null results will be published and affected articles revised publicly.
  • Every framework here sits at level one of the evidence ladder until tested.
Sources
  1. 1Seer Interactive
  2. 2Ahrefs
  3. 3Semrush
  4. 4Conductor
  5. 5Reuters Institute
  6. 6Wix
  7. 7The Age’X practice
M
Mohabbat Khan
The Age’X builds AI search visibility infrastructure. We track the answer engines every week so your brand stays cited.

See how your brand shows up in AI answers.

Get a free GEO audit — the same analysis behind every article here.