Nineteen times in this publication we have said the cause is not established. These are the pre-registered experiments meant to close six of them.
Across this publication we have repeatedly reached the same sentence: we can observe the pattern and we cannot establish the cause. It appears nineteen times, attached to nineteen questions the field answers confidently and nobody has tested. This series is the attempt to close some of them — pre-registered experiments with hypotheses stated before the data is collected, published whether or not the results support what we currently believe.
The evidence base for AI visibility currently has two layers. There is measurement of effects — how click-through changed, how referrals declined, how surfaces differ — which is genuinely good and cited throughout this publication. And there is everything about causes, which is correlational at best and practitioner inference at worst.
That gap would matter less if the industry were honest about it. Instead the inference layer is presented with the confidence of the measurement layer, and brands are making budget decisions on the basis of claims nobody has tested. We have contributed to this — every framework in this publication is a pattern from audit work, and we have labelled them as such, which is a lower bar than testing them.
The specific problem it solves is that outcomes get reinterpreted after the fact. An experiment that fails becomes an exploratory study. A hypothesis that was not supported gets replaced by one the data does support. Success criteria drift toward whatever moved. None of this requires dishonesty; it is what happens when the analysis follows the data.
Publishing the hypothesis, the method, the sample, the analysis plan, and the success criteria before collecting anything removes the flexibility. The result is then interpretable: either the prediction held or it did not, and both outcomes are publishable. That is a low bar in most empirical fields and an unusual one in marketing, which is precisely why it is worth doing.
Four levels of evidence for a causal claim. Almost all AI visibility advice sits at the bottom two while being presented as though it sat at the top.
Practitioner inference
A pattern observed across engagements. Useful, unfalsifiable, and the level most frameworks occupy — including ours.
Claim strength: this is worth trying.
Correlational
Signals that co-occur with the outcome across a population. Directionally informative, cannot distinguish cause from proxy.
Claim strength: cited brands look like this.
Natural experiment
An external change affecting some subjects and not others, observed rather than induced.
Claim strength: this appears to matter.
Controlled intervention
One variable varied deliberately, others held constant, outcome measured against a pre-registered prediction.
Claim strength: this causes that, within stated limits.
The series opens with six, chosen because each isolates a variable this publication treats as consequential and because each is tractable at a scale we can actually run. They are deliberately unglamorous — single-variable tests of things practitioners assert daily.
Each will be published as a pre-registration first, with hypothesis and method, and as results second, whatever those results are. The gap between the two publications is the commitment: once a prediction is public, a null result cannot be reframed as an interesting exploratory finding.
This publication’s strongest claim is that corroboration attaches poorly to an unresolved entity — that fixing entity confidence retroactively increases the value of coverage a brand already has. It is the load-bearing assertion in three articles and it is inferred, not measured.
The test: identify brands with comparable independent mention volume but materially different entity resolution quality, and compare their citation rates. If resolution gates corroboration, the well-resolved cohort should convert mentions into citations at a substantially higher rate. If the conversion rates are similar, the gating claim is wrong and entity work is a parallel investment rather than a precondition — which would require revising several articles here.
An experiment programme that only reports confirmations is marketing. Publishing the pre-registration first, and the result second whatever it says, is the only structure that makes either worth reading.
The most frequently given advice in this discipline is to restructure content into self-contained, answer-first passages. It is asserted everywhere, including here, and rests on a mechanistic argument about how retrieval works rather than on measurement.
The test: within a single site, take matched pages addressing comparable questions with comparable authority, restructure a randomised half, hold everything else constant, and track citation presence on the specific questions those pages address over two quarters. A within-site design controls for entity, domain authority, and corroboration simultaneously, which is what makes it cleaner than a cross-brand comparison.
Our structured data article argues that schema declares rather than persuades, and that most types add little. That is a mechanistic argument with no measurement behind it, and it contradicts a good deal of industry advice, which makes it exactly the sort of claim that should be tested rather than asserted.
The test: matched pages differing only in the presence of a specific markup type, tracked for citation over time. Run separately per type, starting with Organization linking properties — which we claim is the highest-value declaration — and with FAQPage, which we claim adds little. If the results invert our ordering, the article is wrong and will be revised publicly.
The six experiments at a glance
Rows for each experiment. Columns: the claim being tested, the design, the pre-registered prediction, the sample requirement, the timeframe, and which articles in this publication would need revision if the result is null. That last column is the accountability mechanism.
Our audit work suggests that brands described coherently across a dozen sources outperform brands described inconsistently across forty. It is one of the more useful findings we report and it comes from qualitative comparison rather than measurement.
The test: score a cohort of brands on independent mention volume and on description consistency separately, then model citation rate against both. If consistency predicts citation more strongly than volume, the claim holds. If volume dominates once consistency is controlled for, we have been over-weighting a variable because it made a satisfying story.
The ecommerce argument in this publication rests on product data determining whether an item can be represented in a recommendation. It is mechanically plausible and untested, and it asks brands to fund an operations project on the strength of that plausibility.
The test: within a single catalogue, complete the constraint attributes on a randomised subset of comparable products, leave a matched subset unchanged, and track product-level citation on conditional questions over two quarters. A within-catalogue design controls for brand entity, domain, and corroboration entirely, which makes it the cleanest experiment in the series.
Our citation article claims directness usually decides selection among eligible sources — that the passage answering the exact question beats the better treatment of the topic. That is the strongest assertion in that piece and it rests on comparative observation.
The test: construct matched passage pairs differing on exactly one property — directness, specificity, evidence, currency — publish them on comparable pages, and observe which is selected for the same question. This is the hardest experiment to run cleanly because it requires publishing content designed to be tested, and it is the one whose result would be most immediately actionable.
Each experiment publishes its hypothesis before its data. DUNkē supplies the citation measurement these designs depend on — across eight AI engines, per prompt, over time.
Pre-registration means predicting, so: we expect entity resolution to show a substantial gating effect, extractability restructuring to produce a measurable but smaller effect than practitioners claim, and most schema types to show no detectable effect with the linking properties as a possible exception.
We expect consistency to outperform volume, product attribute completeness to show a clear effect on conditional questions and none on general ones, and directness to dominate the passage properties. Writing these down is uncomfortable, which is the point — if three of the six come back null, that is a finding about how much of this discipline rests on plausible reasoning rather than evidence, and it would include our own.
Every one of these is a within-domain or matched-cohort design with a modest sample, which means results will be suggestive rather than definitive. They cannot rule out confounds we have not thought of, they will not generalise cleanly across categories, and a single null result does not disprove a mechanism that might operate at a scale we cannot reach.
What they can do is move specific claims from level one on the evidence ladder to level three or four, which is a meaningful improvement over the current position. We would rather establish six things partially than continue asserting twenty confidently, and we would rather be contradicted by better work than remain untested.
Publishing the pre-registrations makes them checkable and reproducible. It also invites the criticism that improves a design while it can still change, which is the entire function of methodological transparency and is worth more to us than the head start of running quietly.
There is a commercial argument against this: the results might undermine claims we make in engagements, and a competitor could run the same designs. Both are true. The first is the point — a claim that cannot survive testing should not be sold. The second matters less than a field where several parties test openly, which reaches reliable knowledge faster than one where everyone holds private data and asserts from it.
So the commitment is checkable, here is the template each will follow. The claim being tested, stated as a falsifiable proposition rather than a topic. The design, including how subjects are matched and what is held constant. The sample requirement and how it will be assembled. The measurement instrument and sampling cadence.
Then the analysis plan, specified before data exists, including how confounds will be handled. The success criteria, stated as what result would support the hypothesis and what result would contradict it. And the revision commitment, naming which published articles a null result would affect. That last field is unusual and is the one that makes the rest binding.
Four of the six experiments vary something within a single site or catalogue rather than comparing brands, and that choice is doing most of the methodological work. A within-domain design holds entity resolution, domain authority, corroboration, and technical configuration constant automatically — because they are properties of the domain, not of the page.
Cross-brand comparisons cannot control those without matching on variables that are themselves hard to measure. So where a question can be posed within a domain, it should be, and the cross-brand designs are reserved for questions about brand-level properties that a within-domain design structurally cannot address.
Time. Answer engines change during any experiment long enough to be meaningful, which means an observed difference between a treated and untreated group could reflect an engine change interacting with page characteristics rather than the treatment itself.
The mitigation is concurrent controls rather than before-and-after comparison: treated and untreated pages observed in the same periods, so an engine change affects both. That is why every design here is a comparison between groups rather than a comparison against a baseline, and readers should treat any AI visibility experiment using before-and-after without a concurrent control as uninterpretable.
These will be small. Assembling matched brand cohorts is laborious, within-domain experiments are limited by how many comparable pages a site has, and citation is a noisy outcome requiring repeated sampling. We expect to be reporting suggestive differences rather than tight confidence intervals.
Stating that in advance prevents the usual manoeuvre where an underpowered study reports a striking result and the sample size appears in a footnote. Where an experiment is too small to distinguish a moderate effect from noise, we will say so and report it as inconclusive rather than as a finding — which is a third possible outcome alongside support and contradiction.
The pre-registration to publication pipeline
Linear flow: claim identified from published article → pre-registration published with prediction and revision commitment → data collection → analysis per stated plan → result published (supported / contradicted / inconclusive) → affected articles revised. Mark the point after which the hypothesis can no longer change.
Several of the nineteen remain out of reach. Whether the visibility pillars are strictly ordered would require varying three while holding one constant across a large matched cohort. Whether the recommendation ladder’s rungs are genuinely sequential needs a similar design. Whether branded mention volume causes citation or proxies for prominence needs a sample we cannot assemble.
And the population-level questions — how brands distribute across failure modes, what proportion fail each pillar — need a stratified sample of the market rather than of brands we can reach. Those stay at level one on the ladder and will keep being labelled as inference until someone with a larger sample runs them.
Not null results, which are the point. It fails if pre-registrations stop appearing before data, if an experiment is quietly dropped after an unfavourable interim look, if success criteria are restated after the fact, or if a null result is published without the corresponding article being revised.
Each of those is observable from outside, which is deliberate. A reader can check whether the pre-registration predates the result, whether the prediction matches what was later reported, and whether the named article changed. Publishing the failure conditions is the closest thing to a guarantee available when the party making the commitment is also the one running the tests.
The twenty-nine articles preceding this one set out frameworks, protocols, and arguments, and every one of them is honest about resting on observation rather than measurement. Read as a body of work, the publication makes a strong case and repeatedly declines to claim more than it can support.
This article is what that repeated caveat obliges. Having said nineteen times that the mechanism is unestablished, the choice is to keep saying it or to try to establish some of it. The second is slower, riskier, and more likely to prove parts of the publication wrong, which is a reasonable description of what doing this properly looks like.
The most useful response to this article is a criticism of a design before it runs. A confound we have not considered, a matching variable we should control for, an outcome measure that would be more informative — all of those are worth more now than after the data is collected and the analysis is locked.
That is the practical reason for publishing pre-registrations rather than only results. A design improved by criticism produces better evidence than one defended after the fact, and we would considerably rather revise a method than a conclusion.
The immediately usable part is the ladder rather than the experiments. Applied to any claim about AI visibility — ours, a vendor’s, a conference talk’s — ask which level it sits at. Is this someone’s experience, a correlation across a population, an observed natural experiment, or a controlled test with a stated prediction?
Almost everything currently in circulation is level one or two, and almost all of it is delivered with the confidence of level four. Simply noticing that changes how a reader weights advice, and it costs nothing to apply. It is also the reason we put the ladder near the top of this article rather than the experiment list — the classification is more useful than any single result will be.
Two misreadings to head off. This is not an argument for inaction until evidence exists — the frameworks in this publication are unproven and they are also the best available structure for a decision that has to be made now. Waiting for causal evidence in a field this young means waiting years while competitors accrue.
Nor is it an argument that untested claims are worthless. Practitioner inference from many engagements is genuine information, and dismissing it because it sits at level one would discard most of what is currently known. The instruction is to weight claims by their level and to be honest about which level you are on, not to act only on level four.
Publishing this creates an obligation we may come to regret. If experiment one returns null, three articles in this publication need revising, including the entity work we describe as the highest-leverage neglected discipline. If experiment two returns null, the extractability advice that runs through a dozen pieces is overstated.
We think those outcomes are unlikely and we cannot rule them out, which is the definition of a claim worth testing. Publishing the revision commitment in advance means those revisions will happen visibly rather than through quiet edits, and a reader can check. That is the whole mechanism, and it only works because it is inconvenient.
This publication has argued throughout that the field’s central problem is confident claims resting on thin evidence — unfalsifiable audits, ranking reports that stopped predicting, benchmarks with undisclosed frames, movement commentary built on variance. Each of those criticisms applies to us in some form, and the only honest response is to submit our own claims to the standard we have been demanding.
Six experiments will not settle a discipline. What they can do is establish that the claims here are the kind that could be wrong and that we will say so when they are. In a field where almost nothing is testable, being testable is itself a contribution — and it is the one we would rather be judged on than any framework in the twenty-nine articles before this.
Last, and deliberately. The twenty-nine articles before it establish what we believe and label the evidence behind each claim; this one submits a portion of that to testing. Read in order, the publication moves from frameworks to protocols to arguments to the experiments that would validate the frameworks — which is the correct direction and an unusual one to travel publicly.
Readers arriving here first should read the state-of-the-field assessment and the visibility framework next, since those contain most of the claims these experiments target. Readers who have read everything else should treat this as the accountability document for the rest.
Before acting on any claim about what causes AI visibility — from this publication or anywhere else — ask which level of evidence it rests on. Practitioner inference, correlation, natural experiment, or controlled test with a pre-registered prediction.
Almost everything in circulation is one of the first two and is presented as though it were the last. That single classification habit will improve your decisions more than any framework here, and it costs nothing but the willingness to hold good advice and weak evidence in mind at the same time.
What happens if the results contradict your frameworks?
We revise the articles publicly and say what changed and why. That commitment is stated here so it can be checked later, and the comparison table identifies in advance which articles each null result would affect.
Can others participate?
Yes, and replication by parties with different samples is the most valuable contribution available. The designs are published specifically so they can be run independently, and we would publish a contradicting replication alongside our own result.
Why these six first?
Because each isolates a variable this publication treats as load-bearing, and each is tractable at a scale we can reach. The questions we are not testing yet are mostly ones requiring samples or controls we cannot currently achieve.
How long until results?
The within-domain experiments report over two quarters; the cohort comparisons take longer because they depend on assembling matched samples. Pre-registrations publish first, which means the predictions are on record well before any data exists.
This publication has said nineteen times that a pattern is observable and its cause is not established. That is an honest position and a weak one to keep occupying, because the field is making budget decisions on inference presented as mechanism — and our own frameworks sit at the bottom of the evidence ladder alongside everyone else’s.
The series pre-registers six experiments, each isolating one variable we treat as consequential, with hypotheses and success criteria published before data collection and null results committed to in advance. Six partial answers is a better position than twenty confident assertions, and the accountability mechanism is explicit: the comparison table names which articles here would need revising if each result comes back null. Watch which ones do.
Methodology note: this article contains no results, by design. It is a programme statement and a set of commitments: pre-registration before data collection, publication of null results, public revision of affected articles, and open designs so others can replicate. Each experiment will publish its own pre-registration with full method, sample requirements, analysis plan, and success criteria before any data is gathered. Readers should judge this series on whether those commitments are kept rather than on the results themselves.
“We have written nineteen times in this publication that we can see the pattern and cannot establish the cause. That is an honest thing to say once and an embarrassing thing to keep saying.” The Age’X Research Team
Get a free GEO audit — the same analysis behind every article here.