Listicles, articles and product pages dominate AI citations — and search intent decides which format wins the answer.
This study looks at which content formats actually get cited across the major answer engines, and finds three that dominate: listicles, articles, and product pages. The obvious conclusion — publish more listicles — is the wrong one, and the study itself supplies the reason. Format is not what earns the citation. Intent decides which format can answer a given question, and the winning formats dominate because they happen to match the questions people most often ask an assistant.
The research examines citations across several answer engines — Google’s AI Mode, ChatGPT, and Perplexity — and classifies the cited pages by content format, producing a picture of which types of page these systems reach for. Covering multiple engines matters, because it distinguishes patterns common to the answer layer from quirks of a single system.
The headline distribution is clear enough: list-structured content, standard editorial articles, and product pages account for a substantial share of citations. The more useful part of the analysis is the relationship it draws between format and the intent behind the query, which is what turns a descriptive finding into something you can act on.
Listicles perform for a structural reason rather than a stylistic one. A list is pre-chunked: each item is already a bounded, self-contained unit with an implicit label, which is precisely the shape a retrieval system is looking for when it needs a passage it can lift whole. The format does the segmentation work that an engine would otherwise have to do itself.
They also match a very common question shape. A great many queries put to assistants are implicitly enumerative — what are the options, which tools do this, what are the ways to solve that — and a list is the natural form of that answer. The format wins because it is both easy to extract from and aligned with what is frequently being asked, which is a combination rather than a property of lists in themselves.
Standard editorial articles hold their share because most questions are not enumerative. Explanations, causes, comparisons, and judgements need prose, and a well-structured article with clear headings and self-contained sections is entirely extractable — the chunking simply comes from the heading structure rather than from list markup.
This is a useful corrective to the assumption that AI citation rewards fragmented content. It rewards content that is organised, and an article organised into clear sections each answering a specific question is organised. What loses is not prose but undifferentiated prose — long unbroken passages with no signposting and no self-contained units to lift.
The presence of product pages in the top three is the most commercially significant part of the finding. It means answer engines are citing commercial pages, not only editorial ones, when the question calls for it — which contradicts the assumption that the answer layer is purely an informational surface where only guides and explainers compete.
The implication for ecommerce and product-led businesses is direct: product pages are competing for citations and should be built accordingly, with clean structured data, complete specifications, genuine unique description, and real review content. This connects the citation question to catalogue data quality, which is usually treated as an operational concern rather than a visibility one.
Listicles, articles and product pages dominate because they match the questions most often asked. Publishing the format without the matching intent produces content that fits nothing.
The predictable misreading is to treat the distribution as an instruction and convert everything into lists. This fails for a reason the study itself makes explicit: intent determines which format can satisfy a query, so a list written for a question that needs an explanation answers nothing and is cited by no one.
It also produces the recognisable pattern of forced listicles — content bulleted into fragments that would have been clearer as prose, adding structure without substance. Engines are selecting for content that answers the question well and can be cleanly extracted, not for bullet markup. Cargo-culting the format without the fit gets neither.
What it cannot tell you: whether format causes citation or reflects the intent distribution of the queries sampled. Different query mixes would produce different format shares, so read this as a description of what answer engines are currently asked, as much as what they prefer.
The productive translation is a rule for choosing format from the question rather than from a distribution. Enumerative questions — options, tools, ways, examples — want lists. Explanatory questions want structured prose with a clear opening answer. Comparative questions want tables where dimensions are genuinely comparable. Procedural questions want ordered steps. Product questions want well-structured product pages with complete data.
Applied consistently, this produces a content mix that naturally resembles the observed distribution — because the distribution reflects the questions people ask. Arriving at the format through the question is the reliable path; arriving at it by copying the distribution is how sites end up with a library of lists nobody asked for.
The study covers several engines because their behaviour is not identical, and the variation is worth attending to. Engines whose retrieval favours specific extractable passages will lean harder toward structured formats; those weighting broader corroboration may reach further into editorial content. Any single-engine format study risks generalising a local preference.
The practical consequence is that a cross-engine finding is stronger evidence than a single-engine one, and that format decisions should not be tuned to a specific engine’s current behaviour. The stable strategy is matching format to question, which serves every engine because it is what makes the content genuinely answer what was asked.
The only way to know which formats work for your questions is to watch what gets cited. DUNkē tracks your citations across eight AI engines, per prompt and over time, so format decisions rest on evidence.
Look across the three winning formats and the shared property is not that they are lists, articles, or product pages — it is that all three are well-segmented. A listicle has items, an article has headings, a product page has fields. Each arrives pre-divided into units a retrieval system can identify and lift.
That is the durable lesson, and it generalises beyond any format distribution. Whatever you are writing, the property that makes it citable is that it breaks into self-contained, clearly-labelled pieces each answering something specific. Format is one route to that; the underlying discipline is extractability, which applies to everything.
Format share reflects the questions sampled as much as engine preference. A study drawn from a different query mix — more technical, more local, more transactional — would report a different distribution without either being wrong.
Use this to understand why certain formats do well, not as a publishing quota. The transferable finding is the relationship between intent and format, which holds across query mixes; the specific percentages do not.
The rule the study supports is that intent selects format, so it is worth setting out the mapping concretely. Enumerative intent — options, tools, examples, alternatives — wants a list, because the answer genuinely is a set of discrete items. Explanatory intent wants structured prose that opens with a direct answer and develops it under clear headings.
Comparative intent wants a table when the options differ along consistent dimensions, and prose when the comparison turns on judgement rather than specification. Procedural intent wants ordered steps. Commercial intent wants a product page with complete, accurate, structured data. Applied honestly, this mapping produces the format distribution the study observes without anyone having copied it.
Converting explanatory content into list form to chase the citation distribution fails on both dimensions the study identifies. It fails on intent, because a question needing an explanation is not answered by fragments. And it fails on extraction, because bullets carved out of connected reasoning are not self-contained — each depends on the others to make sense, which is precisely what makes a passage unliftable.
The result is content that looks structured and extracts badly, which is worse than the prose it replaced. It also reads poorly, which affects the engagement signals that feed back into how content is assessed. This is the clearest illustration of why the study should be read as explaining a mechanism rather than prescribing a format.
The inclusion of product pages among the dominant cited formats deserves more attention than it usually receives, because it reframes catalogue work as visibility work. If assistants cite product pages when answering commercial questions, then the completeness and accuracy of your product data determines whether your products can be represented in a recommendation.
That connects directly to structured data, feed quality, and product information management — disciplines usually owned by operations rather than marketing. A product page with thin manufacturer-supplied description, missing attributes, and stale availability is not a candidate for citation regardless of how well the rest of the site performs. For product-led businesses this is probably the most actionable finding in the study.
The strong showing of standard articles is worth interpreting carefully, because it is easy to read as reassurance that ordinary content works fine. What is being cited is not ordinary content in general but articles organised well enough to be extracted from — clear headings, self-contained sections, direct answers near the top of each.
An article of the same length and quality written as continuous undifferentiated prose offers a retrieval system nothing bounded to lift, and would not appear in this distribution. So the finding is better read as evidence for structural discipline than as a licence to keep publishing as before. The format is not what earns the citation; the organisation within the format is.
Because the study spans several engines, it supports a claim that single-engine research cannot: that format preference is broadly shared rather than a quirk of one system. That matters, since a format strategy tuned to one engine’s current behaviour would be fragile, while one resting on a pattern common to several is considerably more durable.
It also suggests the underlying cause is architectural rather than editorial. All these systems retrieve passages and compose answers, so all of them favour content that arrives in retrievable passages. That shared mechanism is why the finding generalises, and why the sensible response is structural discipline applied everywhere rather than engine-specific format tuning.
The practical place to apply this is the content brief, before anything is written. If the brief specifies the question the piece answers — drawn from real prompt research rather than a keyword — then the appropriate format follows almost mechanically from the shape of that question.
This is more reliable than deciding format editorially after drafting, which tends to default to house style regardless of fit. It also produces the heterogeneous mix that a well-built topic cluster should have, where each page’s form is determined by what it answers rather than by a template. Making format a required field in the brief is a small process change with a disproportionate effect on citability.
For an established library, the useful exercise is to check format against intent on your most valuable pages. The common failures are recognisable: explanatory content fragmented into bullets, comparison content written as prose where a table would serve, procedures buried in narrative, and product pages carrying no structured data.
Each mismatch is correctable without rewriting the substance, which makes this among the cheaper improvements available. It also tends to surface pages that have been underperforming for reasons nobody could identify — the information was right and the shape was wrong, which is invisible to most content audits because they assess quality rather than form.
A closing caution against overweighting this study: format is a facilitating condition, not a cause of citation. A perfectly formatted page containing nothing worth citing will not be cited, and no amount of structural discipline compensates for having no substance, no evidence, and no authority behind the claims.
Format determines whether your content can be used; relevance, credibility, and originality determine whether it is chosen. The formats identified here dominate because they clear the first hurdle easily and happen to match common question shapes — but every cited page still had to have something worth saying. Read alongside the rest of this hub, format is the cheapest of the citation requirements to satisfy and the least sufficient on its own.
Of everything in this hub, this is the finding most likely to be converted into bad practice, because it appears to license a simple instruction and the instruction is cheap to follow. Publishing lists is easy; determining what question a piece answers and building the right form for it is not.
That asymmetry is exactly why the misuse happens. The disciplined reading requires holding two ideas together — that these formats dominate, and that copying them does not work — which is less satisfying than a directive. Anyone quoting this study as an argument for producing more listicles has taken the half that is easy to act on and dropped the half that makes it work.
Applied across a cluster, the intent-to-format rule produces a deliberately heterogeneous set: a pillar in structured prose, comparison pages as tables, procedural pages as ordered steps, option pages as lists, and product pages built on clean data. Each shaped by the question it owns rather than by a house template.
This is worth stating because clusters are frequently built to a single template for production efficiency, which guarantees format-intent mismatch on a proportion of the pages. The efficiency gain is real and the visibility cost is real too. The compromise that works is templating the structural conventions — heading discipline, answer-first openings, extractable units — while leaving the format itself determined by the question.
Some questions do not map cleanly onto any of the dominant formats — judgement calls, contested topics, questions where the honest answer is that it depends. These are worth attention precisely because the formats that usually win do not apply, and because they are frequently the questions closest to a purchasing decision.
What works here is prose that states the dependency explicitly and then resolves it: naming the conditions under which each answer holds, in self-contained passages an engine can lift with the qualification intact. This is harder to write than a list and considerably more citable than a false enumeration, because it answers a question the enumerative formats cannot. It is also where genuine expertise shows most clearly.
The presence of product pages among the leading cited formats suggests an underexploited position, because commercial pages have historically been optimised for conversion and rarely for extraction. Most product pages carry thin descriptions, incomplete attributes, and no structured data, which makes them poor citation candidates regardless of the underlying product.
A business that treats its product pages as citation assets — complete specifications, genuine unique description, accurate live structured data, real review content — is competing in a category where the average standard is low. Given that these are the pages closest to revenue, the return on closing that gap is unusually direct compared with equivalent effort spent on editorial content further from the purchase.
Because format share depends on query mix, the reliable way to know what works for your questions is to observe your own citations. Tracking which of your pages get cited, and classifying them by format, produces a distribution specific to your category that may differ substantially from the published one.
Where it differs, your data wins — it reflects the actual questions your audience asks rather than a broad sample. Doing this also converts format from an opinion into something testable: change the form of a page that is not being cited, watch whether citation follows, and accumulate evidence about what works in your domain. That is a considerably stronger basis for editorial standards than any published distribution.
Strip away the format labels and this study is making a point about machine legibility. Listicles, structured articles, and product pages share the property that a machine can identify where one unit of information ends and the next begins — through list markup, heading hierarchy, or field structure respectively.
That is the whole mechanism, and it explains why the finding holds across engines with different characters. Any format that makes boundaries explicit performs; any that leaves them implicit does not. Understanding the study at that level makes it durable rather than perishable — format fashions will change, and the requirement that content arrive in identifiable, self-contained pieces will not.
If machine legibility is the mechanism, then it can be stated as a single publishing standard that applies regardless of format: every page should break into units whose boundaries are explicit and whose contents make sense alone. Lists achieve it through markup, articles through headings, product pages through fields — but the requirement is identical.
Written into editorial standards, that requirement is more durable than any format guidance, because it survives changes in what formats are fashionable and in how engines are built. It is also checkable: read the headings alone, read each opening sentence alone, and lift three passages at random to see whether they stand up. A page that passes those checks is citable in whatever form it happens to take, which is the outcome the format distribution is really describing.
The whole finding compresses into a single working rule: let the question decide the format, then make sure every unit inside it can stand alone. The first half prevents the mismatch that makes content unciteable regardless of quality; the second half is what allows a retrieval system to use it at all.
Everything else in this study — which formats dominate, why lists do well, why product pages matter more than expected — is evidence for that rule rather than a set of instructions in itself. Teams that adopt the rule end up producing something close to the observed distribution naturally, because the distribution reflects the questions being asked. Teams that copy the distribution without the rule end up with a library of well-formatted content answering questions nobody posed.
Listicles, editorial articles, and product pages account for a substantial share of citations across AI Mode, ChatGPT, and Perplexity. But the reason is not that engines prefer those formats in themselves — it is that all three arrive pre-segmented into extractable units, and that they match the shapes of question most commonly asked. Search intent decides which format can answer a query, and the observed distribution largely reflects the distribution of questions.
The actionable rule is therefore to choose format from the question rather than from the chart: lists for enumerative questions, structured prose for explanatory ones, tables for genuine comparisons, ordered steps for procedures, and well-built product pages for commercial queries — which the finding confirms are genuinely competing for citations. Underneath all of it, the common factor is segmentation, which is why extractability rather than format is the discipline that actually travels.
“Three formats dominate the citations, and copying them is the wrong lesson. They win because they match the questions being asked — and because all three arrive already broken into pieces a machine can lift.” The Age’X Research Team
Get a free GEO audit — the same analysis behind every article here.