Query fan-out, retrieval-augmented generation, and the signals that decide which source makes the answer.
When an AI engine composes an answer, it draws on some sources and ignores others — and which ones make the answer is not random. A set of signals decides it: relevance to the specific sub-question, whether the source can be retrieved at all, whether a clean point can be extracted, how credible the source is, whether claims are evidenced, whether the entity is clear, and whether the content is current. Understanding these signals is understanding how to be cited — because they are exactly the levers your content can pull.
Every AI answer is built from a selection: out of everything the engine could draw on, it uses and cites a few sources. The practical question for visibility is what makes a source one of the chosen few. The answer is a combination of signals the engine weighs — some about whether it can find and use your content at all, some about whether your content is the best, most trustworthy answer to the specific point it is composing. These signals are the criteria that decide citation.
Knowing the signals matters because they are the levers you can actually pull. Each signal corresponds to a quality you can build into your content — relevance, retrievability, extractability, credibility, evidence, entity clarity, freshness. Understanding them turns "how do we get cited?" from a mystery into a checklist of the qualities that drive selection. This piece walks through the signals that decide which source makes the answer, so you know what to build toward to be the source engines choose.
These signals operate within the pipeline AI engines run: they decompose a question into sub-questions (query fan-out), retrieve relevant sources for each (retrieval-augmented generation), and compose a cited answer (synthesis). Selection happens across this pipeline — some signals determine whether you are retrieved into consideration, others determine whether you are chosen in synthesis to ground and cite the answer. A companion piece covers the pipeline in depth; here, the focus is the signals that decide selection at each point.
The key is that citation requires clearing both hurdles: being retrieved (found and pulled into consideration) and being selected in synthesis (chosen as a source to ground and cite the answer). Some signals — retrievability, relevance — get you into consideration; others — extractability, credibility, evidence — get you chosen. Understanding that selection spans retrieval and synthesis clarifies why a source can be found but not cited, and why the signals below matter at different stages of the engine composing its answer.
The first signal is relevance to the specific sub-question the engine is answering. Because fan-out breaks a query into sub-questions, the engine seeks sources that genuinely address each one — and relevance is semantic, about matching the meaning and intent of the sub-question, not merely its keywords. Content that directly and specifically answers a sub-question is relevant to it; content that is vague, off-topic, or only superficially related is not, and is passed over for sources that address the point squarely.
The practical implication is to make your content directly answer the specific questions your audience asks, in clear, focused sections that address particular sub-questions. Comprehensive content covering a topic’s facets is relevant to more sub-questions across the fan-out, multiplying citation chances. Relevance is the entry criterion for being considered for a citation on a given point — the engine will not cite a source that does not genuinely address the sub-question, so being specifically relevant to the questions in your domain is foundational to being cited.
The second signal is retrievability — whether the engine can find and pull your content at all. An engine can only cite what it retrieves, so being discoverable and accessible to it is a prerequisite: your content must be crawlable to the engine’s systems, present in the index or sources it draws on, and accessible without barriers. A source the engine cannot retrieve is invisible to it, however relevant or excellent, because it never enters consideration for the answer.
The practical work is ensuring you are retrievable: crawlable to AI crawlers, present in the sources engines draw on (including the reputable and community sources they favor), and free of technical barriers that block access. Retrievability is the entry ticket — it does not guarantee citation, but its absence guarantees non-citation. Understanding retrievability as a distinct signal is why the first step in AI visibility is ensuring engines can actually find and pull your content, before any question of whether they choose to cite it.
Some signals get you found; others get you cited. Relevance and retrievability put you in the running; extractability, credibility, evidence, entity clarity and freshness decide whether the model quotes you.
The third signal is extractability — whether the engine can cleanly lift a clear point from your content to use in its answer. In synthesis, the engine draws specific points from sources, and it favors content from which a clear, self-contained answer can be readily extracted, over content where the answer is buried in meandering prose. Answer-first structure — leading with a direct, self-contained, evidenced point — makes your content easy to extract and cite; burying the answer makes it harder and less likely to be chosen.
The practical implication is to structure content so specific points are extractable: leading with clear answers, using self-contained passages, and organizing so the answer to each question is easy to find and lift. This is often the highest-impact change for citability, because it directly serves what synthesis needs. Extractability is why answer-first content wins — it is what the engine can readily use to compose its answer, which makes it the source selected over content whose points are harder to extract.
The fourth signal is credibility and authority — whether the source is trustworthy enough for the engine to rely on and cite. Engines favor sources that are credible, well-regarded, and authoritative on the topic, because they want their answers to be accurate and defensible — and grounded engines especially lean on sources they can trust. Credibility comes from genuine expertise, a strong reputation, links and mentions from other credible sources, and the signals the E-E-A-T framework captures: experience, expertise, authoritativeness, trust.
The practical work is building genuine authority on your topics — through expert, high-quality content, a strong reputation, and recognition from credible sources — so engines trust you enough to cite. Credibility is decisive because an engine will prefer a trustworthy source over a dubious one even for the same point, to keep its answer reliable. Understanding credibility as a signal is why building real authority, not just producing content, is what makes you the source engines choose to ground and cite their answers with.
The signals decide who gets cited; DUNkē shows you the outcome. Track whether you’re the source AI engines choose — across eight engines, per prompt, against competitors — so you can see where your signals are winning and where to strengthen them.
The fifth signal is evidence and specificity — whether claims are backed and concrete rather than vague and unsupported. Engines, especially grounded ones wary of fabrication, favor content that is specific and evidenced: concrete facts, statistics, and claims backed by credible sources are safer to cite than vague assertions, because they can be trusted and attributed. Research on generative-engine visibility finds that adding statistics, quotations, and citations of credible sources can meaningfully lift a source’s chances of being cited.
The practical implication is to make your content specific and evidenced: replace vague claims with concrete detail, support assertions with data and credible references, and be precise rather than general. This is not decoration but direct optimization for what engines reward when choosing sources. Evidence and specificity are why concrete, well-supported content is cited over vague content — it is safer and more useful for the engine to ground its answer in, which makes it the selected source for the points it addresses.
The sixth signal is entity clarity — whether the engine can clearly identify and understand your brand as a source. To cite you, an engine must be able to attribute a point to a clearly-understood entity, which requires that you are consistently defined, recognized, and disambiguated across the web and in the knowledge graph. A clearly-defined entity is easier to attribute facts to and to treat as an authority on its topics; an ambiguous or poorly-understood one is harder to cite confidently.
The practical work is being a clearly-defined entity: consistent naming and description, structured data declaring your entity, accurate representation across your presence, and recognition in the knowledge graph. Entity clarity helps engines attribute answers to you and trust you as a source on your topics. Understanding entity clarity as a signal is why being a well-defined, recognizable entity — not just producing content — is part of being citable, since citation requires an attributable, understood source.
The seventh signal is freshness — whether the content is current, which matters especially for time-sensitive questions. Because engines retrieve at answer time and can pull recent sources, they favor up-to-date content for queries where recency matters — current events, evolving topics, latest information — because a stale answer would be wrong. For such questions, fresh content is retrieved and cited while outdated content is passed over for more current sources.
The practical implication is to keep important content genuinely current where recency matters, treating key pages as living documents with a real updating cadence. Freshness is not equally important for all content — timeless topics are less sensitive — but for anything time-sensitive, it is a real signal in whether you are cited. Understanding freshness as a signal is why maintaining current content, especially in fast-moving domains, is part of being citable, since engines favor sources that reflect the present for questions about the present.
These signals do not operate in isolation — they combine, and being cited generally requires clearing several. A source must be retrievable and relevant to be considered, and extractable, credible, evidenced, entity-clear, and (where relevant) fresh to be chosen. Weakness in one can be partly offset by strength in others, but a serious failure — being unretrievable, irrelevant, or non-credible — can rule you out regardless of other strengths. The signals together determine whether you are the source the engine selects.
The practical takeaway is that citability comes from strength across the signals, not from optimizing one alone. Content that is retrievable, relevant, extractable, credible, evidenced, entity-clear, and fresh is the kind engines cite, because it clears every criterion the selection weighs. Understanding how the signals combine is why AI visibility rewards genuine, all-around quality — being the best available source across the criteria — rather than a single trick, since the engine is weighing the whole picture when it chooses what to cite.
The signals translate into a clear content brief: make your content retrievable (crawlable, present in the sources engines draw on), relevant (directly answering the specific questions in your domain), extractable (answer-first, self-contained), credible (genuinely authoritative), evidenced (specific, backed by data and sources), entity-clear (consistently defined and recognized), and fresh (current where recency matters). Building these qualities is optimizing directly for what decides citation.
This brief is the practical output of understanding how LLMs choose what to cite: rather than guessing, you build the specific qualities the selection signals reward. The rest of the curriculum expands each into detail — extractability, entity SEO, authority, freshness, and the rest — but the signals give the unifying logic. Understanding what gets cited is understanding what to build, and the signals turn AI visibility into a concrete practice of making your content the source that clears every criterion engines weigh when composing their answers.
Engines weight these signals because they serve the engine’s goal: composing accurate, helpful, defensible answers. Relevance ensures the answer addresses the question; retrievability ensures the engine can use the source; extractability lets it compose cleanly; credibility and evidence keep the answer accurate and trustworthy; entity clarity lets it attribute confidently; freshness keeps time-sensitive answers correct. Each signal is not an arbitrary hurdle but a proxy for whether a source helps the engine answer well.
Understanding why engines weight the signals demystifies them: they are what a system trying to give good, grounded, attributable answers would naturally favor. This is why gaming the signals superficially does not work — the engine is selecting for genuine helpfulness and trustworthiness, which the signals approximate. The practical lesson is that being genuinely the best, most trustworthy, most usable source for a question is what the signals reward, because that is exactly what helps the engine achieve its goal of answering well.
In practice, retrievability has several components: being crawlable to AI crawlers (allowing bots like GPTBot, ClaudeBot, and PerplexityBot to access your content), being present in the indexes and sources engines draw on, and being free of technical barriers — broken rendering, blocked resources, access walls — that prevent the engine from reading you. Each component is a potential point of failure that can make excellent content invisible to an engine simply because it cannot be retrieved.
The practical work is to verify each: confirm AI crawlers are permitted and can access your content, ensure you are indexed and present in the sources engines favor, and remove technical barriers to being read. Retrievability is foundational because it gates everything — no citation is possible without it — yet it is often overlooked in favor of content work. Getting retrievability right in practice is the essential first step, ensuring the engine can actually find and read your content before any question of whether it chooses to cite you.
Relevance and comprehensiveness interact in a way that rewards thorough coverage. Because fan-out generates multiple sub-questions, comprehensive content — relevant to many of them — can be retrieved and cited across more of an answer than narrow content relevant to only one. Comprehensiveness thus multiplies relevance’s effect: covering a topic’s facets thoroughly makes you relevant to more sub-questions, and thus present across more of the fan-out, than a thin page addressing a single point.
The practical implication is to build genuinely comprehensive resources that are relevant across a topic’s facets, not thin pages targeting single queries. This is not about length for its own sake but about covering the real questions and sub-questions around a topic, each answered clearly. The interplay of relevance and comprehensiveness is why depth and breadth translate into more citations — comprehensive, relevant content is retrievable and citable across the many sub-questions a generative answer draws on, multiplying your presence.
Credibility is judged from signals engines can detect: links and mentions from other credible sources, consistency and accuracy across your presence, evidence of expertise and authorship, a strong reputation in your domain, and accurate representation on authoritative and community sources. These signals aggregate into an assessment of whether you are a trustworthy source on a topic. Engines cannot directly know your expertise, so they infer it from these detectable signals of credibility and authority.
The practical work is to build and demonstrate the signals of genuine credibility: earn links and mentions from credible sources, be consistent and accurate everywhere you appear, show clear expertise and authorship, and ensure accurate representation across the sources engines draw on. Because engines infer credibility from these signals, building them is how you become a source engines trust enough to cite. Understanding the credibility signals engines can detect is why genuine authority-building — not claims of expertise — is what makes you a citable source.
Vague content loses because it fails multiple signals at once: it is less relevant (not specifically answering sub-questions), less extractable (no clear point to lift), and less credible (unsupported, unspecific claims are harder to trust). An engine composing a grounded, accurate answer has little use for content that is general, unsupported, and unclear, and much use for content that is specific, evidenced, and direct. Vagueness is penalized across the selection criteria simultaneously.
The practical lesson is to be specific and concrete: answer particular questions directly, support claims with evidence and detail, and avoid general, unsupported assertions. Specificity serves relevance, extractability, and credibility together, which is why concrete, evidenced content is cited over vague content. Understanding why vague content loses is a clarifying counterpoint to the signals — it shows that failing to be specific, evidenced, and direct undermines several citation criteria at once, which is why precision and evidence are so consistently rewarded.
The way to know whether your citation signals are working is to test the outcome: whether you are actually being cited in the answers engines give for the questions that matter. Tracking your citation share — per question, per engine, against competitors — reveals whether your content is being selected, on which questions you win or lose, and where your signals may be weak. This closes the loop between building the signals and knowing whether they are producing citations.
The practical discipline is to measure citation outcomes and use them to diagnose: if you are retrieved but not cited, extractability or credibility may be weak; if never retrieved, retrievability or relevance may be the issue. Testing whether your signals are working turns citation from an assumption into a measured result you can improve. Understanding that the signals should be validated by measured outcomes is why AI visibility is a practice of building the signals and tracking whether they produce the citations you are working toward.
While the core signals are shared, engines weight them somewhat differently, reflecting their characters. A grounded, accuracy-focused engine may weight credibility and evidence especially heavily; a real-time-oriented engine may weight freshness more; an engine that leans on community sources may reflect your presence there strongly. These differences mean the same content may fare somewhat differently across engines, and the emphasis of your optimization can be tuned to the engines that matter most to you.
The practical implication is that the shared signals serve all engines, while understanding each engine’s emphasis lets you tune your effort — the engine-specific guides in this curriculum cover these nuances. But the differences are variations on the shared criteria, not different systems, so strong all-around signals serve every engine. Understanding how the signals differ by engine is useful for prioritization and fine-tuning, without changing the core lesson: build genuine strength across the signals, and tune the emphasis toward the engines your audience uses.
Which sources make an AI answer is decided by signals: relevance to the sub-question and retrievability get you into consideration; extractability, credibility, evidence, entity clarity, and freshness get you chosen in synthesis. Understanding these signals is understanding how to be cited, because each corresponds to a quality you can build into your content — they are the levers that decide selection, operating across the retrieval-and-synthesis pipeline engines run.
The practical brief follows directly: make your content retrievable, relevant, extractable, credible, evidenced, entity-clear, and fresh — the qualities that clear the criteria engines weigh. Because the signals combine, citability comes from all-around strength, not a single trick. Understanding how LLMs choose what to cite turns AI visibility from a mystery into a checklist of the qualities that make you the source the machine selects, grounds its answer in, and quotes.
“Which source makes the answer isn’t random — it’s decided by signals. Relevance and retrievability get you in the running; extractability, credibility, evidence, entity clarity and freshness decide whether you’re quoted.” The Age’X Research Team
Get a free GEO audit — the same analysis behind every article here.