- Every AI search vendor publishes access controls. Not one publishes selection criteria.
- Ask whether a system would need your page rather than one of the forty others saying the same thing.
- One test before publishing anything: would this be missing from the internet if we did not publish it?
Ask the current first page of results how to get cited by an AI assistant and you get a formatting exercise. Publish an llms.txt — a text file some vendors say AI crawlers want. Add machine-readable labels to your FAQs. Open every section with a direct answer of forty to sixty words. One widely shared guide states that pages with structured data are cited 3.2 times more often, and sources that figure to a post on Medium.
Google’s own documentation contradicts most of that list, item by item — our companion piece works through what the guidance actually supports. The more interesting point is what happens when you stop reading Google alone.
What the vendors document is access, not selection
Read the primary sources next to each other and a pattern appears.
Google’s Search Central documentation on AI features states it plainly: “There are no additional requirements to appear in AI Overviews or AI Mode, nor other special optimizations necessary.” A page must be indexed and eligible to appear with a snippet. That is the whole technical bar. Its newer guide to optimising for generative AI features adds one editorial instruction with teeth — “create valuable, non-commodity content” and “don’t just recycle what others on the internet have already said, or could easily be produced by a generative AI model.”
OpenAI’s crawler documentation describes four user agents and what each is for. The only lever it gives you is the switch: sites opted out of OAI-SearchBot “will not be shown in ChatGPT search answers”. Perplexity’s documentation covers bot identity, IP ranges and firewall configuration. Neither says a word about how a source gets chosen.
So the published record is a permissions layer. Three vendors, three documentation sets, and every one of them tells you how to be reachable while none tells you how to be worth reaching for. An enormous amount of effort is being spent optimising a selection process nobody has published.
Retrievability is the prior question
The industry framed this as a ranking problem because ranking is the problem it already knew how to talk about. It is the wrong frame.
Retrieval-augmented generation, as Google describes it, retrieves candidate pages using search ranking systems and then generates an answer with supporting links. Query fan-out means the system is not answering your query; it is answering several related ones it composed itself. There is no single ranked list to climb. There is a filter that runs before anything is written, and it asks something simpler than “which page is best” — it asks which sources it needs in order to answer at all.
So stop asking how to rank inside the answer. Ask whether there is any question for which a retrieval system would need your page rather than one of the forty others that say approximately the same thing.
For generic information the answer is no, and it gets worse every quarter. If forty pages carry the same fact, any of them will do, so your odds of being chosen are one in forty and falling. When a model can produce the paragraph itself, it does not need a source for it at all.
Originality is not a virtue signal here. It is arithmetic. Being the only place a piece of information lives collapses the substitute set to one. That is the operational meaning of expertise in an AI-mediated web, and it is where our applied AI work usually starts: not with tooling, but with what the organisation knows that nothing else on the web knows.
What makes a source worth retrieving
Four properties do most of the work, in our reading.
Accountability. A named person, with a real history, who can be found outside your own domain. Not an author box. A human whose judgment is on the line.
Specificity. Narrow enough that a competent general model could not produce an equivalent paragraph from what is already public. Conditions, versions, dates, the case where the advice stops working.
Evidence. The number, where it came from, when it was measured and what it excludes. An unsourced statistic is not evidence; it is a liability that a careful system has every reason to route around.
Structure that serves a reader. Headings that state claims rather than tease them. Not because chunking is a ranking factor — Google explicitly says it is not — but because a claim stated plainly can be checked, and a claim buried in three paragraphs of preamble cannot. Our sibling article on what AI cannot easily summarise works through which categories of knowledge actually survive that test.
The research disagrees with itself, and the reporting cannot settle it
Citation behaviour is moving, it is not controllable, and it is barely measurable. Anyone telling you otherwise is guessing.
Look at the third-party research. Ahrefs analysed 863,000 keywords and four million AI Overview URLs and found 38 per cent of cited pages ranked in the top ten, down from 76 per cent in the previous study. seoClarity, examining 362,000 queries from a single day in October 2025, found 94 per cent of AI Overviews cited at least one top-twenty URL and 56 per cent of citations came from the top twenty. Those are not the same picture. Ahrefs is admirably clear that its own parsing methodology changed between studies and that Google has confirmed no change to fan-out behaviour, so part of that movement may be better detection rather than different behaviour.
Measurement is no better: Search Console’s generative AI report shows impressions but no clicks and no queries, which caps what any of this can be held to.
We cannot claim any of it causes citations. Nobody has that data, and the vendors publishing citation-rate dashboards are inferring from samples of a system being retrained underneath them.
The thesis survives anyway, because it does not depend on the mechanics holding still. If retrieval behaviour changes next quarter, an organisation that has published what it uniquely knows still owns that knowledge in public, and it still converts in a sales conversation and an ordinary search result. An organisation that invested in forty-word answer blocks owns a formatting convention with an expiry date. That asymmetry holds under uncertainty. This is our reading, not a settled question.
The test before you publish anything
One question, applied to every page: would this be missing from the internet if we did not publish it?
If the answer is no, you are producing a substitute for something that already exists, and asking a retrieval system to prefer you for reasons neither you nor the vendor can name. If the answer is yes, you have written something a system has a reason to fetch and a person has a reason to trust — whichever way the citation mechanics move next.
That question is worth more than a scoring rubric because it cannot be gamed. A page either carries something that would otherwise be absent, or it does not, and the person who commissioned it always knows which. Point growth and performance work at the pages that pass, and stop expanding the ones that do not.
Track it over quarters, not weeks, and pay more attention to whether the right conversations are starting than to any impressions figure. That is a slower signal than a citation dashboard. It has the advantage of being real.