How AI engines choose what to cite

AI answer engines cite the pages that are the easiest, most trustworthy answer to a query. Most work in two steps — they retrieve pages relevant to the question, then, from that candidate set, favour pages with a clear extractable answer block, unambiguous entities, structured data that matches the page, and strong source authority (with freshness for time-sensitive topics). To be cited, your page has to be both retrieved and the cleanest option to lift from.

Engines don’t publish their exact ranking logic, and behaviour differs by engine and by whether browsing is on. What follows is the honest, observable pattern — not a claim to read their internals.

Retrieval, then selection

For questions with visible citations, engines overwhelmingly cite pages they retrieve at answer time — not just what they were trained on. There are two gates. First, retrieval: does your page clearly match the query’s phrasing and intent enough to enter the candidate set? Second, selection: from those candidates, which page is the easiest and safest to lift a trustworthy answer from and attribute? Winning a citation means passing both gates — which is exactly why on-page GEO signals move the needle.

The factors that decide it

No single factor guarantees a citation, but the pages that get cited tend to score well across these — the same signals CitePack measures.

  1. 01

    Relevance to the exact query

    The page has to be retrieved first. If it doesn't clearly match the phrasing and intent of the buyer's question, it never enters the candidate set to be cited.

  2. 02

    An extractable answer block

    A self-contained, quotable claim near the top — phrased as the answer to the question — is the fragment an engine lifts and attributes. Buried answers get skipped.

  3. 03

    Entity clarity

    The engine has to resolve who you are and what you do. Plainly named brand, product, and category let it treat you as a distinct, citable entity rather than an ambiguous mention.

  4. 04

    Structured data

    Schema.org markup (Organization, Article, FAQPage, Product) hands the engine a machine-readable map of your content and its entities — as long as it matches what's on the page.

  5. 05

    Source authority

    Signals of expertise and trust — a real author or organization, real sources, corroboration across the web — make a model more willing to lean on you when it decides who to cite.

  6. 06

    Freshness

    For anything time-sensitive, current and dated content is preferred over stale pages. Maintained beats abandoned.

Engine by engine

The core pattern is shared, but each engine grounds its answers a little differently:

Perplexity

Web-grounded by design: it retrieves live sources and shows inline citations, so being in the retrieved, extractable set matters directly. CitePack measures Perplexity natively.

ChatGPT

When browsing / search is active, answers are web-grounded and cite retrieved pages; the same extractability and authority signals apply.

Claude

With web access enabled, Claude retrieves and can attribute sources; clear, well-structured, authoritative pages are easier to cite.

Google (Gemini)

We probe Gemini as a proxy for how Google's generative surfaces read a page. We don't measure AI Overviews directly and don't claim to.

A word on honesty

AI engines don’t disclose their exact citation logic, so nobody can promise a specific engine will cite a specific page. What you can do is add the on-page signals that make citation likelier and measure the change. CitePack proves your score moved before → after deterministically, and probes Gemini as a proxy for Google’s generative surfaces — it does not measure AI Overviews directly.

Frequently asked questions

How do AI engines decide what to cite?

Most AI answer engines work in two steps: retrieval then selection. First they retrieve pages relevant to the query; then, from that candidate set, they favour pages that make it easy to produce a trustworthy answer — a clear, extractable answer block near the top, unambiguous entities, structured data that matches the page, and signals of source authority. Freshness matters for time-sensitive topics. Being cited means being both retrieved and the easiest, most trustworthy option to lift from.

Why does an AI cite my competitor instead of me?

Usually because their page is a cleaner answer to the exact question: a self-contained answer block the engine can lift, plainly named entities, structured data, and stronger authority signals — or simply because they were retrieved for that query and you weren't. It's rarely about your product being worse; it's about their page being easier and safer for the model to cite. The fix is to make your page the better-structured, more extractable answer.

Do AI engines cite pages they were trained on, or pages they retrieve?

For questions with visible source citations, it's overwhelmingly pages the engine retrieves at answer time (web-grounded), not just training data. That's why on-page GEO signals move the needle: you're optimizing to be retrieved and then chosen from the live candidate set. Perplexity is retrieval-grounded by default; ChatGPT, Claude, and Gemini are web-grounded when browsing is enabled.

Does structured data guarantee a citation?

No. Structured data (schema.org markup) makes your page easier for an engine to parse and understand, which improves the odds it gets cited — but only alongside genuine relevance, an extractable answer, and authority. And it must accurately describe what's on the page; mismatched or fabricated markup is a liability, not a boost. No honest tool can guarantee a specific citation.

How can I tell whether an engine cites me right now?

Ask the engine the questions your buyers ask, and check whether your domain appears in the actual cited sources — sampled across several runs, per engine. That's the read CitePack automates across ChatGPT, Perplexity, Claude, and Google (Gemini), so you get a concrete before-and-after instead of guessing.

Find out if engines cite you

CitePack queries all four engines for your buyers’ questions and reports, per engine, whether your domain is in the cited sources — free.

See if AI cites your site.

Paste a URL and get a free AI Authority scan in about 30 seconds — the sealed verdict, the seven signals, and the rivals AI recommends instead.