← Back to the blog
GEO July 16, 2026 · 9 min read Jul 16, 2026 · 9 min

How AI Chooses Sources: The Mechanics Behind ChatGPT Citations

Dawid Walczyk
Dawid Walczyk
Co-Founder
Abstract illustration of how a language model chooses the sources it cites: from a field of documents only selected pages flow into a neural network that produces an answer with citations
The short answer
AI citations come from the grounding layer — a search index (OAI-SearchBot, Bing, Google's core ranking) — not from model training. The strongest correlates of getting cited sit outside your own site: 84% of AI citations come from earned media (Muck Rack, 2026), and citation patterns can shift by tens of percentage points within weeks.

“How AI chooses sources” sounds like a question about machine judgment. It is mostly a question about plumbing. ChatGPT citations — and the links attached by Gemini, Copilot and Google AI Overviews — come out of a retrieval pipeline: a crawler builds a search index, a retrieval layer selects documents for a query, and a large language model (LLM) composes the answer with those documents in its context. This article takes the pipeline apart using platform documentation and 2024–2026 studies; for the strategy on top of it, see our guide to GEO (generative engine optimization) — creating and optimizing content so that it gets cited and used in AI-generated answers.

Do ChatGPT citations come from training data or from search?

Citations in AI answers come from the search layer — grounding — not from model training. Training compresses text into model parameters with no record of where a fact came from, so an answer generated purely from training cannot reliably attach a source. Grounding is the opposite: basing model answers on live search results and an index rather than training data alone. A generative engine — a system that answers with a generated response instead of a list of links (ChatGPT, Perplexity, Google AI Overviews, Copilot) — retrieves specific documents and links them. That link is the citation (in an AI answer): a reference to a page as a source within the generated response.

OpenAI separates the two worlds at the crawler level. It operates separate bots with independent robots.txt controls, three of which matter for content visibility: GPTBot crawls for model training, OAI-SearchBot builds the index behind ChatGPT search, and ChatGPT-User fetches pages when a user asks for them (OpenAI, 2025). Because the controls are independent, a site can block training via GPTBot and remain fully visible in ChatGPT search via OAI-SearchBot — while blocking OAI-SearchBot removes the site from ChatGPT search answers. One control is weaker than it looks: for ChatGPT-User the same documentation states, “Because these actions are initiated by a user, robots.txt rules may not apply”.

robots.txt control What it governs What blocking it does
GPTBot crawling for OpenAI model training excludes content from training; ChatGPT search unaffected (OpenAI, 2025)
OAI-SearchBot the index behind ChatGPT search the site stops appearing in ChatGPT search answers (OpenAI, 2025)
ChatGPT-User page fetches triggered by a user unreliable — “robots.txt rules may not apply” (OpenAI, 2025)
Google-Extended use of content for Gemini training/grounding; a control token, not a separate crawler “does not impact a site’s inclusion in Google Search nor is it used as a ranking signal”; AI Overviews run on the regular Search index (Google, 2026)

The practical consequence: presence in training data reflects years of the web corpus; citations depend on the state of a search index right now. Working on AI citations means working on the grounding layer.

Which search indexes decide what AI can cite?

Three index families feed the major generative engines: OpenAI’s own index crawled by OAI-SearchBot, Microsoft’s Bing index, and Google’s regular Search index. ChatGPT search is no longer a Bing wrapper — OpenAI states it “may share disassociated search queries with the Bing search engine” and with third-party search providers, while building its own index in parallel; only for Enterprise/Edu plans is Bing named as the sole external provider (OpenAI Help, accessed July 2026). Microsoft 365 Copilot “may fetch information from the Bing search service when information from the web helps to provide a better, more grounded response” (Microsoft Learn, 2026) — a page absent from Bing cannot appear in Copilot’s web-grounded answers. Google’s AI Overviews and AI Mode draw on the same index as classic Search (Google, 2026).

The entry ticket to those indexes is technical readability. In a server-log analysis covering, among others, 569 million monthly GPTBot requests, Vercel and MERJ found that “none of the major AI crawlers currently render JavaScript” — GPTBot downloads JS files in 11.5% of requests but never executes them, so client-side-only content is invisible; the exceptions are Google’s rendering infrastructure and AppleBot (Vercel, 2024; the finding held in follow-up sources through 2026).

Being indexed is not the same as ranking, and retrieval does not replay Google’s results. Across 15,000 long-tail queries, only 12% of links cited by chatbots ranked in Google’s top 10 for the query, and about 80% ranked nowhere in Google for it; Perplexity sat closest to Google at 28.6% from the top 10 (Ahrefs, 2025). AI Overviews behave differently, because they share Google’s ranking systems: a page at position 1 appears in AI Overviews 43% of the time versus 7% at position 20, and 94% of AI Overviews contain at least one citation from the top 20 (seoClarity, 2025). How this redefines optimization work is covered in GEO vs SEO — what actually changes.

Diagram showing where AI citations come from: the grounding pipeline crawler → index → retrieval → answer with citations, with model training shown separately as producing no citations

What content traits correlate with getting cited?

The strongest measured correlates of AI citations sit outside a brand’s own website — in what independent sources publish about it. Four findings from 2025–2026 define the current evidence base:

  • Earned media dominates. Third-party sources account for 84% of all citations in AI answers; journalism alone makes up 27% of cited sources, and paid content just 0.3% — based on 25+ million links from ChatGPT, Claude and Gemini answers (Muck Rack, 2026).
  • Mentions beat backlinks. In a study of 75,000 brands, brand mentions on YouTube were the single strongest correlate of AI visibility (Spearman ~0.737), ahead of branded web mentions (0.656–0.709) and far ahead of backlinks (~0.218) or Domain Rating (0.266–0.326) (Ahrefs, 2025). These are correlations on brands with DR above 40, not proof of causation — but the hierarchy is consistent.
  • Freshness helps moderately. ChatGPT, Perplexity, Gemini and Copilot cite content on average 25.7% fresher than Google’s organic results (1,064 vs 1,432 days since publication), yet the average cited page is still about 2.9 years old, and in AI Overviews the freshness effect disappears (Ahrefs, 2025). Relevance outweighs recency.
  • The top of the page works hardest. 44.2% of 18,012 verified ChatGPT citations pointed to the first 30% of a page’s content, versus 24.7% to the final third (Kevin Indig, 2026). Correlational, but it supports stating the answer at the start of every section.

On-page tactics deserve more skepticism than they usually get. The peer-reviewed GEO experiment (KDD 2024) found that adding quotations, statistics and source references raised visibility in generative answers by up to 40% on the Position-Adjusted Word Count metric across a 10,000-query benchmark, while keyword stuffing cut it by about 9% (Aggarwal et al., 2024). An independent replication did not confirm the effect: in C-SEO Bench only 3 of 54 cases showed a statistically significant positive result, adding statistics worsened citation ranking in 19 of 24 settings, and the dominant factor was the document’s position in the model’s context (Puerto et al., NeurIPS 2025). The evidence for editorial tricks is mixed; the evidence for the weight of independent sources is not.

Whether a language model cites your page is decided less by what you publish on your own domain than by what independent sources publish about you. AI citations are largely earned off-site and only collected on-site.

What are ghost citations?

A ghost citation is a citation in which a page’s URL is listed as a source but the brand is never named in the answer text — and it describes 61.7% of citations in AI answers, leaving only 38.3% of citations paired with a brand mention (Semrush and Kevin Indig, 2026; 3,981 domain occurrences across 115 prompts in 14 countries — treat the numbers as an order of magnitude). A citation and a brand mention — the brand name appearing in the AI answer text, regardless of citation — are therefore two separate goals that need separate measurement.

The engines have inverse profiles. Gemini names brands in 83.7% of occurrences but cites their pages in only 21.4%; ChatGPT is the mirror image, citing in 87% of occurrences while naming the brand in just 20.7% (Semrush and Kevin Indig, 2026). Content type shifts the balance too: comparison content produced brand mentions in 43.3% of occurrences — 2.4 times more than informational content (same study). A site can be a heavily used source and remain invisible as a brand, or the reverse; a single “AI visibility” number hides which is happening.

Why do AI citation patterns keep changing?

Source selection can flip by tens of percentage points within weeks, with no change on the publishers’ side. The best-documented case comes from 13 weeks of tracking 230,000 prompts: Reddit appeared in roughly 60% of ChatGPT answers in early August 2025 and about 10% by mid-September, while Wikipedia fell from about 55% to under 20% — a shift confined to ChatGPT, as AI Mode and Perplexity stayed stable over the same period (Semrush, 2025). Google drifts on a slower clock: the share of AI Overviews citations from Google’s top 10 fell from 76.1% in July 2025 to 37.9% in March 2026 — though part of that drop reflects a methodology change between editions (Ahrefs, 2026).

The operational conclusion is blunt: a one-off measurement of AI visibility has a shelf life of weeks, and findings from one engine do not transfer to another. Visibility needs to be measured per engine, on a fixed prompt set, at regular intervals — tracking brand mentions and citations separately. That is what an AI visibility audit does: it measures where a brand appears in ChatGPT, Gemini, Perplexity and Copilot answers and which sources each engine cites in your category — the scope is described in Vistrix Labs services.

FAQ: how AI chooses sources in practice

Does blocking GPTBot remove my site from ChatGPT answers?

No. GPTBot governs model training only; visibility in ChatGPT search depends on a separate crawler, OAI-SearchBot, and the two robots.txt controls work independently (OpenAI, 2025). Blocking OAI-SearchBot is what removes a site from ChatGPT search answers. Note the exception: for user-triggered fetches by ChatGPT-User, OpenAI states that “robots.txt rules may not apply”.

Do I need to rank in Google’s top 10 to be cited by AI?

Not for chatbots: only 12% of links cited by ChatGPT, Gemini and Copilot ranked in Google’s top 10 for the query, and about 80% ranked nowhere for it (Ahrefs, 2025; 15,000 long-tail queries). AI Overviews are the exception — position 1 correlates with 43% presence versus 7% at position 20 (seoClarity, 2025).

Will adding an llms.txt file get my content cited?

No — llms.txt, a proposed standard for a site content map for AI systems, is not used by any major LLM provider as of mid-2026. Among roughly 38,000 domains with a valid file, 97% received zero requests for it in May 2026, and Google officially ignores the standard (Ahrefs, 2026). Indexability, server-rendered HTML and earned media move citations; llms.txt currently does not.

Keep reading

All posts →
AI SEO: a list of search results transforming into an AI-generated answer with a highlighted source citation — hero illustration for an article on AI SEO and GEO

What Is AI SEO and How It Differs From GEO

Abstract illustration of Google AI Overviews: a glowing AI-generated summary above classic search results, linked to three source pages

What Are Google AI Overviews and How Do They Work?

Abstract illustration of the Generative Engine Optimization (GEO) concept: content fragments from three web pages flow through a yellow funnel into an AI answer bubble with marked citations.

Generative Engine Optimization (GEO): The Complete Guide

Find out if AI knows your brand

Book a free consultation