← Back to the blog
GEO October 5, 2026 · 9 min read Oct 5, 2026 · 9 min

Grounding: Where ChatGPT Gets Its Current Information

Michał Rochwerger
Michał Rochwerger
Co-Founder
Abstract illustration of the grounding concept: an AI model pulling fresh pages from a live search index at answer time instead of relying only on frozen training knowledge.
The short answer
ChatGPT knows about fresh events through grounding — basing answers on a live index rather than training data alone. Training and grounding are separate processes: OpenAI runs GPTBot for training and OAI-SearchBot for its search index. A page opted out of OAI-SearchBot will not be shown in ChatGPT search answers, so being in the index is the precondition for visibility.

ChatGPT knows about events that happened long after its training cut-off because of grounding — the practice of basing a model’s answer on live search results and a fresh index rather than on training data alone. Microsoft describes web search in its Copilot as “grounding them in the latest information from the web” (Microsoft 365 Copilot documentation, May 2026). Training and grounding are two separate processes running on two separate pipelines, and the distinction decides whether your page can ever appear in an AI answer. This article explains where ChatGPT gets its current information and what that means for your visibility. Data reviewed as of .

What is grounding, and how is it different from training?

Grounding is when a model composes its answer from a live index or search results, not from the static knowledge frozen into its weights during training. A large language model (LLM) — the system that generates the text you read in ChatGPT — learns a fixed snapshot of the web during training, and that snapshot has a cut-off date. To answer a question about something newer, the model does not “remember” it; it runs a retrieval step, pulls current pages, and writes an answer on top of them. Microsoft’s own definition of web search in Copilot is exactly this: “grounding them in the latest information from the web” (Microsoft 365 Copilot documentation, May 2026).

OpenAI makes the split concrete by running three separate crawlers with three separate jobs. GPTBot fetches content to train models, OAI-SearchBot builds the index that ChatGPT search queries, and ChatGPT-User fetches a page on a user’s live request (OpenAI’s Web Crawlers, July 2026). Training and grounding are not the same event: a page can be excluded from training and still be groundable, or vice versa. That is why “the model was trained before my product launched” is not a reason it can’t cite your product — grounding is the path that reaches new information.

Where does ChatGPT get its current information?

ChatGPT search grounds its answers on more than one index, not on Bing alone. The Bing index still carries significant weight, but OpenAI also maintains its own index built by OAI-SearchBot, uses third-party providers, and shows traces of reaching into Google in paid tiers; for Enterprise and Edu, Bing is declared the only external provider (ChatGPT search, OpenAI Help Center, July 2026). The “ChatGPT equals Bing” shorthand is outdated — but Bing indexation still matters, because it feeds one of the largest grounding sources.

The retrieval query is not your user’s prompt copied verbatim. In Copilot, the system parses the prompt, picks out terms, and generates a short search query from them: “This generated search query is different from the user’s original prompt — it consists of a few words informed by the user’s prompt” (Microsoft 365 Copilot documentation, May 2026). That internal, machine-generated query is what your page has to match to be retrieved — which is why writing for the literal wording of a question matters less than covering the concepts behind it.

Diagram showing where ChatGPT gets its current information: a user question reaches the AI model, which instead of the frozen training snapshot grounds the answer on fresh pages from the search index and returns an answer with citations.

How does grounding work in Copilot, Perplexity, and AI Overviews?

Every major generative engine — a system that answers with generated text instead of a list of links — grounds through an index, but the index differs by engine. Microsoft 365 Copilot is grounded exclusively through Bing: when web search is on, Copilot generates a query, sends it to the Bing service, and composes its answer from the results returned (Microsoft 365 Copilot documentation, May 2026). For Copilot, being indexed in Bing is a hard precondition for visibility.

Google’s AI Overviews (AI-generated summaries above the search results) and AI Mode draw on the ordinary Google Search index. To appear in them, “a page must be indexed and eligible to be shown in Google Search with a snippet” (AI features and your website, Google Search Central, December 2025). Google is also explicit that no special markup is required: “You don’t need to create new machine readable files, AI text files, or markup … There’s also no special schema.org structured data that you need to add” (Google Search Central, December 2025).

Perplexity and Anthropic mirror OpenAI’s training-versus-grounding split with their own bot families. Perplexity runs PerplexityBot to index pages for its search results (declaredly not for model training) and Perplexity-User to fetch a page on a user’s request, which “generally ignores” robots.txt (Perplexity Crawlers documentation, July 2026). Anthropic runs ClaudeBot for training, Claude-SearchBot to index content for Claude’s search, and Claude-User to fetch on demand (Anthropic Support, July 2026).

Engine Grounding index Visibility precondition
ChatGPT search Bing + OpenAI’s OAI-SearchBot index + third-party Not opted out of OAI-SearchBot
Microsoft 365 Copilot Bing only Indexed in Bing
Google AI Overviews / AI Mode Ordinary Google Search index Indexed and snippet-eligible in Google
Perplexity PerplexityBot search index Crawlable by PerplexityBot
Claude Claude-SearchBot index Crawlable by Claude-SearchBot

What does grounding mean for your visibility?

The rule is blunt: to be cited, you have to be in the index the model queries. OpenAI states it directly — “Sites that are opted out of OAI-SearchBot will not be shown in ChatGPT search answers” (OpenAI’s Web Crawlers, July 2026). The same logic holds for Google: blocking Googlebot or failing to get indexed means no grounding in AI Overviews or AI Mode. Note the nuance — the Google-Extended token disables Gemini training and grounding but does not affect Search: it “does not impact a site’s inclusion in Google Search nor is it used as a ranking signal,” so it does not opt you out of AI Overviews (Google crawlers documentation, July 2026).

Being in an index the model queries is not the same as ranking on that engine, though. Search-grounded chatbots mostly cite pages that do not rank in Google’s top 10: only 12% of cited links sat in the top 10, Perplexity came closest at 28.6%, and roughly 80% of citations did not rank anywhere in Google for the original query (Ahrefs AI Search Overlap Study, August 2025). The gap is widening: the share of AI Overviews citations drawn from Google’s top 10 fell from 76.1% in July 2025 to 37.9% in March 2026, as grounding fans out beyond the ranking’s front page — with YouTube the single most-cited domain at 5.6% of all citations (Ahrefs, March 2026). The takeaway: being in the grounding index is necessary, but a top ranking is neither necessary nor sufficient.

What can silently block your content from grounding?

The most common invisible blocker is client-side rendering, because AI crawlers do not run JavaScript. They download JS files but do not execute them: “None of the major AI crawlers currently render JavaScript” (Vercel and MERJ, December 2024). Any content that only appears after client-side JavaScript runs — with the exception of Google’s own ecosystem and AppleBot — is invisible to grounding. The fix is server-side rendering or static generation, so the key facts sit in the HTML the server returns. We walk through this in our complete guide to generative engine optimization.

There is external evidence that grounding really does reach for current pages when the plumbing is right. AI assistants cite content that is materially fresher than organic Google results — on average 1,064 days old for AI citations versus 1,432 days for organic, a 25.7% freshness edge; ChatGPT citations averaged around 958 days from publication (the exception is AI Overviews, whose citations are about as old as the SERP) (Ahrefs, July 2025). Grounding is not a theoretical feature — it demonstrably pulls in newer pages than classic search does.

One more caution about what “cited” buys you. Being in the index and getting cited does not guarantee your brand name appears in the answer text: 61.7% of AI citations are ghost citations — the link is there, the brand name is not — and only 38.3% of citations name the brand (Semrush and Kevin Indig, June 2026). Grounding gets your URL considered; making the model actually say your name is a separate, harder job.

If you want to know which grounding indexes currently reach your site — and where you are already being retrieved versus invisible — that is exactly what a Vistrix Labs AI visibility audit measures across ChatGPT, Perplexity, Copilot, and AI Overviews. For the mechanics of how a model then decides which of those pages to cite, see how ChatGPT chooses the sources it cites, and for the Bing dependency specifically, Bing Webmaster Tools as the gateway to Copilot.

Training gives a model a frozen snapshot of the past; grounding is how it reads today’s web. The consequence is simple and unforgiving: if you are not in the index the model queries at answer time, you do not exist to it — no matter how good your content is.

FAQ

Does ChatGPT get real-time data from the live web?

Yes, through grounding. ChatGPT search bases its answers on a live index rather than training data alone, pulling from Bing, OpenAI’s own OAI-SearchBot index, and third-party providers (ChatGPT search, OpenAI Help Center, July 2026). The training snapshot has a cut-off date; grounding is the separate retrieval step that reaches information published after it. A page excluded from OAI-SearchBot “will not be shown in ChatGPT search answers” (OpenAI’s Web Crawlers, July 2026).

What is the difference between training and grounding for an LLM?

Training bakes a fixed snapshot of the web into the model’s weights, with a cut-off date; grounding retrieves current pages at answer time and composes the reply on top of them. OpenAI keeps them physically separate — GPTBot fetches for training, OAI-SearchBot builds the search index, ChatGPT-User fetches on a user’s request (OpenAI’s Web Crawlers, July 2026). Perplexity and Anthropic use the same split, with distinct bots for indexing versus training.

Do I need special schema or AI files to appear in grounded AI answers?

No. For Google AI Overviews and AI Mode, “You don’t need to create new machine readable files, AI text files, or markup … There’s also no special schema.org structured data that you need to add”; a page simply “must be indexed and eligible to be shown in Google Search with a snippet” (AI features and your website, Google Search Central, December 2025). The real technical requirement is that your content sits in server-rendered HTML, because AI crawlers do not execute JavaScript (Vercel and MERJ, December 2024).

Keep reading

All posts →
Abstract illustration of the GEO glossary as a knowledge graph: a network of connected nodes representing 30 Generative Engine Optimization terms and brand visibility in AI-generated answers.

The GEO Glossary: 30 Terms You Need to Know

Abstract vector illustration of GEO as an umbrella term: a yellow umbrella arc over five smaller acronym badges converging into a single node, symbolizing that GEO and AEO are one phenomenon of optimizing content for AI citations.

GEO vs AEO: The Difference and Which Term to Use

Vector illustration of a large language model (LLM): a neural network turning a stream of scattered data into ordered generated text — the engine behind ChatGPT, Gemini and other AI tools

LLMs for Marketers: What a Large Language Model Is and How It Works

Find out if AI knows your brand

Book a free consultation