TL;DR: AI search research splits into three tiers of evidence. One peer-reviewed benchmark (GEO, KDD 2024) reported up to 40% visibility gains on its own metrics, but the independent C-SEO Bench replication (NeurIPS 2025) found a significant positive effect in only 3 of 54 tested cases. Large citation studies show consistent correlations around brand mentions, earned media, and freshness. Almost nothing is causally proven.
How strong is the evidence behind AI visibility advice?
Most advice about visibility in AI answers rests on correlation, not proof — and the honest way to read the research is to sort every finding into three buckets: proven, correlated, or a bet. Proven means an experimental result, ideally replicated by an independent team. Correlated means an observational study found a pattern, with no demonstration of cause. A bet means a practice with no supporting evidence yet.
The distinction matters because the field is young. GEO (generative engine optimization) — creating and optimizing content so that it gets cited and used in AI-generated answers — got its first peer-reviewed paper in 2024, its first independent replication in 2025, and a wave of industry citation studies in between. As of July 2026, the evidence base is real but thin, and vendors routinely present correlations as causal levers. This overview walks through every major study we consider citable and assigns each conclusion a status. For the practical playbook, see our generative engine optimization guide.
What does the peer-reviewed research on GEO show?
Exactly two peer-reviewed experimental studies exist, and they largely disagree: the original GEO paper reported large gains from content tactics; the independent replication failed to confirm them. Everything else is industry research.
The GEO paper (KDD 2024): up to 40% — on a simulated engine
The paper that coined the term — “GEO: Generative Engine Optimization” by Aggarwal, Murahari and co-authors, accepted at KDD 2024 — states in its abstract that “GEO can boost visibility by up to 40% in generative engine responses” (Aggarwal et al., 2024). The study introduced GEO-bench — 10,000 queries (8K/1K/1K split), 9 optimization methods, two custom metrics: Position-Adjusted Word Count (PAWC) and Subjective Impression (Aggarwal et al., 2024).
The best-performing tactics were adding quotations, adding statistics, and citing sources: relative gains of 30–40% on PAWC and 15–30% on Subjective Impression, with test-set scores of 27.8 (Quotation), 25.9 (Statistics) and 24.9 (Cite Sources) against a 19.5 baseline (Aggarwal et al., 2024, results tables). Effects varied strongly by topical domain. Two more findings matter:
- Keyword stuffing backfired: 17.8 PAWC versus the 19.5 baseline, roughly −9% (Aggarwal et al., 2024).
- Low-ranked sources gained the most: Cite Sources produced +115.1% visibility for the source at position 5 of the underlying results, while the position-1 source lost 30.3% — a zero-sum game inside the answer (Aggarwal et al., 2024, section 5.2, Table 2).
The critical caveat: the engine was simulated — top Google results fed to an LLM — and the metrics were the authors’ own. Proven as a benchmark experiment, not a universal rule about production systems.
C-SEO Bench (NeurIPS 2025): the replication that said no
The only independent large-scale replication did not confirm the GEO paper’s effects. C-SEO Bench by Puerto et al. (NeurIPS 2025 Datasets & Benchmarks) tested content-optimization methods across 1.9K+ queries, 16.3K documents, 6 domains and 4 models — and found a statistically significant positive effect in only 3 of 54 tested cases, with gains around 5.6% (Puerto et al., 2025). Worse: the Statistics Addition method significantly hurt citation ranking in 19 of 24 settings. The authors are blunt: “most current C-SEO methods are not only largely ineffective but also frequently have a negative impact on document ranking” (Puerto et al., 2025).
C-SEO Bench added two strategic findings: traditional SEO — where a document ranks in the context passed to the model — was significantly more effective than any content-rewriting tactic, and under simulated mass adoption gains converge to zero (Puerto et al., 2025). We dissect both papers number by number in our deep-dive on the Princeton GEO paper versus its replication.
One result survived both experiments: keyword stuffing hurts — roughly −9% on the GEO benchmark, about −10% on production Perplexity.ai (PAWC 21.9 vs 24.1), with C-SEO Bench corroborating the negative direction (Aggarwal et al., 2024; Puerto et al., 2025). It is the only tactic-level finding two independent experiments agree on.
As of July 2026, exactly one tactic-level finding in AI search research is corroborated by two independent experiments: keyword stuffing reduces visibility in AI answers. Every positive tactic is unreplicated, correlational, or a bet.

What do large-scale AI citations studies show about who gets cited?
Observational studies of millions of real citations agree on one headline: what AI engines cite overlaps surprisingly little with what ranks in classic Google. Three Ahrefs datasets with published methodology map the pattern.
Across 15,000 long-tail queries, only 12% of links cited by search-enabled chatbots ranked in Google’s top 10 for the same query — Perplexity was the most Google-aligned at 28.6%, ahead of Copilot (8.6%), Gemini (8.2%) and ChatGPT in-text citations (8.0%) — and roughly 80% of cited pages ranked nowhere in Google’s top 100 for the original query (Ahrefs, 2025). Google’s AI Overviews — AI-generated summaries above the search results — sit closer to the classic index, but the share of their citations from top-10 organic results fell to 37.9% in a March 2026 analysis of 863,000 keyword SERPs and 4 million cited URLs (positions 11–100: 31.2%; outside the top 100: 31.0%) (Ahrefs, 2026). The comparable July 2025 figure was ~76% — partly a methodology change, since the earlier study counted only the top 3 visible citations (Ahrefs, 2026). The same dataset found YouTube to be the single most-cited domain in AI Overviews, at 5.6% of all citations.
Freshness shows a moderate, platform-dependent pattern. Across 16.975 million cited URLs from 7 AI platforms, AI assistants cited content averaging 1,064 days since publication versus 1,432 days for organic Google — 25.7% fresher — but AI Overviews showed no freshness preference (an identical 1,432 days), and even for chatbots the average cited page was still ~2.9 years old (Ahrefs, 2025). Everything above: correlational. What changes strategically when rankings and citations decouple is the subject of our comparison of GEO vs SEO.
Do brand mentions predict AI visibility better than backlinks?
Yes — in every correlation study published so far, brand mentions across the web beat backlink metrics as predictors of AI visibility, by a wide margin. Ahrefs ran the largest analysis: 75,000 brands with Domain Rating above 40, Spearman-correlated against visibility in ChatGPT, Google AI Mode and AI Overviews. Brand mentions on YouTube were the strongest single factor at ~0.737; branded web mentions scored 0.656–0.709 and branded anchors 0.511–0.628; Domain Rating — the aggregate backlink-strength metric — managed only 0.266–0.326, and site size just ~0.194 (Ahrefs, 2025).
An earlier Ahrefs study of AI Overviews found the same shape: branded web mentions correlated with brand visibility at 0.664 while backlinks scored 0.218 — roughly three times weaker — with branded anchors at 0.527, branded search volume at 0.392 and organic traffic at 0.274 (Ahrefs, 2025).
Two caveats travel with these numbers. Ahrefs itself states that “correlation isn’t causation” (Ahrefs, 2025) — well-known brands may accumulate both mentions and AI visibility for the same underlying reason. And the sample covered only established domains (DR>40), so the findings say little about new sites. Status: correlated — consistent across two studies and multiple engines, but correlated.
Which sources do AI engines cite — and how fast does that change?
Each AI engine has its own citation profile, and profiles can flip within weeks — which makes any static “most cited domains” list perishable. Profound’s analysis of 680 million citations (August 2024–June 2025) found ChatGPT citing Wikipedia most often (7.8% of all its citations; 47.9% within its top-10 sources), Perplexity favoring Reddit (6.6%), and AI Overviews spreading citations most evenly, with Reddit leading at just 2.2% (Profound, 2025). Treat those numbers as proof that engines differ — not as a current ranking.
How quickly? Semrush tracked 230,000 prompts and 100+ million citations (July–October 2025) and watched ChatGPT’s Reddit citation rate collapse from roughly 60% of answers to roughly 10% after a September 2025 change, while Wikipedia fell from ~55% to ~20% — within weeks (Semrush, 2025). The methodological lesson outweighs either number: AI visibility must be monitored per engine and over time.
One pattern does hold across engines and editions: earned media dominates. Muck Rack’s “What is AI Reading?” analysis of 25+ million links cited by ChatGPT, Claude and Gemini across 17 industries found 84% of citations pointing to earned media — coverage the brand did not own or buy — with a range of 82–89% across three editions since July 2025; journalism alone accounted for 27%, paid content for just 0.3% (Muck Rack, 2026). Muck Rack is a PR company with an interest in that conclusion, but the methodology is public and the direction matches rival studies (5W: 85.5%; Omniscient Digital: earned 48% vs owned 23%). Status: correlated, unusually consistent between competing vendors. The mechanics behind source selection — grounding, indexes, crawlers — are covered in how AI chooses its sources.
Is a citation the same as a brand mention?
No — a citation and a brand mention are two different phenomena, and most citations happen without the brand appearing in the answer at all. The Ghost Citations study by Semrush and Kevin Indig (3,981 domain occurrences, 115 prompts, 14 countries, 4 platforms) found that 61.7% of AI citations are “ghost citations”: the answer links to the page as a source, but the brand name never appears in the answer text (Semrush & Indig, 2026).
Gemini named brands in 83.7% of occurrences but cited their URLs in only 21.4%; ChatGPT was the mirror image, citing in 87% but naming the brand in just 20.7% (Semrush & Indig, 2026). The sample — 115 prompts — is moderate, so read these as orders of magnitude. The operational consequence: measuring AI visibility requires two separate KPIs — brand mentions and URL citations — tracked per engine. A brand can be heavily cited and invisible, or widely named and never linked.
This is exactly the kind of measurement an AI visibility audit exists for: running a fixed prompt set across engines, logging mentions and citations separately, and repeating it monthly — because single snapshots mislead within weeks.
Does content format and structure affect AI citations?
Correlational data says yes: citations skew heavily toward the beginning of a page and toward formats that match the query’s intent — but no experiment has proven either effect causal. Kevin Indig verified 18,012 citations from 1.2 million AI answers against the exact passage cited: 44.2% of verified ChatGPT citations came from the first 30% of the page’s content, 31.1% from the middle third, and 24.7% from the final third (Indig, 2026). That distribution supports answer-first writing — while remaining a correlation, since pages that front-load answers may simply be better pages.
Format matters too. Wix Studio’s AI Search Lab classified 1,056,727 citations from 75,000 answers across ChatGPT, Google AI Mode and Perplexity: listicles led at 21.9%, ahead of articles (16.7%) and product pages (13.7%) — and for commercial-intent queries, listicles alone captured 40.86% of citations (Wix Studio, 2026). Wix’s own interpretation: engines cite the format that structurally answers the question type — intent matching, not listicles as such. Status for both findings: correlated.
Do schema, llms.txt or special files move AI citations?
The technical layer is where AI search research has delivered its clearest negative results: structured data does not cause citations, llms.txt is unused by every major engine, and Google officially requires no special files.
On structured data (schema.org) — machine-readable content markup — Ahrefs ran the field’s closest thing to a controlled experiment: 1,885 pages that added JSON-LD between August 2025 and March 2026, against ~4,000 control pages. The result: no citation gains — AI Overviews visibility dipped 4.6% (a small but statistically significant decline), while AI Mode (+2.4%) and ChatGPT (+2.2%) changed insignificantly, with four statistical tests agreeing (Ahrefs, 2026). The same team’s correlational cut of 6 million URLs found AI-cited pages ~3x more likely to carry JSON-LD — interpreted as co-occurrence, not cause: well-run sites have both (Ahrefs, 2026). Status: correlation confirmed, causation refuted for already-visible pages.
On llms.txt — a proposed standard file mapping a site’s key content for AI systems — the evidence is starker. Across 137,000 domains analyzed by Ahrefs (updated 15 June 2026), 28% publish a valid llms.txt, yet 97% of those files received zero requests in May 2026, and the study states directly: “no major LLM provider currently supports llms.txt. Not OpenAI. Not Anthropic. Not Google.” (Ahrefs, 2026). Google’s John Mueller compared llms.txt to the meta keywords tag in April 2025 (Ahrefs, 2026). Status: a bet — near-zero cost, zero demonstrated payoff.
Google’s official position closes the loop: “You don’t need to create new machine readable files, AI text files, or markup to appear in these features” — AI Overviews and AI Mode run on the same ranking systems as classic Search (Google Search Central, 2025; current as of July 2026). Status: proven, in the narrow sense of official platform documentation.
Study-by-study: methodology, conclusion, evidence status
The table below compresses every study in this overview into one row: what was measured, what it found, and how much weight it can bear.
| Study (source, year) | Methodology and scale | Key finding | Evidence status |
|---|---|---|---|
| GEO paper (Aggarwal et al., KDD 2024) | Experiment on a simulated engine; 10,000 queries, 9 methods, 2 custom metrics | Quotations, statistics, cited sources: +30–40% PAWC; position-5 sources +115.1% | Proven on the authors’ benchmark; not replicated independently |
| C-SEO Bench (Puerto et al., NeurIPS 2025) | Independent replication: 1.9K+ queries, 16.3K documents, 6 domains, 4 models | 3/54 cases with significant positive effect; Statistics hurt in 19/24 settings | Proven (peer-reviewed); key counterpoint to the GEO paper |
| Keyword stuffing (both papers, 2024–2025) | Convergent result of two independent experiments | ~−9% on GEO-bench, ~−10% on production Perplexity.ai | Proven (the only replicated tactic finding) |
| Ahrefs freshness (2025) | 16.975M cited URLs, 7 AI platforms | AI cites content 25.7% fresher than organic Google (1,064 vs 1,432 days); no freshness effect in AI Overviews | Correlated |
| Ahrefs AI search overlap (2025) | 15,000 long-tail queries; ChatGPT, Gemini, Copilot, Perplexity | Only 12% of chatbot citations rank in Google’s top 10; ~80% rank nowhere for the query | Correlated / observational |
| Ahrefs AIO citations (2026) | 863K keyword SERPs, 4M AI Overviews URLs | 37.9% of AIO citations from organic top 10 (vs ~76% in July 2025); YouTube = 5.6% of citations | Correlated / observational |
| Ahrefs brand correlations (2025) | 75,000 brands (DR>40), Spearman; ChatGPT, AI Mode, AI Overviews | YouTube mentions ~0.737, web mentions 0.656–0.709, Domain Rating only 0.266–0.326 | Correlated (sample limited to DR>40) |
| Ahrefs AIO brand study (2025) | 75,000 brands, AI Overviews visibility | Brand mentions 0.664 vs backlinks 0.218 (~3x weaker) | Correlated |
| Profound citation patterns (2025) | 680M citations, Aug 2024–Jun 2025; 3 engines | Engines differ: ChatGPT→Wikipedia (7.8%), Perplexity→Reddit (6.6%), AIO most even (Reddit 2.2%) | Correlated; domain rankings now dated |
| Semrush most-cited domains (2025) | 230K prompts, 100M+ citations, Jul–Oct 2025 | ChatGPT’s Reddit citations fell ~60%→~10% of answers within weeks | Correlated; proves instability over time |
| Muck Rack “What is AI Reading?” (2026) | 25M+ links cited by ChatGPT, Claude, Gemini; 17 industries; 3 editions | 84% of citations from earned media (range 82–89%); journalism 27%; paid content 0.3% | Correlated (vendor interest, but consistent across rival studies) |
| Ghost Citations (Semrush & Indig, 2026) | 3,981 domain occurrences, 115 prompts, 14 countries, 4 platforms | 61.7% of citations omit the brand name; Gemini names but rarely links, ChatGPT links but rarely names | Correlated; moderate sample |
| Indig citation position (2026) | 1.2M answers, 18,012 verified ChatGPT citations | 44.2% of citations come from the first 30% of page content | Correlated |
| Wix AI Search Lab formats (2026) | 1.06M citations from 75,000 answers; 3 engines | Listicles most cited (21.9%; 40.86% for commercial intent); format–intent match is the strongest predictor | Correlated |
| Ahrefs schema experiment (2026) | Quasi-experiment: 1,885 pages adding JSON-LD vs ~4,000 controls; plus 6M-URL correlation | No citation gains from adding schema (AIO −4.6%, others n.s.); cited pages ~3x more likely to have JSON-LD anyway | Correlation confirmed, causation refuted |
| Ahrefs llms.txt (2026) | 137,000 domains, crawler log analysis | 97% of llms.txt files got zero requests in May 2026; no major LLM provider supports the standard | Bet (no engine-side adoption) |
| Google Search Central (2025) | Official platform documentation | No special files or markup needed for AI Overviews / AI Mode; same ranking systems as Search | Proven (official documentation) |
What should you actually do with this evidence?
The research supports a strategy built on measurement and earned authority, not content tricks. The actionable core:
- Stop what is disproven: keyword stuffing, schema deployed purely for AI citations, expectations that llms.txt gets read.
- Invest where correlations converge: brand mentions in independent sources — earned media supplies 84% of AI citations (Muck Rack, 2026), mentions outcorrelate backlinks roughly 3x (Ahrefs, 2025), and YouTube is the strongest correlate of all.
- Write answer-first: the citation skew toward the top of pages is correlational, but the practice costs nothing.
- Measure per engine, monthly, with mentions and citations as separate KPIs.
- Keep classic SEO: document position beat content rewrites in C-SEO Bench, and Google’s AI features run on the same ranking systems as Search.
FAQ
Is GEO scientifically proven to work?
Partially, and the evidence is contested. The GEO paper (KDD 2024) demonstrated up to 40% visibility gains on its own benchmark with a simulated engine (Aggarwal et al., 2024); the independent C-SEO Bench replication (NeurIPS 2025) confirmed a significant positive effect in only 3 of 54 cases (Puerto et al., 2025). The only replicated tactic finding is negative: keyword stuffing hurts.
What is the difference between proven, correlated, and a bet?
Proven means an experimental result, ideally confirmed by an independent team — like keyword stuffing’s negative effect. Correlated means an observational study found a pattern without demonstrating cause — like brand mentions correlating with AI visibility at 0.656–0.709 (Ahrefs, 2025). A bet has no supporting evidence, like llms.txt. Most AI search advice sits in the second and third tiers.
Do backlinks still matter for visibility in AI answers?
Less than brand mentions, according to every correlation study so far. Ahrefs measured backlinks at 0.218 Spearman correlation with AI Overviews brand visibility versus 0.664 for branded web mentions — roughly three times weaker (Ahrefs, 2025). Backlinks still support classic rankings, which feed AI Overviews and grounding indexes — indirectly relevant, not obsolete.
Does llms.txt help you get cited by AI engines?
No evidence says so. Ahrefs analyzed 137,000 domains: 28% publish a valid llms.txt, but 97% of those files received zero requests in May 2026, and no major LLM provider — not OpenAI, Anthropic, or Google — declares support for the standard (Ahrefs, 2026). Publishing one costs almost nothing; treat it as a lottery ticket, not a tactic.