What kind of content does AI actually cite?

AI search

Why discipline beats novelty in brand building

Filed under

AI search

Published

AI engines assemble their answers from a short list of sources they trust. The research on how to get onto that list is clearer than most agencies let on.

Ask ChatGPT, Gemini or Perplexity the questions your buyers ask, and the answer comes assembled from a short list of sources the engine has decided to trust. Getting onto that list is not a dark art, and it is not a separate discipline that needs a new agency. Across the strongest research published through 2025 and 2026, the content AI engines cite most is crawlable, front-loaded, factually specific, self-contained, and backed by original data and a named author. Most of that is good SEO. The rest is a handful of additions most websites still haven't made.

Answer engine optimisation (AEO) is the work of making your content the source an AI engine cites when it answers a question. You'll also see it called generative engine optimisation (GEO). Different labels, same job. Here's what the evidence says that job involves.

How AI engines choose their sources

When an AI engine answers a live question, it runs a search behind the scenes, retrieves passages from pages it can read, and assembles an answer with sources attached. Two findings follow from that.

First, ranking still matters. Ahrefs found around 38% of Google AI Overview citations come from that query's top ten organic results (directional, 2025). If you rank, you're in the room.

Second, page quality is measurable and it pays. A 2025 academic study of 1,702 citations across three AI engines found a high-quality page was roughly four times more likely to be cited than a low-quality one.

The broadest synthesis so far is a meta-analysis by Cyrus Shepard at Zyppy (May 2026), which scored 54 published studies, experiments and patents by strength of evidence. The factors with the strongest support: the page is accessible to crawlers, it already ranks for the query and related queries, the answer sits near the top of the page, the content is structured with clear headings and sections, the writing is factually specific rather than hedged, passages make sense on their own, and the content is fresh. Schema markup showed a small but consistent positive effect. Domain authority was one of the weakest signals tested.

Put the answer first

Analysis of ChatGPT citations by Kevin Indig (reported early 2026) found around 44% came from the first 30% of a page. Pages that open with a plain definition - "X is" - were nearly twice as likely to be cited, and front-loaded pages were cited roughly twice as often overall (directional). The practical rule: give the answer in the first two paragraphs, then earn the depth below it. If a passage can't stand alone and still make sense, an engine can't lift it.

Original numbers beat adjectives

The most-cited academic work in this field is the Princeton-led GEO study (KDD 2024), which tested nine content interventions across 10,000 queries. Adding statistics lifted a page's visibility in AI answers by up to about 40%. Quoting named sources added roughly 28%. Citing sources added 30 to 40%. Keyword stuffing landed about 10% below doing nothing at all. Pages ranked lower in traditional search gained the most from these changes - one mid-ranked page more than doubled its visibility just by adding source citations. The study is from 2024 and tested two engines, so treat the numbers as directional, but the direction matches everything published since: engines reward specific, sourced, verifiable claims over confident adjectives.

Cited is not the same as named

Being used as a source and being named in the answer are different outcomes, and the gap is large. Semrush's ghost citations study found 62% of AI citations produce no brand mention at all. The engines split in opposite directions: Gemini named brands in around 84% of answers but linked them as sources only 21% of the time, while ChatGPT linked sources 87% of the time and named the brand in roughly 21% (directional).

The working interpretation across the research: citations follow topic authority and original content, while mentions follow brand trust and third-party presence. They're two jobs, not one. A related pattern worth knowing: comparison-style questions ("best", "versus", "recommend") produced 2.4 times more brand mentions than informational ones, because the engine has to name the players it's weighing up.

Every engine has its own reading list

A seven-month Conductor study (September 2025 to March 2026) tracked which sources each engine leans on and found each has a persistent editorial identity. ChatGPT is anchored to Wikipedia and rewards citation-grade prose. Perplexity, Google AI Overviews and Gemini lead with YouTube across most query types. Claude, in early data, skipped YouTube, Wikipedia and Reddit entirely and went straight to institutional and primary sources. The practical read: strong foundations get you considered by every engine, but if you know which engine your buyers actually use, you can weight your effort.

Your website is only half the job

A large share of AI answers is built from places you don't control. In a June 2025 Semrush analysis of 150,000 citations, Reddit appeared in around 40% of AI answers, Wikipedia in 26% and YouTube in 24% (directional - other large studies using different counting methods put the shares lower). Businesses with active profiles on two or more review platforms were 3.4 times more likely to be mentioned by ChatGPT (AISOS, 2026). In an Ahrefs study of 75,000 brands (December 2025), YouTube mentions showed the strongest single correlation with AI visibility (about 0.74), and branded mentions across the web correlated more strongly than backlinks did. LinkedIn now appears in roughly 14% of ChatGPT Search responses.

One warning before you rebuild your strategy around any single platform: these patterns move. In September 2025, ChatGPT's Reddit citation share fell from roughly 60% of responses to 10% within a fortnight before recovering. Spread the presence.

What doesn't work

Keyword stuffing actively hurts - it tested below baseline in the Princeton study. Publishing volume without a distinct point of view is expensive and ineffective, because the engines already have generic answers. And llms.txt, the much-promoted "robots file for AI", has no evidence behind it: as of mid-2026 no major AI provider has committed to using it for citations, Google has said publicly that it doesn't support it, and one 90-day study that monitored more than 500 million AI bot visits counted just 408 requests for the file. It costs five minutes, so it's a harmless experiment. It is not a strategy.

The technical prerequisite most sites fail

None of the major AI crawlers execute JavaScript. An analysis of more than 500 million GPTBot requests found no evidence of script execution, and Claude's and Perplexity's crawlers behave the same way - Google's Gemini is the one exception. If your content only appears after scripts run in a browser, most AI engines see an empty page, no matter how well the content is written. The fix is structural: your platform has to serve real content in the raw HTML. We compare how the major website platforms handle this in our guide to the best CMS for AI search (/journal/best-cms-for-ai-search).

What to do with this

Front-load a plain answer on every page. Write sections that stand alone under descriptive headings. Publish your own numbers - client data, benchmarks, anything a generic competitor can't copy. Put a named author with real credentials on your content, because anonymous pages give an engine nothing to trust. Earn presence where the engines read: reviews, YouTube, LinkedIn, the communities your buyers use. And keep doing the SEO, because ranking remains the widest door into AI answers.

A note for Australian businesses

The citation research above is mostly US-based, so treat the specific numbers as directional here. The adoption curve, though, is local and steep. Telsyte's June 2026 study put 17.4 million Australians aged 16 and over on AI tools - 77% of that age group - with ChatGPT at 13.8 million users and daily AI use up 160% year on year. DataReportal's 2026 Australia report found 64% of Australians have already encountered AI Overviews in their search results. The shortlists are being written now, whether or not you're in them.

FAQ

Does schema markup help with AI citations? It shows a small but consistent positive effect across studies. Implement it, but don't expect it to carry a weak page.

Should we publish an llms.txt file? There's no evidence any major engine uses it for citations as of mid-2026. Do it if you like - it takes minutes - but put the real effort into content and structure.

Do we need separate content for Google and for AI? No. The strongest evidence says AI citation is mostly good SEO plus front-loaded answers, original data and named authorship. One foundation serves both.

We're cited but competitors get named. Why? Citations follow content and topic authority; mentions follow brand trust and third-party presence. If you're cited but not named, the content is working and the brand layer - reviews, coverage, community presence - is the gap.

Sources

  • Cyrus Shepard, Zyppy - AI citation ranking factors meta-analysis (May 2026): signal.zyppy.com/p/ai-citation-ranking-factors

  • Aggarwal et al. - GEO: Generative Engine Optimization, KDD 2024: arxiv.org/abs/2311.09735

  • Kumar et al. - citation analysis across three AI engines (September 2025): arxiv.org/abs/2509.10762

  • Semrush - The ghost citations study (2026): semrush.com/blog/the-ghost-citations-study

  • Conductor - How AI citations differ across engines (2026): conductor.com/academy/how-ai-citations-differ

  • BuzzStream - AI mentions vs citations study (2025): buzzstream.com/blog/ai-mentions-vs-citations-study

  • Profound - AI platform citation patterns (2025): tryprofound.com/blog/ai-platform-citation-patterns

  • Vercel and MERJ - The rise of the AI crawler: vercel.com/blog/the-rise-of-the-ai-crawler

  • OrganiKPI - llms.txt adoption and impact (2026): organikpi.com/blog/distribution/llms-txt-adoption-impact

  • Ahrefs - AI Overviews and brand visibility studies (2025); Kevin Indig, Growth Memo - ChatGPT citation position analysis (2026); AISOS - review platform study (2026). Figures hedged as directional.

  • Telsyte - Australian AI study (June 2026): telsyte.com.au; DataReportal - Digital 2026 Australia: datareportal.com

The metrics that actually predict growth

The metrics that actually predict growth

Get In Touch

Let’s build something people remember

Talk to strategists

No account managers - you speak directly with the people running the marketing.

hello@netradigital.com.au

(61) 468 939 110

Melbourne & Berlin

Get In Touch

Let’s build something people remember

Talk to strategists

No account managers - you speak directly with the people running the marketing.

hello@netradigital.com.au

(61) 468 939 110

Melbourne & Berlin

Get In Touch

Let’s build something people remember

Talk to strategists

No account managers - you speak directly with the people running the marketing.

hello@netradigital.com.au

(61) 468 939 110

Melbourne & Berlin