Skip to content

Do schema, llms.txt and Reddit get you cited by AI? What the evidence says

A test on already-cited pages found schema barely moved AI citations, Google says no AI file is needed, and company sites draw most citations overall.

By Published Updated 14 min read
Three translucent glass forms float in a dark blue space while thin beams of light pass through them toward one bright point.

The short answer

Not on their own. In an Ahrefs test on pages that were already heavily cited, adding schema did not measurably raise citations in ChatGPT or Google AI Mode. Google says its AI features need no special files or markup. And company-owned websites drew about 57% of citations across eight engines (47% on ChatGPT). Treat schema and llms.txt as hygiene.

Key takeaways

  • In an Ahrefs test of 1,885 already heavily cited pages (August 2025–March 2026), adding schema did not measurably change citations in ChatGPT or Google AI Mode.
  • Google says its AI Overviews and AI Mode need no special files, AI text files or schema markup to appear.
  • Across 137,210 sites using Ahrefs Web Analytics, 97% of published llms.txt files received no requests in May 2026; this measures requests to the file, not citations.
  • Company-owned websites drew about 57% of citations across eight AI engines from April to July 2026, but 47% on ChatGPT, so outside sources still matter there.
  • Our view: keep schema and llms.txt as cheap hygiene, and put sprint effort into quotable pages and mentions beyond your own site.

Three tactics, one test: what engines cite

Schema markup, an llms.txt file and a presence on Reddit are the three tactics most often sold as ways to get cited by ChatGPT, Gemini and Google's AI features. The evidence for each is thinner than the pitch. Schema did not move citations on pages that were already cited, published llms.txt files are rarely even requested, and company sites, not forums, draw most citations overall.

The test we apply is narrow on purpose. A tactic earns sprint time only if there is evidence it changes what engines cite, or if it is so cheap that the question hardly matters. Adoption rates, crawler hits and ranking signals are interesting. They are not citations.

What is claimed, and what the evidence shows

What is claimed

  • Schema markup gets your pages cited by AI
  • An llms.txt file tells AI engines what to cite
  • Reddit is what AI engines cite most

What the evidence shows

  • On already-cited pages, adding schema made no measurable difference on ChatGPT or AI Mode1
  • Most published llms.txt files received no requests at all in May 20263
  • Company-owned sites drew more than half of citations across eight engines, April to July 20264
Sources: Ahrefs1, Ahrefs3, Profound4. Summary by the Sigzen AI editorial team.

Prior work, and our own interest

Others have covered each myth on its own. Ahrefs wrote up its schema test in detail1, and Search Engine Journal reported Ahrefs' llms.txt data2. This piece sets the three tactics side by side against the primary studies and Google's own documentation, then says what we still ship and why.

A disclosure first. Almost every study below comes from a company that sells SEO or AI-visibility tools or services. So does Sigzen AI: we sell AI visibility audits and sprints, and the sprint includes schema and llms.txt work. Weigh our view with that in mind.

Does schema markup get you cited?

Not measurably, on pages that engines already cite. That is what Ahrefs' controlled test found, and it is the most direct evidence on the question that we have read. Schema still has jobs to do. Getting an already visible page cited more often is not one of them, on this evidence.

What the controlled test found

Ahrefs, which sells SEO and AI-visibility tools, tracked 1,885 pages that were already heavily cited and then added JSON-LD schema between August 2025 and March 2026, against 4,000 matched controls1. Citations moved +2.2% on ChatGPT and +2.4% on Google AI Mode, and neither change was statistically significant. On Google AI Overviews they fell 4.6%, a small decline that was significant1.

Read the sample carefully. These pages were already cited, so the test says little about pages AI does not see at all. Ahrefs says so itself: "for pages that aren't being seen by AI systems at all, schema markup might still play a role."1 Our view: that caveat is fair, but it is a hypothesis, not a result, and it should not be sold as one.

Change in AI citations after adding schema

ChatGPT, not significant1+2.2%
Google AI Mode, not significant1+2.4%
Google AI Overviews, small but significant1−4.6%
Ahrefs, 1,885 already heavily cited pages vs 4,000 controls, Aug 2025–Mar 20261. Bar length shows the size of the change, not its direction; the sign is in the label.

Why most cited pages carry schema anyway

In the same Ahrefs study (August 2025–March 2026), 53% of AI-cited pages carried schema, and Ahrefs reads that as correlation, not cause5. Both findings hold at once. Sites that bother with structured data tend to be the sites that also publish clear pages, keep facts current and earn links. Those habits get pages cited, and the markup comes along with them.

This is why an audit that flags "missing schema" as an AI-visibility problem can mislead. It sees the correlation and presents it as a lever. Our view: when a report shows you how many cited pages carry schema, ask for the controlled result next to it. Without one, the figure describes the pages that win. It does not explain why they win.

Where structured data still earns its place

Structured data still does useful work. It is simply different work from earning citations, and three jobs justify it.

  • Product feeds and merchant data. Shopping answers draw on structured price, availability and variant data. A feed is a separate mechanism from page schema, and for retailers it often matters more. Our e-commerce and D2C page starts with the feed and the Product and FAQ schema for that reason, and our playbook for e-commerce brands sets out the feed checks.
  • Rich results in classic search, which still send clicks to the pages that earn them.
  • Entity consistency: the same name, description, founders and contact details on your site, your profiles and directories, so engines do not confuse you with someone else. Google asks that structured data match the visible text on the page6. If an engine already has your facts wrong, our guide to fixing wrong AI answers starts at the source page.

None of these jobs needs a promise about AI citations to justify it.

Does an llms.txt file help?

There is no evidence that it gets you cited, and Google says its AI features do not need one. The file is cheap, so keeping one is harmless. Paying for one as a citation tactic is not a good use of money.

What Google says

Google's documentation for site owners is unusually direct here. Its page "AI features and your website", last updated in December 2025, says: "You don't need to create new machine readable files, AI text files, or markup to appear in these features."6 It adds: "There's also no special schema.org structured data that you need to add."6

That statement covers AI Overviews and AI Mode only. It says nothing about ChatGPT or Perplexity. What Google asks for instead is ordinary: pages its crawler can reach, important content available as text, and structured data that matches what readers see. The llms.txt proposal is a convention some sites follow. No engine has told site owners that it cites pages because of it.

What the studies found

Of 137,210 sites in Ahrefs Web Analytics with traffic in May 2026, 28% published an llms.txt file, and 97% of those files received no requests that month3. Ahrefs adds that its customers skew technical, so the adoption figure is an upper bound3.

Note what this measures: requests to the file, not citations. A file nobody fetches cannot shape an answer. A fetched file would not prove influence either.

The citation question has been tested too. SE Ranking, in a study of nearly 300,000 domains published in November 2025 (data period not stated), found no link between having the file and being cited7. Our view: keep the file if you have one, because it takes minutes to update. Do not budget for it, and do not let anyone sell it to you as a citation lever.

Is Reddit what AI cites most?

Not overall. Across engines, company websites draw most citations. Reddit's strong showing was specific to ChatGPT at a particular time, and that has shifted too. The larger lesson is that source mixes differ by engine and move from month to month.

A dark sea at night with a scattered archipelago of small islands, each holding a softly glowing lantern whose light reflects across the water toward a larger, brighter island at the centre.
Most of the light engines quote comes from company pages, while outside sources fill the rest in proportions that differ by engine.

Company sites draw most citations overall

Profound, which sells an AI-visibility platform, counted 11.84 billion citations on eight engines in all 8,061 categories it tracks, from 16 April to 16 July 2026. About 57% pointed to company-owned websites: 47% on ChatGPT and 69% on Gemini4.

"Company-owned" means any company's site: corporate pages, product pages, brand blogs and documentation. It does not mean your site. A competitor's comparison page counts. So does a software vendor's help centre that happens to answer your buyer's question.

Our view: this is the strongest single argument against Reddit-first playbooks. If more than half of what engines cite is first-party company content, your own pages are the first place to compete, and the competitor pages being cited instead are the ones to study.

Share of citations to company-owned sites

47%ChatGPT4
69%Gemini4
About 57% across eight engines, all 8,061 categories, 16 Apr–16 Jul 2026 (Profound)4. "Company-owned" means any company's site, not only yours.

In late 2025, Reddit and Wikipedia led ChatGPT's citations

The Reddit story was true for one engine for a while. In Semrush's weekly snapshots from 14 July to 12 October 2025, Reddit and Wikipedia were ChatGPT's two most-cited domains8. That data is now more than a year old, and it described ChatGPT, not AI search as a whole.

Many "Reddit for GEO" playbooks were written in that window. Then the picture moved. Otterly, which sells AI-search monitoring, reported in August 2026 that ChatGPT sharply cut its Reddit citations and shifted toward reference and official pages9. We have not verified Otterly's figures, and it is a single study. It still makes the point: what one engine cited last year is a weak guide to what it cites now.

Engines differ, and they change

Engines also cite at very different volumes. In Semrush's 2026 AI Visibility Index, covering 126 million US prompts from January to April 2026, ChatGPT cited an average of 15 sources per response and Gemini an average of 310. That is US data, and we would not assume it holds elsewhere.

An engine that cites many sources per answer leaves room for forums, reviews and trade press next to company pages. One that cites a handful tends to pick a few authoritative pages and stop. So "should we be on Reddit?" has no single answer. It depends on which engines your buyers use and what those engines cite for your prompts.

Our view: decide per engine, from your own prompt set, not from an average across engines.

What the evidence supports instead

Three things have better support: being cited at all, pages written so an engine can lift a passage, and mentions of your brand on sites you do not own. None of them is a shortcut, and each takes longer than adding markup.

Being cited is what pays

Seer Interactive, a search agency, tracked a panel of 53 brands from January 2025 to February 2026. Pages cited in a Google AI Overview earned 120% more organic clicks per impression than when they were not cited11. The panel is Seer's own, so read it as a strong signal rather than a market average.

That is a Google result, and it measures clicks, not revenue. It still answers the question behind every tactic in this piece. Citation is the outcome worth paying for, and work that does not move it is overhead. If you do not yet know whether engines cite you at all, a free AI Visibility Score shows where you stand before you change anything.

Write pages an engine can quote

Pages that state a clear answer, cite their sources and carry specific facts give an engine something to lift. In a 2024 lab study, Aggarwal and colleagues found that adding citations, quotations and statistics to source pages were the best-performing methods, with relative gains of 30–40% on one of their visibility metrics12. Keyword stuffing did little.

That benchmark ran on engines of its time, so treat the size of the gain as dated and the direction as useful. In practice it means an answer in the first lines of each section, a named source for every figure, and sections short enough to quote whole. This post follows the same pattern for the same reason.

Earn mentions beyond your own site

Your own pages are only part of the job. In May 2025, Ahrefs found that branded web mentions across 75,000 brands correlated more strongly with AI Overview visibility (0.664) than backlinks did (0.218)13. That is a correlation from a dated study, not proof that mentions cause citations.

It fits the rest of the evidence, though. If engines cite company pages for much of each answer, the remainder comes from reviews, trade press, directories and communities. Being described accurately in those places is slower work than markup. It is also where a sprint's effort pays back. Start with the domains engines already cite for your prompts, because those are the pages most likely to be read.

What we still ship on our own site, and why

We publish an llms.txt file and schema on sigzenai.com, and we will keep doing so. Both take minutes to maintain. The schema keeps our organisation and service details consistent for any system that reads them, and the llms.txt file is a plain map of our pages for any tool that asks for one. We do not expect either to get us cited.

Our AEO/GEO sprint (six weeks, from $8,000 (₹3.5L), our published price) lists the same work as "AI-crawler access, schema and llms.txt hygiene". It is there because blocked crawlers and broken markup are worth fixing, not because we promise citations from them. The sprint's effort goes into answer-first rewrites of your top pages, entity consistency and the first ten third-party citations.

What we sell in place of a promise is measurement: a prompt set run before and after, so each tactic is kept or dropped on evidence. Our playbooks describe the checks we run. To see what other providers put in a sprint, and at what price, read our dated list of published GEO prices.

Where to spend your next sprint

Spend it on pages and mentions, in that order, and keep the hygiene items to an afternoon.

  1. Fix access first. Check that robots rules let search and AI crawlers in, and that key copy is in the HTML your server sends, not added later by scripts.
  2. Rewrite the pages your buyers' prompts should land on, so each section opens with its answer and names its sources.
  3. Find which outside domains engines cite for your prompts, then earn accurate mentions there: reviews, directories, trade press, and communities where your buyers actually talk.
  4. Keep schema accurate and matched to visible text, and keep llms.txt current if you have one.

How the tactics compare

Plotting the common tactics by effort and by strength of evidence makes the trade-off plain. The cheap ones are worth doing for reasons other than citations. The ones with better support cost real time.

Effort vs strength of evidence

Effort →

Costly, weakly supportedReddit seeding campaigns; mass-produced AI pages
Where sprint effort goesAnswer-first rewrites of key pages; earning mentions on sites engines cite
Cheap, no citation evidencellms.txt as a citation lever; adding schema to pages already cited
Cheap, do it anywayCrawler access and server-rendered copy; checking which domains your prompts cite

Strength of evidence →

Illustrative: our assessment of eight tactics, not a measured result.

Our view: if a proposal leads with schema or llms.txt as its citation strategy, ask what evidence it rests on, and for which engines.

How do you test a tactic on your own prompts?

Run a fixed prompt set before and after the change, several times each, on every engine that matters to you. Then compare the share of answers that name or cite you, and change one thing at a time.

AI answers vary from run to run, so a single before-and-after check proves little. You need repeated runs and a margin of error on each number. Our measurement spec, Can you trust your AI visibility score?, sets out what to ask any vendor for, including us.

If the change is schema or llms.txt, expect the honest result to be no detectable difference. That is still worth knowing, because it tells you where not to spend the next sprint.

Sources

  1. Ahrefs, schema and AI citations test: 1,885 already heavily cited pages, Aug 2025–Mar 2026 (May 2026)
  2. Search Engine Journal, 97% of llms.txt files got no requests, Ahrefs data shows (Jun 2026)
  3. Ahrefs, llms.txt study: 137,210 sites, May 2026 data (Jun 2026)
  4. Profound, where AI citations come from: 11.84B citations, 8 engines, 16 Apr–16 Jul 2026 (Jul 2026)
  5. Ahrefs, schema and AI citations: correlation across cited pages (May 2026)
  6. Google Search Central, AI features and your website (updated Dec 2025)
  7. SE Ranking, llms.txt study: nearly 300,000 domains (Nov 2025); data period not stated
  8. Semrush, most-cited domains in AI search: weekly snapshots 14 Jul–12 Oct 2025 (Nov 2025)
  9. Otterly, ChatGPT and Reddit citations (Aug 2026)
  10. Semrush, 2026 AI Visibility Index: 126M US prompts, Jan–Apr 2026 (Jun 2026)
  11. Seer Interactive, AI Overviews and CTR, 2026 update: 53 brands, Jan 2025–Feb 2026 (Apr 2026)
  12. Aggarwal et al., GEO: Generative Engine Optimization, KDD 2024(dated)
  13. Ahrefs, brand mentions and AI Overview visibility, correlational (May 2025)(dated)

Questions readers ask

  • No. It costs little to keep, and some tools do read it. Just do not expect citations from it. Ahrefs found that almost all published llms.txt files received no requests in May 2026, and Google says its AI features need no such file. Keep it accurate if you have one, and put the time you would spend expanding it into the pages it points to.

Keep reading

Free, in two minutes. Enter your domain and we'll score it against three competitors across six engines.

No account needed. The report is emailed within 24 hours.