One web page in ten now shows AI authorship. Does AI-written content get cited, or penalised?
Pew finds AI authorship on one sampled page in ten. Google penalises scaled, low-value pages, not the tool. What gets reviewed pages cited, and what is risky.

The short answer
AI-written content is neither cited nor penalised by default. Google judges helpfulness, not how a page was made, and its spam policies target pages generated at scale without adding value. AI engines cite pages with specific, checkable facts. A page drafted with AI, then sourced, reviewed and owned by a person, is safe. Scaled, unreviewed output is not.
Key takeaways
- In a random sample of 10,000 English web pages from Pew Research Center's July 2026 snapshot, 10% showed significant signs of AI authorship.
- Google's spam policy on scaled content abuse names generative AI used to make many pages without adding value, not AI use as such.
- Google says there are no special requirements for AI Overviews or AI Mode beyond its existing Search fundamentals.
- Pages get cited for what they contain: a direct answer, named sources, dates and facts an engine can check against other pages.
- Our view: the risk sits in volume without review, so measure it per page by who checked each fact before publishing.
In this article
How much of the web is now written with AI?
About one sampled page in ten, by Pew's count. Pew Research Center collected about 490,000 English-language pages from Common Crawl snapshots between January 2021 and July 2026. In a random sample of 10,000 pages from the July 2026 snapshot, 10% showed significant signs of AI authorship.1 Among pages in that snapshot published after ChatGPT's release in November 2022, the share was over one-third.1
The split by domain is wide. In Pew's 2026 samples, about 10% of .com pages showed those signs, against 4.6% of .org pages and about 1% each for .edu and .gov.1 Commercial sites are where the drafting has moved fastest.
Two cautions before anyone quotes this in a board deck. Pew used a machine-learning detector that looks for patterns in wording, so each label is a statistical estimate, not proof. And Pew itself warns that single habits, such as a particular punctuation mark, do not show that a person used AI.1
AI authorship on the web
Common on commercial pages, rare on .edu and .gov
Does Google penalise AI-written pages?
Not for being AI-written. Google's guidance on AI-generated content, published in February 2023, says it rewards helpful content however it was produced. The same guidance says that using automation mainly to manipulate rankings breaks its spam policies.2 The line is purpose and value, not the tool.
The spam policy that matters here is scaled content abuse. Google defines it as many pages generated mainly to manipulate rankings rather than help users. Its list of examples names generative AI tools used to produce many pages without adding value.3 Scraped feeds and stitched-together pages sit on the same list, so the policy is about the pattern, not the model.
Enforcement arrives in updates. Google released its August 2026 spam update on 18 August, applying globally and to all languages.4 Google did not say it targeted AI content. It said sites that break its policies may rank lower or drop out, and that recovery can take months once the problem is fixed.4
Our view: a site that published hundreds of near-identical AI pages this year should read that recovery note twice. The cost of a spam action is measured in months, which is longer than the drafting time it saved.
Do AI engines cite AI-written content?
They cite pages, and the evidence we can point to is about what pages contain, not who drafted them. For Google's own AI features the guidance is short. Google says there are no additional requirements to appear in AI Overviews or AI Mode, and that its existing Search fundamentals still apply.5 A page has to be indexed and eligible for a snippet; snippet controls and noindex limit how it can be shown.5
None of the policy pages we cite here says an AI engine down-ranks text because a model wrote it. What engines reward, in practice, is retrievable and checkable material: a direct answer near the top, named sources, dates, prices and specifications that agree with other pages. Generic AI drafts tend to lack exactly those things, which is why they are easy to ignore rather than penalised.
The tactics debate is covered in our review of schema, llms.txt and Reddit as citation levers, and the date question in what the freshness studies actually show. Neither found a shortcut that replaces substance.
Two ways to use AI in publishing
Scaled, unreviewed output
- Many pages from one template and a keyword list
- Facts nobody checked against a source
- No named owner, no update date
- Fits Google's description of scaled content abuse
Drafted with AI, owned by people
- One page per real question a buyer asks
- Every figure linked to a primary source
- A reviewer who checked facts before publishing
- Specifics an engine can match to other pages
What makes an AI-assisted page citable?
Four things, and none of them depends on whether a model wrote the first draft. Each one is something a reviewer can check in minutes.
- An answer first. The opening paragraph answers the page's question in plain words, so a retrieval system can lift it without the rest of the page.
- Primary sources. Every figure links to the study, filing or policy page it came from, with the data period in the sentence.
- Facts only you hold. Your prices, your process, your product specifications and your published terms. Generic drafts cannot invent these without inventing errors.
- A reviewer and a date. A named team or role that checked the page, and a last-checked date that changes only when the substance does.
Our view: the third item does most of the work. A model can draft a competent overview of any topic, which is why overviews are crowded. The pages that get cited carry something the engine cannot find elsewhere, such as a published price list or a specification table.
Disclosure is the honest default. This blog says on every post that it is drafted with AI from sources our editors select, then fact-checked against each source. We do it because readers deserve to know, not because any engine asked.
Where scaled AI content goes wrong
It goes wrong on volume, sameness and errors. Each is a signal a search system can see, and each is avoidable with a review step.
Volume without demand
The classic pattern is a keyword list fed to a template: one page per city, product variant or long-tail phrase, with nobody asking for most of them. That is close to Google's own description of scaled content abuse.3 Publish fewer pages, each tied to a question your sales or support team actually hears.
Sameness
Pages generated from one prompt tend to say the same thing in the same order. An engine choosing a source has no reason to prefer one copy over another, so it prefers neither. Distinct facts break the tie.
Confident errors
The most expensive failure is a wrong fact stated well: an old price, a discontinued product or an invented certification. AI engines then repeat it. Our guide to fixing what ChatGPT gets wrong about your company traces how one bad page spreads. A reviewer who checks each figure against its source stops that at the point of publishing.
A review rule you can run this week
Our view: treat AI as a drafting tool and keep publishing a human decision. Set one rule and audit against it monthly.
- No page goes live without a named reviewer who checked every figure, price and claim against a source or an internal record.
- No template publishes more than a handful of pages until the first ones show demand in Search Console or in AI answers.
- Every page answers one question someone asked, and says so in its first paragraph.
- Pages that fail the rule get merged, rewritten or removed, starting with the thinnest.
Then check the outcome where it shows. Search Console's Performance report includes traffic from Google's AI features in the overall web figures,5 and since 31 August its AI reports show impressions by AI surface, as covered in our note on the new reports. For other engines, run the questions your buyers ask and record which pages are cited. That is the measurement behind our AI visibility audit, which runs each prompt on three separate days across six engines.
Sources
- Pew Research Center, AI authorship of web pages: ~490,000 English Common Crawl pages, Jan 2021–Jul 2026; headline share from a 10,000-page July 2026 snapshot (Aug 2026)
- Google Search Central, Google Search's guidance about AI-generated content (Feb 2023)
- Google Search Central, spam policies for Google web search: scaled content abuse (updated Aug 2026)
- Search Engine Roundtable, Google August 2026 spam update (Aug 2026)
- Google Search Central, AI features and your website (updated Dec 2025)


