Thin content and the scaled content abuse policy
Google's spam policies name a specific offense: scaled content abuse, meaning the production of many pages primarily to manipulate search rankings rather than to help people. The policy is deliberately method-agnostic. It applies whether the pages are generated by AI, written by freelancers, spun from templates, or assembled by any combination. The question is never who or what wrote the content. The question is whether each page adds value.
What the policy actually names
The policy targets volume without substance: paraphrasing existing content while adding nothing, stitching together feeds and scraped snippets, and mass-producing pages that technically match queries but do not help the person asking. Google's AI optimization guide adds a warning that matters for anyone doing programmatic SEO: creating separate content for every query variation violates the scaled content abuse policy. One strong page that covers a topic beats fifty near-clones targeting fifty phrasings, not because fewer is better, but because the fifty near-clones are the pattern the policy describes.
The cluster fingerprint
How does detection work in practice? Google Research's S-CTS work is instructive: it targets templated narratives produced from shared infrastructure at scale, not AI authorship as such. The fingerprint has three parts:
- Templated narratives. The same article skeleton repeated with entities swapped in: same structure, same transitions, same conclusions, different keywords.
- Shared infrastructure. The same hosting, boilerplate, themes, and technical footprints repeated across a network of sites.
- Publish frequency. Bursts of near-identical pages appearing at a rate no editorial process would produce.
Each element alone is innocent. Plenty of legitimate sites use templates, share hosting, or publish often. It is the combination, sameness repeated across infrastructure at high frequency, that makes a cluster look like an operation rather than a publication.
AI authorship is not the trigger
A study across 1,000,000 pages and 100,000 search results pages found that top-10 positions consistently contain 8.4 to 11.7 percent pages that are at least 80 percent AI generated, with stable impressions over twelve months. AI-assisted pages rank, and they keep ranking. If AI authorship itself were the target, those numbers would not hold steady for a year. What gets removed is substance-free sameness at scale, whoever or whatever produced it.
What volume without substance risks
The risk accrues at the cluster and site level, not per article. A thousand interchangeable articles create a statistical fingerprint that one weak article never would, and the consequences arrive at the same level: spam enforcement, whether algorithmic or through manual action, can suppress a whole site, taking the good pages down with the templated ones. Google's helpful content guidance points the same direction, noting that unhelpful content can weigh on a site as a whole. To be precise about the framing: the risk does not come from publishing a lot. It comes from publishing a lot of the same thing.
What safe scaling looks like
Publishing at volume is compatible with the policy when each page carries its own weight:
- Varied templates. Let structure follow the subject. A comparison, a how-to, and a reference page should not share one skeleton with nouns swapped.
- Real sourcing. Cite primary sources, use real data, include first-hand perspective. Google's own guidance says unique, non-commodity content is what gets selected for AI answers.
- Per-article substance. Every page should answer a question a human actually has, and answer it better than a paraphrase of the current top ten would. Consolidate query variations into one strong page instead of splitting them across clones.
This is the standard Vupie holds its own autopilot output to: varied structures, cited sources, and a reason for each article to exist. Start from the beginning of this series with what SEO looks like in 2026 or how Google ranks content.