Skip to main content
All pages in this guide
แหล่งความรู้

Last updated

Programmatic SEO done legitimately

Programmatic SEO is generating many pages from one template plus a data source. A row in the data becomes a page, and the template decides how every page is laid out. It is how product catalogs, location pages, specification pages and directory listings have been built for years.

It is neither inherently spam nor inherently fine. Google's spam policies define scaled content abuse as generating many pages for the primary purpose of manipulating search rankings and not helping users, focused on "creating large amounts of unoriginal content that provides little to no value to users, no matter how it's created". The last clause settles it. The policy is about substance and intent, not production method, which is the same reason AI authorship on its own is not the trigger, as set out in does Google penalize AI content and thin content and the scaled content abuse policy. A template is a production method. It decides nothing by itself.

Where the line runs

Google draws it in one sentence in its AI features guide: creating separate content for every possible variation of how people might search, primarily to manipulate rankings or generative AI responses in Google Search, violates the scaled content abuse spam policy. The same passage adds that "a high quantity of pages doesn't make a website higher quality or more relevant to users", and that its systems can judge a page relevant without an exact match between the query and the page's primary content.

That converts into one question for a build: does each generated page carry information the other pages do not? If yes, the program publishes data that happens to be templated. If no, it publishes query variations, and the template is how it does so at scale.

The test to apply before building

  1. Does each page have unique data a reader would actually want? Not a unique title and a unique noun in the second paragraph. Numbers, availability, measurements, terms, photographs, hours, results: information that differs per row and that somebody had a reason to look up.
  2. Would a human find this page useful landing on it directly? Most visitors arrive from a search result without the context of the index page. If a page only makes sense as one of ten thousand, it does not make sense.
  3. Could two of these pages be merged without losing anything? If yes, they should be. This question is answerable by inspection: put two generated pages side by side and mark what actually differs. If the difference is one field, the two pages were one page.
  4. Does the data source have enough depth per row to fill a page honestly? Count populated fields on the thinnest row, not the richest. A source with three fields per row supports a table, not ten thousand articles.

What legitimate programmatic looks like

The working pattern is that the template supplies structure and the data supplies substance. The pages stand on information that is genuinely per row: real inventory with real availability and prices, measurements that were taken rather than estimated, locations the business actually operates in with their own addresses, staff and hours, specifications that differ between models, records that exist independently of the page.

Two properties follow. The page could not have been written without the data, and the data would be worth publishing even if search did not exist. The pages are many because the underlying reality has many items, not because a plan called for many URLs.

What illegitimate looks like

  • The swap. The same paragraphs with a city, an industry or an adjective substituted. Google's spam policies name a neighboring case under doorway abuse: pages targeted at specific regions or cities that funnel users to one page. If the service is identical everywhere and no page carries anything local, those pages are addresses for a single page.
  • Combinatorial coverage. Every product crossed with every use case crossed with every audience, generated because the grid had cells rather than because anyone asked the resulting question. Most cells have no answer behind them.
  • Pages with no data. A template filled by a model with plausible sentences, where the per-row facts were invented instead of retrieved. This is the worst version, because the pages are not only thin but wrong, and wrong is discoverable by any reader who checks one number.

Controls that make a program survivable

  • A minimum-data threshold. Decide, before building, which fields a row must have populated for its page to exist at all. Rows below the threshold generate nothing. This is the control that separates a data-driven program from a page-count program, because it puts a floor under the output that the template cannot supply.
  • Deduplication. Compare generated pages against each other, not only against the rest of the web. If two outputs are near-identical once the templated frame is removed, one of them should not ship.
  • Exclusion for thin rows. Rows that fall below the threshold later, or that were published before a threshold existed, should be noindexed or removed rather than left as filler. Google's spam policies give the same instruction for content of this kind: exclude it from Search.
  • Monitoring at the cluster level. A generated set behaves as a group. Track impressions, clicks and indexed count for the whole URL path in Search Console rather than reading individual URLs, because a program fails as a set and the per-page view hides it.

The risk is site-wide, which is why the threshold matters

Enforcement does not land only on the pages that earned it. Spam action, algorithmic or manual, operates above the level of a single URL and can suppress a site rather than a page, taking the good pages down with the generated ones. A large thin program therefore does not merely fail to rank; it puts at risk the pages that were already working. The cluster-level fingerprint that makes such a program recognizable is described in thin content and the scaled content abuse policy.

The template is therefore the least important part of the decision. Templates are how a good data source becomes many good pages, and equally how a poor one becomes many poor ones. The threshold decides which happened. Coverage in the sense worth pursuing, as opposed to coverage measured in URLs, is the subject of topical authority.

Vupie's generation pipeline holds a floor of the same kind and refuses to publish an article when the sourcing behind it is too thin, which is this control applied to editorial content rather than to rows in a table.

ให้ทำงานแบบออโตไพลอต

Vupie ทำทุกอย่างในคู่มือนี้ให้คุณ แปดบทความพรีเมียมต่อเดือน

เข้าร่วมเบต้า