Marketing · 03 · Search, answers & analyticsLesson 2 of 4
Programmatic SEO with Claude Code
- Generate template-driven pages at quality, not spam
- Build the keyword -> page pipeline with review gates
Programmatic without the spam
This lesson produces one of two things: a 10-page pilot generated from your own entity data - validated, similarity-checked, staged behind a review gate - or the honest verdict that you don't have a programmatic play, which is just as valuable. Programmatic SEO means generating many pages from structured data: comparisons, integrations, glossaries, location pages, 'X for Y' plays. Done badly, it's the thin-content spam that search engines have spent a decade learning to bury - and AI made generating it free, which made the bar higher, not lower.
Done right, it's the same engine you've built all course - grounded generation, deterministic validation, human gates - run in a loop. The mature 2026 pattern: real entity data in, structured content per entity out, similarity-checked, validated, and monitored as a cohort. Each page must deserve to exist for the person who lands on it. That's the whole test.
Entity data first - or no play at all
The make-or-break input is the entity dataset: one record per page in JSON or SQLite. The working threshold is 6-8 genuinely different fields per entity - facts that differ in substance between entities, not just labels swapped in a template.
{
"workflow": "weekly client reporting",
"category": "agency operations",
"who_feels_it": "marketing agencies, 5-15 people",
"manual_cost": "3-4 hours per client per month",
"inputs": ["analytics export", "campaign notes", "last report"],
"failure_modes": ["numbers without narrative",
"copy-paste errors at 6pm Friday"],
"automation_shape": "scheduled skill renders an HTML report
from cached weekly data pulls",
"proof_point": "reporting day became a half-hour review"
}- Pick the play where you have real data advantage: things you know from actual work that a generic writer would have to guess.
- Define the schema: which 6-8 fields will make every page substantively different?
- Build the dataset from real sources - your own research, public APIs, documented facts. Have Claude help compile, but every field traces to something true.
- Run a completeness check: entities missing fields get completed or cut before generation. No padding a thin record with adjectives.
The generation loop
The defining rule: Claude generates structured content per entity - reasoning over that entity's actual fields - rather than substituting names into one master template. Each page gets its own angle, examples, and FAQ derived from its own data. Output is structured JSON per entity (sections as fields), which your site renders; content and presentation stay separate. pSEO skills circulate in the open catalogs too - install the library's version if one fits, then adapt it to your entity schema and the grounding rule before writing your own.
- Write the generation skill: input one entity record plus voice.md and the page spec, output structured JSON - intro, body sections, FAQ entries - grounded only in the record's fields.
- Add the grounding instruction with teeth: any claim not derivable from the entity record is a defect, same as the writer agent's TODO rule.
- Run the loop headless, one entity per run - a small script that calls
claude -p --model sonnet --output-format jsonwith each record - because 200 entities in one chat session blows the context window. Sonnet 5 is the right default for reader-facing bulk pages; validate 10 against Opus before deciding. The JSON output carriestotal_cost_usd: log it per entity, and multiply the pilot's average by the entity count before you run the full set. Alternatively, ask Claude for a dynamic workflow: it fans the entities out to parallel agents in the background (up to 1,000 per run) and you watch progress in/workflows. - Have Claude write you a small similarity script - TF-IDF or shingle overlap with cosine on token vectors; no embeddings API needed - then fingerprint every page and compute pairwise similarity. Flag any pair above 0.92 as near-duplicates: differentiate or cut.
Near-dup flags are diagnostic gold. If your comparison pages keep colliding, the entities aren't different enough in the fields the page actually uses - fix the dataset, not the prose.
Validation, schema, and the review gate
Between generation and launch sits the validation pass - deterministic checks every page must clear, plus a human sampling gate. The same lint-not-vibes philosophy from module 1, in a loop.
- Required sections present and non-trivial - no empty FAQ stubs.
- Headline under 60 characters; meta description under 160.
- Grounding check: claims in the page trace to fields in the entity record.
- Near-duplicate check passed (the 0.92 flag resolved).
- JSON-LD generated per page and validated via the Rich Results Test API.
- Run validation across the full set; pages fail loudly with reasons, and failures route back to generation or to the dataset.
- Human gate, sampled: read 10-15 pages across the quality spectrum, every page for sets under 50. Ask the only question that matters: would a person landing here be glad they did?
- Stage the survivors through your normal publishing path - for git-native sites that's one reviewable PR for the whole set, which is the ideal case.
Launch, monitor as a cohort, prune
Two hundred pages can't be watched one at a time. Monitor the play as a cohort in your search console, grouped by URL pattern - impressions, clicks, and average position for the set as one curve. Expect indexing to take 2-8 weeks; do not panic-edit in week one.
- Submit the sitemap; confirm the URL pattern is crawlable.
- Add a cohort section to your weekly report (next lesson): the pattern's impressions and clicks over time.
- At 90 days, prune: pages with zero demand - no impressions - get cut or consolidated. Pruning is part of the play, not an admission of failure.
- Feed winners back: top pages by clicks tell you which entities your audience cares about - often the seed of editorial content worth writing by hand.
Do this now
Sources and further reading
Want us to set it up with you, end to end?
Three one-on-one sessions. We train you on your real stack and build your first agents together, until you can run it yourself. You keep everything.