Marketing · 01 · The content engineLesson 2 of 4
Research that doesn't hallucinate
- Build research workflows with sources and receipts
- Add the fact-check pass that catches made-up claims
Why research is where AI content dies
Voice gets you content that sounds like you; it does not get you content that is true. This hour builds the part that makes it true: a research pipeline that can't ship a made-up number - a brief format where every claim carries a receipt, a researcher subagent that fills it, and a fact-check pass that audits it, run on your next real topic, not a sample one. Models state plausible falsehoods with total confidence: invented statistics, misattributed quotes, product claims that were never made. In marketing, one fabricated number in a published post costs more trust than a hundred good posts earn.
The rule this lesson installs: receipts or it doesn't ship. Every claim, number, and quote in anything you publish must trace back to a source file or URL saved in the brief.
The brief: where receipts live
A brief is one file per piece, living in content/briefs/. It is the contract between research and writing: the writer may only use what's in the brief, and the brief must carry its evidence. The non-negotiable part is the sources section, where quotes and numbers are saved verbatim - not paraphrased - so they can be string-matched later.
# Brief: Why we publish our pricing
## Angle
Contrarian take: hiding pricing costs B2B companies more
deals than it protects.
## Target query / audience
"should b2b agencies publish pricing" - agency founders.
## Claims to make (each must cite a source below)
- C1: Buyers shortlist before talking to sales. [S1]
- C2: Our own close rate rose after publishing pricing. [S2: internal]
## Sources (verbatim - never paraphrase here)
S1: https://example.com/b2b-buyer-report
QUOTE: "..." (copied exactly, with surrounding sentence)
PULLED: 2026-06-18
S2: internal - data/analytics/close-rates-2026Q1.csv, row 14
## Open questions for the writer
- Do we name the competitor example or anonymize it?- Every claim gets an ID (C1, C2...) and points at a source ID. Unsourced claims are opinions and must read as opinions.
- Quotes are copied exactly, with enough surrounding text to verify context. Paraphrases cannot be string-matched, so they are banned in the sources section.
- Internal numbers cite the file and location they came from, so the fact-checker can open the CSV and look.
- Each source notes the date it was pulled. Stale claims get re-verified, not trusted.
The researcher subagent
Research and writing should not share a context window. A model that just read ten web pages starts absorbing their style; a model drafting in your voice starts treating its own fluent sentences as facts. Claude Code subagents fix this: each one runs in its own isolated context with its own restricted tools.
---
name: researcher
description: Gathers sources for a content brief. Searches corpus/
first (calls, stories, posted), then the web. Fetches pages and saves
claims with verbatim quotes and URLs into the brief's Sources
section. Never writes prose. Never invents a number - if a claim
can't be sourced, it goes under Open questions.
tools: WebSearch, WebFetch, Read, Grep, Glob, Write
model: opus
---Two choices in that frontmatter are deliberate. Internal first: the researcher searches your own corpus before the web, because a quote from your own call or a number from your own data is the source nobody else can cite. And the model pin: research decides what the post is allowed to say, so it runs on the strongest everyday model, Opus 5.5 (opus). For a research question that takes dozens of fetches across many sources, step up to Fable 5.1 (fable), Anthropic's model for long multi-step autonomous work, at 2.5x the Opus price per token as of September 2026.
- Create .claude/agents/researcher.md with web tools and file access only - no Bash, no publishing tools.
- Give it the brief template as its output contract: it fills in Claims and Sources, verbatim quotes required.
- Hard-code its prime directive in the agent body: a claim without a source goes in Open questions, never in Claims.
- Run it on a real topic: 'Research the brief at content/briefs/<file>. Fill in claims and sources.'
- Spot-check three sources yourself: open the URL, find the quote. If any quote isn't on the page, tighten the agent instructions and rerun.
The fact-check pass
The researcher saves receipts; the fact-checker audits them. It is a separate subagent that reads a finished draft, extracts every factual claim, and verifies each against the brief's saved sources. It has no write access to the draft - it produces a verification report, pass or fail, claim by claim.
- Create .claude/agents/fact-checker.md: read-only tools, separate context, and a single job - trace every claim, number, and quote in the draft to a source in the brief.
- Its output format: a table of claim, source ID, verdict (VERIFIED / UNSUPPORTED / CONTRADICTED), and evidence. Anything UNSUPPORTED blocks the draft from moving to approved/.
- For quotes and numbers, add a deterministic check: a small script that string-matches every quoted passage and digit sequence in the draft against the saved source text. Scripts don't get tired and don't get charmed by fluent prose.
- Wire both into your flow: no draft moves to content/approved/ without a clean fact-check report next to it.
Level up: briefs from live keyword data
Once the receipts discipline is in place, upgrade the input. Instead of starting a brief from a hunch, start it from search data you already have: what people actually query, and what currently ranks. No new tools required - your search console and whatever keyword tool you already use both export CSVs, and the pages that currently rank are public pages Claude can read. The data lands in the repo as files; Claude does the rest.
- Export the raw material: a queries report from your search console, or a keyword list from your keyword tool, saved as CSV into data/keywords/.
- Seed with one target query. Have Claude read the export and shortlist what's worth a page - use volume where you have it, but the real filter is: would a buyer actually ask this?
- Capture what currently ranks: open the top 5 ranking pages for the target query and save each one's content as a file in data/keywords/ for gap analysis.
- Have Claude produce the brief: target queries, entities to cover, FAQ candidates, and a requirement to answer the core query in the first 200 words (that's a GEO rule - more in module 3).
- Save it as a normal brief in content/briefs/ and encode the procedure as the seo-brief skill - SEO-brief skills exist in the open catalogs, so install one if it matches and adapt it to your brief format and receipts rule.
Ten minutes, one file, fully sourced. That brief now feeds the drafting system you build in the next lesson.
Do this now
Sources and further reading
Want us to set it up with you, end to end?
Three one-on-one sessions. We train you on your real stack and build your first agents together, until you can run it yourself. You keep everything.