anfloy.AcademyBook a call

Sales · 02 · Outbound at scaleLesson 2 of 4

Personalization that isn't cringe

75 min working time · Weeks 7-8

By the end of this lesson you can
  • Generate first lines grounded in verifiable facts only
  • Build the personalization skill with a fact-check pass

The reply-rate ladder

The deliverable here: grounded first lines generated for 25 of your real A-band leads, by a /first-line skill that cites evidence or refuses to write - with a separate verify pass that catches anything invented.

Personalization has a measurable price ladder. The 2026 vendor benchmark data (treat the exact bands as directional ranges, not gospel) stacks up like this:

  • No personalization: 1-3% reply rate
  • Basic merge fields ({{first_name}}, {{company}}): 5-9%
  • Pain-point or news-based relevance: 9-15%
  • Signal-based trigger plus a tailored value prop: reported 15-25% - the upper band comes from a single source family, so treat it as a ceiling claim, but the direction across all sources is consistent: each rung roughly doubles the one below.

For calibration: the average cold-email reply rate sits around ~3.5% (vendor-reported, 2026). Most senders are on rung one or two. The signal work you did last lesson bought you a ticket to rung four - this lesson is about cashing it without writing slop.

The failure mode: confident fabrication

Ask a model to "write a personalized first line for Jane at Acme" with no source material and it will - by inventing a plausible detail. "Loved your recent post about scaling SDR teams" reads great until Jane, who wrote no such post, replies asking what you're talking about. One fabricated line costs more trust than a hundred generic ones.

The fix is a discipline, not a better prompt: separate gathering from writing. First, fetch real evidence about the lead. Then write ONLY from that evidence. If the evidence is thin, skip the lead rather than pad the line. Every sentence must survive the reply "how do you know that?" with a URL.

The two-pass pipeline: gather, then write

The pipeline formalizes the separation. In practice it's four passes, and the order is the whole trick - the writing pass never runs against an empty evidence file.

  1. Pass 1 - gather. For each lead, fetch the available evidence: company homepage and one or two key pages, the LinkedIn profile data you already have from enrichment, recent news mentions, the signal that flagged them (job posting, funding item, social engagement). Store it per lead: evidence/{lead_id}.json, each fact with its source URL.
  2. Pass 2 - judge. Does the evidence support a specific, relevant observation? Score evidence quality 0-2. Zero means skip - the lead gets template treatment, no shame in it.
  3. Pass 3 - write. Generate 1-2 lines that cite ONLY gathered facts, in your voice, connecting the observation to the problem you solve. The signal is usually the strongest material: "saw you're hiring two SDRs" beats "love what you're building."
  4. Pass 4 - verify. A separate check pass re-reads each line against the evidence file: every claim must trace to a stored fact. Lines that don't trace get rejected back to template. This pass uses a fresh context so it can't rationalize the writer's choices.

Build the personalization skill

The skill encodes your taste: voice, banned phrases, and the evidence-or-skip rule. The scripts fetch and loop; the skill is what keeps a thousand generated lines sounding like your team on its best day.

.claude/skills/first-line/SKILL.md
---
name: first-line
description: Write grounded first lines for outbound emails.
  Requires an evidence file per lead - never writes without one.
argument-hint: <evidence-file or leads-folder>
---

Write 1-2 opening lines for a cold email.

Hard rules:
1. Every claim must come from the evidence file. No outside
   knowledge, no inference beyond what a fact states.
2. If evidence quality is 0 (no specific, recent, relevant
   fact), output SKIP with a one-line reason. Do not write.
3. Under 25 words for the opener. Plain words. No flattery.
4. Connect the observed fact to the problem we solve in the
   reader's terms - see ./voice.md for tone and examples.
5. Banned openers and phrases in ./banned.md - includes
   "I noticed", "Congrats on", "I came across", "impressive",
   "quick question", and anything a thousand other senders
   wrote this morning.

Output per lead: { line, facts_used: [urls], confidence }
  1. Write voice.md with 5 real first lines your best rep actually sent (and got replies to). Examples teach tone better than adjectives.
  2. Write banned.md - start with the list above, add every cliche you personally hate receiving.
  3. Run the full pipeline on 25 A-band leads from your enriched list.
  4. Read every line. You're the editor of record on this first batch: mark each line keep / edit / reject, and note WHY for the rejects.
  5. Fold the reject reasons back into banned.md and voice.md. Two or three iterations of this loop is what separates your output from generic AI slop.

Cost and scale math

Grounded personalization is token-hungry by design - you're fetching and reading real pages per lead. Production systems budget roughly 500K to 1.2M tokens per 1,000 records with personalization. Plan for it rather than discovering it on an invoice:

  • Personalize only A-band leads. B-band gets the rung-two template - the math rarely justifies more, and the ladder says template beats skipped.
  • Route by pass. Gather and extract (pass 1) is bulk, schema-checked and read by nobody: Haiku 4.5 with the extraction schema cached, or Batch at half price when same-day is not needed. Judge and write (passes 2-3) carry your voice and reach a prospect: Sonnet 5 as the daily driver, Opus 5.5 when the voice is hard to hit. Verify (pass 4) is judgment about money and trust: Claude, fresh context. Never route a prospect-facing line to a cheap open model; the honest-routing rule from lesson 4 puts writing and verifying firmly on the keep-on-Claude side.
  • Compact before you prompt. The rubric and the writer read a handful of fields; a raw profile carries 15KB of job history and 8KB of recommendations they do not. Strip to what the pass reads, and put the stable part (voice.md, banned.md, the rules) in a cached system prompt. For bursty runs use the 1-hour cache TTL: a 5-minute cache expires between batches and you pay the write premium every time.
  • Cache evidence files - a re-run on the same leads should cost near zero gathering tokens.
  • Track skip rate per batch. Under 10% means your evidence bar is probably too loose; over 40% means your gather pass needs more sources.
  • Log tokens per call (input, cache read, cache write, output) with provider and model, and roll them up per batch. Output is the expensive token; a model that writes long costs more than its rate card says.

Do this now

Sources and further reading

We set it up with you

Want us to set it up with you, end to end?

Three one-on-one sessions. We train you on your real stack and build your first agents together, until you can run it yourself. You keep everything.