anfloy.AcademyBook a call

Sales · 03 · Pipeline ops & reportingLesson 4 of 4

Sales capstone: the full engine

120 min working time · Weeks 9-10

By the end of this lesson you can
  • Run the complete flow: signal -> list -> enrich -> send -> reply -> report
  • Document it as your team's playbook

The engine you've already built

Look at what the last six weeks produced: a rubric, a list builder, a waterfall, a scorer, signal watchers, a personalization skill, sending infrastructure, a sequencer, a triage agent, call extraction, and a self-writing report. Today you wire them into one engine, run it end to end on your real ICP with your real accounts, and document it so it outlives this course.

The canonical pipeline - every stage is now yours
SEED -> LIST BUILD -> FREE GATES -> ENRICH (waterfall) -> VERIFY
  -> SCORE/QUALIFY -> DEDUPE + BLOCKLIST -> PERSONALIZE
  -> PRE-LAUNCH AUDIT -> SEND -> MONITOR
  -> REPLY HANDLING -> CRM SYNC -> REPORT

L1 rubric        gates  SCORE (one file, zero copies)
L2 apollo pull   feeds  LIST (api_search, silent-filter check)
L3 waterfall     owns   ENRICH + VERIFY (valid-only, $/verified lead)
L4 score+sync    owns   SCORE + CRM (model routing, fail-closed guard)
L5 signals       feed   SEED (priced before they run)
L6 first-line    owns   PERSONALIZE (evidence-or-skip, Claude only)
L7 instantly     owns   SEND (20-30/inbox/day, audit, per-domain monitor)
L8 sequencer     owns   multichannel timing (20 invites/day cap)
L9 triage        owns   REPLY (human approves every send)
L10 call-extract feeds  CRM + REPORT
L11 report       owns   Monday 8am (cost line priced per model)

The repo structure that runs in production

This layout is what GTM agencies actually run client pipelines on. One repo, one folder per client (or per segment, if you're in-house), shared logic factored out, state and logs disciplined from day one:

sales-engine/ - the production layout
sales-engine/
  CLAUDE.md              <- the playbook (see final section)
  .env                   <- keys (gitignored)
  .claude/
    settings.json        <- permissions, deny Read(./.env*)
    skills/              <- icp-qualify, first-line, sequence-rules,
                            triage, call-extract, weekly-report,
                            crm-hygiene
  clients/
    acme/                <- or segments/ if in-house
      config.json        <- ICP params, budgets, campaign ids
      input/             <- raw pulls land here
      output/            <- enriched, scored, clean CSVs
      state/             <- engine.db (SQLite), caches, report state
      logs/              <- JSON lines, one file per run
  shared/
    waterfall.py  score.py  sync.py  handoff.py
    pull_report_data.py  morning_check.py
  tasks/                 <- cron entries, one documented job each
  icp-rubric.md  voice.md  banned.md
  deliverability-defaults.md  field-mapping.md
  • State in SQLite once a pipeline passes ~10k records - CSVs stop being honest about what's been processed, retried, or spent. One engine.db per client tracks every lead's stage, enrichment attempts, and costs.
  • Structured JSON logs, one line per enrichment attempt and per decision - when a run misbehaves, you grep the log instead of guessing. You've been doing this since lesson 3; now it's a repo-wide rule.
  • config.json per client/segment carries everything that varies: ICP thresholds, budget caps, campaign IDs, sequence variant. The shared scripts read config; no per-client forks of code.
  • Claude Code is not a daemon: everything fires from cron (or Routines), runs, logs, exits. The only always-on piece is the small webhook receiver from lesson 9.
  • Every new scheduled job ships with its schedule OFF behind an env var (SCHEDULE_ENABLED=0) until one manual run has been read end to end. The first unattended run is the one that finds the missing config folder.

The run: end to end, for real

Now run it. Real ICP, real accounts, real sends - a controlled cohort of about 100 leads so every gate is checkable in one sitting.

  1. SEED: pull this week's signal hits (hiring sweep + any social/visitor hits) and merge with a fresh Apollo pull against the rubric. Target ~150 raw leads.
  2. ENRICH + VERIFY: run the waterfall. Gate: 90%+ verified coverage on A-band. Record the cost report - blended cost per verified email goes in the playbook.
  3. SCORE + DEDUPE + SYNC: score.py, dedupe against the CRM (domain-level included), upsert the clean set. Gate: zero already-known contacts survived dedupe.
  4. PERSONALIZE: evidence gathering plus first lines on the A-band (~60-80 leads). Gate: every line traces to evidence; skip rate noted and the skips fall back to template.
  5. AUDIT: run prelaunch_audit.py against the created campaign. Gate: every step has a real body (and step 1 a subject), campaign cap equals mailboxes x per-mailbox cap, every removal verified by GET, every merge tag has a value, the blocklist guard read a non-empty list.
  6. SEND: load into the campaign via create_campaign.py - caps from deliverability-defaults, sequence from sequence-rules, LinkedIn drafts for the top slice (human-gated start, 20 invites a day at most). Gate: no inbox exceeds its daily cap; bounce check at 48 hours must hold under 2%, joined by sending domain so a young domain is throttled rather than the list blamed.
  7. REPLY: confirm the triage agent catches the first real replies; approve responses from Slack. Gate: median reply-to-response under your SLA.
  8. REPORT: trigger the weekly report manually at the end of the run window. It should narrate this cohort accurately - that's your verification that every pipe is connected.

Document it: the playbook is the deliverable

The engine isn't done until someone who isn't you can run it. The repo's CLAUDE.md is the playbook: what this system is, how each stage runs, what the gates are, and what to do when something breaks. Write it for the teammate who joins in six months.

  1. Write CLAUDE.md with: the pipeline diagram, a one-paragraph description per stage pointing at its script and skill, the gate numbers (90% coverage, <2% bounce, 0.10% spam rate, caps, SLAs), the model per step with its price, and the cron schedule with what runs when. Run /doctor prompt-audit on it once; it flags instructions written for older models.
  2. Add the numbers YOU measured: $/verified lead by source, clean rate, skip rate, classification accuracy, tokens per 1,000 records by model, reply rate vs your benchmarks.json baseline. Your playbook should brag in specifics.
  3. Add the failure playbook: bounce spike -> join to sending domain, throttle young domains, re-verify only if every domain bounces; report missing -> check cron log and the heartbeat's build marker; classification drift -> re-run calibration; a metric at exactly zero or 100% -> suspect an ignored filter or a swallowed exception before believing it. Each known failure, one prescribed response.
  4. Make the repo a company brain, not a project: CLAUDE.md is the index that says what exists and which build to copy for what, every folder carries its own short README so it self-describes, and the skills live in the repo so the next project inherits them on install. Keep a hand-maintained "current best" list at the top: the one file you update when a newer build does something better.
  5. Commit everything. Tag it v1. This repo - not any single campaign - is the asset the track was building.

Do this now

Sources and further reading

We set it up with you

Want us to set it up with you, end to end?

Three one-on-one sessions. We train you on your real stack and build your first agents together, until you can run it yourself. You keep everything.