anfloy.AcademyBook a call

Sales · 01 · Your GTM data machineLesson 3 of 4

The enrichment waterfall

90 min working time · Weeks 5-6

By the end of this lesson you can
  • Build a cost-ordered waterfall: cheap first, expensive last
  • Wire each finder in as an API step (Prospeo as the example)
  • Dedupe and validate before anything hits the CRM

Why waterfalls beat any single provider

The build this time: a waterfall script running on your own list - 90%+ of A-band rows with a verified email, a cost report per run, and your API keys handled like a professional.

No single email-finding provider covers everyone. Vendors quote 55-62% coverage for one provider and 80-98% for a chain; our own measured version is plainer. On a population of LinkedIn engagers at small agencies, one pay-on-hit provider found emails for 40%. Adding a second provider on the misses only took the yield to 80%. Same list, same day, double the sendable rows, and the second provider was paid only for what it found. That chain is called a waterfall, and it is the standard pattern across every serious GTM team.

Spreadsheet enrichment platforms sell waterfalls as a feature behind a per-seat fee (as of September 2026 the entry paid tier of the best-known one starts at $167 a month billed monthly, for 3,000 data credits and 15,000 actions). You're going to build the same thing as a 30-line script you own. Same coverage, no seat, full cost visibility per row. The pattern matters more than any provider name. Here is the order we run and why:

  • Free to exhaustion first, ICP gate before the first credit. Website check, title rules, dedupe and the rubric all run before any finder sees a row. Every row you drop for free is a credit you did not spend on someone you would never email.
  • The database you already pay for, second. Apollo bulk_match costs 1 credit per MATCH (not per email), up to 10 people per request, and the company switchboard phone comes free with a match. You are already paying the seat; use its credits before anyone else's.
  • A pay-on-hit finder on the misses, third. Prospeo, the example this lesson builds on, charges 1 credit per verified email, 0 on a miss, and re-enrichment of the same record within 90 days is free. Misses cost nothing, so the second provider only ever sees leftovers and only bills for wins. Confirm the billing model for any provider before it joins your chain; some waterfall vendors charge even on a miss.
  • The verifier last, on everything. It is the only step that touches every found address, so it is priced per address, and it is the gate that decides what may enter a campaign (the 2% rule, below). Aggregators that waterfall 15-20 sources themselves cost several times more per email; if you use one, it sees only the stubborn rows.

Which providers? Whichever you already have accounts with, in that order: flat-rate or already-paid, then pay-on-hit, then verify. This lesson uses Prospeo so the code is real; every other finder slots into the same chain the same way - an API key, a find call, a billing model to confirm.

.env discipline before the first call

You're about to hold three or four API keys. The rule is absolute: keys live in a .env file that is gitignored, scripts read them from the environment, and a key never appears in code, in chat, or in a commit. Set this up once and the habit carries through every build in this course.

.env (gitignored - this file never leaves your machine)
# Enrichment providers in waterfall order. Prospeo is the
# course example - add a line per provider YOU use, same shape.
APOLLO_API_KEY=...          # already-paid database, runs first
PROSPEO_API_KEY=pk_...      # pay-on-hit finder, runs on misses
VERIFIER_API_KEY=...        # runs last, on every found address

# Per-run caps - the waterfall script refuses to start a run
# estimated above these, and aborts if it reaches them
ENRICH_BUDGET_USD=15
MAX_CREDITS=500
  1. Create .env in your sales workspace root and add the keys you have.
  2. Confirm .env is in .gitignore. If the repo doesn't have one, add it now: echo ".env" >> .gitignore
  3. Add a permission deny rule so Claude can't read the file either: in .claude/settings.json, add "Read(./.env*)" to permissions.deny.
  4. Test: ask Claude to read .env and confirm it refuses. The deny rule covers Claude's built-in file tools and the file commands it recognizes in Bash (cat, head, tail, sed).

Hello world: one Prospeo call

Prospeo is the textbook first API for a sales team: simple auth, cheap, and you're only charged when it finds something. Make one call by hand before building the waterfall, so the machinery underneath never feels like magic. As of September 2026 the endpoint is POST /enrich-person (the older /email-finder path in old tutorials is not the documented one; verify against the docs before building).

One enrich-person call (docs: prospeo.io/api-docs)
curl -s -X POST https://api.prospeo.io/enrich-person \
  -H "Content-Type: application/json" \
  -H "X-KEY: $PROSPEO_API_KEY" \
  -d '{
    "only_verified_email": true,
    "data": {
      "first_name": "Jane",
      "last_name": "Doe",
      "company_website": "acme.com"
    }
  }'

# Billing (Prospeo help center, Sept 2026): 1 credit per
# verified email, 10 per verified mobile (email included),
# 0 when nothing is found, 0 to re-enrich the same record
# within 90 days. Add "enrich_mobile": true only for rows
# a rep will call.

Run it with a real lead from your list. You'll get back an email plus its verification status. That's the entire atomic unit - the waterfall is just this call, repeated per provider, per row, with bookkeeping. only_verified_email: true is the setting that matters: it tells the provider to return nothing rather than a guess, which is exactly the trade you want a pay-on-hit provider to make.

Build the waterfall script

Now the real build. Describe the logic to Claude and let it write the script; your job is to specify the order, the stop condition, and the bookkeeping. This is the moment that replaces the rented enrichment spreadsheet.

waterfall.py - the core logic (sketch)
# Providers in waterfall order. Each returns (email, status) or None.
# ANY finder slots in the same way: a find function, a cost per HIT,
# a confirmed billing model. Costs are yours to fill in from your
# plan; the shape is what matters.
PROVIDERS = [
    ("apollo",  find_apollo,  COST["apollo"]),   # already paid, 1 credit/match
    ("prospeo", find_prospeo, COST["prospeo"]),  # pay-on-hit, 0 on miss
    # ("aggregator", find_agg, COST["agg"]),     # stubborn rows only
]

def enrich_row(row: dict, ledger: list, cache: dict) -> dict:
    key = (row.get("linkedin_url") or row.get("apollo_id")
           or (row["first_name"], row["company"]))
    if key in cache:                       # one lookup per person, ever
        return {**row, **cache[key], "email_source": "cache"}
    for name, finder, cost in PROVIDERS:
        if ledger_total(ledger) >= MAX_CREDITS:
            raise SystemExit("MAX_CREDITS reached - stopping, not eating the month")
        result = finder(row)
        log_attempt(row, provider=name, hit=bool(result))  # JSON line
        if result:
            email, status = result
            ledger.append({"row": row["id"], "provider": name, "cost": cost})
            found = {"email": email, "email_status": status,
                     "email_source": name, "email_cost": cost}
            cache[key] = found
            return {**row, **found}
    return {**row, "email": None, "email_source": "none", "email_cost": 0}

# After the run: defect guards, then verify every found email (below),
# then print the cost report - rows, hits per provider, total spend,
# effective cost per VERIFIED email (never per found email).
  1. In Claude Code: "Write waterfall.py implementing my providers in waterfall order: the database I already pay for first (Apollo bulk_match), then the pay-on-hit finder on the misses (Prospeo enrich-person), pointing at each one's API docs. Read keys from the environment. Stop at the first hit per row. Keep a persistent cache keyed on LinkedIn URL so a person is never looked up twice across runs. Log every attempt as a JSON line to logs/. Track cost per provider, support a --dry flag that prints the estimated bill and exits, refuse to start if the estimate exceeds ENRICH_BUDGET_USD, and abort if the ledger reaches MAX_CREDITS. Read input CSV, write output CSV with email, email_status, email_source and email_cost columns."
  2. Review the plan before it writes code - check the provider order, the cache key, and both guards especially.
  3. Run the dry mode first: python waterfall.py input/leads_raw.csv --dry. Read the bill. Then run it for real on 50 rows.
  4. Read the cost report. Expect the first provider to find 40-60% and the second to roughly double the yield on the misses. If the second provider adds almost nothing, your population is one it does not cover; try a different second provider before adding a third.

The defects a hit rate hides

A found email is not a sendable email, and not only because it might bounce. On a 12,000-row run of our own, every one of these guards fired. Put them between the finder and the verifier, in code, and report the clean rate they leave behind:

  • Wrong company: about 20% of found addresses were at a DIFFERENT company from the one on the row (a previous employer, a namesake). Compare the address domain to the company domain; allow a legitimate sister brand only when the two domains share a distinctive token.
  • One address, two people: the same email returned for two different rows. Neither is sendable. Drop both and log it.
  • Same person twice under name variants (Jon / Jonathan, maiden name, initials). Dedupe on LinkedIn URL and on email before counting anything.
  • Report clean rate, not raw hit rate. 20% raw became 12% clean on one pool; 63% raw became 16% clean on another. The vendor's number is the first one; your budget runs on the second.

Verify everything: the 2% rule, and the verifier that fits your list

Finding an email and being able to send to it are different things. Every found email goes through a verification pass before it's allowed near a campaign, because bounces poison the asset you'll spend weeks warming up in module 2: your domain reputation.

  • Verification is cheap relative to what it protects. As of September 2026, pay-as-you-go runs about $0.008 per check at one mainstream verifier and under $0.002 per check in 50k blocks at another: a fraction of an enrichment credit.
  • The gate is valid-only, and it never loosens. Only valid or deliverable may send. catch-all, unknown and error do not, and an error fails closed (the row waits; it does not send). The day the numbers look bad is the day someone proposes relaxing this. Don't.
  • Pick the verifier by population, not by brand. On LinkedIn-engager leads at small agencies (lots of catch-all domains) one well-known verifier passed 1 of 8 addresses; the other 7 came back unknown or catch-all. Some verifiers cannot judge a catch-all domain at all. The fix was a catch-all-capable verifier, not a looser gate. And never re-verify a good list with the wrong tool: one such pass wrote 176 deliverable leads as bounced.
  • One verification record per address, reused across every source. Verification is cached like the finder results: the same address arriving from a second list, a signal feed or a CRM import is never paid for twice. The cache is the cost control.
  • Rows that fail verification go back through the remaining providers once, then get flagged unverified and stay out of every campaign. Catch-all rows sit in their own band; some aggregators do not charge on a catch-all result, so if you run one, it sees those rows, but the send gate still applies to whatever comes back.

Target after verification: 90%+ of your A-band rows with a verified email. Then compute the number that runs the whole engine: dollars per VERIFIED lead, all-in (scrape or database, finder credits, verifier, LLM). On our engager pipeline that came to about $0.10 per verified lead across a first full run, with individual sources ranging from $0.02 to $0.60, a 30x spread on the same rubric. Budget in $/verified lead, never in rows. That measured number, not a vendor claim, is what you carry into the team budget conversation.

For teams: budgets, caps, and one shared waterfall

Enrichment is the spendiest stage of the data machine, and it's also the most centralizable - nothing about a waterfall needs to be per-rep except the budget.

Do this now

Sources and further reading

We set it up with you

Want us to set it up with you, end to end?

Three one-on-one sessions. We train you on your real stack and build your first agents together, until you can run it yourself. You keep everything.