anfloy.AcademyBook a call

Sales · 01 · Your GTM data machineLesson 2 of 4

List building: Apollo & the sources

90 min working time · Weeks 5-6

By the end of this lesson you can
  • Pull targeted lists via the Apollo API from Claude Code
  • Know the current source landscape and when to use which

The 2026 source map

Here's what you'll walk away with: 200 leads matching your real ICP pulled from Apollo, scored, the top 50 enriched, and the whole pull saved as a one-command skill you run again next week.

List building in 2026 is not one tool - it's a small map of source TYPES, all reachable from a Claude Code session. This lesson uses Apollo as its worked example because the code has to be real and runnable, but hold the transferable lesson the whole way through: your data tool already has an API. The skill is pointing Claude Code at that API and letting it do the work. The same move works for whatever database you use.

  • A B2B contact database - the example here is Apollo, a filter-driven database (title, industry, headcount, geo, funding stage). If you already pay for a different one, keep it - everything below maps 1:1.
  • Lookalike engines - feed in your best customers, get similar companies back. Useful when filters can't describe your ICP (several vendors offer this; any of them slots in here).
  • Signal sources (lesson 5 covers these in depth) - funding, hiring, technographics, social engagement, website visitors. These tell you WHEN, the databases tell you WHO.
  • Your own CRM - closed-lost from 12+ months ago and stale leads are the cheapest list you'll ever build.

Connect the Apollo MCP

Apollo's official MCP server lives at https://mcp.apollo.io/mcp. It authenticates with OAuth - no API key to manage - and exposes people and company search, enrichment, sequence management, job postings and credit stats. The tool list comes from the server and changes, so run /mcp and read what is there today rather than trusting a count from a blog post.

  1. In your terminal: claude mcp add --transport http apollo mcp.apollo.io/mcp
  2. Inside a Claude Code session, run /mcp - you'll see apollo listed. Select it to trigger the OAuth flow in your browser.
  3. Log in with your Apollo account and approve. The connection persists across sessions.
  4. Sanity check: ask Claude to "search Apollo for 5 heads of growth at US marketing agencies, 10-50 employees" and confirm results come back.

If your team already added Apollo as a claude.ai connector, it flows into Claude Code automatically when you're logged in with your subscription - check /mcp before adding it twice.

Conversational prospecting: your first 200-lead list

With the MCP connected, list building becomes a conversation grounded in your rubric. This is the recipe to run today.

  1. Start a fresh session in your sales workspace (so icp-rubric.md is in scope).
  2. Prompt: "Read icp-rubric.md. Using the Apollo tools, build me a list of 200 people matching the A-band profile. Search only - do not enrich anything yet."
  3. Review the first page of results together. Tighten filters where Apollo's taxonomy disagrees with your rubric (industry labels rarely match 1:1).
  4. Have Claude run /icp-qualify over the results and sort by score.
  5. Enrich only the A-band rows: "Enrich the top 50 by score." Watch the credit usage - Apollo's MCP exposes credit stats, ask for them before and after.
  6. Export to CSV: leads_YYYY-MM-DD.csv with score and band columns included.

The API route for bulk and scheduled pulls

The MCP is for interactive work. When a pull becomes recurring or large, move to the REST API in a script - deterministic, loggable, schedulable. This is the track's rule, and it recurs by name: MCP is hands, scripts are muscle, skills are taste. Hands explore, muscle repeats, taste decides - and you just used all three in one lesson. Apollo's API is documented at docs.apollo.io; the endpoints you'll live in:

  • POST /api/v1/mixed_people/api_search - people search with the same filters as the UI. 0 credits, no emails or phones, 50k display cap. The older /mixed_people/search path answers 422 "deprecated for API callers" (we hit it on 2026-09-14); any tutorial still using it is stale.
  • POST /api/v1/mixed_companies/search - company-level search (funding stage and date filters live here). 1 credit per page of 100, thin fields: discovery only.
  • POST /api/v1/people/match and people/bulk_match (up to 10 per request) - enrichment, 1 credit per match, +8 for a mobile, returns emails.
  • POST /api/v1/organizations/enrich - company enrichment, 1 credit per company.
  • GET /api/v1/organizations/{id}/job_postings - hiring signals, 1 credit per page (lesson 5 prices this before it runs).
pull_list.py - the shape of a scripted pull (a sketch, not gospel)
# Paths and field names drift; have Claude verify each
# endpoint against docs.apollo.io before building on this.
# api_search: 0 credits, no emails/phones, last names masked,
# 50k display cap. The old /mixed_people/search path returns
# 422 "deprecated for API callers". Search returns an Apollo
# person id, first name, title and company name - no LinkedIn
# URL, domain or email. Those come from people/match, on the
# rows that pass your ICP gate (lesson 3).
import os, csv, requests

API = "https://api.apollo.io/api/v1"
HEADERS = {"X-Api-Key": os.environ["APOLLO_API_KEY"]}

def search_people(page: int, titles: list[str]) -> dict:
    payload = {
        "person_titles": titles,
        "organization_num_employees_ranges": ["10,50"],
        "person_locations": ["United States"],
        "page": page,
        "per_page": 100,
    }
    r = requests.post(f"{API}/mixed_people/api_search",
                      headers=HEADERS, json=payload, timeout=30)
    r.raise_for_status()
    return r.json()

# Silent-filter check (see next section): the count must MOVE
# when the filter changes, or the vendor is ignoring your key.
TITLES = ["Head of Growth", "Founder"]
def total(resp: dict) -> int:
    # top-level in the current docs; older responses nested it
    return resp.get("total_entries") or \
        (resp.get("pagination") or {}).get("total_entries", 0)

n_all = total(search_people(1, TITLES))
n_one = total(search_people(1, TITLES[:1]))
assert n_one < n_all, "person_titles filter had no effect - stop"

rows = []
for page in range(1, 3):  # 200 leads
    for p in search_people(page, TITLES)["people"]:
        rows.append({
            "apollo_id": p.get("id"),  # the key people/match takes
            "first_name": p.get("first_name"),
            "title": p.get("title"),
            "company": (p.get("organization") or {}).get("name"),
            "source": "apollo:api_search",
        })

os.makedirs("output", exist_ok=True)
with open("output/leads_raw.csv", "w", newline="") as f:
    w = csv.DictWriter(f, fieldnames=rows[0].keys())
    w.writeheader()
    w.writerows(rows)
print(f"wrote {len(rows)} leads")

You don't write this by hand - you ask Claude to write it, pointing it at docs.apollo.io. Using a different database? Point Claude at YOUR tool's API docs instead and describe the same outcome - the script that comes back differs in field names, nothing else. Your job is to review the filters and run it. Note the API key comes from the environment, never the file; lesson 3 makes that discipline concrete. Note also the source column: every row you ever build carries where it came from, because in lesson 3 it also carries which provider found the email and what that cost.

The silent filter: the first check you write against any vendor API

Here is the most transferable verification habit in this track, and the one no vendor documents. Several GTM APIs silently ignore a filter key they do not recognize and return the UNFILTERED set with HTTP 200. Three vendors in our own stack do it: a company database, an enrichment provider, and a sender's lead-list endpoint where campaign_id is a no-op and the working key is campaign. Nothing errors. The response looks like a result.

  • What it looks like: a 17-country market-sizing sweep returned an identical 66,502 companies for every country. An overlap scan reported "100% of your list was already contacted"; the real figure was 4.6%. Both were reported as findings before anyone noticed they were bugs.
  • The habit: after writing any filter, change its value and confirm the count moves. If a title filter of two titles and one title return the same total, the filter is not being applied. The assert in the script above is this habit in four lines.
  • The tell: treat an implausibly clean number - 100%, 0%, identical across segments - as evidence of an ignored filter first and a finding second.
  • Variants to expect: a malformed filter that is ignored AND charged; not-found delivered as {"status": 404} inside an HTTP 200 body; a Cloudflare 403 error code: 1010 to a default Python urllib User-Agent, which looks like an auth failure and is not.

What a list build looks like when the credits are yours

The conversational pull above is the classroom version. When we build a market list for money, it runs as nine stages, free ones to exhaustion before the first metered call, with the ICP gate before the first credit. The order is the lesson:

Our list-build waterfall (measured Aug-Sep 2026)
#  Stage                                  Tool type            Cost
1  Market harvest                          company database     flat
2  Is it really this kind of company?      regex on name/site   free
3  Does it exist?                          HTTP fetch of site   free
4  Tier by buying signal                   site + firmographics free
5  Contacts, pass 1                        people search        free
6  Contacts, gap fill (native titles)      people search        free
7  Emails, pass 1                          bulk_match           1 credit / MATCH
8  Emails, pass 2 on the misses            pay-on-hit finder    1 credit / hit
9  Deliverability                          verifier             1 credit / address
  • Stage 3 earns the list. Roughly a third of database companies fail a plain website fetch: expired domains resold to gambling sites, "Index of /" servers, pivots into a different business. A gentler re-check recovered about 1.5%, which says the first pass was fair. Every one of those would have cost a credit at stage 7.
  • Titles: recall from the API, precision locally. Providers match loosely ("VP Sales" returns "VP, Sales Enablement"). Write the title rules as code and unit-test them, including must-REJECT cases. That test is what caught \bpresident\b matching "Vice President".
  • Use the vertical's own titles. SaaS-style titles covered 29% of target companies in a services vertical; that vertical's native titles (Branch Manager, Practice Director, Partner, Principal) took coverage past 50%. Your rubric's persona list should read like their org chart, not yours.
  • Report the clean rate, never the raw hit rate. One pool came back 20% raw and 12% clean after the defect guards in lesson 3; another 63% raw and 16% clean. Raw hit rate is a vendor metric.
  • Every metered script has a hard MAX_CREDITS cap in its config and aborts when it is reached. Eating the month's credits because a loop did not stop is a solved problem: solve it before the first run, not after.

For teams: shared searches, separate credit budgets

Solo, you feel a wasted credit immediately. On a team, waste hides: five reps each running slightly different searches, enriching overlapping lists, on one shared credit pool. The fix is the same move as the rubric - make the searches shared, versioned assets.

Do this now

Sources and further reading

We set it up with you

Want us to set it up with you, end to end?

Three one-on-one sessions. We train you on your real stack and build your first agents together, until you can run it yourself. You keep everything.