Sales · 01 · Your GTM data machineLesson 2 of 4
List building: Apollo & the sources
- Pull targeted lists via the Apollo API from Claude Code
- Know the current source landscape and when to use which
The 2026 source map
Here's what you'll walk away with: 200 leads matching your real ICP pulled from Apollo, scored, the top 50 enriched, and the whole pull saved as a one-command skill you run again next week.
List building in 2026 is not one tool - it's a small map of source TYPES, all reachable from a Claude Code session. This lesson uses Apollo as its worked example because the code has to be real and runnable, but hold the transferable lesson the whole way through: your data tool already has an API. The skill is pointing Claude Code at that API and letting it do the work. The same move works for whatever database you use.
- A B2B contact database - the example here is Apollo, a filter-driven database (title, industry, headcount, geo, funding stage). If you already pay for a different one, keep it - everything below maps 1:1.
- Lookalike engines - feed in your best customers, get similar companies back. Useful when filters can't describe your ICP (several vendors offer this; any of them slots in here).
- Signal sources (lesson 5 covers these in depth) - funding, hiring, technographics, social engagement, website visitors. These tell you WHEN, the databases tell you WHO.
- Your own CRM - closed-lost from 12+ months ago and stale leads are the cheapest list you'll ever build.
Connect the Apollo MCP
Apollo's official MCP server lives at https://mcp.apollo.io/mcp. It authenticates with OAuth - no API key to manage - and exposes people and company search, enrichment, sequence management, job postings and credit stats. The tool list comes from the server and changes, so run /mcp and read what is there today rather than trusting a count from a blog post.
- In your terminal: claude mcp add --transport http apollo mcp.apollo.io/mcp
- Inside a Claude Code session, run /mcp - you'll see apollo listed. Select it to trigger the OAuth flow in your browser.
- Log in with your Apollo account and approve. The connection persists across sessions.
- Sanity check: ask Claude to "search Apollo for 5 heads of growth at US marketing agencies, 10-50 employees" and confirm results come back.
If your team already added Apollo as a claude.ai connector, it flows into Claude Code automatically when you're logged in with your subscription - check /mcp before adding it twice.
Conversational prospecting: your first 200-lead list
With the MCP connected, list building becomes a conversation grounded in your rubric. This is the recipe to run today.
- Start a fresh session in your sales workspace (so icp-rubric.md is in scope).
- Prompt: "Read icp-rubric.md. Using the Apollo tools, build me a list of 200 people matching the A-band profile. Search only - do not enrich anything yet."
- Review the first page of results together. Tighten filters where Apollo's taxonomy disagrees with your rubric (industry labels rarely match 1:1).
- Have Claude run /icp-qualify over the results and sort by score.
- Enrich only the A-band rows: "Enrich the top 50 by score." Watch the credit usage - Apollo's MCP exposes credit stats, ask for them before and after.
- Export to CSV: leads_YYYY-MM-DD.csv with score and band columns included.
The API route for bulk and scheduled pulls
The MCP is for interactive work. When a pull becomes recurring or large, move to the REST API in a script - deterministic, loggable, schedulable. This is the track's rule, and it recurs by name: MCP is hands, scripts are muscle, skills are taste. Hands explore, muscle repeats, taste decides - and you just used all three in one lesson. Apollo's API is documented at docs.apollo.io; the endpoints you'll live in:
POST /api/v1/mixed_people/api_search- people search with the same filters as the UI. 0 credits, no emails or phones, 50k display cap. The older/mixed_people/searchpath answers 422 "deprecated for API callers" (we hit it on 2026-09-14); any tutorial still using it is stale.POST /api/v1/mixed_companies/search- company-level search (funding stage and date filters live here). 1 credit per page of 100, thin fields: discovery only.POST /api/v1/people/matchandpeople/bulk_match(up to 10 per request) - enrichment, 1 credit per match, +8 for a mobile, returns emails.POST /api/v1/organizations/enrich- company enrichment, 1 credit per company.GET /api/v1/organizations/{id}/job_postings- hiring signals, 1 credit per page (lesson 5 prices this before it runs).
# Paths and field names drift; have Claude verify each
# endpoint against docs.apollo.io before building on this.
# api_search: 0 credits, no emails/phones, last names masked,
# 50k display cap. The old /mixed_people/search path returns
# 422 "deprecated for API callers". Search returns an Apollo
# person id, first name, title and company name - no LinkedIn
# URL, domain or email. Those come from people/match, on the
# rows that pass your ICP gate (lesson 3).
import os, csv, requests
API = "https://api.apollo.io/api/v1"
HEADERS = {"X-Api-Key": os.environ["APOLLO_API_KEY"]}
def search_people(page: int, titles: list[str]) -> dict:
payload = {
"person_titles": titles,
"organization_num_employees_ranges": ["10,50"],
"person_locations": ["United States"],
"page": page,
"per_page": 100,
}
r = requests.post(f"{API}/mixed_people/api_search",
headers=HEADERS, json=payload, timeout=30)
r.raise_for_status()
return r.json()
# Silent-filter check (see next section): the count must MOVE
# when the filter changes, or the vendor is ignoring your key.
TITLES = ["Head of Growth", "Founder"]
def total(resp: dict) -> int:
# top-level in the current docs; older responses nested it
return resp.get("total_entries") or \
(resp.get("pagination") or {}).get("total_entries", 0)
n_all = total(search_people(1, TITLES))
n_one = total(search_people(1, TITLES[:1]))
assert n_one < n_all, "person_titles filter had no effect - stop"
rows = []
for page in range(1, 3): # 200 leads
for p in search_people(page, TITLES)["people"]:
rows.append({
"apollo_id": p.get("id"), # the key people/match takes
"first_name": p.get("first_name"),
"title": p.get("title"),
"company": (p.get("organization") or {}).get("name"),
"source": "apollo:api_search",
})
os.makedirs("output", exist_ok=True)
with open("output/leads_raw.csv", "w", newline="") as f:
w = csv.DictWriter(f, fieldnames=rows[0].keys())
w.writeheader()
w.writerows(rows)
print(f"wrote {len(rows)} leads")You don't write this by hand - you ask Claude to write it, pointing it at docs.apollo.io. Using a different database? Point Claude at YOUR tool's API docs instead and describe the same outcome - the script that comes back differs in field names, nothing else. Your job is to review the filters and run it. Note the API key comes from the environment, never the file; lesson 3 makes that discipline concrete. Note also the source column: every row you ever build carries where it came from, because in lesson 3 it also carries which provider found the email and what that cost.
The silent filter: the first check you write against any vendor API
Here is the most transferable verification habit in this track, and the one no vendor documents. Several GTM APIs silently ignore a filter key they do not recognize and return the UNFILTERED set with HTTP 200. Three vendors in our own stack do it: a company database, an enrichment provider, and a sender's lead-list endpoint where campaign_id is a no-op and the working key is campaign. Nothing errors. The response looks like a result.
- What it looks like: a 17-country market-sizing sweep returned an identical 66,502 companies for every country. An overlap scan reported "100% of your list was already contacted"; the real figure was 4.6%. Both were reported as findings before anyone noticed they were bugs.
- The habit: after writing any filter, change its value and confirm the count moves. If a title filter of two titles and one title return the same total, the filter is not being applied. The
assertin the script above is this habit in four lines. - The tell: treat an implausibly clean number - 100%, 0%, identical across segments - as evidence of an ignored filter first and a finding second.
- Variants to expect: a malformed filter that is ignored AND charged; not-found delivered as
{"status": 404}inside an HTTP 200 body; a Cloudflare403 error code: 1010to a default PythonurllibUser-Agent, which looks like an auth failure and is not.
What a list build looks like when the credits are yours
The conversational pull above is the classroom version. When we build a market list for money, it runs as nine stages, free ones to exhaustion before the first metered call, with the ICP gate before the first credit. The order is the lesson:
# Stage Tool type Cost
1 Market harvest company database flat
2 Is it really this kind of company? regex on name/site free
3 Does it exist? HTTP fetch of site free
4 Tier by buying signal site + firmographics free
5 Contacts, pass 1 people search free
6 Contacts, gap fill (native titles) people search free
7 Emails, pass 1 bulk_match 1 credit / MATCH
8 Emails, pass 2 on the misses pay-on-hit finder 1 credit / hit
9 Deliverability verifier 1 credit / address- Stage 3 earns the list. Roughly a third of database companies fail a plain website fetch: expired domains resold to gambling sites, "Index of /" servers, pivots into a different business. A gentler re-check recovered about 1.5%, which says the first pass was fair. Every one of those would have cost a credit at stage 7.
- Titles: recall from the API, precision locally. Providers match loosely ("VP Sales" returns "VP, Sales Enablement"). Write the title rules as code and unit-test them, including must-REJECT cases. That test is what caught
\bpresident\bmatching "Vice President". - Use the vertical's own titles. SaaS-style titles covered 29% of target companies in a services vertical; that vertical's native titles (Branch Manager, Practice Director, Partner, Principal) took coverage past 50%. Your rubric's persona list should read like their org chart, not yours.
- Report the clean rate, never the raw hit rate. One pool came back 20% raw and 12% clean after the defect guards in lesson 3; another 63% raw and 16% clean. Raw hit rate is a vendor metric.
- Every metered script has a hard
MAX_CREDITScap in its config and aborts when it is reached. Eating the month's credits because a loop did not stop is a solved problem: solve it before the first run, not after.
For teams: shared searches, separate credit budgets
Solo, you feel a wasted credit immediately. On a team, waste hides: five reps each running slightly different searches, enriching overlapping lists, on one shared credit pool. The fix is the same move as the rubric - make the searches shared, versioned assets.
Do this now
Sources and further reading
Want us to set it up with you, end to end?
Three one-on-one sessions. We train you on your real stack and build your first agents together, until you can run it yourself. You keep everything.