Rollout · 01 · Team setup & governanceLesson 3 of 3
Usage, cost & model routing
- Monitor team usage and set expectations
- Write the team's model-routing and spend policy
- Right-size plans as adoption grows
What teams actually spend
Numbers replace vibes in this lesson. You finish with a recorded usage baseline, a cost per run for every automation, a written routing and spend policy, and a one-sentence budget for leadership. Anchor on Anthropic's published benchmarks, unchanged as of September 2026: roughly $13 per developer per active day, $150-250 per developer per month, and 90% of users under $30 per active day. Operator workloads - reports, list work, document generation - typically land below the engineering-heavy averages.
- Pro: $20/month, or $17/month billed annually.
- Max 5x: $100/month. Max 20x: $200/month.
- Team Standard seat: $25/month, or $20/month billed annually, minimum 2 seats.
- Team Premium seat: $125/month, or $100/month billed annually, 5x the usage of a standard seat.
- Enterprise: talk to sales before you put a number in a budget.
Those are claude.com list prices as of September 2026. Every plan from Pro up includes Claude Code, and the free plan does not. Usage runs on a rolling 5-hour window plus a weekly limit, shared across Claude chat and Claude Code. Frame it against what it replaces and the conversation gets short: a seat plus usage costs less per month than most single per-seat GTM tools, and you have been replacing several of those since week five.
Seeing usage: three levels
- Individual, in-session: /usage shows session tokens and your plan-limit status, broken down by skill, subagent, plugin, and MCP server, with a Loops line and a prompt-cache line. /skill-doctor shows each skill's context cost and how often it runs. /usage-credits manages the extra-usage spend cap.
- Team: the analytics dashboard at claude.ai/analytics/claude-code (Team and Enterprise) shows who is adopting, sessions, and the trend; the spend report in org analytics shows estimated usage-credit spend per user and model. This is the rollout's scoreboard.
- Automations:
claude -p --output-format jsonreturnstotal_cost_usdplus a per-model breakdown for every run, so every scheduled job already produces its own cost ledger. API workspaces get spend and rate limits in the Console.
For orgs that want usage data in their own dashboards, Claude Code exports OpenTelemetry metrics (set CLAUDE_CODE_ENABLE_TELEMETRY=1) and the monitoring docs cover the setup. If you pay contracted rates, the modelPricing admin setting makes /cost and telemetry use them. File under nice-to-have until someone in finance asks.
- This week: every team member runs /usage and reads their own breakdown once.
- The admin opens the analytics dashboard and records the baseline: active users, spend per user, trend.
- Add a cost line to every scheduled automation's log review: cost per run times runs per month, per system.
- Put a ten-minute usage review into the monthly audit from last lesson.
The routing table: which model for which job
Opus 5.5 is the default model in Claude Code on every plan since v2.1.280 (September 22, 2026), replacing Sonnet on Pro and Team Standard. That is a good default for judgment work and an expensive one for bulk work. The current lineup, with API list prices per million input / output tokens as of September 2026:
- Fable 5.1 (
fable, $10 / $50): the hardest long autonomous work, or when Opus 5.5 at higher effort still falls short. Not a default on any plan; pick it with/model fable. On some plans it bills to usage credits, and Claude Code asks for consent first. - Opus 5.5 (
opus, $4 / $20): architecture, reviews, judgment calls, debugging, anything a client or prospect reads. Anthropic's own guidance: start here for most workloads. - Sonnet 5 (
sonnet, $2 / $10, now the permanent price): the fast daily driver for routine edits, internal drafts, and reports, and the cost saver when someone keeps hitting limits. Theopusplanalias plans on Opus and executes on Sonnet. - Haiku 4.5 (
haiku, $1 / $5): bulk mechanical work - tagging, extraction, formatting, file triage, search subagents. Its retirement floor is "not sooner than October 15, 2026", so reference it by thehaikualias in Claude Code and keep the model ID in one config value in your scripts.
/effort is the second dial, and on the newest models it is the main one: thinking cannot be turned off on Opus 5.5 or Fable, so you lower effort instead. Levels are low, medium, high, xhigh, and max. Opus 5.5 defaults to medium, most others to high. /effort saves per model (press s for this session only), /effort auto clears it, and admins can cap it with maxEffortLevel. Haiku 4.5 does not support effort.
---
name: tag-titles
description: Normalize job titles in a CSV into our seniority and function buckets. Use when a list file needs title tagging before scoring.
model: haiku
allowed-tools: Read, Write, Bash(python3 scripts/check_tags.py *)
---
THE RULE: tag from the title only, never guess from the company.
...---
name: list-researcher
description: Reads company websites and returns a one-line verdict per row. Use for bulk site checks.
tools: Read, WebFetch
model: sonnet
---
Return one JSON object per company. Never draft outreach.Pipelines outside Claude Code: Batch, caching, cheap models
Your scheduled systems call the API directly, and there the levers are bigger than model choice:
- Batch: 50% off every model for anything that can wait (Haiku 4.5 drops to $0.50 / $2.50, Opus 5.5 to $2 / $10). Nightly scoring, enrichment passes, and report generation almost always can.
- Prompt caching: cache reads cost 0.1x input on most models, 0.05x on Opus 5.5, and 0.025x on Fable 5.1. Put the long system prompt and rubric first and cache them. For bursty jobs use the 1-hour TTL: its write costs 2x, but a 5-minute cache expires between bursts and every burst pays the write premium again.
- Shrink the input: strip what the rubric does not read. Our qualifier drops full job history (about 15KB per profile, the single biggest cost driver), recommendations, and "people also viewed" before the call.
- Log four numbers per call:
input_tokens,cache_read_input_tokens,cache_creation_input_tokens,output_tokens. A cache that never hits shows up here and nowhere else.
Cheap open models through OpenRouter (DeepSeek, GLM-5.3, MiniMax) have a place, and it is narrow: bulk, stateless, schema-checked steps with no tools and nothing a prospect reads - classification, ICP tagging, field extraction, title normalization, internal summaries. Enforce a JSON schema, validate every row, and fail the row loudly on a parse error. Keep on Claude: agent loops with tools, client-facing text, judgment calls that decide money or relationships, and the verification pass.
Compare against the cheapest Claude option, not the default one. As of September 2026, GLM-5.3 at $1.40 / $4.40 plus OpenRouter's 5.5% fee costs about the same as live Haiku 4.5 and more than Batch Haiku. The genuinely cheap tier is the Flash class, such as DeepSeek Flash V4.1 or MiniMax-M3 at $0.30 / $1.20. Our own outbound qualifier moved from Sonnet 4.6 to GLM 5.2 via OpenRouter in June and saved a lot; on today's price cards, Haiku 4.5 with a cached system prompt or a Flash model would be cheaper still. Our newer engager qualifier runs on cached Haiku: $0.80 to judge 2,603 people.
- A/B 50-100 real rows against Claude and compare verdicts before switching anything.
- Pin OpenRouter's
orderandquantizations: its default routing favors the cheapest host, which may serve a quantized model. - Privacy: DeepSeek's first-party API stores personal data in China and trains on it unless you opt out. For client data, use a US host of the same open weights or OpenRouter with
zdr: trueplus anonlyallowlist, and check the client's contract first. - Claude Code itself stays on Claude. Anthropic's docs say they do not support routing Claude Code to non-Claude models through any gateway, and a gateway credential switches you from subscription limits to per-token billing.
The spend policy, written down
Routing is half of it. The other half is who may spend what, and when. Commit this file to the brain and install it into every project with your foundation skills, so every session and every teammate follows it:
# Spend policy
1. A key being available is not permission to spend it.
Keys load into every shell. Before any run that costs money,
the owner asked for that run, or you ask first.
2. Before any batch or loop: state per-unit cost x units = total.
"0.4 cents x 12,000 rows = $48. Proceed?" Stating is not asking.
3. Every spending script has --dry (prints the bill and stops)
and a hard MAX_CREDITS cap. Abort rather than eat the month.
4. New schedules ship OFF behind an env var until one manual run
is verified end to end.
5. Two wallets: in-session work runs on the subscription;
the API key is only for unattended systems.
6. Price model swaps on measured output tokens, not the rate card.
7. Price the whole pipeline first. Optimize the line that dominates.
8. Every usage row records provider + model + the four token counts.Spending less without getting less
- Context hygiene is cost hygiene: /clear between tasks, skills instead of a fat CLAUDE.md, path-gated rules. Every junk token is paid for, usually more than once.
- Disable unused MCP servers. Tool search already keeps connected-but-idle tools cheap, but a server nobody uses is still surface area.
- Subagents for heavy research keep your main session lean, and they can be pinned to Sonnet or Haiku.
- Dynamic workflows (
/workflows) can fan out to dozens of agents. Keep the size guideline small for routine work and check the run's cost before saving it for reuse.
Right-sizing plans as adoption grows
Plan fit drifts as adoption grows - in both directions. The signals and the moves:
- Someone hitting Pro limits weekly: try Sonnet 5 as their daily driver first, then upgrade to Max or a Team Premium seat. The upgrade usually costs less than the productivity lost to waiting on limits.
- Someone barely using their seat: pair them with a champion before downgrading - it is almost always an enablement gap, not a need gap.
- Automations growing:
claude -pand Agent SDK runs on a subscription still draw from the same subscription limits. Anthropic announced a separate Agent SDK credit for June 15, 2026, then paused it that day, and it has not taken effect. Heavy scheduled workloads are cleaner on an API workspace with explicit spend limits, where each system's cost is its own line item. - Review plan fit quarterly against the dashboard, not annually against a guess.
Do this now
Sources and further reading
Want us to set it up with you, end to end?
Three one-on-one sessions. We train you on your real stack and build your first agents together, until you can run it yourself. You keep everything.