Foundation · 02 · How Claude Code thinksLesson 2 of 5
Models, effort and cost
- Route each task to the right model and effort level, and know what each costs
- Pin the model where the work is defined: skills, subagents, workflows
- Price a batch job before it runs and keep the receipt honest
Four models, one routing rule
You leave this lesson with your defaults set on purpose, one subagent pinned to a cheap model, and a habit: price the job before you run it. Start with the lineup. Anthropic's own guidance, verbatim in spirit: start with Opus 5.5 for most workloads; use Fable 5.1 for demanding reasoning and long-horizon agentic work, or when Opus 5.5 at higher effort still falls short. Prices below are per million tokens, input / output, on the Claude API as of September 2026.
- Fable 5.1 (
claude-fable-5-1) - $10 / $50, 1M context, cache reads $0.25. The hardest long-horizon autonomous work: a multi-hour build, deep research, a plan that must survive contact with fifty files. Never a default;/model fableopts in, and where it bills to usage credits Claude Code asks for consent first. - Opus 5.5 (
claude-opus-5-5) - $4 / $20, 1M context, cache reads $0.20. The default on every plan and the right starting point: architecture, judgment calls, planning, anything a client reads. Anthropic's line: it performs at Fable 5.1 level on most work and costs 40% less to run than Opus 5. - Sonnet 5 (
claude-sonnet-5) - $2 / $10, 1M context. The fast daily driver when speed matters more than depth, and the cost saver for high-volume interactive work. The price started as a promo and became permanent in August 2026. - Haiku 4.5 (
claude-haiku-4-5) - $1 / $5, 200K context, cache reads $0.10. Bulk mechanical work: classification, extraction, formatting, the engine inside cheap subagents. It is the only current Haiku and its retirement floor is October 15, 2026, so pin it by thehaikualias, not a dated ID, and expect a successor.
Aliases move with the lineup: opus means Opus 5.5 today and the next Opus tomorrow; sonnet is Sonnet 5; best picks Fable where it is available; opusplan plans on Opus and executes on Sonnet. Skills and scripts should name aliases so they never pin a retired model. Billing records should carry the exact model ID so the cost is right.
/effort: the dial most people never touch
/effort sets how much reasoning the model spends before answering: low, medium, high, xhigh, max. The default is high on most models and medium on Opus 5.5. Thinking cannot be switched off on Opus 5.5 or the Fable models; lowering effort is the lever. The setting saves per model (press s for session-only), /effort auto clears your saved level, and admins can cap the ceiling with maxEffortLevel.
- low - renames, formatting, moving files, anything a rewind fixes in two keystrokes.
- medium (the Opus 5.5 default) - the daily driver setting for operator work.
- high - planning a multi-file change, analysis you will act on, anything with a number in it a client sees.
- xhigh and max - a plan that must be right the first time, a hard debugging session. Expensive and slow; use it on purpose.
/effort ultracode- xhigh plus automatic dynamic workflows (next section): Claude plans a multi-agent workflow for every substantive task. Only for sessions that are all big tasks.
/fast toggles fast mode: the same Opus 5.5, faster output, billed at $8 / $40 instead of $4 / $20 (as of September 2026, still a research preview; also on Opus 5 and Opus 4.8 at $10 / $50). On a subscription plan fast mode bills to usage credits only, never to your plan limits. It is a latency dial, not a quality dial. Use it when you are waiting on the screen, never for background jobs.
Dynamic workflows: when one session is not enough
Put the word ultracode in a prompt, or ask for "a workflow", and Claude writes a small orchestration script (agent(), pipeline(), parallel(), phase()) and runs dozens of agents in the background: up to 1,000 per run, 16 at a time by default. /workflows monitors runs (p pauses, x stops, s saves the script to .claude/workflows/ so the next run is one command). The bundled /deep-research <question> is a workflow you can try today. Runs pause at usage limits and resume on their own.
Available on every paid plan; on Pro it is off until you turn it on in /config, and the size guideline there is small. The default guideline elsewhere is medium, under ten agents. Plugins can ship saved workflows, which is how a team turns 'research every account on this list' into a repeatable command.
Pin the model where the work is defined
Switching /model by hand is for exploration. Production routing lives in files: a skill's frontmatter (model and effort), a subagent's frontmatter, and permission rules. That way the cheap mechanical step is always cheap and the judgment step is never accidentally downgraded, whoever runs it.
---
name: lead-tagger
description: Tag one lead export against the ICP rules in brain/icp.md.
Bulk and mechanical. Use for any CSV of leads that needs an in/out column.
model: haiku
tools: Read, Write
---
Read brain/icp.md first. Then for every row in the CSV you are given,
add an `icp_fit` column with one of: fit, no-fit, unclear, and a
`reason` column under ten words. Write <name>-tagged.csv alongside.
Report row counts in, out, and per label. Never delete a row.- Subagents run in the background by default and report back when done. A fork subagent (on by default since August 2026) inherits your conversation and prompt cache, so it starts already knowing the task.
- The
/agentswizard was removed in July 2026. Create a subagent by editing.claude/agents/or by asking Claude: "write me a subagent that ...". It knows the format. CLAUDE_CODE_SUBAGENT_MODELsets a default for subagents that do not pin one; addCLAUDE_CODE_SUBAGENT_MODEL_FORCE=1and that model overrides every pin, useful for a cost-capped shared machine.- Deny and ask rules can match a tool parameter:
Agent(model:opus)indenyblocks Agent calls that request theopusalias. It matches the model Claude asks for in the call, not a subagent's frontmatter pin, so pair it with the env vars above to cap cost. - Rule of thumb: Haiku for subagents that read and tag, Opus for the reviewer, and never let a Haiku-pinned agent write anything a client reads.
Outside Claude Code: cheap open models in your own scripts
Claude Code and the Agent SDK run on Claude models only. Anthropic's docs are explicit that non-Claude models are not supported through any gateway, so this section is about the pipelines you write yourself in week four and the vertical tracks, where you pick the model per step. There, cheap open models through OpenRouter (DeepSeek, GLM-5.3, MiniMax) have a real place, and a narrow one.
- Send to a cheap open model: work that is bulk, stateless and schema-checked, with no tools and no text a prospect reads. Classification and ICP tagging, field extraction, title normalization, dedup hints, internal summaries. Enforce a JSON schema, validate every row, and fail the row loudly on a parse error.
- Keep on Claude: every agent loop with tools, anything a prospect or client reads, judgment calls that decide money or relationships, and the verification pass.
- Compare against the cheapest Claude option, not the default one. As of September 2026 Haiku 4.5 is $1 / $5, $0.50 / $2.50 on the Batch API (50% off, results within 24 hours), with cache reads at $0.10. GLM-5.3 at $1.40 / $4.40 plus OpenRouter's 5.5% fee is roughly live Haiku and dearer than Batch Haiku. The truly cheap tier is the Flash class: DeepSeek Flash and MiniMax-M3 around $0.30 / $1.20.
- Our own lesson: the outbound qualifier moved from Sonnet 4.6 to GLM 5.2 via OpenRouter in June 2026 and saved real money against Sonnet. Re-priced in September, Haiku 4.5 with a cached system prompt (about $0.30 per 1,000 people on our competitor-gate call, $0.80 for a 2,603-person run) or a Flash-class model would beat it. Prices move; re-run the comparison quarterly.
- Price the whole pipeline first. On our LinkedIn signals engine the LLM step cost about $0.06 per run against about $60 a month of scraping: 0.1% of the bill. Downgrading the model there saves nothing and costs judgment quality.
- Privacy is part of the price. DeepSeek's first-party API stores data in the PRC and trains on it unless you opt out; Z.ai's API runs from Singapore. For client data use a US host of the same open weights, or OpenRouter with
zdr: trueand anonlyallowlist, and read the client's contract before any of it.
Keep the receipt honest
/usageshows your plan's usage bars plus what consumed them: skills, subagents, MCP servers, cache misses and your heaviest loops (/costis now an alias for it)./insightsbuilds a report over your recent sessions.claude -p ... --output-format jsonreturnstotal_cost_usdplus a per-model breakdown, which is how a scheduled job reports what it cost (week three).- Record provider and model on every usage row your own systems write. Our dashboard priced every qualifier call at Sonnet rates for weeks while the default route was GLM, so the LLM line was wrong by the ratio of the two price cards. When you add a route, the accounting follows it, in the same change.
- Measure swaps, do not estimate them. Moving one of our reasoning pipelines from Opus 5 to Sonnet 5 was estimated at -40% and measured at -10% ($0.177 vs $0.196 per company) because the cheaper model wrote far more output tokens. Run 50-100 real rows through both and compare the bill and the verdicts.
- Prompt caching in your own scripts: a bursty workload lets a 5-minute cache expire between bursts, so every burst paid the write premium. A 1-hour cache write costs 2x base but survives across bursts and redeploys. Log
cache_read_input_tokensandcache_creation_input_tokensper call, or you cannot tell whether the cache is working.
Do this now
Sources and further reading
Want us to set it up with you, end to end?
Three one-on-one sessions. We train you on your real stack and build your first agents together, until you can run it yourself. You keep everything.