Internal Ops · 03 · Scheduled ops & reportingLesson 3 of 4
When things break
- Build error alerts and recovery into every job
- Debug a failed run without an engineer
Failure is a design input
Two artifacts come out of this lesson: every scheduled job you run announcing its own failures, and a one-page runbook anyone on the team can work. Every scheduled job will eventually fail: a connector token expires, a source API has a bad morning, your laptop was closed, a domain gets blocked. None of that is a crisis. The crisis is failing silently - the report that quietly stopped arriving three weeks ago and nobody noticed until the numbers mattered.
So the design goal is not "never fails." It is "never fails silently, and a non-engineer can diagnose it in five minutes." Both are achievable with patterns you can copy directly.
- Pattern one: the dead-man's switch - every run announces itself, even when it fails.
- Pattern two: the diagnosis checklist - a fixed order of checks that finds the cause without guesswork.
- Pattern three: the audit trail - a log that wrote itself, for the day you need to know exactly what happened.
The dead-man's switch
This costs one sentence per prompt and converts every failure mode into a visible event. It also pairs with the stale-data guard from the KPI lesson: the as-of timestamp catches the subtler failure where the job ran fine but on old data.
Go add the line to every scheduled job you have right now, before reading on. It is the highest-value sixty seconds in this lesson, and the checklist below assumes the signal exists.
Counts that must move
The dead-man's switch catches a job that did not run. It does not catch the worse case: a job that ran, reported success, and did zero work. Three of ours, all real, all green.
- A signals handler caught each row's exception and answered HTTP 200 with
{"received": 21, "promoted": 0}. A gitignored config folder meant the deployed container had no ICP file, so every row crashed. For a day, 77 people were marked handled with no tier and would never have been offered again. Zero promoted looks exactly like a strict gate doing its job. - A dead API key returned 401 on every scoring batch. Every person fell through as unscored, the competitor gate rejected nobody, and the pipeline posted its success summary to Slack.
- A prune reported
deleted=753having deleted nothing. Every DELETE returned 400, and the script counted attempts, not confirmations.
- Put the counts in the heartbeat, not just the verdict: rows read, rows judged, rows failed, and the ids to retry. A batch handler that swallows per-row errors must return the failures in its response, and the caller records only rows that were actually judged.
- When a success metric is zero, ask which path produced it, the success path or the crash path. They render identically. Log a 100% failure rate at ERROR, naming the ids.
- After any delete or write, read back a sample: fetch about 15 removed rows and expect 404, about 15 kept rows and expect 200. A count of attempts is not a count of effects.
- Prove a check can fail before you trust it. Inject the fault (rename the config file, revoke the test key) and watch the alert fire. A check that has never fired is an assumption.
- A rejected row is not an error. Keep it out of the retry list, or every disqualified record is re-pushed forever.
The 'report didn't run' checklist
When the Monday report is missing, work this list in order. It resolves the vast majority of incidents without an engineer.
- Identify the tier from your jobs registry: Routine, Desktop task, or CI. The diagnosis differs by tier.
- Routine? Open the run transcript at claude.ai/code/routines. Remember green status only means a clean exit - the transcript is where blocked domains, missing connector tools, and task failures actually show up. Read what it tried.
- Desktop task? Was the machine awake at trigger time? Missed runs get exactly one catch-up run on wake (covering the last 7 days) - so the report may arrive late rather than never. Check whether the catch-up already fired.
- Connector failure in the transcript? Run /mcp and reconnect - expired OAuth is the most common single cause. Then re-run the job manually.
- Cron or Actions? Check the exit code and the JSON output; your wrapper should alert on nonzero exit. While you are there, scan total_cost_usd both ways - a spike is a job stuck in a loop, and a near-zero cost on a job that normally costs something is a job that did nothing.
- Fixed it? Re-run manually, confirm the output lands, and add one line to the job's prompt or skill so this failure mode self-reports next time.
Step 6 is the one people skip. Every diagnosed failure should make the job more diagnosable - that is how a flaky script becomes an operating layer.
The audit trail
When something did run and you need to know exactly what it did, you want a log that wrote itself. A PostToolUse hook gives you one in five lines: every tool call, appended as a JSON line with a timestamp.
{
"hooks": {
"PostToolUse": [
{
"matcher": "*",
"hooks": [
{
"type": "command",
"command": "jq -c '{ts: now | todate, tool: .tool_name, input: .tool_input}' >> ~/.claude/audit-log.jsonl"
}
]
}
]
}
}- Hooks run deterministically - unlike CLAUDE.md guidance, a hook fires every time, guaranteed. That is why the audit trail is a hook and not an instruction.
- The one-liner uses jq, which a stock Mac does not ship with: brew install jq, or ask Claude to swap in a three-line python3 fallback (python3 is already on every Mac) doing the same append.
- The log captures the full tool_input of every call - which can include message bodies, customer data, anything the agent touched. Treat audit-log.jsonl as sensitive: keep it out of git and store it where your other sensitive files live.
- A ConfigChange hook can log edits to settings and skills the same way - tamper-evidence for your guardrails (this mirrors an official docs example).
- The brain repo's git history is itself an audit trail of every knowledge change: who, what, when, with diffs.
- Team and Enterprise plans add OpenTelemetry export and usage dashboards on top, but the JSONL file is enough to answer "what did the agent do on Tuesday?"
Escalate or handle?
The last skill is knowing which failures are yours and which ones need an engineer or a security review. The split is cleaner than you would expect.
- Handle yourself: expired OAuth (reconnect via /mcp), a missed run (check catch-up), a wrong output (edit the skill, re-run), a stale source (the guard already flagged it).
- Handle yourself: a routine missing a connector tool - re-check the routine's connector list; someone probably trimmed too far.
- Escalate: the same job failing three runs in a row after reconnects - something structural changed in a source system.
- Escalate: anything security-shaped - a key that may have leaked, an audit line you cannot explain, a hook that stopped firing.
Do this now
Sources and further reading
Want us to set it up with you, end to end?
Three one-on-one sessions. We train you on your real stack and build your first agents together, until you can run it yourself. You keep everything.