anfloy.AcademyBook a call

Rollout · 02 · The 30-day rolloutLesson 1 of 3

Capstone demo day

120 min working time · Week 12

By the end of this lesson you can
  • Every participant demos their shipped system
  • Peer review against the quality checklist

What demo day is for

Every participant walks out of demo day having shown a running system, holding a score, three concrete improvements, and a named reviewer - and the team has voted two systems into the shared library. Demo day is not a presentation exercise. It is the moment individual capability becomes team culture: everyone sees what everyone else built, the quality bar becomes shared and visible, and the best systems get claimed by the whole team instead of staying one person's trick.

The bar for what counts as a demo: a real system, on real company data, that ran at least once without its builder touching it. Slideware does not qualify. A scheduled job that fired this morning does.

The format

Tight structure is what keeps two hours useful for fifteen people. Run it like this:

  1. Ten minutes per person: five to demo, three for review, two for the improvement list.
  2. Demo structure (same as the foundation capstone): the chore it replaced, the system in one sentence, the live proof, the guardrails, the monthly cost.
  3. Reviewers score against the checklist below while watching - on paper, before discussion, so the loudest voice does not set the grade.
  4. Each builder leaves with three concrete improvements and one named reviewer who will re-check in two weeks.
  5. Close by voting two systems into the org skills library - the ones most worth generalizing beyond their builder.

The quality checklist

Score each system one point per line, ten possible. Eight or above is library-candidate quality.

  • Solves a real recurring chore - the builder can say how many hours per month it returns.
  • Runs without conversation: inputs from files or connections, outputs to fixed paths.
  • Verification is built in: the system checks its own work and shows evidence in its output.
  • Logs every run: timestamp, inputs seen, outputs produced, checks passed or FAILED, and the model and cost of the run.
  • Fails loudly and safely: missing input means a clear failure line, never a silent guess.
  • Permissions are minimal: dontAsk with a narrow allowlist, not a broad mode for convenience.
  • Side effects are gated: anything that sends, posts, or spends requires human invocation, and anything that leaves the building stops at a draft for a named approver.
  • Secrets are clean: keys in .env or the vault, nothing key-shaped in the repo or the skill.
  • It is committed: skill, settings, and README in git - it survives the builder's laptop.
  • Another person ran it cold, from the README alone, successfully.

Reviewing well: scope it to correctness

A known failure mode of review - human and AI alike - is that reviewers asked to find problems will always find some, and unscoped review feedback pushes systems toward over-engineering. So scope it: reviewers flag correctness, safety, and checklist gaps. They do not redesign the architecture or pile on nice-to-haves.

  • Good review note: "The dedupe step doesn't log how many rows it dropped - add the count to the run log."
  • Bad review note: "I would have built this as three separate skills with a coordinator." Maybe true; not demo day's job.
  • The question every reviewer asks last: "Would I let this run against MY data unattended?" If no, the reason is the review.

Do this now

Sources and further reading

We set it up with you

Want us to set it up with you, end to end?

Three one-on-one sessions. We train you on your real stack and build your first agents together, until you can run it yourself. You keep everything.