Legacy systems rarely fail in exciting ways. They fail in quiet, expensive ways: a stuck job queue, a certificate that expired, a disk that slowly filled, a partner API that changed its behavior. When only one person knows the fix, you do not have a system, you have a single point of failure.
A runbook is simply the set of instructions that turns “we think we know” into “we can do it again.” The problem is that teams often picture a runbook as a big documentation project, and then never start.
The minimum viable runbook (MVR) is a lighter approach. You document only the operations that actually keep the service alive: how to recognize common failures, how to restore service safely, and how to confirm it is healthy afterward.
Why a minimum viable runbook works
Runbooks fail when they aim for completeness. A legacy application has years of tribal knowledge, half-remembered “do not touch that” rules, and configuration that has drifted over time. Trying to document everything is a guarantee that nothing gets finished and what does get written is outdated.
An MVR works because it is optimized for two real moments: an incident and an on-call handoff. In both cases, the reader needs a small set of high-confidence steps, not a history lesson. The MVR is also easy to maintain because it is short enough to review after each meaningful change.
Think of it as a contract between the system and the team: if the system emits these symptoms, the team takes these actions, and the outcome is measured with these checks.
What to include (and what to skip)
A practical runbook has a predictable shape. Readers should be able to scan it quickly during an incident and find the one section that matches what they are seeing.
The 15-minute rule
Include anything that would take a capable engineer more than 15 minutes to figure out from scratch under pressure. Skip anything that is obvious, discoverable, or already documented elsewhere in a reliable place. The goal is to remove ambiguity, not to duplicate your repository README or cloud provider docs.
Here is a compact outline that works well for small teams:
Runbook: Service Name
1) What "healthy" looks like (2-3 signals)
2) How to access it (dashboards, logs, shells)
3) Top alerts and what they usually mean
4) Safe recovery playbooks (step-by-step)
5) Data safety notes (what not to delete, invariants)
6) Escalation + owners (who to call, when)
7) Post-incident checklist (confirmations, follow-ups)
Some guidance on each part:
- Healthy signals: Choose a few signals you can quickly check. Examples include “requests succeed,” “queue depth stable,” “scheduled jobs complete,” or “error rate below a threshold.”
- Access: Write down where logs live, how to get a shell if needed, and which environment is which. Include the “least privilege” path first.
- Top alerts: Focus on the handful that pages people or causes customer-facing impact. If an alert has paged twice, it deserves a runbook entry.
- Recovery playbooks: Prefer reversible actions. If a step could delete data, make it explicit and add verification steps before and after.
- Data safety notes: Document invariants like “do not rerun job X without setting flag Y” or “table Z is append-only.”
- Escalation: Include “when to wake someone” conditions, not just names. Conditions reduce hesitation and prevent unnecessary escalations.
Key Takeaways
- Keep the runbook short enough that it gets used during real incidents.
- Document the “healthy” baseline first, then add only the alerts and fixes that repeatedly matter.
- Write recovery steps as reversible actions plus explicit verification.
- Make ownership and escalation rules unambiguous so incidents do not stall.
How to write it in one afternoon
You do not need a perfect understanding of the system to write an effective MVR. You need a focused process that captures what is already known, tests it, and leaves clear TODOs where certainty is missing.
Use this sequence:
- Pick one service boundary. Choose a single application or job family, not “the entire platform.” If you are unsure, pick the component that pages people.
- Write “healthy” checks. Spend 15 minutes listing how you know it is working. If you cannot define health, the rest of the runbook will be vague.
- List the top five failures. Pull from memory, tickets, and on-call notes. If you have none, ask: “What breaks after deploys?” and “What breaks on Mondays?”
- Draft recovery steps. For each failure, write the smallest safe action that moves the system toward healthy. Add a verification step immediately after each action.
- Do a desk-check. Have someone who did not write it read it and point out missing assumptions. Fix anything that requires mind-reading.
- Run a low-risk rehearsal. You are not creating an outage. Instead, verify access steps, confirm dashboards exist, and validate that the described commands or UI paths are real.
Copyable checklist for your MVR draft
- Service purpose in one sentence
- Links or pointers to logs and metrics (names matter)
- Three “healthy” signals with how to check each
- Five common alerts with a one-line meaning
- For each alert: steps, rollbacks, and verification
- Two “do not do this” warnings
- Escalation conditions and primary owner
- Where to record incident notes for follow-up
A concrete example: the nightly billing job
Imagine a small SaaS company with a legacy billing pipeline. Every night, a scheduled job aggregates usage, generates invoices, and posts them to a payment processor. It is business-critical, but only one engineer has touched it in years.
An MVR for this job might define health like this:
- The job starts within 10 minutes of schedule.
- It processes all accounts (count roughly matches active customers).
- It ends with “COMPLETED” status and a stable error count (near zero or within known bounds).
Then it documents the top failure modes with crisp actions:
- Symptom: job status is “RUNNING” for more than 2 hours. Likely cause: deadlock in a reporting query. Action: pause new billing runs, take a snapshot of job parameters, then restart the job with a “resume” option if supported. Verify: processed account count increases steadily after restart.
- Symptom: spike in “payment API 429” errors. Likely cause: rate limits. Action: switch to a smaller batch size and longer delay setting for retries. Verify: error rate falls and completion estimate stabilizes.
- Symptom: invoices generated but not emailed. Likely cause: email queue disconnected. Action: reroute to a “send later” queue and notify support to set expectations. Verify: email backlog drains and invoice status updates.
Notice what is not included: detailed schema docs, the full history of why the pipeline exists, and a full tutorial on the payment processor. The runbook is about restoring service and protecting data integrity.
Common mistakes (and quick fixes)
Most runbooks are not wrong, they are unusable under stress. Here are frequent failure points and how to correct them.
- Too many links, not enough steps. Replace “see dashboard” with “check metric X; if above Y, do Z.” Links can stay, but they cannot be the only content.
- No verification steps. Every action needs a “how you know it worked” check. Otherwise responders over-correct and create a second incident.
- Assumes privileged access. If the instructions require an admin account, include an alternative path or document the access request process and who approves it.
- Unclear stop conditions. Document when to stop trying and escalate. Example: “If two retries fail or errors involve data corruption, escalate to owner.”
- Stale screenshots or UI paths. Prefer stable identifiers like system names, log streams, and alert titles. If you must reference UI, write the text labels, not pixel-perfect directions.
When NOT to rely on a runbook
Runbooks are a tool, not a substitute for engineering investment. There are cases where writing more documentation is not the best next step.
- When the system is actively changing daily. If the architecture is in motion, pause and define stable interfaces first. Otherwise the runbook becomes stale immediately.
- When failures indicate deeper design flaws. If you constantly “fix” issues caused by missing backpressure, no isolation between tenants, or unsafe deployments, prioritize technical remediation. Document the emergency steps, but do not accept them as normal operations.
- When the team lacks basic observability. If you cannot tell healthy from unhealthy, add minimal monitoring first. A runbook that says “check the logs” without saying which logs or what patterns to look for is not a runbook.
A helpful framing: the MVR reduces operational risk quickly, but it should also surface where the system needs stabilization work. Treat the “frequent incident” section as an input to your engineering roadmap.
Conclusion
The minimum viable runbook is a small investment that pays off every time someone new joins the on-call rotation, every time the primary expert is unavailable, and every time an incident hits at the worst possible moment. Keep it short, focused on recovery, and anchored in verification.
If you want a simple maintenance routine, add one line to your change checklist: “Does this alter the runbook?” That single habit keeps the document alive.
FAQ
Where should we store the runbook?
Put it where responders already look during incidents. For many teams, that is a repository folder near the service, an internal wiki with good search, or a ticketing knowledge base. The best location is the one that is easy to update and easy to find.
How detailed should recovery steps be?
Detailed enough that a capable engineer can follow them without guessing. Include prerequisites, safe defaults, and verification. Avoid long narratives. If a step is risky, say why and provide a safer alternative if possible.
How do we keep it from becoming stale?
Attach runbook updates to real events: incidents, alert changes, and deployments that modify operational behavior. A small rule helps: if an incident required a Slack thread longer than a page, update the runbook the same week.
Do we need a runbook for every service?
No. Start with the services that page humans, handle money, or have high customer impact. Expand only when you feel the operational pain, and keep each runbook scoped to a clear service boundary.