1. Introduction
A runbook is a documented set of procedures for handling a specific operational task or incident. In theory, every team has them. In practice, most teams have a folder full of documents that nobody reads during an incident because they're too long, too vague, too out of date, or too hard to navigate under pressure.
The goal of a runbook is not completeness. It's usability at 2am by someone who may not be the most experienced person on the team. That changes how you write it — shorter, more direct, more command-heavy, less background explanation.
This guide covers what makes a runbook actually useful, what sections to include, what to leave out, a full template with annotated examples, and a process for keeping runbooks alive over time rather than letting them go stale.
2. When and Why to Write a Runbook
Write a runbook for any operational task that:
- Happens infrequently enough that the steps aren't memorised
- Has enough steps that improvising under pressure introduces real risk
- Is owned by a rotation of engineers, not just one person
- Has caused an incident before due to a missed step or wrong assumption
- Requires coordination across teams or involves escalation paths
Common runbook categories in DevOps teams:
| Category | Examples |
|---|---|
| Incident response | Service down, high error rate, database failover, DDoS response |
| Deployment procedures | Production release process, hotfix deployment, rollback procedure |
| Infrastructure tasks | Scaling a node group, rotating credentials, certificate renewal |
| Routine maintenance | Database vacuuming, log rotation, index cleanup, cache flush |
| On-call procedures | Alert triage, escalation paths, stakeholder communication templates |
| Recovery procedures | Restoring from backup, DR failover, data recovery steps |
3. Prerequisites
Before writing a runbook, you need:
- A real scenario to document — write from a post-incident review, not speculatively
- Access to the actual commands and steps, verified to work
- A designated location where runbooks are stored and searchable (Confluence, Notion, GitHub wiki, PagerDuty)
- Agreement on the template format across your team so runbooks are consistent and predictable
4. The Runbook Structure
A practical runbook has eight sections. Each one serves a specific purpose for the person executing it under pressure. Here is the complete structure, followed by a detailed explanation of each section:
# Runbook: [Name of Issue or Task]
- ## Overview
- One paragraph: what this runbook covers and when to use it.
- ## Severity
- & Impact Who is affected, how many users, what SLO/SLA is at risk.
- ## Prerequisites
- Access, tools, and permissions required before starting.
- ## Symptoms
- What the engineer will see: alert name, error messages, metrics.
- ## Diagnosis Steps
- Ordered commands and checks to identify the root cause.
- ## Resolution Steps
- Ordered commands and actions to resolve the issue.
- ## Verification
- How to confirm the issue is resolved. Specific checks.
- ## Escalation
- Who to contact if this runbook doesn't resolve the issue.
Section 1: Overview
Two to four sentences. What is this runbook for? What system or service does it cover? What is the condition that triggers its use? Keep it to the minimum context a new on-call engineer needs to confirm they are looking at the right runbook.
## Overview This runbook covers the response procedure for the payment-api service reporting a high error rate (>5% 5xx responses over 5 minutes). It applies to the production environment only. For staging issues, see: runbook-payment-api-staging-errors.md Related alert: payment-api-high-error-rate (PagerDuty policy: payments-team)
Section 2: Severity and Impact
State the business and user impact clearly. Who is affected? What functionality is degraded or unavailable? What SLO or SLA is at risk and over what time horizon? This section helps the on-call engineer decide how urgently to escalate and whether to wake up additional people immediately.
## Severity & Impact Severity: SEV-2 (escalate to SEV-1 if error rate exceeds 20% for > 10 minutes) User impact: Payment processing is degraded or unavailable. Checkout flow fails for affected users. SLO at risk: payment-api availability SLO (99.9% / 30-day window). Each minute of >5% error rate consumes ~0.07% of monthly error budget. Stakeholders: Notify #incidents Slack channel immediately. Page payments-team-lead if not resolved within 15 minutes.
Section 3: Prerequisites
List everything the engineer needs before they can start. Be specific about access levels, tools, and credentials. An engineer who discovers mid-incident that they don't have the right kubectl context or Vault access has wasted critical minutes. Better to surface that before the steps begin.
## Prerequisites - kubectl access to the production cluster (context: eks-prod-us-east-1) - AWS CLI configured with the ops-readonly role (for CloudWatch/RDS access) - Access to #incidents Slack channel for communication - PagerDuty responder access (to acknowledge and reassign alerts)
# Verify your cluster context before starting: kubectl config current-context
# Expected: eks-prod-us-east-1
Section 4: Symptoms
Describe exactly what the engineer will observe. If your team runs Loki or ELK, include the exact log query the engineer should run — not just "check the logs". Include the alert name, the specific metric thresholds that triggered it, example error messages they will see in logs, and any dashboard panels that are relevant. This helps the engineer confirm they are responding to the right incident and not something unrelated.
## Symptoms Alert: payment-api-high-error-rate fires when: rate(http_requests_total{status=~'5..', job='payment-api'}[5m]) > 0.05 Grafana: Dashboard 'Payment API Overview' — panel 'Error Rate %' showing spike Dashboard 'Payment API Overview' — panel 'P99 Latency' may also be elevated Logs: Loki query to confirm: {namespace="production", container="payment-api"} | = "error" | rate [5m] Common error signatures: "connection refused" — downstream DB or dependency issue "context deadline exceeded" — timeout, check latency metrics "OOMKilled" — memory issue, check pod restarts
Section 5: Diagnosis Steps
This is the most important section. Write ordered steps with exact commands. Do not write prose — write commands. Every step should produce observable output that confirms whether the problem is in that layer or not. The engineer should be able to follow these steps without knowing the system in depth.
## Diagnosis Steps 1. Check pod status and recent restarts kubectl get pods -n production -l app=payment-api kubectl describe pod <pod-name> -n production | tail -30 2. Check pod logs for the error kubectl logs -n production -l app=payment-api --tail=100 --previous # Look for: error message, stack trace, timestamp of first failure 3. Check if the issue is a single pod or all pods kubectl get pods -n production -l app=payment-api # If one pod is bad: likely a code or config issue on that pod # If all pods are bad: likely a downstream dependency or config change 4. Check downstream dependencies # Test DB connectivity from inside a pod: kubectl exec -it <pod-name> -n production -- \ nc -zv payment-db.internal 5432 # Test Redis: kubectl exec -it <pod-name> -n production -- \ redis-cli -h redis.internal ping 5. Check for recent deployments kubectl rollout history deployment/payment-api -n production # If a recent deployment is visible, proceed to Resolution Step 1 (rollback)
Section 6: Resolution Steps
Write the resolution steps in the same way as diagnosis — ordered, command-first, with expected output. Where there are multiple possible resolutions depending on the diagnosis, number them and reference which diagnosis step leads to each. A branching structure is fine as long as the branch conditions are explicit.
## Resolution Steps --- If caused by a bad deployment (from Diagnosis Step 5) --- 1. Rollback to the previous deployment kubectl rollout undo deployment/payment-api -n production kubectl rollout status deployment/payment-api -n production # Wait for rollout to complete — watch error rate in Grafana --- If caused by DB connectivity failure (from Diagnosis Step 4) --- 2. Check RDS instance status aws rds describe-db-instances \ --db-instance-identifier payment-db-prod \ --query 'DBInstances[0].DBInstanceStatus' # If 'available': check security group rules and VPC routing # If 'rebooting' or 'failing-over': wait and monitor — RDS will recover 3. If connection pool exhaustion is suspected kubectl rollout restart deployment/payment-api -n production # Restarting pods forces new DB connections — short-term mitigation # File a follow-up ticket to investigate pool sizing --- If caused by memory (OOMKilled from Diagnosis Step 2) --- 4. Temporarily increase memory limit and redeploy kubectl set resources deployment/payment-api -n production \ --limits=memory=1Gi # IMPORTANT: file a ticket to update the manifest in git # kubectl set resources does not persist — it will be overwritten on next deploy
Section 7: Verification
Tell the engineer exactly how to confirm the issue is resolved. Name the specific metric, dashboard panel, or command output that represents a healthy state. 'Monitor it for a few minutes' is not a verification step — it leaves the engineer guessing about when to stand down.
## Verification 1. Error rate has returned below threshold Grafana: Dashboard 'Payment API Overview' > panel 'Error Rate %' Should show < 1% for 5 consecutive minutes 2. Pods are Running and not restarting kubectl get pods -n production -l app=payment-api Expected: all pods READY 1/1, RESTARTS count not increasing 3. Confirm with a synthetic transaction (if available) curl -X POST https://api.example.com/v1/payments/health-check Expected: HTTP 200, {"status": "ok"} 4. Confirm error budget impact Grafana: Dashboard 'SLO Overview' > 'payment-api availability' Note the error budget remaining and record in the incident ticket
Section 8: Escalation
Define clearly who to contact when this runbook does not resolve the issue. Include the contact method (Slack, PagerDuty, phone), not just a name. People change roles and leave teams — use a role or on-call rotation name rather than an individual wherever possible.
## Escalation If this runbook does not resolve the issue within 30 minutes: 1. Payments team lead PagerDuty: escalate incident to 'payments-team-lead' policy Slack: @payments-team-lead in #incidents 2. Database team (if DB issue is suspected) PagerDuty: page 'database-oncall' policy Slack: @database-team in #incidents 3. AWS Support (if RDS infrastructure issue confirmed) Support plan: Business (4hr SLA for production issues) Case URL: https://console.aws.amazon.com/support Incident commander: first responder owns the incident until handoff. Communication: post updates to #incidents every 15 minutes.
5. Complete Worked Example
Here is a condensed but complete runbook for a real-world scenario, showing how all eight sections work together. This is what a finished runbook looks like — not a template, but an actual document ready for use.
# Runbook: Kubernetes CrashLoopBackOff — payment-api Last updated: 2025-04-15 | Owner: payments-team | Severity: SEV-2
## Overview Use this runbook when the payment-api container is in CrashLoopBackOff in the production namespace. The container starts and exits repeatedly. This runbook does not apply to ImagePullBackOff — see: runbook-image-pull.md
## Severity & Impact SEV-2. Payment processing unavailable for affected pods. Escalate to SEV-1 if all pods are crashing simultaneously.
## Prerequisites - kubectl access: context eks-prod-us-east-1 - Read access to AWS Secrets Manager (for secret validation steps)
## Symptoms Alert: payment-api-pod-crash-loop (threshold: restarts > 3 in 10 min) Visible in: kubectl get pods -n production | grep payment-api
## Diagnosis Steps 1. kubectl logs <pod> -n production --previous → Read last 50 lines. Note the error and exit code. 2. kubectl describe pod <pod> -n production | grep -A5 'Last State' → Note exit code. Exit 137 = OOMKilled. Exit 1 = app error. 3. kubectl get secret payment-api-config -n production → Confirm secret exists. Missing secret = common crash cause.
## Resolution Steps Exit code 1 / app error in logs: → Check if caused by recent deploy: kubectl rollout history ... → If yes: kubectl rollout undo deployment/payment-api -n production Exit code 137 (OOMKilled): → kubectl set resources deployment/payment-api -n production \ --limits=memory=1Gi → File ticket: update memory limit in git and merge to main Missing secret: → Confirm in AWS Secrets Manager. Re-sync with External Secrets Operator. → kubectl rollout restart deployment/payment-api -n production
## Verification kubectl get pods -n production -l app=payment-api → All pods: READY 1/1, RESTARTS not increasing for 5 minutes Grafana: 'Payment API Overview' > 'Error Rate %' < 1%
## Escalation Not resolved in 20 min → page payments-team-lead via PagerDuty
6. What to Leave Out
Runbooks fail not because they are missing information but because they contain too much of the wrong kind. These things actively make runbooks worse:
- Background explanation of how the system works — this belongs in architectural documentation, not a runbook. During an incident, nobody reads it
- Step-by-step instructions for things the engineer already knows — don't document how to run kubectl or how to log into AWS. Document the specific commands for this specific issue
- Multiple alternative procedures in one document — if there are two different scenarios, write two runbooks
- Speculative steps — 'you might also want to check X' adds cognitive load without adding value. Every step should be there because it is needed for this specific issue
- Out-of-date screenshots — they become misleading faster than text. Use commands and expected output text instead
- Contact details for specific individuals — use on-call rotation names or PagerDuty policy names, not personal phone numbers
7. Keeping Runbooks Current
A runbook that is six months out of date is worse than no runbook. An engineer who follows stale steps and takes a wrong action because the commands or service topology changed has been actively harmed by the documentation.
Embed runbook review into your team's existing processes
- Post-incident reviews: update the runbook immediately after every incident that used it — add new steps, correct wrong assumptions, update commands
- Quarterly rotation: dedicate 30 minutes per quarter to reviewing runbooks for services your team owns — check that commands still work and contact details are accurate
- Deployment checkpoints: when a significant infrastructure change ships — for example, adding a new node group (see Pod Pending: No Nodes Available for what changes can break scheduling) — identify which runbooks are affected and update them (new service version, dependency change, credentials rotation), identify which runbooks are affected and update them
Metadata that runbooks must always have
| Field | Why it matters |
|---|---|
| Last updated date | Tells the reader whether to trust the content or treat it with caution |
| Owner / team | Who to ask if the runbook is wrong or incomplete |
| Reviewed by | Confirms another engineer has validated the steps |
| Related runbooks | Links to connected procedures — avoids dead ends mid-incident |
| Related alerts | Maps the runbook to the alert that triggers it for fast lookup |
| Change log | Brief history of what changed and why — critical for post-incident context |
8. Summary
A runbook that gets used during incidents shares the same characteristics: it is short, specific, command-first, and written for someone who is stressed and working fast. The structure is consistent across every runbook so engineers know where to find what they need without reading linearly.
| Section | Purpose | Keep it to |
|---|---|---|
| Overview | Confirm this is the right runbook | 2–4 sentences |
| Severity & Impact | Calibrate urgency and escalation timing | 5–8 lines |
| Prerequisites | Prevent mid-incident access failures | Bullet list |
| Symptoms | Confirm the incident matches the runbook | Alert name + key signals |
| Diagnosis Steps | Identify the root cause systematically | Numbered, commands only |
| Resolution Steps | Fix the issue with exact actions | Numbered, branched by cause |
| Verification | Confirm resolution with specific checks | 3–5 concrete checks |
| Escalation | Define next steps if the runbook doesn't work | Roles + contact method |
Start with your most-paged alert. Write its runbook the day after the next incident that triggers it. Repeat for the next most-paged alert. Within a quarter, you will have covered the scenarios that matter most — and your on-call rotation will be noticeably less stressful.