Skip to content

Incident Response and Chaos

  • Incidents are learned by causing them on purpose: a drill injects a known fault into your own running system and grades how you detect, mitigate, and record it.
  • Mitigate first (roll back, fail over, shed load), find the root cause second, write it down third.
  • A runbook is written before the incident, for the person who did not build the system; a postmortem is written after, without blame.
  • Detection moves from a person noticing (Pass 1) to burn-rate alerts (Pass 7), and ol drill end measures time to detect and time to mitigate from Prometheus.
  • Do each drill when its pass reaches it, against your own kind cluster, with the time limit on.
  • Write the runbook during the drill, not after: the commands you actually typed are the Diagnosis section.
  • Read one chapter of the SRE book after each drill and compare its advice with what you did.

Every system fails; what differs is how long it stays failed. Time to mitigate is mostly time spent working out what is wrong, so the work that pays is making the next diagnosis shorter: telemetry that shows the cause, runbooks that say where to look, and rollouts that can be undone in one command.

Key ideas:

  • Outside in: which workload is unhealthy, why its last container exited (exit code and reason), what its previous log says, what changed (rollout history).
  • Exit codes: below 128 the program chose to exit; 128 + n means signal n (137 SIGKILL, usually OOMKilled; 143 SIGTERM).
  • CrashLoopBackOff: restarts wait 10 s, doubling to 5 minutes.

Key ideas:

  • Roll back the change that started it (kubectl rollout undo, helm rollback), then fix forward through the chart.
  • Readiness and rollout strategy decide whether a bad rollout becomes an outage.

Key ideas:

  • Blameless postmortems: Summary, Impact, Timeline, Root cause, Detection, Resolution, Action items.
  • Action items that prevent, detect, or mitigate the next occurrence, each with an owner.
#ModuleChapterKindPass
1ops.00First drill: the engine crashloopsdrill1
2ops.01A decode worker dies mid-streamdrill7
3ops.02The durable server is SIGKILLed in a loop during a corpus builddrill8
4ops.03A poison task lands in the dead-letter queuedrill8
5ops.04Drill: KV v2 migrationdrill11
6ops.05Drill: API v2 migrationdrill11
7ops.06Drill: breaking dependency upgradedrill11
8ops.07Drill: performance regression bisectdrill11
9ops.08Drill: data incidentdrill11
10ops.09Drill: noisy neighbordrill11
11ops.10Drill: runaway agentdrill11
12ops.11Drill: event log disk fulldrill11
13ops.12Drill: Raft partition (optional)drill11
TrackConnection
Observabilitythe SLOs and burn-rate alerts that detect what these drills inject
Cloud Nativethe Kubernetes objects every drill acts on
Software Craftsmanship: documentationrunbooks and decision records
CompanyPractice
GoogleSRE incident command and blameless postmortems
Netflixchaos engineering against production
Any on-call teamrunbooks linked from alerts; game days