Incident Response and Chaos
Overview
Section titled “Overview”- Primary reference: Site Reliability Engineering (Google, free): ch. 12 Effective Troubleshooting, ch. 14 Managing Incidents, ch. 15 Postmortem Culture
- Supplementary: The Site Reliability Workbook (free), ch. 9 Incident Response and ch. 10 Postmortem Culture; Chaos Engineering by Rosenthal and Jones (O’Reilly); Kubernetes docs, Debug Applications (free)
- Prerequisites: Cloud Native, Observability
- Estimated time: one drill per pass of the course, 1 to 2 h each; the full set in Pass 11
Key Takeaways
Section titled “Key Takeaways”- Incidents are learned by causing them on purpose: a drill injects a known fault into your own running system and grades how you detect, mitigate, and record it.
- Mitigate first (roll back, fail over, shed load), find the root cause second, write it down third.
- A runbook is written before the incident, for the person who did not build the system; a postmortem is written after, without blame.
- Detection moves from a person noticing (Pass 1) to burn-rate alerts (Pass 7), and
ol drill endmeasures time to detect and time to mitigate from Prometheus.
How to Study
Section titled “How to Study”- Do each drill when its pass reaches it, against your own kind cluster, with the time limit on.
- Write the runbook during the drill, not after: the commands you actually typed are the Diagnosis section.
- Read one chapter of the SRE book after each drill and compare its advice with what you did.
Concepts & Techniques
Section titled “Concepts & Techniques”Core Insight
Section titled “Core Insight”Every system fails; what differs is how long it stays failed. Time to mitigate is mostly time spent working out what is wrong, so the work that pays is making the next diagnosis shorter: telemetry that shows the cause, runbooks that say where to look, and rollouts that can be undone in one command.
1. Triage on Kubernetes
Section titled “1. Triage on Kubernetes”Key ideas:
- Outside in: which workload is unhealthy, why its last container exited (exit code and reason), what its previous log says, what changed (rollout history).
- Exit codes: below 128 the program chose to exit;
128 + nmeans signaln(137 SIGKILL, usually OOMKilled; 143 SIGTERM). - CrashLoopBackOff: restarts wait 10 s, doubling to 5 minutes.
2. Mitigation
Section titled “2. Mitigation”Key ideas:
- Roll back the change that started it (
kubectl rollout undo,helm rollback), then fix forward through the chart. - Readiness and rollout strategy decide whether a bad rollout becomes an outage.
3. Learning from incidents
Section titled “3. Learning from incidents”Key ideas:
- Blameless postmortems: Summary, Impact, Timeline, Root cause, Detection, Resolution, Action items.
- Action items that prevent, detect, or mitigate the next occurrence, each with an owner.
Chapters
Section titled “Chapters”| # | Module | Chapter | Kind | Pass |
|---|---|---|---|---|
| 1 | ops.00 | First drill: the engine crashloops | drill | 1 |
| 2 | ops.01 | A decode worker dies mid-stream | drill | 7 |
| 3 | ops.02 | The durable server is SIGKILLed in a loop during a corpus build | drill | 8 |
| 4 | ops.03 | A poison task lands in the dead-letter queue | drill | 8 |
| 5 | ops.04 | Drill: KV v2 migration | drill | 11 |
| 6 | ops.05 | Drill: API v2 migration | drill | 11 |
| 7 | ops.06 | Drill: breaking dependency upgrade | drill | 11 |
| 8 | ops.07 | Drill: performance regression bisect | drill | 11 |
| 9 | ops.08 | Drill: data incident | drill | 11 |
| 10 | ops.09 | Drill: noisy neighbor | drill | 11 |
| 11 | ops.10 | Drill: runaway agent | drill | 11 |
| 12 | ops.11 | Drill: event log disk full | drill | 11 |
| 13 | ops.12 | Drill: Raft partition (optional) | drill | 11 |
Connections to Other Tracks
Section titled “Connections to Other Tracks”| Track | Connection |
|---|---|
| Observability | the SLOs and burn-rate alerts that detect what these drills inject |
| Cloud Native | the Kubernetes objects every drill acts on |
| Software Craftsmanship: documentation | runbooks and decision records |
Company Relevance
Section titled “Company Relevance”| Company | Practice |
|---|---|
| SRE incident command and blameless postmortems | |
| Netflix | chaos engineering against production |
| Any on-call team | runbooks linked from alerts; game days |