Skip to content

First drill: the engine crashloops

Moduleops.00 · drill · ops, docs · Pass 1 · 1 to 2 h
You builddocs/runbooks/engine-crashloop.md (Symptoms, Diagnosis, Mitigation) and docs/postmortems/<date>-engine-crashloop.md, and you restore the service
Contractthe drill spec course/drills/engine-crashloop/drill.toml (what is injected and what ol drill end checks: section 4)
Testsgraded by ol drill end; no course tests to read (section 4 lists every check)
Needsdep.00 (your engine and gateway on kind); you triage what L10.0 and gw.00 built, with the kubectl of lang.07
Used byno call site: a drill. Its runbook is the first of the runbooks craft.10 requires per alert
MilestoneMS-P1
Optional depthSite Reliability Engineering, ch. 12 Effective Troubleshooting, ch. 14 Managing Incidents, ch. 15 Postmortem Culture (free); Kubernetes docs, Debug Running Pods (free)
  • A crashlooping pod tells you why it died in three places: Last State and the exit code in kubectl describe, and the log of the previous run in kubectl logs --previous.
  • Mitigate first, then find the root cause: kubectl rollout undo restores service in seconds; the durable fix goes through the chart.
  • Whether users see an outage depends on your rollout strategy and readiness probe, not only on the bug.
  • A runbook is a how-to guide written for the next person at 3 a.m.: symptoms they will see, commands that narrow the cause, steps that restore service.
  • A blameless postmortem has fixed sections, and ol drill end checks that yours does.
Terminal window
ol drill list # engine-crashloop is ops.00
ol drill start engine-crashloop # injects the fault, prints only the page
# ... triage with kubectl, restore service, write the runbook and the postmortem ...
ol drill end # grades: resolved, runbook, postmortem, time limit
ol drill reset # replays the undo journal; always run it after `end`

Your tracer runs on kind: one engine pod behind one gateway pod, deployed by your Helm charts (dep.00). There is no Prometheus and no alert yet (they arrive in Pass 7), and the engine has a single replica. So when a rollout goes wrong, nothing pages you except a user, and nothing tells you why except kubectl. This drill makes that happen on purpose: a change to the engine Deployment points --model-dir at a directory that does not exist. The engine exits on startup, Kubernetes restarts it again and again, and depending on how your chart rolls out, the gateway’s /readyz turns 503 and every completion fails. You will find the cause by hand, restore the service, and write the runbook and the postmortem that make the second occurrence a five-minute fix.

The objects involved. A Deployment declares “run N copies of this pod template”. It owns ReplicaSets, one per version of the template; each ReplicaSet creates Pods. Changing anything in the template (an image, an arg, an env var) creates a new ReplicaSet: that is a rollout, and kubectl rollout history lists each one as a revision. A Service sends traffic only to pods that are Ready.

Restart policy and back-off. A Deployment’s pods have restartPolicy: Always. When the container exits, the kubelet restarts it in place, but waits longer each time so a broken container does not spin the node. The pod’s status shows CrashLoopBackOff while it waits.

SymbolMeaningUnit
kkthe restart number, k=1,2,…k = 1, 2, \dotscount
dkd_kthe wait before restart kkseconds
rrhow long the container runs before it diesseconds
sks_kwhen attempt kk starts; s0=0s_0 = 0 is the first startseconds after the fault
tinj,tdet,tmitt_{inj}, t_{det}, t_{mit}when the fault was injected, noticed, and mitigatedclock time

With the default kubelet settings, the wait doubles from 10 s and is capped at 5 minutes, and it resets once a container has run for 10 minutes:

dk=min⁡(10⋅2k−1, 300),sk=sk−1+r+dk.d_k = \min(10 \cdot 2^{k-1},\ 300), \qquad s_k = s_{k-1} + r + d_k .

So the RESTARTS column grows fast at first and then about once every 5 minutes. A restart count that is still climbing slowly does not mean the problem is going away.

Exit codes say who stopped the process. A code below 128 is the program’s own choice: 1 usually means it detected a problem and quit, and its last log lines say which. A code of 128 + n means it died from signal nn: 137 is SIGKILL (9), which in a pod almost always means the kernel killed it for exceeding its memory limit (Reason: OOMKilled); 143 is SIGTERM (15), a normal shutdown.

Logs of the dead. kubectl logs shows the current attempt, which in a crashloop is often empty or not started. kubectl logs --previous shows the attempt that just died, which is where the reason is.

Readiness decides whether users notice. The default Deployment strategy is RollingUpdate with maxSurge: 25% and maxUnavailable: 25%, rounded up and down. With one replica that is one extra pod and zero unavailable: the new pod is created first and the old one is removed only when the new one is Ready. So:

Your engine chart hasDuring a bad rolloutUsers see
a readiness probe on /healthz (health port)the new pod never becomes Ready; the old pod keeps serving; the rollout stallsnothing yet, but the next restart of the old pod is an outage
no readiness probea container counts as Ready the moment it starts, so the old pod is removed during the new pod’s first second alive503 no_capacity on every completion
strategy: Recreatethe old pod is removed first, always503 no_capacity on every completion

Either way the Deployment is broken until you act. The drill grades the service you restore, not which row you were in.

Mitigate, then fix, then learn. During an incident the first goal is to stop user impact, even if you do not yet understand the cause. Rolling back the change that started it is the fastest safe action, because the previous version is known to work. The root cause and the durable fix come after, and the postmortem comes last.

The cluster drifts from the chart. Helm renders your chart into objects and applies them. A kubectl patch or kubectl edit changes the live object without changing the chart, so the cluster no longer matches what you would deploy. The durable fix is always in the chart, applied with helm upgrade.

Runbooks and postmortems are different documents. A runbook is a how-to guide for one symptom: what you see, how to narrow the cause, how to restore service. It is written before the next incident and kept current. A postmortem is the record of one incident: what happened, when, why, and what will change. It is blameless: it names systems and processes, never a person at fault, so people report what really happened. Its sections are fixed: Summary, Impact, Timeline, Root cause, Detection, Resolution, Action items.

Time to detect, time to mitigate. TTD=tdet−tinj\text{TTD} = t_{det} - t_{inj} and TTM=tmit−tinj\text{TTM} = t_{mit} - t_{inj}. Pass 1 has no alerts, so TTD is however long it takes a person to notice. From Pass 7, burn-rate alerts (obs.03) cut it to minutes, and ol drill end measures both from Prometheus.

A different incident with the same shape, so you can check every number: someone lowers the engine’s memory limit to 8 MiB. The engine is killed while loading its model, after running r=2r = 2 s each time. Your chart has no readiness probe, so users see the outage at once. The fault lands at 10:00:00.

The restart schedule. Apply dk=min⁡(10⋅2k−1,300)d_k = \min(10 \cdot 2^{k-1}, 300) and sk=sk−1+r+dks_k = s_{k-1} + r + d_k:

kkdkd_k (s)sks_k (s)how sks_k was computedclock time
00first start10:00:00
110120+2+100 + 2 + 1010:00:12
2203412+2+2012 + 2 + 2010:00:34
3407634+2+4034 + 2 + 4010:01:16
48015876+2+8076 + 2 + 8010:02:38
5160320158+2+160158 + 2 + 16010:05:20
6300622320+2+min⁡(320,300)320 + 2 + \min(320, 300)10:10:22

At 10:03:00 (t=180t = 180 s) kubectl get pods shows RESTARTS 4: attempts 1 to 4 started before 180 s, attempt 5 starts at 320 s.

The timeline you would write.

TimeEventEvidence
10:00:00memory limit lowered to 8 MiB by a manual kubectl patchkubectl rollout history: new revision
10:00:02first killLast State: Terminated, Reason: OOMKilled, Exit Code: 137
10:01:10a user reports 503 no_capacitygateway response
10:03:00triage: engine pod CrashLoopBackOff, RESTARTS 4; exit 137 points at memory, not at the programkubectl get pods, kubectl describe pod
10:04:40kubectl rollout undo deploy/<system>-engine
10:05:00new pod Ready; completions stream againkubectl rollout status, curl

The two numbers. tinjt_{inj} = 10:00:00, tdett_{det} = 10:01:10, tmitt_{mit} = 10:05:00, so TTD=70\text{TTD} = 70 s and TTM=300\text{TTM} = 300 s. Note what dominated TTM: not the fix (20 s) but the time between the report and reading Exit Code: 137. That is what a runbook removes.

What differs in your drill. Your fault is not memory: expect Exit Code: 1 and a log line about the missing model directory. The method is the same.

Before you start. dep.00 passes, and your system.toml has a [deploy] table:

[deploy]
kube_context = "kind-<system>" # ol refuses any context that is not kind- or k3d-
namespace = "<system>"
gateway_url = "http://127.0.0.1:30080" # the gateway NodePort
traces = "http://127.0.0.1:30686" # Jaeger query
services = { engine = "deploy/<system>-engine", gateway = "deploy/<system>-gateway" }

Export your gateway key in the variable [endpoints].api_key_env names (default TL_API_KEY); the second resolve check calls the gateway with it.

Inject. ol drill start engine-crashloop checks the safety gate (current context equals kube_context, which starts with kind- or k3d-; the namespace exists), then patches the first container of services.engine: its args become the tracer engine’s flags (spec/cli-roles.md) with --model-dir /missing, at the Kubernetes ports 8000 and 9464. The undo (your original args) goes to .ol/drills/<run>/journal.jsonl. You see only the page.

Detect. There is no alert in Pass 1. Work outside in: which workload is unhealthy, why its last container exited, what its previous log says, and what changed. These commands answer the four questions in that order; your runbook’s Diagnosis section says what each output means, in your own words:

Terminal window
kubectl get deploy,pods -n <system> -o wide # which workload: the engine pod is not 1/1 Running
kubectl describe pod -n <system> <engine-pod> # why it exited: Last State, Exit Code, Restart Count, Events
kubectl logs -n <system> deploy/<system>-engine --previous # what the crashed container said before it died
kubectl rollout history deploy/<system>-engine -n <system> # what changed: a new revision
kubectl get deploy <system>-engine -n <system> -o jsonpath='{.spec.template.spec.containers[0].args}'

Mitigate. Restore the last revision that worked, wait for it, then make the chart the truth again:

Terminal window
kubectl rollout undo deploy/<system>-engine -n <system>
kubectl rollout status deploy/<system>-engine -n <system> --timeout=120s
helm upgrade --install <system>-engine deploy/helm/<system>-engine -n <system>

Write the runbook. docs/runbooks/engine-crashloop.md, with at least these headings (any level, any extra sections you like):

# Runbook: engine crashloop
## Symptoms
What the person paged will see: the client error, the gateway's /readyz,
the pod status and restart count.
## Diagnosis
Numbered steps, each a command and what its output means, from "which
workload" to "what changed".
## Mitigation
Numbered steps that restore service, how to confirm it, and the durable fix.

Write the postmortem. docs/postmortems/<YYYY-MM-DD>-engine-crashloop.md (the date in UTC), with the headings Summary, Impact, Timeline, Root cause, Detection, Resolution, Action items. The timeline is a table like the one in section 3; the action items say how the next rollout could not cause this (a readiness probe, maxUnavailable: 0, changes only through helm upgrade).

Verify. ol drill end grades, then prints the seed and what was injected:

CheckPasses whenWhy it matters
detectedalways, with “graded manually”: this drill has no [detect] blockPass 1 has no Prometheus; you are the detector
resolve 1the engine rollout (services.engine) has converged within 60 s: every replica updated and availableKubernetes, not you, now keeps a Ready engine running
resolve 2ol conform openapi:v0:gateway:smoke against [deploy].gateway_url passes: the v0 smoke cases (completion schema, SSE framing, and 401 without a key; the health case is pending unless [deploy].gateway_health_url is set) through your gateway NodePortcompletions flow again, so a Ready engine is behind the gateway
postmortemthe newest docs/postmortems/*-engine-crashloop.md has all seven headingsthe incident is recorded in the fixed shape
runbookdocs/runbooks/engine-crashloop.md has Symptoms, Diagnosis, Mitigation headings (the drill’s [[doc]] table)the next person needs minutes, not your 45
time limitend within 45 minutes of startkeeps the drill honest

ol drill end grades the runbook and the postmortem the same way: a missing file or heading fails the drill. It records the verdict under ops.00, which MS-P1 requires. Then run ol drill reset: it replays the journal in reverse (after your rollback the replayed undo changes nothing).

#PitfallSymptomCaught by
1Reading kubectl logs without --previousan empty log, or “container is waiting to start”no check; the runbook’s Diagnosis step 3
2Deleting the crashing pod to “restart it”the ReplicaSet makes an identical pod that crashes the same wayresolve 1 and resolve 2 stay red
3Scaling the engine to zero to stop the restartsthe restarts stop and so does the serviceresolve 2 stays red
4Fixing the args with kubectl edit and never touching the chartit works until the next helm upgrade reapplies whatever the chart saysthe postmortem’s action items; review
5Running ol drill reset before ol drill endthe run is marked reset and never gradedol drill end refuses: no active drill
6A postmortem that names who broke it, or skips a headingthe team stops reporting honestly; end fails on the missing headingthe postmortem check
DirectionModuleHow it uses this
Backdep.00the Deployment, Service, and NodePort the drill injects into and checks through
BackL10.0the engine whose startup fails on a missing model directory
Backgw.00the gateway whose /readyz and 503 no_capacity are the symptom
Backlang.07kubectl get, describe, logs, and Helm
Forwardobs.03burn-rate alerts that page you before a user does
Forwardops.01the first drill with a [detect] block, graded on TTD from Prometheus
Forwarddep.03chart policy: readiness probes and disruption budgets, the action items here
Forwardcraft.10a runbook for every required alert, in the format you start here
Your pieceProduction equivalentWhat it addsWhere to look
manual rollout undoprogressive delivery (Argo Rollouts, Flagger)automatic rollback when the new version’s health or SLO checks failArgo Rollouts analysis templates
reading exit codes by handstartupProbe, Kubernetes events, kubectl debugseparates slow starts from crashes; ephemeral debug containers in a broken podKubernetes docs, Configure Liveness, Readiness and Startup Probes
the fixed 10 s to 5 min back-offKEP-4603, tunable CrashLoopBackOfffaster restarts for quick-failing containerskubernetes/enhancements, keps/sig-node/4603
docs/runbooks/*.mdincident tooling (PagerDuty, incident.io, Rootly)runbooks linked from the alert, incident roles, automated timelinesGoogle SRE book ch. 14 (free)
your postmortemblameless postmortem programsreview meetings, tracked action items, shared learningGoogle SRE book ch. 15 (free); Etsy, Debriefing Facilitation Guide (free)
CompanyPracticeWhy it matters
GoogleSRE incident management and blameless postmortemsthe origin of most of this chapter’s vocabulary
Netflixchaos engineering in productiondrills are how a team learns its system before a real outage teaches it
Any on-call teamrunbooks linked from every alertthe difference between a 5-minute and a 45-minute mitigation