Skip to content

The durable server is SIGKILLed in a loop during a corpus build

Moduleops.02 · drill · ops, docs · Pass 8 · 2 to 3 h
You builddeploy/observability/rules/durable.yaml (the alerts DurableRedeliveriesSpike and DurableDeadLetters, shared with ops.03), docs/runbooks/DurableRedeliveriesSpike.md (Symptoms, Diagnosis, Mitigation), and docs/postmortems/<date>-durable-kill9.md; and you show that a corpus build survives five SIGKILLs of its server with the output of a clean run
Contractthe drill spec course/drills/durable-kill9/drill.toml (what is injected, detected, and checked: section 4); tl_durable_redeliveries_total, tl_durable_task_queue_depth, and tl_durable_dlq_size in otel/metrics.yaml; the ctl verbs fixed by MS-durable (course/milestones/MS-durable.toml)
Testsgraded by ol drill end; no course tests to read (section 4 lists every check)
Needsdep.06 (the durable StatefulSet and the worker Deployment it kills pods of) and obs.05 (the scrapes and the control-plane dashboard you diagnose on); reading: ops.01 (how a drill runs), dur.09 (how a killed activity resumes), and the log, lease, and replay chapters of this pass
Used byno call site: a drill. MS-durable requires it; MS-ops reruns the drills
MilestoneMS-durable (drill ops.02 resolved)
Optional depthSite Reliability Engineering, ch. 22 Addressing Cascading Failures and ch. 26 Data Integrity (free); Temporal: activity timeouts and retries (free); Kleppmann, Designing Data-Intensive Applications, ch. 8 (process pauses and leases)
  • Killing the server of a durable engine must cost time, never correctness: every acknowledged event is in the fsynced log, and recovery truncates only the torn record that was never acknowledged.
  • Leases turn a dead worker or a dead server into a timeout: work in flight is delivered again after the visibility timeout, so each kill shows up as a redeliveries spike, which is what pages.
  • Redelivery is safe only because activities are idempotent: a rerun with the same key finds its work directory and resumes or re-emits DONE.json. The proof is the output, not the logs: the corpus equals a clean run’s and the ledger has every row once.
  • Mitigation for a self-healing fault is to confirm, not to act: the StatefulSet restarts the pod on the same claim. Deleting the WAL claim “to reset” is the one action that loses data.
  • The postmortem’s action items are about exposure: one replica, and long activities without checkpoints that redo minutes of work.
Terminal window
kubectl apply -f deploy/observability/rules/durable.yaml # your two alerts (section 4)
{ctl} data build --config course/fixtures/MS-corpus/corpus.toml --id drill-kill9 --durable 127.0.0.1:30733 --json
ol drill start durable-kill9 # five SIGKILLs of the durable pod, 30 s apart, then only the page
# ... dashboard, kubectl, the runbook; wait for the build ...
{ctl} data build --config course/fixtures/MS-corpus/corpus.toml --id drill-kill9 --durable 127.0.0.1:30733 --wait --json
TL_ARTIFACTS=artifacts {corpus} ledger verify --config course/fixtures/MS-corpus/corpus.toml
ol drill end # grades detection, resolution, evidence, runbook, postmortem
ol drill reset

{ctl} and {corpus} are your [entry].ctl and [entry].corpus (spec/cli-roles.md); the ctl verbs are fixed in MS-durable. ol drill start blocks while it kills (about two and a half minutes), so the build starts first. The two commands after the wait are the pass condition ol drill end also runs itself (section 4).


Pass 8 built a durable execution engine: a write-ahead log (dur.01), leases with visibility timeouts and a dead-letter queue (dur.03), workers (dur.04), idempotent activities (dur.05, dur.09), and replayed workflows (dur.06), deployed on kind as a StatefulSet with its log on a persistent volume (dep.06) and observed by obs.05. Every one of those modules was tested with kills in a test process. This drill does it to the deployed system, during a real corpus build, the way a node failure or an OOM kill would: SIGKILL, five times, with no chance to flush or shut down cleanly. The question is the one the whole pass exists to answer: after it, is the result exactly what an undisturbed run produces?

SymbolMeaningType
kknumber of kills5 here
Δ\Deltatime between kills30 s
rrtime for the pod to restart and recover the logseconds (about 8)
VVvisibility timeout of an activity lease30 s ([durable].visibility_timeout_ms)
aia_iactivity tasks whose lease was live at kill iicount

SIGKILL cannot be caught: the process ends between two instructions. What survives is what was on disk: the log up to the last fsync. An append is acknowledged only after its fsync (dur.01), so every acknowledged event survives; the record being written at the instant of the kill may be torn, and recovery truncates it, which is correct because nobody was told it happened.

A worker holds each activity task under a lease that expires VV after its last heartbeat. While the server is down (rr seconds), workers can neither heartbeat nor complete. After the restart the server’s queue state is rebuilt from the log with the leases’ deadlines, and every lease whose deadline passes without a completion makes its task visible again: a redelivery. So each kill produces about aia_i redeliveries roughly VV after it, and the counter tl_durable_redeliveries_total climbs in steps.

2.3 Why the second delivery has no second effect

Section titled “2.3 Why the second delivery has no second effect”

A redelivered activity runs again with the same idempotency key (TL_IDEMPOTENCY_KEY), in the same work directory (spec/subprocess-activity.md). Outputs are written as <name>.tmp and renamed, DONE.json last; a rerun that finds DONE.json re-emits its result and exits. So whether the first attempt finished, half-finished, or never started, the second leaves the same files. The ledger (data.08) records each source document and shard by content hash, so a duplicate would show up as a second row, and {corpus} ledger verify fails on it.

2.4 Detection: alert on the symptom you can see

Section titled “2.4 Detection: alert on the symptom you can see”

The page comes from the redeliveries rate, not from pod restarts: a restart that causes no redelivery hurts nobody, and redeliveries also catch dying workers and activities that outlive their lease. Your rule (section 4) fires when one queue sees more than 5 redeliveries in 2 minutes for 20 s. Normal churn (a deploy, a preempted node) stays under it.

k=5k = 5 kills at t=20,50,80,110,140t = 20, 50, 80, 110, 140 s after ol drill start, r=8r = 8 s, V=30V = 30 s. The build is in its dedup and shard stages, with about 3 activity tasks leased at any moment.

KillServer downLeases that expire unansweredRedeliveries (cumulative)
1 at 20 s20 to 28 sthe 3 live at 20 s, at about 50 s3
2 at 50 s50 to 58 s3 more, at about 80 s6
3 at 80 s80 to 88 s3 more, at about 110 s9
4 at 110 s110 to 118 s2 (a quieter stage), at about 140 s11
5 at 140 s140 to 148 s3, at about 170 s14

The alert’s expression is the increase over 2 minutes. At t=80t = 80 s the window holds 6, more than 5; with for: 20s it fires at about t=100t = 100 s, 80 s after the first kill, inside the drill’s 300 s. Without the kills the build runs about 150 s; with them each kill costs about r+V≈38r + V \approx 38 s for the tasks it hit, and the stages overlap, so expect roughly 3 to 4 extra minutes, not 5×385 \times 38.

Then check the output by hand: the clean numbers of the MS-corpus fixture are 169 documents, 3 shards, 75164 training tokens, and 6868 validation tokens. The same start run again (--wait) must report exactly those, with started: false (no second run was created).

StepWhat ol drill doesWhat you do
inject20 s after start, kubectl delete pod --grace-period=0 --force on a Ready pod of {deploy.services.durable}; four more, each 30 s after the previousstart the build first; then nothing: the page is your first signal
detectat end, reads the ALERTS series: DurableRedeliveriesSpike must have fired within 300 s of the first killopen the control-plane dashboard (obs.05) and your runbook
mitigateconfirm the server recovers each time (logs: torn tail truncated, no ErrCorrupt), the workers reconnect, and the build progresses; touch nothing that holds state
verifyat end: both rollouts complete; dead letters at 0 and every queue at depth 0, each held 60 s; a workflow CorpusBuild trace in the windowrun the two pass-condition commands below; write the runbook and the postmortem

Your system.toml needs [deploy].services.durable (statefulset/<system>-durable), [deploy].services.worker (deploy/<system>-worker), [deploy].prometheus, [deploy].traces, and [entry].ctl and [entry].corpus.

Your rules file, a PrometheusRule with the label release: observability (contracts/helm/observability.md), holds both durable alerts:

- alert: DurableRedeliveriesSpike
expr: sum by (namespace, queue) (increase(tl_durable_redeliveries_total{namespace="<system>"}[2m])) > 5
for: 20s
labels: {severity: page}
annotations: {runbook: docs/runbooks/DurableRedeliveriesSpike.md}
- alert: DurableDeadLetters
expr: sum by (namespace, queue) (tl_durable_dlq_size{namespace="<system>"}) > 0
for: 15s
labels: {severity: page}
annotations: {runbook: docs/runbooks/DurableDeadLetters.md}
CheckPasses when
detectedDurableRedeliveriesSpike fired within within_s = 300 of the first kill (TTD is reported)
resolvedrollout status of the durable StatefulSet and the worker Deployment; tl_durable_dlq_size and tl_durable_task_queue_depth summed to 0, each held 60 s
trace evidencea trace in the window has the span workflow CorpusBuild from <system>-worker
runbookdocs/runbooks/DurableRedeliveriesSpike.md has Symptoms, Diagnosis, Mitigation
postmortemdocs/postmortems/<date>-durable-kill9.md has the seven sections
time limit45 minutes from start to end (not graded for the CI responder)

The drill’s pass condition also holds two commands, [[resolve.run]] in the spec, that ol drill end runs and grades: the same start with --wait returns COMPLETED with 169 documents, 3 shards, 75164 and 6868 tokens; and {corpus} ledger verify over ./artifacts exits 0. Run both yourself first, and put their output in your postmortem’s Resolution.

PitfallSymptomCaught by
1. Alerting on restarts instead of redeliveries, or no durable rule loadedno page, or a page for restarts that hurt nobodyol drill end, check detected
2. Deleting the WAL claim or scaling the StatefulSet to 0 to “reset” the serverevery workflow’s history is gone; the build never completesthe rollout and queue-depth checks; the build command fails
3. An activity that writes outputs in place, without a temporary name and a renamea redelivered attempt sees a half-written shard and keeps it; the token counts differthe pass-condition commands (clean-run numbers, ledger verify)
4. A rerun that ignores the idempotency keya shard written twice, two ledger rowsledger verify
5. Declaring the incident over while the build still runsthe queue-depth check fails its holdresolve check 4
DirectionModuleHow it uses this
Backdep.06the StatefulSet, its WAL claim, and the worker Deployment the kills land on
Backobs.05the dashboard and scrapes the diagnosis reads
Forwardops.03the second alert of the same rules file, and a fault that does not heal itself
Forwardops.11the same log under a full disk instead of a kill
Your pieceProduction equivalentWhat it addsWhere to look
one durable replicaTemporal’s history service on Cassandra or Postgres; Raft-replicated logsa kill is not an outagedur.10 (optional), etcd
a kill loop in a drillJepsen, Chaos Mesh PodChaos on a schedulerandomized faults with a checker over the historyjepsen.io
a fixed visibility timeoutheartbeat timeouts per activity, sticky queuesa dead worker is noticed in seconds, not a full timeoutTemporal docs
a ledger check after the factend-to-end checksums and exactly-once sinksduplicates caught at write timeKafka transactions, Flink two-phase commit