Skip to content

A poison task lands in the dead-letter queue

Moduleops.03 · drill · ops, docs · Pass 8 · 2 to 3 h
You builddocs/runbooks/DurableDeadLetters.md (Symptoms, Diagnosis, Mitigation), docs/postmortems/<date>-poison-task.md, and the alert DurableDeadLetters in deploy/observability/rules/durable.yaml (written with ops.02); and you find the cause of a task that fails every attempt, remove it, redrive the task, and see the build complete
Contractthe drill spec course/drills/poison-task/drill.toml (section 4); tl_durable_dlq_size and tl_durable_task_queue_depth in otel/metrics.yaml; the failpoint hook your worker evaluates (section 2.4); the ctl verbs wf dlq list, wf dlq redrive, wf describe fixed by MS-durable
Testsgraded by ol drill end; no course tests to read (section 4 lists every check)
Needsdep.06 (the worker Deployment the fault is shipped to) and obs.05 (the scrapes and the dead-letter panel); reading: ops.02 (the rules file and the drill routine), dur.09 (exit codes and retryable failures), the queue chapter of this pass (retries, the dead-letter queue, redrive)
Used byno call site: a drill. MS-durable requires it; MS-ops reruns the drills
MilestoneMS-durable (drill ops.03 resolved)
Optional depthAWS: Amazon SQS dead-letter queues (free); Site Reliability Engineering, ch. 14 Managing Incidents (free); Hohpe and Woolf, Enterprise Integration Patterns, “Dead Letter Channel” and “Invalid Message Channel”
  • A poison task fails the same way on every attempt. Retries cannot fix it; they only decide how long it takes to stop trying. After dlq_after_attempts the queue moves it aside so the rest of the work keeps flowing.
  • A dead letter is a paused workflow, not a finished one: the build waits for that one shard forever, so the alert pages on the first dead letter.
  • Diagnosis asks two questions in order: what failed (the dead letter’s last failure) and what changed (rollout history, the environment). A deterministic failure that appeared without a code change came with a configuration change.
  • Fix, then redrive. Redriving first buys five more failures and a second page. Redrive only the tasks you fixed.
  • The postmortem’s best action item usually prevents the configuration from reaching production at all (a policy check on the chart), not just a better alert.
Terminal window
kubectl apply -f deploy/observability/rules/durable.yaml # DurableDeadLetters (written with ops.02)
ol drill start poison-task # ships the fault to the workers, prints the page
kubectl -n <system> rollout status deploy/<system>-worker
{ctl} data build --config course/fixtures/MS-corpus/corpus.toml --id drill-poison --durable 127.0.0.1:30733 --json
# ... wait for the page; dashboard, wf dlq list, the runbook; fix; redrive ...
{ctl} wf describe drill-poison --durable 127.0.0.1:30733 --json # COMPLETED, dead_lettered 0
ol drill end
ol drill reset

{ctl} is your [entry].ctl; its wf and data build verbs are fixed in MS-durable. The build starts after the poisoned rollout, as a scheduled build would.


ops.02 killed the server and the system healed itself: leases expired, tasks were delivered again, every effect happened once. This drill is the other kind of failure: one task that can never succeed, however often it is retried. Your queue (dur.03) handles it by moving the task to a dead-letter queue after a bounded number of attempts, and your dashboard (obs.05) shows the dead-letter count. That keeps one bad task from spinning forever, but it also means a human has to act: the workflow waiting on the task is stuck until someone understands why it failed, fixes that, and redrives it.

SymbolMeaningType
AAattempts before the dead-letter queue ([durable].dlq_after_attempts)5
bnb_nbackoff before attempt n+1n + 1: min⁡(b0βn−1,bmax)\min(b_0 \beta^{n-1}, b_{max}), then full jitter (RetryPolicy, dur.05)seconds
b0b_0, β\beta, bmaxb_{max}initial interval, multiplier, cap1 s, 2, 60 s
TdlqT_{dlq}time from the first attempt to the dead letterseconds

Without jitter the waits are b1=1b_1 = 1, b2=2b_2 = 2, b3=4b_3 = 4, b4=8b_4 = 8 s, so TdlqT_{dlq} is at most 1515 s plus five attempts’ run time. Full jitter draws each wait uniformly in [0,bn][0, b_n], so on average it is half that. The point of the budget: a transient failure (a flaky download) succeeds within a few attempts; a deterministic one is set aside in seconds instead of blocking a worker slot forever.

The activity’s failure type decides what the queue does (spec/subprocess-activity.md, dur.09): exit 65 (bad input) is non-retryable and fails at once; a crash, a panic, or exit 75 is retryable. A panic is classified retryable because most panics are not deterministic, which is exactly why a deterministic one walks the whole retry budget before anyone notices. Only the dead-letter queue stops it.

Redrive (RedriveDeadLetter) puts a dead-lettered task back in its queue with a fresh attempt budget. It is safe because the task is idempotent (the same key, the same work directory), and it is useless until the cause is gone. The order is: find the cause, remove it, confirm the removal is live (a completed rollout), then redrive the specific task ids.

A bad rollout ships a test-only setting to production: the worker Deployment’s environment gains TL_FAILPOINTS=dur/activity/poison/shard-00001=panic. Your worker (dur.04) evaluates the failpoint dur/activity/poison/<activity id> before it runs each activity task (the TL_FAILPOINTS syntax of the course testkit: crash, panic, error(msg), sleep(d)), and recovers a panic in an activity as a retryable failure of type Panic. CorpusBuild (data.09) names its per-shard tasks shard-00000, shard-00001, and so on. So exactly one task panics, on every attempt, on every worker. That is how real poison looks from the outside: one input, one code path, always the same failure.

Backoff without jitter (the worst case), each attempt failing in about 0.1 s, the scrape every 15 s, and the rule tl_durable_dlq_size > 0 for 15 s.

Time after the build reaches shard 1EventDead letters
0 sattempt 1 panics0
1 sattempt 2 (after b1=1b_1 = 1 s) panics0
3 sattempt 3 (after b2=2b_2 = 2 s) panics0
7 sattempt 4 (after b3=4b_3 = 4 s) panics0
15 sattempt 5 (after b4=8b_4 = 8 s) panics: A=5A = 5 reached, the task is dead-lettered1
at most 30 sthe next scrape sees tl_durable_dlq_size = 11
at most 45 sthe rule has held for 15 s: DurableDeadLetters fires1

So the page arrives at most about 45 s after the first failure, well inside the drill’s 600 s. Meanwhile shards 0 and 2 completed, and the build reports RUNNING with dead_lettered: 1. After the fix and the redrive, the one task runs once more and succeeds; the build completes with the same numbers as a clean run (169 documents, 3 shards).

StepWhat ol drill doesWhat you do
injectadds TL_FAILPOINTS=dur/activity/poison/shard-00001=panic to the environment of {deploy.services.worker} (a JSON patch; the undo restores the old environment)wait for the rollout, start the build drill-poison
detectat end, reads the ALERTS series: DurableDeadLetters must have fired within 600 s of the injectionread the page, open the dead-letter panel and your runbook
mitigatewf dlq list data (what failed and how), rollout history and the environment (what changed), remove the setting, wait for the rollout, wf dlq redrive data <task id>
verifyat end: the worker rollout is complete; dead letters at 0 and every queue at depth 0, each held 60 swf describe drill-poison shows COMPLETED and dead_lettered: 0; write the runbook and the postmortem
CheckPasses when
detectedDurableDeadLetters fired within within_s = 600 of the injection
resolvedrollout status of the worker Deployment; tl_durable_dlq_size and tl_durable_task_queue_depth summed to 0, each held 60 s
runbookdocs/runbooks/DurableDeadLetters.md has Symptoms, Diagnosis, Mitigation
postmortemdocs/postmortems/<date>-poison-task.md has the seven sections
time limit45 minutes from start to end (not graded for the CI responder)

The design’s last condition, that the workflow completes, is a command check ol drill end runs and grades ([[resolve.run]] in the spec): {ctl} wf describe drill-poison --json must report COMPLETED with dead_lettered: 0. Run it yourself first and quote it in your Resolution.

PitfallSymptomCaught by
1. Redriving before fixingfive more panics, a second dead letter, the alert fires againthe dead-letter check fails its hold
2. Deleting the dead letter to clear the alertthe alert clears, the build waits foreverthe queue-depth check; the workflow never completes
3. Restarting the workers without removing the settingthe new pods read the same environment and panic the same waythe dead-letter check after the redrive
4. No alert on dead letters, or on the raw count without a forno page, or a page on a scrape blipol drill end, check detected
5. A postmortem that stops at “removed the variable”the next test setting reaches production the same wayol drill end, check postmortem; the self review against course/rubrics/postmortem.md
DirectionModuleHow it uses this
Backdep.06the worker Deployment the fault is shipped to and rolled back on
Backobs.05the dead-letter panel and the scrapes the alert reads
Forwardops.10a runaway agent stopped by a budget and resumed after a grant: the same pause, fix, resume shape
Forwardops.08a data incident traced through the ledger, the other kind of bad input
Your pieceProduction equivalentWhat it addsWhere to look
one dead-letter queue per queueSQS redrive policies, RabbitMQ dead-letter exchangesper-queue budgets, automatic redrive to the sourceAWS SQS docs
redrive by handTemporal’s failed-activity handling in the workflow; DLQ consumersthe workflow decides: skip, compensate, or faildur.08 (sagas)
a failpoint in the environmentfeature flags with audit trails, admission policiestest settings cannot reach productionOPA Gatekeeper, Kyverno
an alert on the dead-letter countalerts on failure rate per activity typethe page comes before the budget is spentobs.05’s dashboard