Skip to content

Resilience tests (R10)

Modulecraft.21 · practice · Go · Pass 8 · 4 to 6 h
You buildprimers/craft.21/taskq.go (a kata written by ol start: a durable task queue and its worker, the log, leases, and idempotent effects of your durable engine in one file) and primers/craft.21/resilience_test.go (your resilience tests, package craft21_test)
Contractnone of its own: the kata’s API is in its file, and the course’s fault kit is course/tests/craft.21/faults/faults.go (copied to primers/craft.21/faults/ by the first check)
Testscourse/tests/craft.21/check: the course’s suite (course/tests/go/craft_21/) on your kata, then the grade of your tests by 12 planted faults (section 4)
Needsreading: craft.03 how tests are graded · craft.04 properties · craft.07 reading survivors · lang.06 Go testing · dur.09 subprocess activities; the engine pieces the kata models: the event log (dur.01), fenced leases (dur.03), idempotent activities (dur.05)
Used byno call site (a practice): rung R10 grades resilience tests in the durable engine (dur.*) and the drills (ops.*) from here on
MilestoneMS-durable (its kill loop is the system-scale version of your tests)
Optional depthKyle Kingsbury, Jepsen analyses (free); Pillai et al., “All File Systems Are Not Created Equal” (OSDI 2014, free); Martin Kleppmann, “How to do distributed locking” (fencing tokens, free); FoundationDB, “Testing Distributed Systems w/ Deterministic Simulation” (free)
  • A resilience test kills the system at its worst moment and checks a promise, not a return value: effects exactly once, nothing acknowledged lost, one owner per task.
  • A SIGKILL loses nothing a write returned. To catch a missing fsync the test must lose the page cache: the fault kit’s CrashFS.Crash() keeps only synced bytes.
  • Time is an input. With a manual clock a test can hold a lease for 90 s, pause a worker past its lease, and expire it, in microseconds and the same way every run.
  • Exactly once is “at least once delivery” plus an idempotency key that is the same on every attempt. The test that proves it must crash between the effect and the acknowledgement.
  • The grade is rung R10: your tests must fail on at least 90% of 12 planted faults, and on all three resilience faults: the idempotency key ignored, the lease not renewed, fsync skipped.
Terminal window
ol start craft.21 # writes primers/craft.21/taskq.go with stub bodies
ol tests craft.21 # read the test catalog first
ol check craft.21 # first run: writes go.mod and faults/faults.go, then fails (no tests yet)
# write resilience_test.go (section 4), against the stub first: every test must fail
(cd primers/craft.21 && go test ./...)
# then make the kata pass your tests and the course's
ol check craft.21 # grades your tests by planted faults

ol check never shows a planted fault’s code. A surviving resilience fault prints its one-line name (they are public: section 2.5); any other survivor prints only “a planted pitfall”.


Your durable engine now makes three promises that no unit test has ever checked: an acknowledged event survives a crash (dur.01), a task has one owner at a time (dur.03), and an activity’s effect happens once however often it is retried (dur.05, dur.09). Unit tests run the happy path in one process that never dies, so a log that forgets to fsync, a worker that forgets to renew its lease, or an activity keyed by its attempt number passes all of them. These bugs appear only when a machine dies at the wrong instant, which in production is “eventually, at 3 a.m., during the biggest training run”. MS-durable will kill your real server and worker 500 times; this module teaches you to write the small, deterministic version of that test first, on a kata small enough to hold in your head, and grades it by whether it catches the faults that matter.

A resilience test is a loop: put the system in a state, inject a fault, let it recover, check an invariant. The invariants are the system’s promises, stated so a test can check them after the fact:

PromiseChecked asThe kata’s mechanism
durability: a call that returned nil is on diskafter Crash() and Open, every acknowledged enqueue, lease, and completion replaysappend one JSON line, Sync, only then update memory
exclusivity: one owner per task at a timewhile a lease is live, Lease gives ErrNoTask; a stale token’s Renew and Complete get ErrLeaseLostleases with a deadline and a fencing token that grows on every delivery
exactly-once effectthe external system’s applied keys equal the task ids, with deliveries counted separatelythe effect is keyed by the task id, the same on every attempt
idempotent submissionenqueueing an id twice is one taskEnqueue of an existing id is a no-op

Delivery is at least once: a worker can die after applying an effect and before completing, and the retry applies it again. That is unavoidable (the Two Generals problem), so the effect must be idempotent: the external system deduplicates by key. The fault kit’s Effects does exactly that, and counts deliveries, so a test sees both the state (Keys) and the retries (Deliveries).

2.2 Crashes: what survives and what does not

Section titled “2.2 Crashes: what survives and what does not”
FaultWhat is lostWhat a test needs
SIGKILL of the processmemory; nothing a write returned (the page cache belongs to the kernel)a child process, or abandoning the goroutine
power loss, kernel panicmemory and every byte written but not synceda disk model: CrashFS.Crash()
a crash during a writethe end of that write: a torn recordCrashFS.CrashTorn(n) keeps n bytes of the unsynced tail
a full disknothing yet; the write failsCrashFS.FailWrites(err)

Two rules follow. First, sync before acknowledging, and update memory only after the sync succeeded: a queue that changes memory first serves a completion that vanishes on restart. Second, replay must survive a torn tail: the last line may be half a record. Its call never returned, so it is safe to ignore; but the next append lands right after it, so Open must seal the fragment with a newline, or the first record acknowledged after the restart is glued to garbage and lost on the next replay.

A lease is ownership with a deadline: worker w may work on task t until now + TTL. If w dies, the task comes back after the deadline. Two things go wrong:

  • A task that takes longer than the TTL loses its lease mid-run and a second worker starts it too. The worker must renew (heartbeat) while it works, every TTL/3 in the kata, so two missed renewals still leave time.
  • A worker that was paused (GC, a stopped VM, a partition) wakes up after its lease expired and someone else owns the task. It must not finish: every lease carries a fencing token, a number that grows with every delivery, and Renew and Complete with a stale token are ErrLeaseLost. When a renewal fails, the worker cancels its work and does not apply the effect.

The kata takes its disk (FS) and its time (Clock) as interfaces, so a test owns both. With the fault kit’s Clock, time moves only on Advance, and BlockUntil(n) waits until the code under test is actually waiting, so a test never sleeps and never races. With CrashFS, a “crash” is one call at an exact point. The test is the same on every run, which is what makes a failure a bug report instead of a flake. A kill loop then explores many crash points with a seeded generator: each round picks where to die (before the effect, after it, after Complete), crashes the disk, reopens the queue, expires the leases, and continues until no task is pending; at the end the promises of 2.1 must hold.

Your tests run against the course’s correct kata (they must pass, twice), then against 12 copies of it, each with one planted fault. A fault your tests make fail is killed. Rung R10 needs a score of at least 0.90 and every required fault killed. Three of them are the resilience faults of the testing ladder, and their names are public:

FaultWhat it breaksA test that catches it
idempotency key ignoredthe effect’s key changes per attempt, so a retry applies it twicecrash after the effect, retry, AssertExactlyOnce
lease not reneweda long task is handed to a second worker mid-runhold a task for 3 x TTL, try to lease it from another worker
fsync skippedacknowledged records live only in the page cacheCrash(), reopen, check acknowledged state

The other nine are pitfalls (section 5) and stay hidden until you pass.

Clock at t0t_0 = 1760000000 s (Unix), TTL 30 s, empty disk.

StepActionDurable log after it
1Enqueue("t1","a"), Enqueue("t2","b")two lines: {"op":"enq","id":"t1","payload":"a"}, {"op":"enq","id":"t2","payload":"b"}
2w1 leases t1: token 1, attempt 1, deadline t0t_0 + 30 s{"op":"lease","id":"t1","worker":"w1","token":1,"deadline":1760000030000000000}
3w1 applies the effect t1 = done:a (deliveries of t1: 1)unchanged: the effect is outside the log
4power loss before Completeunchanged: everything above was synced
5restart: Open replays three records; t1 is leased to w1 until t0t_0 + 30 s
6clock moves to t0t_0 + 31 s; w2 leases t1: token 2, attempt 2lease t1 w2 token 2
7w2 applies t1 again: the external system has the key, so it keeps done:a (deliveries 2, applied keys {t1})
8w2 completes t1 with token 2, then runs t2done t1, lease t2, done t2
9w1 wakes up and calls Complete with token 1ErrLeaseLost: fenced

Every promise holds: applied keys are exactly {t1, t2}, t1 was delivered twice and applied once, the zombie’s completion was refused. This is the course’s TestHandExampleCrashBeforeComplete; with the idempotency key changed to t1#<token>, step 7 applies a second key t1#2 and AssertExactlyOnce fails.

Write primers/craft.21/resilience_test.go in package craft21_test (black box: only the kata’s exported API and craft21/faults). It must at least crash the disk, move time past a lease, restart the queue with craft21.Open, and assert on effects; the check rejects a suite that never does. Cover each promise of 2.1 with at least one test, and write one kill loop (2.4).

// primers/craft.21/taskq.go (the kata)
func Open(fs FS, clock Clock) (*Queue, error)
func (q *Queue) Enqueue(id, payload string) error
func (q *Queue) Lease(worker string, ttl time.Duration) (Lease, error) // ErrNoTask
func (q *Queue) Renew(l Lease, ttl time.Duration) (Lease, error) // ErrLeaseLost
func (q *Queue) Complete(l Lease, result string) error // ErrLeaseLost
func (q *Queue) Result(id string) (string, bool)
func (q *Queue) Pending() []string
type Worker struct { Q *Queue; ID string; TTL time.Duration; Clock Clock; Work func(context.Context, string) (string, error); Apply Effect }
func (w *Worker) RunOne(ctx context.Context) (bool, error)
// primers/craft.21/faults/faults.go (the course's kit)
fs := faults.NewCrashFS() // Crash(), CrashTorn(n), FailWrites(err), Durable(name)
clk := faults.NewClock(t0) // Now, After, Advance(d), BlockUntil(n)
fx := faults.NewEffects() // Apply(key, v), Deliveries(key), Keys(), FailNext(key, n), AssertExactlyOnce(t, keys)

The check (course/tests/craft.21/artifacts.py):

TestKINDChecks
test_files_presentunitthe kata and your tests exist; writes go.mod and faults/faults.go once
test_your_tests_inject_faultsunitblack-box package; uses a crash, a clock advance, a restart, and an effect assertion
test_your_kata_passes_the_course_suitefaultthe course’s suite passes on your kata
test_your_tests_pass_on_your_kataunityour tests pass on your kata
test_your_tests_pass_on_the_course_kataunityour tests pass on the course’s kata, twice (baseline A, and no flakes)
test_your_tests_catch_the_planted_faultsfaultscore >= 0.90 over 12 faults, every required one killed

The course’s suite (course/tests/go/craft_21/), run on your kata:

TestKINDChecks
TestHandExampleCrashBeforeCompletefaultsection 3, line by line
TestEnqueueIsIdempotentunita duplicate enqueue is one task, even after it is done
TestAcknowledgedSurvivesCrashfaultacknowledged enqueues and completions replay after Crash()
TestTornTailIsIgnoredfaulta torn record is ignored and sealed; the next record survives the next crash
TestFailedWriteChangesNothingfaulta failed write leaves memory as it was
TestLeaseIsExclusiveUntilExpiryboundaryno second lease until the deadline, then a higher token and attempt 2
TestStaleTokenIsFencedfaulta zombie’s Renew and Complete are ErrLeaseLost
TestLongTaskKeepsItsLeasefaulta 90 s task keeps its 30 s lease by renewing
TestLostLeaseCancelsWorkfaulta failed renewal cancels the work and skips the effect
TestCrashAfterApplyIsExactlyOncefaultcrash between effect and completion: delivered twice, applied once
PitfallSymptomCaught by
1. no fencing on Complete or Renewa paused worker overwrites the current owner’s resultTestStaleTokenIsFenced (mutants s01, s06)
2. lease deadlines checked wrongan expired task is never redelivered, or a live one is delivered twiceTestLeaseIsExclusiveUntilExpiry (mutants s02, s03)
3. torn tailsthe queue refuses to start, or loses the first record after a restartTestTornTailIsIgnored (mutants s04, s09)
4. enqueue not idempotenta retried submission runs a done task againTestEnqueueIsIdempotent (mutant s05)
5. memory updated before the diska completion served from memory vanishes on restartTestFailedWriteChangesNothing (mutant s07)
6. a lost lease does not stop the worktwo workers apply the same task’s effectTestLostLeaseCancelsWork (mutant s08)
7. idempotency key per attempt (resilience fault)a retry applies the effect twiceTestCrashAfterApplyIsExactlyOnce (mutant s10)
8. lease not renewed (resilience fault)long tasks are stolen mid-runTestLongTaskKeepsItsLease (mutant s11)
9. fsync skipped (resilience fault)acknowledged work disappears after a power cutTestAcknowledgedSurvivesCrash (mutant s12)
10. testing crashes with SIGKILL alonethe fsync fault survives your suite: the page cache outlives the processyour grade: s12 survives
11. real sleeps in resilience testsflaky timing, slow suites, a check that fails “only sometimes”test_your_tests_pass_on_the_course_kata (two runs)
DirectionModuleHow it uses this
Backcraft.07reading survivors: a surviving resilience fault names the promise your suite never checked
Backdur.09the kata’s worker is the shape of a subprocess activity: lease, heartbeat, idempotent outputs
ForwardMS-durablethe same promises checked by a kill loop over your real {durable} and {worker} with 500 activities
Forwardops.02, ops.03drills durable-kill9 and poison-task check these promises on kind
Your pieceProduction equivalentWhat it addsWhere to look
CrashFSALICE, CrashMonkeyrecord every syscall and replay every crash state the file system allows, reordering includedPillai et al., OSDI 2014
the kill loopdeterministic simulation (FoundationDB, TigerBeetle VOPR)the whole cluster in one thread with a seeded scheduler: every interleaving is reproducible from a seedFoundationDB “Simulation” docs
Effects checksJepsen and Ellea history checker that proves linearizability or serializability over a recorded runjepsen.io, elle
fencing tokensChubby sequencers, etcd lease revisionsthe storage service checks the token, not the clientBurrows, “The Chubby lock service”