Skip to content

Durable and worker images and charts: WAL PVC, StatefulSet, KEDA autoscaling of workers

Moduledep.06 · practice · ops · Pass 8 · 4 to 6 h
You builddeploy/docker/durable.Dockerfile and deploy/docker/worker.Dockerfile; the charts deploy/helm/<system>-durable (a StatefulSet whose WAL is on a PersistentVolumeClaim, with the WAL quota rendered into runtime.toml) and deploy/helm/<system>-worker (a Deployment that drains before it is killed, writes /artifacts, and is scaled by a KEDA ScaledObject); deploy/keda pinning KEDA
Contractcourse/contracts/helm/durable.values.schema.json, course/contracts/helm/worker.values.schema.json, the KEDA pin in course/contracts/helm/observability.md, the gauge tl.durable.task_queue.depth in course/contracts/otel/metrics.yaml, and [durable]/[worker] of course/contracts/config/runtime.schema.json
Testscourse/tests/dep.06/ (check runs artifacts.py: static, docker, and cluster tiers; section 4)
Needsdep.03 (the chart policy and the stack these charts join, with the images of dep.01 and the cluster of dep.02), dur.09 (the Python units the worker image carries); your durable server (dur.02) and worker (dur.04) entry points, which the images build
Used bydep.07 extends the worker chart for agent workers; MS-durable deploys these charts on kind, obs.05 traces and scrapes them, and drills durable-kill9, poison-task, and eventlog-disk-full run against them
MilestoneMS-durable (its kind steps install these charts)
Optional depthKubernetes: StatefulSets (free); KEDA: Prometheus scaler (free); Kubernetes: pod termination (free)
  • The durable server’s state is its WAL. It runs as a one-replica StatefulSet whose volumeClaimTemplates gives the pod a PersistentVolumeClaim that outlives it; an fsGroup lets the non-root server write the fresh volume.
  • kind’s local volumes ignore their size, so walMaxBytes (rendered as an integer [durable].wal_max_bytes) is the real disk quota, and it sits below the claim’s size.
  • A worker gets terminationGracePeriodSeconds long enough to drain (30 s) and let its Python child checkpoint (30 s more); the default 30 s kills training between checkpoints on every rollout.
  • KEDA scales workers on queue depth, not CPU, through a ScaledObject that owns the replica count: the Deployment sets none, and no second HPA competes.
  • The worker image is two worlds: a static Go binary copied into a Python base that can import the units its activities exec.
Terminal window
ol start dep.06 # records the start; there are no stubs
ol tests dep.06 # read the test catalog first
# write the Dockerfiles, both charts, and deploy/keda (section 4), then:
helm dependency update deploy/keda # writes deploy/keda/Chart.lock
OL_SMOKE=1 ol check dep.06 # static and docker tiers, no cluster
helm upgrade --install keda deploy/keda -n keda --create-namespace
helm upgrade --install forge-durable deploy/helm/forge-durable -n forge
helm upgrade --install forge-worker deploy/helm/forge-worker -n forge
ol check dep.06 # every tier, against kind-<system>

Pass 8 gave you a durable server and workers that survive kills on your laptop. On kind they need three things your Pass 7 charts never had to think about. The server has state: a Deployment pod writes its WAL to the container filesystem, and the first reschedule erases every workflow. The worker’s load is a queue, not requests: a corpus build enqueues 200 shard tasks in a second, the CPU of two idle workers says nothing about it, and a fixed replica count either wastes the cluster or leaves the queue growing. And the worker’s work is long: a training step runs for minutes inside a Python child, so the way Kubernetes stops a pod (SIGTERM, then SIGKILL 30 s later) decides whether a rollout loses a checkpoint. This module packages both components so they keep the promises of dur.01 to dur.09 when Kubernetes, not you, decides when pods start and stop.

Both Dockerfiles keep the dep.01 rules: multi-stage, every base pinned by digest, a numeric non-root USER, an exec-form ENTRYPOINT (the process is PID 1 and receives SIGTERM, which the worker needs to drain), a HEALTHCHECK in exec form on the health port 9464, and the OCI labels. The durable image is a static Go binary on alpine, like the gateway. The worker image is different: its runtime stage is a Python base (glibc, so numpy and pyarrow install as wheels) that also carries the static Go worker and your python/ tree, because the worker execs {tinyllm} train and {corpus} run --stage as subprocess activities (dur.09).

Each chart carries its values.schema.json copied unchanged from contracts/helm/, so Helm refuses unsafe values before an install: image.tag=latest, testClock=true (a test-only clock in a cluster), persistence.enabled=false, a missing or zero walMaxBytes, two durable replicas without Raft, a worker with no task queue. The dep.03 policy applies to what helm template renders: probes, CPU and memory requests and limits, a numeric non-root user without privilege escalation, a pinned image, a PodDisruptionBudget for every workload, and secrets only by reference.

2.3 The durable server: StatefulSet, WAL claim, quota

Section titled “2.3 The durable server: StatefulSet, WAL claim, quota”
ObjectWhy
StatefulSet, replicas: 1, podManagementPolicy: OrderedReadyone writer per WAL; a stable pod name (<system>-durable-0) that Raft peers can address later (dur.10)
volumeClaimTemplates wal, mounted at /var/lib/durablethe claim wal-<system>-durable-0 survives pod deletion and rescheduling; [durable].wal_dir points inside the mount
securityContext.fsGroup: 10001a freshly provisioned volume belongs to root; fsGroup makes it group-writable for the server’s gid
runtime.toml wal_max_bytesthe quota at which appends fail with RESOURCE_EXHAUSTED and nothing acknowledged is lost (dur.01, drill eventlog-disk-full)
Service NodePort 7233 to 30733, plus 9464workers in the cluster use <system>-durable:7233; a training worker on your Mac uses 127.0.0.1:30733 (the port dep.02 maps)

One Helm detail bites here: Helm reads YAML numbers as float64, so toToml writes wal_max_bytes = 1.073741824e+09, which runtime.schema.json (an integer) rejects. Cast every integer with int64 in the template.

2.4 The worker: drain time and a writable artifact root

Section titled “2.4 The worker: drain time and a writable artifact root”

Kubernetes stops a pod by sending SIGTERM to PID 1 and, after terminationGracePeriodSeconds (default 30), SIGKILL. Your worker on SIGTERM stops polling and lets in-flight activities finish for up to 30 s (dur.04), and the subprocess runner gives a Python child 30 s between SIGTERM and SIGKILL to checkpoint (dur.09). Both must fit: the grace period is at least 60 s (the reference uses 75). The worker writes checkpoints, shards, and DONE.json under /artifacts, so it mounts that hostPath read-write (the engine mounts it read-only), while its root filesystem stays read-only with an emptyDir at /tmp.

KEDA is an operator that turns a ScaledObject into a HorizontalPodAutoscaler fed by an external metric. Here the metric is the durable server’s gauge tl_durable_task_queue_depth{queue, kind}, scraped by Prometheus, and the rule is: replicas =⌈depth/threshold⌉= \lceil \text{depth} / \text{threshold} \rceil, clamped to [minReplicaCount, maxReplicaCount].

SymbolMeaningReference value
ddpending activity tasks on the worker’s queuefrom Prometheus
ttthreshold: pending tasks one worker should absorb4
r=min⁡(max⁡(⌈d/t⌉,rmin⁡),rmax⁡)r = \min(\max(\lceil d/t \rceil, r_{\min}), r_{\max})desired replicasrmin⁡r_{\min} = 1, rmax⁡r_{\max} = 6

The ScaledObject owns the count, so the Deployment sets no replicas (a helm upgrade would otherwise reset it) and the chart renders no HPA of its own. cooldownPeriod is long (120 s): removing a worker mid-activity costs a 75 s drain and a redelivery, more than an idle pod. KEDA itself is pinned in deploy/keda/Chart.yaml and Chart.lock to the version in contracts/helm/observability.md.

The reference values: claim 2Gi, walMaxBytes: 1073741824, worker threshold 4, replicas 1 to 6, grace 75 s.

The quota. 2 Gi=2×230=21474836482\,\text{Gi} = 2 \times 2^{30} = 2147483648 bytes; the quota is 230=10737418242^{30} = 1073741824, half of it, which leaves room to raise the quota during an incident (drill eventlog-disk-full) without resizing the claim. helm template must render wal_max_bytes = 1073741824 (an integer), and with --set walMaxBytes=123456789, wal_max_bytes = 123456789.

A burst. CorpusBuild enqueues 17 shard tasks on queue data. Prometheus scrapes d=17d = 17; KEDA computes ⌈17/4⌉=5\lceil 17/4 \rceil = 5, within [1,6][1, 6], so the worker Deployment goes from 1 to 5 replicas. As the queue drains to d=3d = 3, ⌈3/4⌉=1\lceil 3/4 \rceil = 1; after the 120 s cooldown KEDA scales back to 1, and each removed pod gets SIGTERM, drains, and exits within 75 s.

A rollout. helm upgrade of the worker during a training activity: the old pod gets SIGTERM at t=0t = 0, stops polling, the runner SIGTERMs the Python child, which checkpoints at its next step (say t=8t = 8 s) and exits 130; the activity is reported retryable (WorkerShutdown), the pod exits at t≈9t \approx 9 s, far inside 75 s, and the new pod resumes from the checkpoint.

These are test_wal_quota_is_rendered, test_keda_scales_on_queue_depth, and test_worker_drains_before_it_is_killed; the burst itself runs on kind in MS-durable.

FileWhat it holds
deploy/docker/durable.Dockerfilestatic build of ./cmd/durable onto a digest-pinned alpine, USER 10001, ports 7233, 7234, 9464
deploy/docker/worker.Dockerfilestatic build of ./cmd/worker copied onto a digest-pinned Python base with numpy, pyarrow, zstandard, and python/
deploy/helm/<system>-durable/values.schema.json (the contract), StatefulSet, Service and headless Service, ConfigMap with runtime.toml, PDB
deploy/helm/<system>-worker/values.schema.json, Deployment, ConfigMap, PDB, ScaledObject
deploy/keda/Chart.yaml with the pinned keda dependency, Chart.lock, values.yaml
TestKINDChecks
test_layout_presentunitevery file above exists
test_dockerfiles_follow_the_image_rulesunitdep.01 rules, digest pins, exec-form HEALTHCHECK, labels; the worker image carries python/ and builds ./cmd/worker
test_values_schemas_are_the_contractconformanceboth schemas equal the contracts
test_charts_lintconformancehelm lint of both charts
test_schema_rejects_unsafe_valuesboundaryeight unsafe values each fail helm template
test_policy_over_rendered_chartsunitprobes, resources, non-root, pinned images, PDBs, no literal secrets
test_durable_is_a_statefulset_with_a_wal_volumeunitone-replica StatefulSet, a claim template mounted where wal_dir points, fsGroup
test_wal_quota_is_renderedunitsection 3’s quota as an integer, following --set, below the claim size
test_durable_service_portsunitNodePort 30733 for gRPC 7233, port 9464
test_worker_drains_before_it_is_killedboundarygrace period at least 60 s
test_worker_writes_artifactsunit/artifacts read-write, not an emptyDir; read-only root filesystem
test_keda_scales_on_queue_depthunitthe ScaledObject’s target, bounds, query, and threshold; no replicas, no HPA
test_keda_is_pinnedconformancedeploy/keda Chart.yaml and Chart.lock match the pin
test_images_buildunitboth images build from your entry points
test_images_run_as_numeric_non_rootunitimage USER numeric and not 0
test_worker_image_has_the_python_unitsunitthe worker image imports numpy and tinyllm.io.activity
test_keda_installedconformanceKEDA’s CRD and release on kind
test_server_side_dry_runconformanceboth charts pass kubectl apply --dry-run=server
test_releases_readyconformancedurable ready with a bound WAL claim, worker available, KEDA’s HPA exists
PitfallSymptomCaught by
1. the server as a Deployment with an emptyDirevery workflow vanishes on the first rescheduletest_durable_is_a_statefulset_with_a_wal_volume
2. no fsGroupthe non-root server cannot open its WAL on a fresh volume: CrashLoopBackOfftest_durable_is_a_statefulset_with_a_wal_volume
3. toToml of a float quotathe server rejects wal_max_bytes = 1.073741824e+09 at starttest_wal_quota_is_rendered
4. quota above the volume sizethe disk fills before RESOURCE_EXHAUSTED trips; the WAL tearstest_wal_quota_is_rendered
5. the default 30 s grace on workersevery rollout and scale-down kills training between checkpointstest_worker_drains_before_it_is_killed
6. /artifacts read-only in the workerevery subprocess activity fails writing its first outputtest_worker_writes_artifacts
7. replicas set beside a ScaledObject, or a second HPAeach helm upgrade resets the count; two controllers flaptest_keda_scales_on_queue_depth
8. scaling on CPUa 200-task burst waits on two idle-looking workerstest_keda_scales_on_queue_depth
9. --test-clock or :latest in a charta test seam in production; kind runs whatever it cachedtest_schema_rejects_unsafe_values
10. a worker image without the Python unitsevery activity exits 1 with an ImportError, only on kindtest_worker_image_has_the_python_units
DirectionModuleHow it uses this
Backdep.03the chart policy, the observability stack, and the Prometheus that KEDA queries
Backdur.09the worker image carries the units the subprocess runner execs
Forwardobs.05control-plane dashboards read the durable server’s queue and DLQ gauges and the worker traces
ForwardMS-durablethe kind steps install these charts and run the kill loop against them
Forwardops.02, ops.03, ops.11drills durable-kill9, poison-task, eventlog-disk-full target these releases
Forwarddep.07the agent chart follows the worker chart
Your pieceProduction equivalentWhat it addsWhere to look
one-replica StatefulSet on local-pathTemporal on Cassandra or PostgreSQL with history shardsstateless frontends, shard ownership over a replicated storeTemporal docs, “Temporal Platform: persistence”
walMaxBytes as quotaa CSI driver that enforces capacity, volume expansionreal per-volume limits and online growthKubernetes “Volume expansion”
KEDA on queue depthTemporal worker autoscaling on schedule-to-start latencyscales on how long tasks wait, not how manyTemporal docs, “Worker performance”
hostPath /artifactsobject storage (S3, GCS) or a ReadWriteMany volumemany nodes, durable outputs, no node affinityCSI drivers for S3, ReadWriteMany access modes