Skip to content

ModelRelease: gates, canary, PromQL burn check, rollback

Moduledur.12 · build · Go · Pass 9 · 5 to 7 h
You buildgo/workflows/model_release.go: ModelRelease, ReleaseSpec (Validate, ExportSpec, EvalInput), CheckGates, ModelCardProblems, LedgerProblems, and two activities, Preflight (release.preflight) and ReadResults (release.results); go/activities/routes.go: the admin route client Routes (Get, Put, Snapshot, Ready, Canary, Promote, Restore, BackendsOf, CanaryBackends, SameBackends); go/activities/promql.go: Prom.Query, Decode, ParseSample
Contractthe Go API of section 4 (fixed by this chapter and its tests, DEVIATIONS B112-01), speaking openapi/admin.v1.yaml (routes and workers), formats/export-spec.schema.json, formats/eval-results.schema.json, formats/ledger.schema.json, and templates/MODEL_CARD.md
Testscourse/tests/go/dur_12/: the workflow under a replaying fake of the durable SDK, the route client against a fake gateway, the query against a fake Prometheus; what they check: section 4
Needsdur.08 the workflow Runtime seam and Saga · dur.11 EvalSuite, run as the evaluation step (or --ref-deps); reading: dur.09 (export and eval run as subprocess activities) · data.08 the data ledger · gw.05 and gw.07 the admin route API · obs.03 burn-rate alerts
Used byC1 (the capstone ships through your ModelRelease; its check runs your preflight on your release); later ops.08 (the retrain decision after a data incident)
MilestoneMS-C1 ({ctl} release --spec specs/c1/release.json promotes or rolls back your capstone model)
Optional depthBeyer et al., The Site Reliability Workbook, ch. 5 “Alerting on SLOs” (burn rates, free); Prometheus HTTP API (free); Argo Rollouts analysis and Flagger (canary analysis, free); RFC 9110 section 13.1.1 (If-Match)
  • A release is a durable workflow, not a script: every step is an activity or a timer recorded in history, so a worker that dies mid-canary resumes where it was, and the 10-minute bake does not start over (TestWorkerKilledMidCanary).
  • Gates fail closed and early. The model card and the data ledger are checked before the export; a missing eval row fails its gate; a burn query that returns nothing, NaN, or an error rolls back (TestMissingEvalRowFailsTheGate, TestNoDataFailsClosed).
  • No traffic moves before a human approves, and a release nobody approves expires on the durable clock instead of holding a worker forever. Once traffic has moved, every way out but a promotion (a burn, a cancel, a failed step) puts the snapshot back through a saga compensation (TestCancelRestoresTheRoute).
  • Route changes are read-modify-write with If-Match on one route: on 412 read again and redo the change, and send every other route back byte for byte.
  • Every activity is idempotent by desired state: a retried canary finds its weight in place and does nothing, because activities run at least once.
Terminal window
ol start dur.12 # stubs go/workflows/model_release.go, go/activities/{routes,promql}.go
ol tests dur.12 # read the test catalog first
ol check dur.12 # the workflow (replayed), the route client, the PromQL query
ol diff dur.12 # after passing: your code against the reference

Then wire it in your worker’s composition root (go/cmd/worker, learner territory): register the workflow as ModelRelease through the same workflow.Context-to-Runtime adapter as TrainRun (dur.08), and register the activities under the Act* names: Preflight, ReadResults, the subprocess runner of dur.09 for export and eval, and the methods of one activities.Routes and one activities.Prom. Your umbrella CLI gets {ctl} release --spec <file> (start) and <system> wf signal <id> approve.


Your capstone model exists as a checkpoint under runs/, and your gateway serves whatever the route table says. Between the two there is only you, typing: copy the files, hope the model card is current, edit the route by hand, watch a dashboard, and undo it if something looks wrong. Every one of those steps fails in a known way. A model trained on a source whose license forbids training ships because nobody re-read the ledger. A route edit erases a colleague’s cascade because you PUT a table you read an hour ago. The canary runs for ten minutes, your laptop sleeps, and nobody checks the error budget. Your durable engine (Pass 8) already survives crashes and keeps timers and signals; this module turns the release itself into a workflow on it: gated, approved, canaried, measured, and reversible.

SymbolMeaningType
oothe SLO objective, the fraction of requests that must succeed (0.99)float64
b=1−ob = 1 - othe error budget as a fraction of requests (0.01)float64
eWe_Wthe canary’s error ratio over a window WW: 5xx responses / all responsesfloat64
BW=eW/bB_W = e_W / bthe burn rate: how many times faster than allowed the budget is spentfloat64
Bmax⁡B_{\max}the ceiling (burn.max): roll back when BW>Bmax⁡B_W > B_{\max}float64
wwthe canary’s share of the route’s traffic, 0<w<10 < w < 1float64
TTthe bake time (canary_wait_s)duration

ModelRelease is deterministic Go that decides; activities do (DESIGN 2.7). The workflow may branch only on its input and on what its Env recorded: activity results, timer firings, signals, and Env.Now(). It never reads a file, the wall clock, or the network itself, because on a restart the durable server replays the history through the same code, and any other input could take a different path (ErrNondeterminism). That is why the model card check is an activity (release.preflight) and not an os.ReadFile in the workflow.

The workflow is a function of dur.08’s Runtime, the SDK calls it needs (ExecuteActivity, Sleep, Now, AwaitSignal, Detached), like TrainRun and EvalSuite. Your worker adapts its workflow.Context to it once for all of them, and the course tests drive it with a fake that records a history, crashes and cancels on cue, and replays.

#StepHowFails as
1validate the specpureGateError{spec} before any activity
2preflight: model card, ledgeractivity release.preflightGateError (non-retryable)
3export the checkpointactivity export ({tinyllm} export --spec)activity failure
4evaluate the exported model, then gateEvalSuite (one eval activity per suite), release.results reads the reports, CheckGatesGateError{eval}
5wait for approve or rejectAwaitSignal, timeout approve_timeout_soutcome rejected or expired
6snapshot the route, then canaryroutes.snapshot, routes.canaryactivity failure
7bakeSleep(T)(durable: survives restarts)
8burn checkpromql.query at Now()no data or error: roll back
9promote or restoreroutes.promote, or the saga’s routes.restoreoutcome promoted or rolled_back

Every step carries a fixed activity id (release-preflight, release-export, eval-release-<suite>, release-results, routes-snapshot, routes-canary, promql-burn, routes-promote, routes-restore), so its idempotency key <workflow_id>/<id> is the same on every replay and retry.

Gates first because they are cheap and permanent: a forbidden license does not become allowed on retry, so a gate failure is non-retryable and the workflow fails before exporting gigabytes. The eval runs on the exported directory, not on the checkpoint: the export may quantize (quant: int4-g32), and the artifact you serve is the one you must measure.

  • Model card. MODEL_CARD.md exists, has the five sections of templates/MODEL_CARD.md, and holds no template placeholder (<model_id>, <you>) outside HTML comments. A < followed by a digit or a space (p < 0.05) is prose, not a placeholder.
  • Ledger. Every line of LEDGER.jsonl is a ledger row whose allowed_uses includes train and that is not revoked; an empty ledger fails (a model trained on undocumented data cannot ship).
  • Eval rows. Each gate names (suite, task, metric, op, value). It looks for the released model’s row in that suite’s report and passes when value op threshold holds, with op either >= or <= (inclusive). No row, a row whose status is not ok, or a row with a null value fails: a suite that did not run proves nothing.

A signal can arrive before the workflow waits for it (your operator approves while the eval still runs); the SDK buffers it (dur.08). The wait has a timeout on the durable clock; when it passes, the release ends as expired with no traffic moved.

A route either has no backends (all traffic to the model of its own name) or a list of {served_model, weight} summing to 1. The canary gives the new model the share ww and scales every other backend by 1−w1 - w. The gateway replaces the whole table on PUT and guards it with an ETag: GET returns the table and ETag: "routes-7"; PUT with If-Match: "routes-7" succeeds and returns "routes-8"; a PUT with a stale ETag is 412. So every change is:

  1. GET the table; keep every route as the raw JSON the gateway sent.
  2. Change one route’s backends (or replace it, for a restore).
  3. PUT the table with If-Match. On 412, start again at 1 (at most MaxConflicts times).

Before the canary, Ready checks that a live, non-draining worker reports the new served model (<model_id>-<version>): routing a share to a model nobody serves fails that share of requests. Until one appears the canary returns a retryable ErrNoWorkers, and the activity’s retries wait.

Right after the snapshot, the workflow registers one compensation with a dur.08 Saga: “restore the route to the snapshot”. If the run is canceled during the bake, or the canary or promotion fails for good, saga.Fail runs the compensation on rt.Detached() (so the cancel cannot interrupt the undo) and returns the original error, which still matches ErrCanceled. A burn above the ceiling runs the same compensation and ends as rolled_back. A cancel before the snapshot (during the approval wait) has nothing to undo.

An activity runs at least once: the worker can die after the PUT and before the server records the completion, and then the activity runs again. Each route activity therefore compares the table with the state it wants and does nothing when it is already there. Restoring needs the state from before the canary, so the workflow snapshots the route as its own recorded activity first; the snapshot is in history, and a replay restores exactly that.

The error budget over a 30-day SLO period of 720 h is bb of all requests. A burn rate B=1B = 1 spends it exactly in 30 days; B=14.4B = 14.4 for one hour spends 14.4⋅1/720=2%14.4 \cdot 1 / 720 = 2\% of it, the page threshold of the SRE Workbook. The release asks Prometheus once, at the end of the bake, for BWB_W of the canary’s traffic (an expression that aggregates to one number), evaluated at Env.Now() so that a retried query asks the same question. It promotes when BW≤Bmax⁡B_W \le B_{\max} and rolls back otherwise. An empty vector (no canary request matched) and NaN (0/00/0) are no data; no data, and a query that failed for good, roll back: the canary did not prove itself.

The route tinystories serves tinystories-10m-v1 with weight 1. You release v2 with w=0.1w = 0.1, T=600T = 600 s, o=0.99o = 0.99, Bmax⁡=14.4B_{\max} = 14.4; the approval arrives; the workflow clock reads t0t_0 = 2026-10-09 12:00:00 UTC (Unix 1791547200).

Canary weights. Other backends total r=1r = 1. v1 gets 1/r⋅(1−0.1)=0.91 / r \cdot (1 - 0.1) = 0.9, v2 gets 0.10.1; sum 11. (With an even split a 0.5, b 0.5: each gets 0.5/1⋅0.9=0.450.5 / 1 \cdot 0.9 = 0.45, and v2 0.1.) The gateway had ETag "routes-7"; the canary PUT sends If-Match: "routes-7" and the table becomes "routes-8".

Bake. The timer fires at t0+600t_0 + 600 s, so the query carries time=1791547800.000.

Burn, healthy. In the last 5 minutes the canary answered 1,000 requests with 8 errors: e=8/1000=0.008e = 8 / 1000 = 0.008, B=0.008/0.01=0.8≤14.4B = 0.008 / 0.01 = 0.8 \le 14.4. Promote: backends = [{tinystories-10m-v2, 1}], ETag "routes-9". Two PUTs in all. This is TestHandExampleCanaryPromotes (and the weights are TestHandExampleCanaryWeights, the query TestHandExampleBurnRate).

Burn, too fast. 1,000 requests with 205 errors: e=0.205e = 0.205, B=20.5>14.4B = 20.5 > 14.4. Restore the snapshot: the route is again exactly {"model":"tinystories","aliases":["ts"],"backends":[{"served_model":"tinystories-10m-v1","weight":1}],"x-owner":"ml-team"}, and the cascade route smart never changed. This is TestFastBurnRollsBackToTheSnapshot.

A crash. The worker dies right after the canary PUT; another picks the run up 2 minutes later, replays steps 1 to 6 from history, and runs the canary activity again: the table already says v2 0.1, so no PUT. The timer starts at t0+2t_0 + 2 min. A second worker dies with the timer pending and comes back 2 minutes later; the timer keeps its deadline t0+2t_0 + 2 min +600+ 600 s, so the query is at Unix 1791547920. This is TestWorkerKilledMidCanary.

// go/workflows/model_release.go (Runtime, StepOptions, StepFailure, Saga: dur.08; EvalSuite: dur.11)
func ModelRelease(rt Runtime, spec ReleaseSpec) (ReleaseResult, error) // *GateError, a *StepFailure, or ErrCanceled
func (s ReleaseSpec) Validate() error
func (s ReleaseSpec) ExportSpec() map[string]any // formats/export-spec.schema.json
func (s ReleaseSpec) EvalInput(modelDir string) EvalSuiteInput // the release's EvalSuite
func CheckGates(gates []Gate, reports []EvalReport, subject string) []GateResult
func ModelCardProblems(text string) []string
func LedgerProblems(ledger []byte) []string
func Preflight(ctx context.Context, in PreflightInput) (PreflightResult, error) // activity release.preflight
func ReadResults(ctx context.Context, in ResultsInput) (EvalResult, error) // activity release.results
// go/activities/routes.go
func (r Routes) Get(ctx) ([]json.RawMessage, string, error) // the routes as sent, and the ETag
func (r Routes) Put(ctx, routes []json.RawMessage, etag string) (string, error) // ErrStale on 412
func (r Routes) Snapshot(ctx, SnapshotInput) (RouteSnapshot, error) // activity routes.snapshot
func (r Routes) Ready(ctx, served string) (bool, error) // GET /admin/v1/workers
func (r Routes) Canary(ctx, CanaryInput) (RouteResult, error) // activity routes.canary
func (r Routes) Promote(ctx, PromoteInput) (RouteResult, error) // activity routes.promote
func (r Routes) Restore(ctx, RestoreInput) (RouteResult, error) // activity routes.restore
func BackendsOf(route map[string]any) ([]Backend, error)
func CanaryBackends(current []Backend, served string, w float64) []Backend
func SameBackends(a, b []Backend) bool
// go/activities/promql.go
func (p Prom) Query(ctx, QueryInput) (Sample, error) // activity promql.query
func Decode(body []byte) (Sample, error)
func ParseSample(pair []json.RawMessage) (float64, bool, error) // (value, isNaN, err)

An error that a retry cannot fix implements NonRetryable() bool returning true: *GateError and activities.Permanent (a route that does not exist, a weight outside (0,1)(0, 1), a table the gateway rejects with 400, a query Prometheus rejects). Your composition root maps it to Failure.non_retryable. Transport errors, ErrStale after MaxConflicts, ErrNoWorkers, and a 5xx from Prometheus are retryable.

A release spec (specs/c1/release.json in the capstone):

{"model_id": "tinystories-10m", "version": "v2", "from": "runs/c1/ckpt",
"model_card": "docs/MODEL_CARD.md", "ledger": "corpus/LEDGER.jsonl",
"suites": ["quality", "safety"], "eval_seed": 0,
"gates": [{"suite": "quality", "task": "ts-val", "metric": "bpb", "op": "<=", "value": 1.30},
{"suite": "safety", "task": "refusal", "metric": "score", "op": ">=", "value": 0.9}],
"route": "tinystories", "canary_weight": 0.1, "canary_wait_s": 600, "approve_timeout_s": 86400,
"burn": {"query": "sum(rate(...{model=\"tinystories-10m-v2\",http_response_status_code=~\"5..\"}[5m])) / sum(rate(...{model=\"tinystories-10m-v2\"}[5m])) / (1 - 0.99)", "max": 14.4}}
TestKINDChecksWhy it matters downstream
TestHandExampleCanaryPromotesunitsection 3: the step order and activity ids, 0.9/0.1 during the bake, the query at t0+600t_0 + 600 s, promotion, two PUTs, the eval runs on the exported directoryyou and the test agree on the release
TestFastBurnRollsBackToTheSnapshotfaultburn 20.5: the route is restored to its snapshot, the other route untoucheda bad model leaves no trace in routing
TestBurnAtTheCeilingPromotesboundaryburn equal to the ceiling promotesthe ceiling is inclusive
TestNoDataFailsClosedfaultan empty vector and NaN roll backa canary with no traffic proves nothing
TestFailedBurnQueryFailsClosedfaulta rejected query rolls back after one trya broken check is not a pass
TestUnlicensedSourceFailsBeforeExportfaultan eval-only source fails the gate once, before exportlicenses are checked before any work
TestRevokedSourceFailsfaulta revoked source fails the gatethe ops.08 data incident blocks releases
TestModelCardGatefaultmissing card, missing section, leftover placeholder; p < 0.05 is fineethics.03’s card is a real gate
TestMissingEvalRowFailsTheGatefaultno safety row: the eval gate refuses; no signal consumed, no PUTethics.04’s rows must exist to pass
TestGateDirectionsboundary<= and >= inclusive; another model’s row ignoredlower-is-better and higher-is-better metrics
TestRejectStopsBeforeTrafficunitreject ends the release with no route accesshumans can stop a release
TestApprovalTimesOutunitno answer: expired after exactly approve_timeout_s on the durable clocknothing waits forever
TestWorkerKilledMidCanaryfaultcrash after the canary PUT and during the timer: one canary PUT, the timer keeps its deadlinereleases survive worker loss
TestCancelRestoresTheRoutefaulta cancel in the approval wait touches nothing; a cancel in the bake restores the snapshot and never queriesan operator can stop a release safely
TestReplayIsDeterministicregressionreplaying a finished history with no live activities gives the same resultthe workflow obeys the replay rules
TestSpecIsValidatedFirstboundaryweight 1, an empty query, op >: refused before any activitya bad spec moves nothing
TestHandExampleCanaryWeightsunit1 to 0.9/0.1; 0.5/0.5 to 0.45/0.45/0.1; re-weighting a canarythe gateway’s sum-to-1 rule
TestImplicitBackendIsTheRouteModelboundarya route without backends keeps 0.75 for its own modelthe old model keeps its traffic
TestPutCarriesTheETagconformanceone PUT with If-Match: "routes-7"admin.v1 optimistic concurrency
TestStaleETagRereadsAndKeepsTheOtherChangefaulta concurrent edit: 412, reread, both changes surviveoperators and releases can overlap
TestOtherRoutesAreUntouchedregressionthe cascade route is byte-identical; unknown fields kepta release cannot erase routing it did not own
TestCanaryIsIdempotentpropertya second canary changes nothingat-least-once activities
TestCanaryWaitsForAServingWorkerfaultno live worker (or a draining one): retryable ErrNoWorkers, no PUTno traffic to a model nobody serves
TestUnknownRouteAndBadInputArePermanentboundaryunknown route, weight 1, a 400 table: non-retryablefail fast on what retries cannot fix
TestPromoteAndRestoreunitpromote to weight 1; restore exactly and idempotently; another route’s snapshot refusedthe two ends of a release
TestHandExampleBurnRateunitsection 3’s query: value 20.5, query and time sentthe measurement the decision rests on
TestScalarAndInfinityunita scalar result; +Inf is a value, not no datascalar(...) expressions; all-error canaries
TestNoSeriesAndNaNAreNoDataboundaryempty vector and NaN are no datafail closed (pitfall 11)
TestSeveralSeriesAreRefusedboundarytwo series or a matrix: permanent ErrQueryan unaggregated query checks a random pod
TestErrorStatusIsPermanentAndServerErrorsRetryfaultbad_data is permanent; a 503 is retryableone Prometheus blip must not roll back

The workflow tests run ModelRelease under a fake Runtime that keeps a history, crashes and cancels on cue, and replays from the top after every crash, re-executing only activities whose completion was not recorded. The clock is virtual: no test sleeps. Its eval activity writes one tl.eval-results.v1 report per suite, which your ReadResults reads.

PitfallSymptomCaught by
1. gate results computed but not enforceda model that failed its safety suite takes trafficTestMissingEvalRowFailsTheGate (mutant s01)
2. a gate with no row passesthe safety suite crashed and the release went outTestMissingEvalRowFailsTheGate (mutant s02)
3. >= written as >a model exactly at its threshold is refusedTestGateDirections (mutant s03)
4. the ledger check reads eval as enough, or skips revokeda non-commercial or revoked source shipsTestUnlicensedSourceFailsBeforeExport, TestRevokedSourceFails (mutants s04, s05)
5. a template copied but never filled passes<model_id> in a published model cardTestModelCardGate (mutant s06)
6. a gate failure returned as a plain errorthe gate is retried three times, then reported as an outageTestUnlicensedSourceFailsBeforeExport (mutant s07)
7. evaluating the checkpoint, not the exported modelthe int4 export ships unmeasuredTestHandExampleCanaryPromotes (mutant s08)
8. reject, timeout, or the wait itself missinga rejected model ships; a release waits forever; traffic moves unapprovedTestRejectStopsBeforeTraffic, TestApprovalTimesOut (mutants s09, s10, s11)
9. the snapshot taken after the canary, a rollback branch that never restores, or a restore that mergesrollback keeps the canary’s backendsTestFastBurnRollsBackToTheSnapshot, TestPromoteAndRestore (mutants s12, s27, s37)
10. no bake, or a query at no fixed timethe burn is measured before the canary served anything; a retry measures another windowTestHandExampleCanaryPromotes, TestBurnAtTheCeilingPromotes, TestWorkerKilledMidCanary (mutants s13, s33)
11. no data or a failed query read as healthya canary that served nothing, or that Prometheus could not see, is promotedTestNoDataFailsClosed, TestFailedBurnQueryFailsClosed (mutants s15, s16, s29)
12. a strict ceilinga canary exactly at the budget line rolls backTestBurnAtTheCeilingPromotes (mutant s17)
13. PUT without If-Match, or giving up on 412428 on every change; a concurrent edit kills the releaseTestPutCarriesTheETag, TestStaleETagRereadsAndKeepsTheOtherChange (mutants s18, s19)
14. the table rebuilt instead of returned as read>= comes back as >=; other routes vanishTestOtherRoutesAreUntouched (mutants s20, s21)
15. a canary that is not idempotenta retried activity makes 0.9/0.1 into 0.81/0.09/0.1TestCanaryIsIdempotent, TestWorkerKilledMidCanary (mutant s22)
16. weights not rescaled, or a missing backends read as nonethe gateway rejects the table (sum not 1); the old model loses its shareTestHandExampleCanaryWeights, TestImplicitBackendIsTheRouteModel (mutants s23, s24)
17. no check that a worker serves the new model10% of requests fail with no backendTestCanaryWaitsForAServingWorker (mutant s25)
18. permanent errors returned as retryablean unknown route or a bad expression retries for minutesTestUnknownRouteAndBadInputArePermanent, TestFailedBurnQueryFailsClosed (mutants s26, s34)
19. the first of several series usedthe check reads one random podTestSeveralSeriesAreRefused (mutant s30)
20. scalars refused, or a 503 treated as permanentscalar(...) queries fail; one overload rolls back a good modelTestScalarAndInfinity, TestErrorStatusIsPermanentAndServerErrorsRetry (mutants s31, s32)
21. a canary weight of 1 acceptedthe “canary” is a full cutover with no way back but a rollbackTestSpecIsValidatedFirst (mutant s35)
22. a cancel that does not compensatea canceled release leaves 10% of traffic on an unapproved modelTestCancelRestoresTheRoute (mutant s36)
23. activity ids left to the runtime’s counteran added step renumbers every later idempotency keyTestHandExampleCanaryPromotes, TestCancelRestoresTheRoute (mutant s39)
24. time.Now(), rand, or map order in the workflowreplay diverges after a restart (ErrNondeterminism)TestReplayIsDeterministic
DirectionModuleHow it uses this
Backdur.08the Runtime the workflow runs on, and the Saga that undoes the canary
Backdur.11EvalSuite is the evaluation step
Backdur.09export and eval are subprocess activities run by your runner
Backdata.08the ledger the preflight gate reads
Backgw.05, gw.07the admin route and worker API this client speaks
Backobs.03the burn-rate expressions and thresholds of your SLO rules
ForwardC1your capstone model is released through this workflow; C1’s check runs your Validate and Preflight on its release
Forwardops.08the data-incident drill revokes a source and expects the next release to fail its ledger gate
Your pieceProduction equivalentWhat it addsWhere to look
ModelReleaseArgo Rollouts, Flaggerprogressive steps (1%, 5%, 25%, …), several metric analyses per step, automatic abort, traffic mirroringargo-rollouts/rollout/canary.go, Flagger pkg/controller/scheduler.go
the burn checkKayenta (Spinnaker)statistical canary analysis: the canary compared with a baseline started at the same time, Mann-Whitney per metrickayenta-judge
RuntimeTemporal Go SDKworkflow.Context, workflow.NewSelector over signals and timers, versioning with GetVersiongo.temporal.io/sdk/workflow
gatesMLflow model registry, Vertex AI Model Registrystage transitions with approval, lineage to data and runsMLflow model_registry docs
RoutesEnvoy AI Gateway, Gateway API HTTPRoute weightsweighted backends as Kubernetes objects; the API server’s resourceVersion is the ETagGateway API backendRefs[].weight