Skip to content

Durable agent runs (AgentRun)

Moduleag.05 · build · Go · Pass 10 · 4 h
You buildgo/agent/durableagent/journal.go (Record, Store, FileStore, Journal, ErrIndeterminate, Reconcile), go/agent/durableagent/run.go (Config, Run, Signal, Outcome)
Contractthe step names and the StepRunner seam of ag.01; the Go API is section 4
Testscourse/tests/go/ag_05/ (what they check: section 4)
Needsag.01 (StepRunner, StepToolName), ag.02 (tool.Result), ag.03 (the loop), ag.04 (the gate and IsWrite) · reading: Durable Orchestration & Workers
Used byag.09 runs suite subjects as durable runs (DurableSubject) · dep.07 deploys the agent worker when it lands
MilestoneMS-agent
Optional depthTemporal, Workflow execution: event history and replay (free); Helland, Life beyond Distributed Transactions (ACM Queue, 2016, free)
  • A durable run is a journal of steps: each model call and tool call records its start, then its result, before the result is used (TestHandExample).
  • On restart, completed steps are replayed from the journal: the model is not called again and no tool runs twice (TestKillResumesWithoutRecall).
  • A step that started and never completed is re-run when it only reads (a model call, a lookup) and never guessed at when it writes: the run stops with ErrIndeterminate until a human reconciles it (TestIndeterminateWrite, TestReadStepRerunsAfterCrash).
  • Approvals and reconciliations are signals: records appended to the journal, applied at the next attempt (TestHandExample, TestIndeterminateWrite).
  • A crash can tear the last record: the journal ignores a torn tail and cuts it off before the next append (TestFileStoreTornTail).
Terminal window
ol start ag.05
ol tests ag.05
ol check ag.05 # ag.01 to ag.04 smoke tests run first
ol diff ag.05

An agent run that files a ticket, waits an hour for a human to approve it, and then reports back outlives any single process: the worker is redeployed, evicted, or OOM-killed in the middle. Started again from scratch, the run pays for every model call a second time and, worse, repeats side effects: two tickets, two emails, two refunds. You built a durable engine in Pass 8 for exactly this problem; this module applies its core idea (record every step, replay instead of re-executing) to the agent loop, through the StepRunner seam the loop already has.

A run is identified by a run id and owns an append-only list of records:

RecordWrittenHolds
run.startedonce, firstthe input messages
step.startedbefore a step runsthe step name and kind
step.completedafter it returns, before its result is usedthe result bytes
signalby Signal, any timeapprove with a marker, or reconcile with a step and an outcome
run.completedwhen the loop answersthe final message

The run’s state is a fold over its records. That is event sourcing, the same idea as the durable engine’s history (Pass 8): nothing else is stored, so nothing else can disagree.

2.2 Durable means fsynced, and crashes tear

Section titled “2.2 Durable means fsynced, and crashes tear”

Append returns only after the record is on disk (f.Sync()). A crash in the middle of a write can leave half a line at the end of the file. That record was never acknowledged, so Load ignores a final line without its newline, and the next Append truncates the file to the last newline before writing. A corrupt line in the middle (not a torn tail) is an error: something other than a crash damaged the file.

The run id becomes a file name (<dir>/<run id>.jsonl), and it arrives from a CLI flag or an API request, so it is validated: letters, digits, _, ., -, starting with a letter or digit, at most 128 characters. ../escape never reaches the file system.

SymbolMeaning
S(s)S(s)a step.started record exists for step ss
C(s)C(s)a step.completed record exists for ss
W(s)W(s)ss is a tool step whose tool the gate classifies as a write (IsWrite)

RunStep(s) decides:

CaseAction
C(s)C(s)return the recorded result; do not call fn
S(s)∧¬C(s)∧W(s)S(s) \wedge \neg C(s) \wedge W(s), reconciled donerecord the human’s result as completed; do not call fn
S(s)∧¬C(s)∧W(s)S(s) \wedge \neg C(s) \wedge W(s), reconciled retryrun it again
S(s)∧¬C(s)∧W(s)S(s) \wedge \neg C(s) \wedge W(s), not reconciledErrIndeterminate
otherwiserecord start, call fn, record result, return it

The crash between “the tool did its thing” and “the result was recorded” is the hard case. For a read (a model call, a lookup) running it again is harmless. For a write nobody knows whether it happened: re-running could file the ticket twice, skipping could lose it. The journal refuses to guess. The run reports Indeterminate with the step’s name, and a human checks (is there a ticket?) and sends reconcile with done and the result, or retry. The step.started record must therefore be written before fn runs; written after, a crash inside the write leaves no trace and the write is repeated.

An fn error is not recorded: the next attempt runs the step again (a transient model error should not become permanent). The loop (ag.03) ends the run on any StepRunner error, so ErrIndeterminate reaches Run instead of being shown to the model.

Run(ctx, cfg, id, input) loads the records. A completed run returns its recorded answer. A new run records run.started; an existing run with different input is ErrInputMismatch (a retried start with the same input is harmless). Then it invokes the loop from the recorded input with the journal as StepRunner, the gate, and every approved marker from approve signals. Completed steps replay; the first step without a result runs. The loop’s outcome becomes the run’s: Completed (recorded), AwaitingApproval with the pending markers, or Indeterminate with the step.

The model looks up the access policy, then files a ticket (a write), then answers. Run a1:

AttemptRecords appendedModel calls (total)Outcome
Run 1run.started; start and result of llm/1 (call lookup), tool/1/0/lookup, llm/2 (call create_ticket)2AwaitingApproval: the gate wants a human for create_ticket
Signal(approve, marker)signal approve2
Run 2start and result of tool/2/0/create_ticket, llm/3; run.completed3Completed: “Ticket filed.”
Run 3none3the recorded answer

After Run 1 the journal holds exactly seven records (the input, then two per step). In Run 2, llm/1, tool/1/0/lookup, and llm/2 come back from the journal (no model call, no second lookup), the approved marker lets the ticket through, and only llm/3 calls the model. This is TestHandExample.

Now kill the worker inside create_ticket after the ticket exists. The journal ends with step.started tool/2/0/create_ticket and no result. The next worker replays to that step and stops: Indeterminate, step tool/2/0/create_ticket, still one ticket. reconcile with done and TICKET-7 records the result, the next worker replays it as the tool’s answer, and llm/3 sees “TICKET-7”. This is TestKillResumesWithoutRecall, with real SIGKILLs of a child process.

package durableagent // import "tinyllm/agent/durableagent"
type Record struct { Seq int64; Type, Step string; Kind types.StepKind; Data json.RawMessage }
type Store interface {
Load(ctx context.Context, runID string) ([]Record, error)
Append(ctx context.Context, runID string, r Record) error // durable when it returns
}
func ValidRunID(id string) bool
func OpenFileStore(dir string) (*FileStore, error)
func NewJournal(store Store, runID string, records []Record, isWrite func(string) bool) *Journal
func (j *Journal) RunStep(ctx context.Context, name string, kind types.StepKind, fn func(context.Context) ([]byte, error)) ([]byte, error)
var ErrIndeterminate error
type IndeterminateError struct{ Step string } // Unwrap() is ErrIndeterminate
type Reconcile struct { Step, Outcome, Result string } // Outcome "done" or "retry"
type Config struct { Agent loop.Config; Gate *gate.PolicyGate; Store Store; Options []loop.Option }
type Outcome struct { Status Status; Message types.Message; Pending []loop.Pending; Step string }
var ErrInputMismatch error
func Run(ctx context.Context, cfg Config, runID string, input []types.Message) (Outcome, error)
func Signal(ctx context.Context, store Store, runID, name string, payload any) error // "approve" {"Marker": m}, "reconcile" Reconcile

Use only the standard library and the ag.01 to ag.04 packages.

TestKINDChecksWhy it matters downstream
TestHandExampleunitsection 3: the seven records, approval, replay with no new model calls, a completed run calls nothingyou and the tests agree on the journal
TestIndeterminateWritefaulta crash between the write and its record: Indeterminate until reconciled; done never re-runs the write, retry runs it onceno double tickets, no lost ones
TestReadStepRerunsAfterCrashfaultan unrecorded read is re-run; the model call before it is notreads do not block the run
TestInputMismatchboundarysame id and input resumes, other input is refused, an unknown run cannot be resumeda run id names one run
TestFileStoreTornTailfaulta half-written last line is ignored and cut by the next appenda crash mid-write does not brick the run
TestRunIDValidationboundarypath-like ids are refused; nothing is written outside the store; signals need an existing runrun ids come from users
TestHelperAgentWorkerfaultthe child process of the kill test (skipped when run alone)
TestKillResumesWithoutRecallfaultSIGKILL inside a read and inside a write across five workers: 3 model calls, 1 ticketthe module’s promise under a real kill
PitfallSymptomCaught by
1. completed steps run againevery restart pays for every model call; tools repeatTestHandExample, TestKillResumesWithoutRecall (mutant s01)
2. an unrecorded write re-run, an unrecorded read treated as a write, or the start recorded after the steptwo tickets after a crash; runs stuck on harmless readsTestIndeterminateWrite, TestReadStepRerunsAfterCrash, TestKillResumesWithoutRecall (mutants s02, s03, s14)
3. a torn last record kept, or treated as corruptionthe next record is glued onto garbage; the run cannot loadTestFileStoreTornTail (mutants s05, s06)
4. a done reconciliation that runs the write anywaythe human said it happened; it happens againTestIndeterminateWrite (mutant s07)
5. approve signals recorded but not appliedthe run waits for approval foreverTestHandExample (mutant s08)
6. a reused run id answering the old questiona client retrying with a new prompt gets the previous run’s answerTestInputMismatch (mutant s10)
7. run ids used as paths unchecked../../somewhere writes outside the storeTestRunIDValidation (mutant s11)
DirectionModuleHow it uses this
Backag.01StepRunner, StepKind, StepToolName
Backag.02a reconciled write is recorded as a tool.Result
Backag.03Run invokes the loop with the journal as its step runner and resumes from the recorded input
Backag.04IsWrite decides which unrecorded steps are writes
Forwardag.09DurableSubject runs each case as a durable run, so a killed eval worker resumes without re-calling the model
Forwarddep.07the deployed agent worker runs durable agent workflows

The catalog also places AgentRun on the durable engine (dur.06 replay, dur.08 signals); until those modules are registered the run keeps its own journal, and adapting Store to the engine’s event log is the step that joins them (DEVIATIONS B121-04).

Your pieceProduction equivalentWhat it addsWhere to look
Journalsaige durable agentlocal and DBOS backends, WAL stores, budget reservations restored on replaysaige agent/durable/{local,dbos}
ErrIndeterminate and reconcilesaige ErrIndeterminate + Reconcilethe same refusal to guess, with fenced leases for remote workerssaige agent/durable/local
approvals as signalsTemporal signals and updatesvalidated synchronous updates; the worker is released while waiting (ErrSuspended)Temporal message passing (free)
replayTemporal workflow replaydeterminism checks on code changes, GetVersion, history pagingTemporal docs (free)