Skip to content

Platform workflows TrainRun and EvalSuite

Moduledur.11 · build · Go · Pass 8 · 4 to 6 h
You buildgo/workflows/train_run.go: TrainRun, Segments, TrainOptions; go/workflows/eval_suite.go: EvalSuite, EvalOptions, MarshalSpec; then the activity form of your {tinyllm} CLI (train --spec and eval --spec, entry-point territory) and {ctl} train --spec, {ctl} eval --suite
Contractcourse/contracts/formats/train-spec.schema.json, course/contracts/formats/eval-spec.schema.json, course/contracts/formats/eval-result.schema.json, and course/contracts/spec/subprocess-activity.md
Testscourse/tests/go/dur_11/ (TrainRun and EvalSuite through the replaying simulator, over fake train and eval activities with a deterministic loss; section 4)
Needsdur.08 (the Runtime seam, ErrCanceled), data.09 (TrainRun builds its corpus with CorpusBuild); reading: dur.09 (the subprocess runner and --resume), L0.6 (atomic checkpoints, the token cursor), L6.7 (the eval suites)
Used bydur.12 (ModelRelease evaluates the candidate with EvalSuite); ag.12 upgrades EvalSuite with paired experiments; C1 runs the capstone as a TrainRun through {ctl}
MilestoneMS-durable ({ctl} train --spec specs/tiny.json survives a worker kill with a bitwise-equal final loss)
Optional depthTemporal: long-running activities and heartbeating (free); PyTorch: saving and loading a general checkpoint (free)
  • A training run is a workflow of segments: train to step ee, evaluate that checkpoint, train to 2e2e, and so on. Evaluation happens at the same steps on every run, and each evaluation runs once.
  • Every segment’s activity id names its end step (train-000500), so a kill replays finished segments and redelivers the one in flight, which resumes from its last heartbeated checkpoint.
  • Resuming from a checkpoint that holds the model, the optimizer, the RNG state, and the data cursor gives exactly the loss of an uninterrupted run: the property MS-durable checks bit for bit.
  • EvalSuite runs one activity per suite: a crash in the middle of the zoo suite does not rerun the finished ppl suite.
  • A cancel stops the run after the activity in flight checkpointed; nothing is deleted.
Terminal window
ol start dur.11 # stubs go/workflows/{train_run,eval_suite}.go
ol tests dur.11
ol check dur.11
ol diff dur.11

Then wire your entries: {tinyllm} train --spec <dir>/spec.json --progress <file> [--resume <ckpt>] under tinyllm.io.activity.run (resume from --resume, else from <run>/ckpt/LATEST), {tinyllm} eval --spec, the activities train and eval in your worker over the dur.09 runner, and {ctl} train --spec <file> starting TrainRun with workflow id train/<name>.


You can train (L0.5, L0.6) and evaluate (L6.7) from the command line, and your durable engine can run Python as activities (dur.09). But a training run started by hand still dies with the terminal, evaluation is something you remember to run, and nothing records which checkpoint was evaluated with which result. The capstone (C1) trains for hours and is released by a workflow (dur.12) that must find evaluated checkpoints. TrainRun and EvalSuite make training and evaluation platform workflows: started by {ctl}, resumed after any kill, cancellable, and with evaluation at fixed points of the run.

SymbolMeaning
NNthe spec’s steps
eethe spec’s eval_every (0: no periodic evaluation)
bi=min⁡(i⋅e,N)b_i = \min(i \cdot e, N)the end step of segment ii, for i=1,2,…i = 1, 2, \dots until bi=Nb_i = N

Segments(N, e) is [e,2e,…,N][e, 2e, \dots, N], ending exactly at NN even when ee does not divide it; e=0e = 0 or e≥Ne \ge N is one segment. Segment ii is one train activity whose spec is the run’s spec with steps set to bib_i and nothing else changed; it resumes from the run directory’s LATEST checkpoint (the previous segment’s) and stops at bib_i. Then, when the input names suites, EvalSuite evaluates the checkpoint at bib_i. The optional eval_interval_seconds field durably waits between that evaluation and the next segment, through Runtime.Sleep; it has no effect after the last segment. The run is optionally preceded by CorpusBuild (data.09) in the same workflow, so training never starts on a missing corpus.

What diesWhat happens
the worker between activitiesthe workflow replays: finished segments and evaluations come from history
the worker during trainthe activity times out (heartbeat) and is redelivered with attempt 2 and --resume the last heartbeated checkpoint (dur.09)
the serverit recovers every run from the WAL (dur.01, dur.02) and re-enqueues what was pending

A checkpoint holds everything the next step depends on: weights, optimizer moments, the shuffle RNG state, and the token stream’s cursor (L0.6). So the resumed run takes the same steps on the same batches, and its final loss equals the uninterrupted run’s in every bit. The only waste of a kill is the steps since the last checkpoint.

CancelWorkflow (dur.08) tells the train activity in flight to stop; the runner SIGTERMs it, it checkpoints and exits 130, and TrainRun returns an error wrapping ErrCanceled. TrainRun has nothing to compensate: the checkpoints are work done, and a later run with the same spec continues from them.

One activity per suite, in the order given, each with the id eval-<tag>-<suite> (TrainRun’s tag is the segment’s end step) and a spec of that one suite. The first suite that fails for good fails the step. Each activity writes evals/<suite>/<run>/results.jsonl and summary.json (formats/eval-result.schema.json) and returns them as its outputs.

Spec {"name": "tiny", "steps": 1200, "eval_every": 500, "eval_interval_seconds": 30, "seed": 7, ...} with suites ["ppl"]. Segments(1200, 500): b1=500b_1 = 500, b2=1000b_2 = 1000, b3=min⁡(1500,1200)=1200b_3 = \min(1500, 1200) = 1200. The workflow records a 30 second durable wait after the first two evaluations. The commands, in order:

train-000500 spec steps 500 -> runs/tiny/ckpt/step-000500
eval-000500-ppl subject tiny@500
sleep 30s durable timer before the next segment
train-001000 spec steps 1000 -> resumes from step-000500
eval-001000-ppl
sleep 30s durable timer before the next segment
train-001200 spec steps 1200
eval-001200-ppl

1200 optimizer steps in all, each checkpoint evaluated once. If the worker dies during a timer wait, replay reads the recorded sleep command and resumes with the same remaining timer, without repeating the evaluation. If it dies at step 740 of train-001000 (last checkpoint 700), the activity is redelivered with --resume runs/tiny/ckpt/step-000700; steps 701 to 740 are redone, the run ends at the same final loss, and eval-000500-ppl is not rerun.

This is TestHandExampleSegmentsAndEvals.

const (
ActivityTrain = "train"
ActivityEval = "eval"
)
type TrainRunInput struct { Spec, Corpus json.RawMessage; EvalSuites []string; EvalSeed int64 }
type TrainRunResult struct { Name string; Steps int; Checkpoint string; Segments []Segment; Corpus *CorpusBuildResult }
type Segment struct { Until int; Checkpoint string; Evals *EvalSuiteResult }
func Segments(steps, evalEvery int) []int
func TrainOptions(until int) StepOptions
func TrainRun(rt Runtime, in TrainRunInput) (TrainRunResult, error)
type EvalSubject struct { ID, Model string }
type EvalSuiteInput struct { Name, Tag string; Suites []string; Subjects []EvalSubject; Seed int64 }
type EvalSuiteResult struct { Suites []SuiteOutputs }
func EvalOptions(tag, suite string) StepOptions
func EvalSuite(rt Runtime, in EvalSuiteInput) (EvalSuiteResult, error)
func MarshalSpec(in EvalSuiteInput, suite string) ([]byte, error)
TestKINDChecksWhy it matters downstream
TestHandExampleSegmentsAndEvalsunitsection 3’s ids, checkpoints, 1200 steps, one evaluation per checkpointdur.12 finds every checkpoint’s evaluation
TestEvalIntervalUsesDurableSleepAndReplaysOncefaultthe configured gap is recorded between evaluations and is not repeated on replaydur.07 fires the timer once
TestSegmentsMathboundaryboundaries end at NN, no repeats, e=0e = 0 and e≥Ne \ge Nevaluation points are the same on every run
TestEachSegmentGetsTheSpecWithItsEndunitonly steps changes between segmentsreproducibility lives in the spec
TestKillLoopResumesBitwisefaulta crash at every position: same result and exactly the same final lossMS-durable’s bitwise check
TestRetryResumesFromLatestfaulta segment that dies halfway resumes, no steps redone beyond the last checkpointhours of training are not repeated
TestCancelStopsTheRunfaultcanceled after the current activity; no further segment; nothing deleteda cancelled capstone keeps its checkpoints
TestCorpusFirstunitCorpusBuild runs first; its failure stops trainingno training on a half-built corpus
TestBadSpecFailsFastboundaryno name, no steps, negative eval_every or eval_interval_seconds: non-retryable, no activitya typo does not burn a GPU-hour
TestEvalSuiteOneActivityPerSuiteunitone activity per suite, replay after a crash, the spec of one suite, a failed suite stops the restthe zoo suite’s crash keeps ppl’s results
TestOptionsIDsunittrain-000500, eval-000500-ppl, eval-zoo; heartbeat timeouts setidempotency keys are unique per segment and suite
PitfallSymptomCaught by
1. the last, shorter segment dropped or a boundary repeatedthe run stops at 1000 of 1200 steps, or evaluates step 1200 twiceTestHandExampleSegmentsAndEvals, TestSegmentsMath (mutants s01, s04)
2. activity ids that do not name the segment or suitesegments share one work directory; a replay answers the wrong oneTestHandExampleSegmentsAndEvals, TestOptionsIDs (mutants s02, s03)
3. every segment trains to the endthe first segment runs the whole run; evaluations see the final modelTestEachSegmentGetsTheSpecWithItsEnd (mutant s05)
4. no retries for trainone OOM ends a day-long runTestRetryResumesFromLatest (mutant s06)
5. a cancel treated as “move on”the cancelled run trains every remaining segmentTestCancelStopsTheRun (mutant s07)
6. a failed corpus build ignoredtraining starts on missing token filesTestCorpusFirst (mutant s08)
7. no spec validationa typo fails after the corpus build, or neverTestBadSpecFailsFast (mutant s09)
8. all suites in one activitya crash in the last suite reruns every suiteTestEvalSuiteOneActivityPerSuite (mutant s10)
9. continuing after a failed suitea release gate sees partial results as completeTestEvalSuiteOneActivityPerSuite (mutant s11)
10. no heartbeat timeout on traina dead trainer is noticed after a dayTestOptionsIDs (mutant s12)
11. using a wall-clock sleep or skipping the durable cadencea worker restart repeats or loses the evaluation waitTestEvalIntervalUsesDurableSleepAndReplaysOnce (mutant s13)
DirectionModuleHow it uses this
Backdur.08Runtime, ErrCanceled
Backdata.09CorpusBuild before training
Backdur.09, L0.6the subprocess runner’s --resume and the checkpoints it resumes from
ForwardC1the capstone run is a TrainRun on the TinyStories config
Forwarddur.12ModelRelease evaluates the candidate with EvalSuite
Forwardag.12paired experiments extend EvalSuite with scorers and A/B options
Forwardobs.05one TrainRun is one trace down to train.step
Your pieceProduction equivalentWhat it addsWhere to look
segments with evaluationKubeflow Pipelines, Flyte, MetaflowDAGs of steps with caching by input hash, artifact lineageFlyte “cached tasks”
resume from LATESTPyTorch Lightning fault-tolerant training, TorchElasticrestart on node loss, rendezvous of many workers, sharded checkpointstorch.distributed.elastic
one eval activity per suitelm-evaluation-harness, HELMthousands of tasks, cached model outputs, aggregate leaderboardsEleutherAI lm-evaluation-harness