Skip to content

Checkpoints and released models

A model directory is what the engine serves and what {tinyllm} loads: model.safetensors (safetensors.md), config.json, and, when config.json says "tl_tokenizer": "file", tokenizer.json (tokenizer.md). A checkpoint and a released model are model directories with more files. All paths are relative to /artifacts.

runs/<run_id>/ckpt/step-<nnnnnn>/ (six digits, the optimizer step), holding:

FileContent
model.safetensorsthe parameters, F32 (the master weights, also under bf16 training)
optimizer.safetensorsper parameter <name>: AdamW <name>.exp_avg and <name>.exp_avg_sq; SGD with momentum and Muon <name>.momentum_buffer; all F32, metadata {"format": "tinyllm", "optimizer": "<name>", "step": "<step>"}
config.jsonthe model config, verbatim from the train spec
tokenizer.jsonwhen the model uses one
generation_config.jsonoptional (generation-config.schema.json)
trainer_state.jsontrainer-state.schema.json: step, tokens seen, the data cursor, the generator state, the learning rate, the spec’s sha256, the git sha
MANIFEST.jsonmanifest.schema.json, kind: checkpoint: every other file with its sha256 and size

runs/<run_id>/ckpt/LATEST is a text file holding the name of the newest complete step directory and a newline, for example step-000500\n.

A crash at any instant leaves either the previous checkpoint or the new one, never a mix:

  1. Write every file into step-<nnnnnn>.tmp/, fsync each file.
  2. Write MANIFEST.json last, fsync it and the directory.
  3. rename step-<nnnnnn>.tmp to step-<nnnnnn>.
  4. Write LATEST.tmp, fsync, rename it to LATEST, fsync the ckpt/ directory.
  5. Emit {"kind": "ckpt", ...} on the progress file (spec/subprocess-activity.md) only now.
  6. Delete step directories beyond the spec’s keep_ckpts newest, and any stale *.tmp.

--resume <step dir> (or, without a path, the directory LATEST names) loads the parameters, optimizer state, and trainer_state.json, after checking every file against MANIFEST.json. A directory whose manifest is missing or does not verify is skipped for the newest older one that verifies. Resuming continues bitwise as the uninterrupted run would have: the data cursor names the next window, the shuffle generator state continues its stream (spec/pcg32.md), and the step count drives the schedule. dur.11’s kill test compares the final loss of a resumed run with an uninterrupted one on the smoke config.

models/<model_id>/<version>/, written by ModelRelease (dur.12) from one checkpoint:

FileContent
model.safetensorsweights, F32 or BF16, or int4 (<name>.qweight and <name>.scales, safetensors.md)
config.json, generation_config.json, tokenizer.json, tokenizer_config.jsonas in the checkpoint; tokenizer_config.json when a chat template is needed
MODEL_CARD.mdfrom templates/MODEL_CARD.md; the release gate fails without it
ledger.jsona JSON array of the ledger rows of every source the training data came from; the gate fails when any lacks train in allowed_uses or is revoked
evals/summary.jsonthe EvalSuite summary (eval-result.schema.json $defs/summary) the gate read
heads/<name>.jsonoptional linear heads over this model’s embeddings (linear-head.schema.json)
MANIFEST.jsonkind: release, with model_id and version

A released directory is immutable: a new version is a new directory.