Checkpoints and released models
A model directory is what the engine serves and what {tinyllm} loads: model.safetensors (safetensors.md), config.json, and, when config.json says "tl_tokenizer": "file", tokenizer.json (tokenizer.md). A checkpoint and a released model are model directories with more files. All paths are relative to /artifacts.
Checkpoint step directory
Section titled “Checkpoint step directory”runs/<run_id>/ckpt/step-<nnnnnn>/ (six digits, the optimizer step), holding:
| File | Content |
|---|---|
model.safetensors | the parameters, F32 (the master weights, also under bf16 training) |
optimizer.safetensors | per parameter <name>: AdamW <name>.exp_avg and <name>.exp_avg_sq; SGD with momentum and Muon <name>.momentum_buffer; all F32, metadata {"format": "tinyllm", "optimizer": "<name>", "step": "<step>"} |
config.json | the model config, verbatim from the train spec |
tokenizer.json | when the model uses one |
generation_config.json | optional (generation-config.schema.json) |
trainer_state.json | trainer-state.schema.json: step, tokens seen, the data cursor, the generator state, the learning rate, the spec’s sha256, the git sha |
MANIFEST.json | manifest.schema.json, kind: checkpoint: every other file with its sha256 and size |
runs/<run_id>/ckpt/LATEST is a text file holding the name of the newest complete step directory and a newline, for example step-000500\n.
Writing atomically
Section titled “Writing atomically”A crash at any instant leaves either the previous checkpoint or the new one, never a mix:
- Write every file into
step-<nnnnnn>.tmp/,fsynceach file. - Write
MANIFEST.jsonlast,fsyncit and the directory. renamestep-<nnnnnn>.tmptostep-<nnnnnn>.- Write
LATEST.tmp,fsync,renameit toLATEST,fsynctheckpt/directory. - Emit
{"kind": "ckpt", ...}on the progress file (spec/subprocess-activity.md) only now. - Delete step directories beyond the spec’s
keep_ckptsnewest, and any stale*.tmp.
Resuming
Section titled “Resuming”--resume <step dir> (or, without a path, the directory LATEST names) loads the parameters, optimizer state, and trainer_state.json, after checking every file against MANIFEST.json. A directory whose manifest is missing or does not verify is skipped for the newest older one that verifies. Resuming continues bitwise as the uninterrupted run would have: the data cursor names the next window, the shuffle generator state continues its stream (spec/pcg32.md), and the step count drives the schedule. dur.11’s kill test compares the final loss of a resumed run with an uninterrupted one on the smoke config.
Released model
Section titled “Released model”models/<model_id>/<version>/, written by ModelRelease (dur.12) from one checkpoint:
| File | Content |
|---|---|
model.safetensors | weights, F32 or BF16, or int4 (<name>.qweight and <name>.scales, safetensors.md) |
config.json, generation_config.json, tokenizer.json, tokenizer_config.json | as in the checkpoint; tokenizer_config.json when a chat template is needed |
MODEL_CARD.md | from templates/MODEL_CARD.md; the release gate fails without it |
ledger.json | a JSON array of the ledger rows of every source the training data came from; the gate fails when any lacks train in allowed_uses or is revoked |
evals/summary.json | the EvalSuite summary (eval-result.schema.json $defs/summary) the gate read |
heads/<name>.json | optional linear heads over this model’s embeddings (linear-head.schema.json) |
MANIFEST.json | kind: release, with model_id and version |
A released directory is immutable: a new version is a new directory.