Pass 5 milestones
Pass result: transformer, GPT/BERT/ELECTRA/LoRA, the model zoo, modern block; SmolLM2-135M loads and matches HF.
The path includes the milestone stages below. Run each component milestone after its modules pass, then run the pass gate. The component gates run before the pass gate, which also reruns the smoke steps of earlier passes.
| Gate | Specification | What it covers |
|---|---|---|
MS-L5 | MS-L5.toml | Your transformer learns addition, and pre-LN trains where post-LN does not. Requires L5.1, L5.2, L5.3, L5.4, L5.5. |
MS-L6 | MS-L6.toml | GPT, BERT, ELECTRA, and LoRA through your CLI, scored by your model zoo. Requires L6.1, L6.2, L6.3, L6.5, L6.6, L6.7. |
MS-L7 | MS-L7.toml | Your modern decoder loads Hugging Face checkpoints and matches them. Requires L7.1, L7.2, L7.3, L7.4, L7.5, L7.6, L7.7, L7.8, L7.9. |
MS-P5 | MS-P5.toml | Transformers: the 2017 model, its objectives, and a modern decoder that matches SmolLM2. Requires M01.4, M07.5, M07.7, S-M07d, L5.1, L5.2, L5.3, L5.4, L5.5, L6.1, L6.2, L6.3, L6.6, L6.5, L6.7, M05.1, L7.1, L7.2, L7.3, L7.4, L7.5, L7.6, L7.7, L7.8, L7.9, craft.05. |
Component gate details
Section titled “Component gate details”MS-L5: Your transformer learns addition, and pre-LN trains where post-LN does not
Section titled “MS-L5: Your transformer learns addition, and pre-LN trains where post-LN does not”Your encoder-decoder Transformer (L5.1 to L5.5), trained and decoded through your own entry point: it learns three-digit addition to an exact-match rate of 0.98 with beam search, and the LayerNorm-placement demo shows why Part 7 uses pre-LN: with no warmup, a deep post-LN stack does not train while the same stack in pre-LN does.
The task (course/fixtures/small-corpora/add3.tsv, 4000 pairs, and add3-test.tsv, 500 unseen pairs, from course/oracle/MS-L5/add3.py) writes every number least significant digit first: a = 123, b = 456 is the source “321+654” and the target “9750” (579 reversed, zero padded to 4 digits). DEVIATIONS B72-01 says why it is add3, not the catalog’s add5.
This file fixes the Pass 5 transformer verbs of your tinyllm role
(spec/cli-roles.md, “Verbs of later passes”). Every verb keeps the rules of
that page: exit 2 on a usage error, the last stdout line is one JSON object.
{tinyllm} train transformer —task <pairs.tsv> —cfg <config.json> —out
source<TAB>target pairs; builds L1.1’s CharTokenizer over every
source and target with the specials steps updates of batch pairs,
shuffled by PCG32(S).substream(“shuffle”), at the Noam rate
factor * noam(step + 1, d_model, warmup) and label smoothing
smoothing. The flags override the config’s keys. —warmup 0 means no
warmup: a constant rate, given with —lr (exit 2 without it). Writes
the model directory (save_transformer) and tinyllm_char.json into {tinyllm} translate —model
ol milestone MS-L5 --smoke runs every smoke step (no services, about a
minute on a laptop). The full run trains 3500 updates (about 30 s to 2 min).
MS-L6: GPT, BERT, ELECTRA, and LoRA through your CLI, scored by your model zoo
Section titled “MS-L6: GPT, BERT, ELECTRA, and LoRA through your CLI, scored by your model zoo”Your objectives and your evaluation harness, run through your own entry point: a GPT trained on the byte corpus to the reference’s calibrated validation loss; a BERT and an ELECTRA pretrained on the same in-domain text, each then fine-tuned with LoRA (r = 8) into a sentiment classifier that trains under 5% of its parameters and reaches the reference’s calibrated accuracy; a perplexity with its confidence interval; and the model zoo scoring every family trained here in one report.
Data (stand-ins, DEVIATIONS B73-05 and B73-06): the byte corpus of MS-L2 (course/fixtures/MS-L2/{train,val}.bin) for GPT; the synthetic SST-2-shaped set course/fixtures/small-corpora/sst2-2k.tsv (split, sentence, label; 1600 train, 400 val) for the classifiers, and its train sentences as a token shard, course/fixtures/MS-L6/sst2-train.bin, for BERT and ELECTRA. Configs: course/fixtures/configs/{gpt-tiny,bert-tiny}.json.
This file fixes the Pass 5 objective verbs of your tinyllm role
(spec/cli-roles.md, “Verbs of later passes”). Every verb keeps the rules of
that page: exit 2 on a usage error, the last stdout line is one JSON object.
The byte tokenizer (D32) with four control bytes as specials: 0 [PAD],
1 [CLS], 2 [SEP], 3 [MASK]. Seeds: PCG32(S).substream(“init”) builds a
model, .substream(“shuffle”) draws windows and batches, .substream(“sample”)
masks and samples.
{tinyllm} train gpt —cfg <config.json> —data <train.bin> —val <val.bin> —out
{tinyllm} train bert —cfg <config.json> —data <train.bin> —out
batch windows of window bytes at
shuffle.below(n - window), wraps each as [CLS] window [SEP], masks with
L6.2’s mlm_mask (p, the specials never picked, mask id 3), and takes
one M10.3 AdamW step (lr, weight_decay). save_bert (tokenizer bytes).
Final line: {“out”, “arch”: “bert”, “steps”, “mlm_loss” (mean of the
last 20), “params”}.
{tinyllm} train electra —cfg <config.json> —data <train.bin> —out
{tinyllm} finetune classify —base
{tinyllm} eval ppl —model
{tinyllm} zoo add —manifest <zoo.json> —id
{tinyllm} eval —suite zoo —manifest <zoo.json> [—out <report.json>] [—seed S] L6.7’s run_zoo with PCG32(S).substream(“sample”); writes the formats/eval-results.schema.json report (default: zoo-report.json next to the manifest) and prints one line per row. Final line: {“suite”, “report”, “rows”, “ok”, “errors”, “skipped”}.
ol milestone MS-L6 --smoke runs every step with 10% of the updates (about
a minute on a laptop). The full run takes about 4 minutes.
MS-L7: Your modern decoder loads Hugging Face checkpoints and matches them
Section titled “MS-L7: Your modern decoder loads Hugging Face checkpoints and matches them”Your modern decoder, run through your own entry point: a Llama-family checkpoint exactly as transformers saves it loads into your L7.9 model (built from L7.1 to L7.8), reports Hugging Face’s parameter count, gives HF’s next-token logits, and decodes HF’s greedy continuation token for token. PR CI runs it on the committed tiny model; the nightly run pulls SmolLM2-135M-Instruct with your own downloader and checks it against HF.
This file fixes the Pass 5 Llama verbs of your tinyllm role
(spec/cli-roles.md, “Verbs of later passes”). Every verb keeps the rules of
that page: exit 2 on a usage error, the last stdout line is one JSON object.
—model names a Llama directory, or its config.json (the committed tiny
model is referenced by its config.json, a fixture file with a MANIFEST row).
A “Llama directory” holds config.json (formats/config.schema.json: tl_arch
“llama”, or a plain HF model_type llama, mistral, or qwen2) and
model.safetensors (or shards with model.safetensors.index.json). Its
tokenizer is the byte tokenizer when config.json says tl_tokenizer “bytes”
(D32), else tokenizer.json read with your L1.2 BPETokenizer.from_hf_json.
{tinyllm} pull <owner/name> [—revision R] [—cache-dir DIR] [—files a,b,…]
L7.9’s hf_download of the repository’s files (default: config.json,
generation_config.json, model.safetensors, tokenizer.json,
tokenizer_config.json, special_tokens_map.json) into
DIR/
{tinyllm} info —model
{tinyllm} logits —model
{tinyllm} generate —model
Fixtures (course/oracle/L7.9/llama_hf.py, course/oracle/MS-L7/smollm2_parity.py): tiny-llama-2l is a random LlamaForCausalLM (2 layers, BF16 weights, byte tokenizer) saved by transformers 5.19.0; its logits and greedy ids are HF’s float32 forward. Every prompt’s greedy continuation keeps a top-2 logit margin >= 1e-3 at all 32 steps (DESIGN 5.7 near-tie rule), and the expected files carry that margin per step. The nightly model is SmolLM2-135M-Instruct at the revision pinned in course/fixtures/ASSETS.tsv, not the base SmolLM2-135M of DESIGN 4.3 (DEVIATIONS B75-07).
ol milestone MS-L7 --smoke runs the four tiny-model steps (seconds).
MS-P5: Transformers: the 2017 model, its objectives, and a modern decoder that matches SmolLM2
Section titled “MS-P5: Transformers: the 2017 model, its objectives, and a modern decoder that matches SmolLM2”Pass 5 builds the transformer three times over: the 2017 encoder-decoder (MS-L5), the objectives and the model zoo (MS-L6), and the modern decoder that loads SmolLM2 (MS-L7). The gate is those three plus the spiral invariant: the smoke steps of every earlier pass gate rerun, so your engine still streams the Pass 1 bigram and your earlier models still train.
requires is every Pass 5 stage of DESIGN 7.4 (the optional L6.4 is not
required); each needs a fresh pass of your own. craft.05 grades your oracle
tests (rung R5) with its own artifact check.
Run a gate with practice/bin/ol milestone <ID> --smoke; omit --smoke for its full local and cluster steps.