Skip to content

Capstone: TinyStories, owned end to end

ModuleC1 · practice · Python, Go, and ops · Pass 9 · 2 to 3 days of work plus 4 to 12 h of laptop training (the short config: under 1 h; the smoke tier: minutes)
You buildnothing new in the library: you run your system on one goal and write up what it measured. Artifacts: specs/c1/ (the full, short, and release specs), docs/capstone/c1/ (report.json, zoo.json, samples.jsonl, the model card and data ledger the release names), and two ADRs in docs/adr/
Contractformats/train-spec.schema.json, formats/eval-results.schema.json (the zoo table), formats/ledger.schema.json, templates/ADR.md, templates/MODEL_CARD.md, and the report schema course/tests/C1/capstone-report.schema.json
Testscourse/tests/C1/: the artifact check (section 4), which recomputes what can be recomputed and runs your dur.12 preflight; the run itself is MS-C1’s
Needsdur.12 (release workflow preflight), L11.1 (bf16, accumulation, activation recompute in the full and short train specs), ethics.04 (safety suite in the release report); reading: every pass before this one, above all data.07 and data.08, L1.2, L1.4, L1.6, L2.1, L6.7, L7.4, L7.6, L7.8, L7.9, M03.5, M05.1, M07.5, M10.3, M10.4, dur.09, craft.22
Used byno call site (the capstone): C2 post-trains its release, and the ops.08 data-incident drill retrains it
MilestoneMS-C1 (part of MS-P9)
Optional depthEldan and Li, TinyStories (2023, free); Kaplan et al., Scaling Laws for Neural Language Models (2020, free); Hoffmann et al., Training Compute-Optimal Large Language Models (2022, free); DeepSeek-AI, DeepSeek-V2 (MLA, free)
  • The capstone is your system doing one real job end to end: clean data through CorpusBuild, a trained tokenizer, a Llama trained as a durable TrainRun with mixed precision and resume, evaluation against the zoo, a gated release through ModelRelease, and serving through your engine and gateway.
  • An ablation is an experiment: change one thing, hold the budget equal (the same vocabulary size, equal KV bytes per token, equal active parameters), pair the measurements by held-out document, and name a winner only when the 95% CI of the paired difference excludes zero (test_ablations_are_fair_and_paired).
  • Every number in the report is checkable: the parameter count follows from the config (test_hand_example_parameter_count), the scaling fit from its points (test_scaling_fit_reproduces), the zoo’s Llama row is the report’s bpb (test_zoo_table).
  • Decisions are recorded with their evidence: the vocabulary ADR cites the tokenizer ablation, the architecture ADR cites the attention and MLP ablations (test_adrs_cite_the_ablations).
  • The model ships only through your release workflow (test_release_passes_your_preflight) and is then served like any other model.
Terminal window
ol check C1 # tells you which artifact is missing
<system> data build --dataset tinystories --version v1 # CorpusBuild (data.09)
{tinyllm} tok train --algo bpe --vocab 4096 --sample 16MB # and --algo unigram, for the ablation
<system> train --spec specs/c1/tinystories-short.json # TrainRun (dur.11); the full spec when you can spare a night
<system> train --spec specs/c1/abl-mla.json # one short run per ablation arm, same seed
<system> eval --suite zoo # EvalSuite (dur.11) over every family you trained
# write docs/capstone/c1/report.json from the runs' outputs, then the two ADRs
ol check C1 # the artifact check
<system> release --spec specs/c1/release.json && <system> wf signal <id> approve
ol milestone MS-C1 # val loss at the calibrated bar, samples, conformance, artifacts, a trace

ol milestone MS-C1 --smoke runs the same pipeline on the smoke config (about 1M parameters, 200 steps) in minutes; that is what CI runs.


Every part of your system has passed its own tests. None of them has had to work with all the others for hours, on data you did not choose, with numbers nobody gave you in advance. That is where the remaining bugs live: a tokenizer whose vocabulary disagrees with the model config, a checkpoint whose data cursor resumes one window late, an eval that reads the wrong split, a release spec your gateway cannot route. The capstone runs the whole system on one goal (a model that writes small stories) and asks you to defend each choice with a measurement. It is also the first time the course cannot hand you an oracle: the report is the evidence, and the check verifies that the evidence agrees with itself.

SymbolMeaningType
NNtrainable parameters of the modelint
DDtraining tokensint
C≈6NDC \approx 6NDtraining compute in FLOP (2 for the forward pass, 4 for the backward)float
bpb\mathrm{bpb}held-out bits per byte: total NLL in bits over the held-out bytesfloat
u=1,…,nu = 1, \dots, na paired unit: one held-out document
δu=bpbb(u)−bpba(u)\delta_u = \mathrm{bpb}_b(u) - \mathrm{bpb}_a(u)the ablation’s paired difference on document uufloat
δˉ\bar\delta, [ℓ,h][\ell, h]the mean difference and its 95% bootstrap CIfloat
aa, α\alphathe scaling fit bpb=aN−α\mathrm{bpb} = a N^{-\alpha}float
StageWhat runsYour modules
DataCorpusBuild: fetch, filter (with the KN perplexity filter), dedup, decontamination, PII, shards; the tokenizer trained on a fixed 16 MB sample; encode to uint16 .bindata.01 to data.09, L1.2, L1.4, L1.5, L2.1
ModelLlamaConfig(vocab=4096, d=320, layers=8, heads=8, kv_heads=4, d_ff=864, ctx=512, tied)L7.1 to L7.9, M05.1
TrainAdamW, WSD, clipping, bf16 emulation, accumulation, recompute; TrainRun segments with resumeM10.3, M10.4, L11.1, L0.6, dur.09, dur.11
Ablationstokenizer, attention, MLP, one short run per armL1.4, L7.6, L7.8, M07.5
Evaluateheld-out bpb with a CI, seeded samples, the scaling fit, the zoo tableL6.7, M03.5, M11.2, craft.22
Releaseexport, eval gates, model card, ledger, canarydur.12, data.08
Servethe Rust engine behind your gatewayL10.*, gw.*
TierModelTokensLaptop time (estimate)Accepted by
fullabout 10.4M parametersabout 10810^84 to 12 hol check C1, MS-C1
shortabout 2.4Mabout 2×1072 \times 10^7under 1 hol check C1, MS-C1 at a looser bar
smokeabout 0.8M, 200 stepsabout 4×1054 \times 10^5minutesCI only (OL_SMOKE=1, MS-C1 --smoke)

C=6NDC = 6ND sizes the run before you start it: the full tier is 6×1.04×107×108≈6.2×10156 \times 1.04 \times 10^7 \times 10^8 \approx 6.2 \times 10^{15} FLOP. At an effective 0.2 to 0.4 TFLOP/s for numpy with Accelerate on an M-series laptop that is 4 to 9 hours (uncertain: measure your own throughput with ol bench first).

One ablation changes one thing and holds the budget equal, or the comparison measures the budget:

  • tokenizer: BPE vs Unigram at the same vocabulary size on the same sample. Compare in bits per byte, never per token: a tokenizer that splits text into more, easier tokens has a lower per-token loss and no better model.
  • attention: GQA vs MLA at equal KV bytes per token. GQA caches L⋅2⋅nkv⋅dhL \cdot 2 \cdot n_{kv} \cdot d_h values per token; MLA caches the latent rr plus the shared rope key drd_r, L⋅(r+dr)L \cdot (r + d_r). In bf16 each value is 2 bytes.
  • MLP: dense SwiGLU vs a mixture of experts at equal active parameters per token (the router plus kk experts, not all EE).

Measure both arms on the same held-out documents and pair by document: δu\delta_u removes how hard each document is. The verdict follows the 95% bootstrap CI of δˉ\bar\delta: b wins when h<0h < 0, a when ℓ>0\ell > 0, and otherwise it is a tie, which is a result, not a failure.

Three sizes of the short family give three points (Ni,bpbi)(N_i, \mathrm{bpb}_i). In log space the power law is a line, log⁡bpb=log⁡a−αlog⁡N\log \mathrm{bpb} = \log a - \alpha \log N, fitted by least squares (M03.5’s lstsq). Three points do not prove a law; they tell you whether the next size is worth training.

The model-zoo suite (L6.7, run by EvalSuite) scores every family you trained on the same held-out text: KN-4, NPLM, LSTM, GPT, and the capstone Llama in bits per byte, plus the seq2seq (EM), classification (accuracy), and word-similarity (Spearman) rows. At the full and short tiers the capstone must beat every baseline.

Parameters of the full config (V=4096V = 4096, d=320d = 320, 8 layers, 8 heads of dh=40d_h = 40, 4 KV heads, SwiGLU f=864f = 864, tied embeddings):

PartCount
embedding VdV d (shared with the output)4096⋅320=1,310,7204096 \cdot 320 = 1{,}310{,}720
per layer: WqW_q, WoW_o (d⋅8⋅40d \cdot 8 \cdot 40 each)2⋅102,400=204,8002 \cdot 102{,}400 = 204{,}800
per layer: WkW_k, WvW_v (d⋅4⋅40d \cdot 4 \cdot 40 each)2⋅51,200=102,4002 \cdot 51{,}200 = 102{,}400
per layer: gate, up, down (3df3 d f)3⋅320⋅864=829,4403 \cdot 320 \cdot 864 = 829{,}440
per layer: two RMSNorm gains640640
8 layers8⋅1,137,280=9,098,2408 \cdot 1{,}137{,}280 = 9{,}098{,}240
final norm320320
total10,409,280

KV bytes per token in bf16: 8⋅2⋅4⋅40⋅2=5,1208 \cdot 2 \cdot 4 \cdot 40 \cdot 2 = 5{,}120. An MLA arm at equal KV bytes needs r+dr=5120/(8⋅2)=320r + d_r = 5120 / (8 \cdot 2) = 320, for example r=288r = 288, dr=32d_r = 32.

Compute: 6⋅10,409,280⋅108≈6.2×10156 \cdot 10{,}409{,}280 \cdot 10^8 \approx 6.2 \times 10^{15} FLOP.

An ablation verdict, from the reference smoke report: Unigram minus BPE is δˉ=+0.032\bar\delta = +0.032 bits per byte over 32 documents, CI (0.014,0.059)(0.014, 0.059). The CI excludes zero and δ=b−a>0\delta = b - a > 0, so a (BPE) wins. MLA minus GQA is −0.011-0.011 with CI (−0.042,0.008)(-0.042, 0.008): it includes zero, a tie.

The first two blocks are test_hand_example_parameter_count; the verdict rule is test_ablations_are_fair_and_paired.

ArtifactContent
specs/c1/tinystories-10m.json, specs/c1/tinystories-short.jsontrain specs (formats/train-spec.schema.json) of the two tiers
specs/c1/release.jsonthe ModelRelease spec of your capstone model (dur.12 section 4): suites, a quality and a safety gate, the route, the canary, the burn query of <model_id>-<version>
docs/capstone/c1/report.jsonthe report (course/tests/C1/capstone-report.schema.json): tier, spec, tokenizer, model config and parameters, the training run, held-out bpb with its CI, the three ablations, the scaling points and fit, and the paths below
docs/capstone/c1/zoo.jsonthe zoo suite’s report (formats/eval-results.schema.json)
docs/capstone/c1/samples.jsonl20 samples, seeds 0 to 19: {"seed", "prompt", "text", ...}
the model card and ledger your release spec namesMODEL_CARD.md from the template (ethics.03), the release’s LEDGER.jsonl
two ADRs in docs/adr/the vocabulary (citing the tokenizer ablation) and the architecture (citing the attention and MLP ablations)

The report is written by your tooling from the runs’ outputs (a few lines of Python over the progress files and eval reports), never by hand. The course’s reference report is a smoke-tier run, course/oracle/C1/smoke_capstone.py.

The check (ol check C1, course/tests/C1/check):

TestKINDChecks
test_hand_example_parameter_countunitsection 3, then your report’s parameter count against its own config
test_specs_are_capstone_train_specsunitboth specs validate, are Llamas of the tier’s size, and train on enough tokens
test_report_is_completeunitthe report schema; a full or short tier (smoke only with OL_SMOKE=1); its spec, zoo, samples, and release exist; bpb inside its CI
test_ablations_are_fair_and_pairedunitthree ablations; equal vocab, KV bytes, active parameters (recomputed); mean inside its CI; the verdict follows the CI
test_scaling_fit_reproducesunitthree or more sizes; the log-log least-squares fit of the points is the reported fit
test_zoo_tableunitthe zoo report validates; KN-4, NPLM, LSTM, GPT, Llama on one text; the other tasks scored or skipped with a reason; the Llama row is the report’s bpb; the capstone wins at full and short
test_samples_pass_the_quality_scorersunit20 seeds; at least 20 words; repeated 8-grams at most 0.25 per sample and 0.10 on average; at least 30% distinct words
test_adrs_cite_the_ablationsunitthe two ADRs, their sections, and the cited mean differences
test_release_passes_your_preflightconformanceyour Go ReleaseSpec decodes the release spec strictly, Validate and Preflight pass on its card and ledger; quality and safety gates; the burn query selects the served model
PitfallSymptomCaught by
1. an ablation at unequal budgetMLA “wins” because it caches more; the MoE “wins” because it computes moretest_ablations_are_fair_and_paired
2. a winner declared inside the noisean ADR built on a CI that includes zerotest_ablations_are_fair_and_paired
3. comparing tokenizers per tokenthe tokenizer with more, easier tokens looks bettertest_ablations_are_fair_and_paired (bits per byte only)
4. a report that describes another modelparameters, config, and zoo row disagreetest_hand_example_parameter_count, test_zoo_table
5. a scaling fit drawn by eye or in linear spaceaa and α\alpha that no least-squares fit givestest_scaling_fit_reproduces
6. baselines scored on different texta zoo table whose rows cannot be comparedtest_zoo_table
7. a looping or degenerate samplerthe same 8-gram again and again; a handful of wordstest_samples_pass_the_quality_scorers
8. decisions without evidenceADRs that state a choice and cite nothingtest_adrs_cite_the_ablations
9. a release your own workflow refusesa template placeholder in the card, a non-train source, a burn query of another modeltest_release_passes_your_preflight
10. a smoke run presented as the capstone200 steps of a 0.8M model reported as “the model”test_report_is_complete
DirectionModuleHow it uses this
Backdur.12the release workflow the capstone ships through
BackL11.1the full and short training specs use bf16, accumulation, and checkpointed activations
Backethics.04the release report carries the capstone’s safety suite rows
ForwardC2 (optional)post-trains the released model into an instruction follower
Forwardops.08revokes a source of this model’s data and drives the retrain decision through ModelRelease
Your pieceProduction equivalentWhat it addsWhere to look
the capstone pipelineKarpathy’s nanoGPT and llm.cthe same loop on GPUs, with careful throughput accountingllm.c/train_gpt2.c
the ablationsthe DeepSeek-V2 and Mixtral reportsMLA and MoE ablations at scale, with compute-matched baselinesthe papers above
the scaling fitChinchilla’s three approachesfits over hundreds of runs, isoFLOP profilesHoffmann et al., section 3
the zoo tableHELM, the Open LLM Leaderboardmany tasks, standardized prompts and seedscrfm-helm