Pass 6 milestones
Pass result: Python inference stack (sampler, KV, paged KV, quant, spec decoding, constrained decoding), plus optional standalone C exercises.
The path includes the milestone stages below. Run each component milestone after its modules pass, then run the pass gate. The component gates run before the pass gate, which also reruns the smoke steps of earlier passes.
| Gate | Specification | What it covers |
|---|---|---|
MS-L8 | MS-L8.toml | Your Python inference stack matches no-cache output token for token. Requires ds.07, L8.1, L8.2, L8.4, L8.5, L8.6, L8.7, M09.4. |
MS-P6 | MS-P6.toml | Inference: cached, quantized, speculative, and constrained decoding. Requires M07.6, M09.3, S-M09b, craft.06. |
MS-L9 | MS-L9.toml | Optional C kernels, runtime utilities, and data structures. Requires ds.01, ds.02, ds.03, ds.04, rt.02, rt.03, rt.04, M09.7, M09.5, M09.6, L9.1, L9.2, L9.3, L9.4, L9.5, L9.6, lang.03, craft.06. |
Component gate details
Section titled “Component gate details”MS-L8: Your Python inference stack matches no-cache output token for token
Section titled “MS-L8: Your Python inference stack matches no-cache output token for token”Your Python inference stack, run through your own entry point, on the Llama model of MS-L7: the same greedy tokens with no cache, your contiguous cache (L8.2, with float16 keys and values); quantized weights (L8.5, using M09.4’s FP8 and microscaling conversions) that keep the model’s predictions; speculative decoding (L8.6) that changes the speed and never the greedy output; and constrained decoding (L8.7) whose every document parses and validates. PR CI runs it on the committed tiny model; the nightly run uses SmolLM2-135M-Instruct pulled by MS-L7; the cache speedup is a local perf step.
This file fixes the Pass 6 inference forms of your tinyllm role
(spec/cli-roles.md, “Verbs of later passes”: generate with —cache and
—spec, eval ppl, bench decode). Every verb keeps the rules of that page:
exit 2 on a usage error, the last stdout line is one JSON object.
{tinyllm} generate —model
{tinyllm} eval ppl —model
{tinyllm} bench decode —model
DESIGN 4.3 step (5), “the L8.4 course tests pass from cargo test”, and the
rt.04 block pool are this milestone’s requires: a fresh pass of L8.4 ran
its Rust course tests with cargo, as MS-L9 treats its sanitizer step
(DEVIATIONS B83-08). The tiny model is random, so its perplexity says
little: the PR quantization steps bound top-1 agreement and the NLL shift;
the nightly steps keep DESIGN 4.3’s perplexity budgets on SmolLM2.
ol milestone MS-L8 --smoke runs the tiny-model steps (seconds).
MS-P6: Inference: cached, quantized, speculative, and constrained decoding
Section titled “MS-P6: Inference: cached, quantized, speculative, and constrained decoding”MS-L8 verifies the core Python inference stack. Standalone C exercises are optional and deliberately do not gate this pass. Earlier pass smoke steps rerun through the spiral invariant.
MS-L9: Optional C kernels, runtime utilities, and data structures
Section titled “MS-L9: Optional C kernels, runtime utilities, and data structures”requirement of MS-P6 or any production runtime milestone. Each required module’s C check compiles its own test binary, and numerical parity uses the checked-in vectors emitted by the corresponding Python reference.
Run a gate with practice/bin/ol milestone <ID> --smoke; omit --smoke for its full local and cluster steps.