LLM Foundations
Where language models came from and the math that makes them work: the history from n-grams to reasoning models, transformer math, and the architecture variants you will meet in every config.json.
For: Anyone new to LLMs, and every role that has to explain a model to someone else.
Included in: Superstar FDE. Progress made here counts there too.
The stages, their chapters, and what “done” means for each are in path.tsv.
practice/bin/ol learn llm-foundations # stages and progresspractice/bin/ol learn llm-foundations next # read the next unfinished stageStages
Section titled “Stages”Generated from path.tsv. Track progress locally with practice/bin/ol learn.
| Stage | Read | Done when |
|---|---|---|
| 1 | History of language models Neural Network Architectures & the History of Deep Learning | You can explain, in order, why n-grams gave way to RNNs and LSTMs, why seq2seq needed attention, and why the Transformer, scaling laws, RLHF, MoE, and reasoning models each followed. |
| 2 | Deep learning and transformer math Deep Learning | You can derive backprop through softmax attention and compute, by hand, the FLOPs per token and the weight plus KV-cache memory of a 70B model. |
| 3 | Architecture variants Foundation Models & Architectures | Given any config.json you can name the variant: dense or MoE, MHA, GQA, or MLA, RoPE scaling, context length, and say what each choice costs at inference. |