Skip to content

LLM Foundations

Where language models came from and the math that makes them work: the history from n-grams to reasoning models, transformer math, and the architecture variants you will meet in every config.json.

For: Anyone new to LLMs, and every role that has to explain a model to someone else.

Included in: Superstar FDE. Progress made here counts there too.

The stages, their chapters, and what “done” means for each are in path.tsv.

Terminal window
practice/bin/ol learn llm-foundations # stages and progress
practice/bin/ol learn llm-foundations next # read the next unfinished stage

Generated from path.tsv. Track progress locally with practice/bin/ol learn.

StageReadDone when
1History of language models
Neural Network Architectures & the History of Deep Learning
You can explain, in order, why n-grams gave way to RNNs and LSTMs, why seq2seq needed attention, and why the Transformer, scaling laws, RLHF, MoE, and reasoning models each followed.
2Deep learning and transformer math
Deep Learning
You can derive backprop through softmax attention and compute, by hand, the FLOPs per token and the weight plus KV-cache memory of a 70B model.
3Architecture variants
Foundation Models & Architectures
Given any config.json you can name the variant: dense or MoE, MHA, GQA, or MLA, RoPE scaling, context length, and say what each choice costs at inference.