Optimization
Overview
Section titled “Overview”- Primary references: Boyd and Vandenberghe, Convex Optimization (free PDF); Nocedal and Wright, Numerical Optimization (Springer, 2nd ed.)
- Supplementary: Kingma and Ba, Adam (free); Loshchilov and Hutter, Decoupled Weight Decay Regularization (free); Goh, Why Momentum Really Works (free); Cohen et al., Gradient Descent on Neural Networks Typically Occurs at the Edge of Stability (free); Jordan, Muon (free)
- Prerequisites: Calculus 3 (gradients, Hessians), Linear Algebra (eigenvalues), Precalculus (geometric series for schedules)
- Estimated time: 3 weeks at 10 to 12 h/week; in the course, Pass 2 (M10.1 to M10.4), with optional curvature and Muon in Pass 9
Key Takeaways
Section titled “Key Takeaways”- Gradient descent converges at a rate set by the condition number : the gap shrinks by about per step on a quadratic.
- Momentum and Nesterov average past gradients to move faster along shallow directions; Adam adds a per-parameter scale from the second moment, with bias correction for early steps.
- Weight decay is not L2 regularization under Adam: AdamW decouples it from the adaptive scale.
- The schedule is part of the optimizer. Warmup, cosine or warmup-stable-decay, and gradient clipping decide whether a run trains at all.
- Optimizers are stateful: resuming a run bitwise needs the moments, the step count, and the schedule position in the checkpoint.
How to Study
Section titled “How to Study”Read Boyd and Vandenberghe chapters 2, 3, and 9 for convexity and descent methods, then the Adam and AdamW papers in full. Plot every optimizer on a 2D ill-conditioned quadratic before you trust it on a network. In the course, M10.1 to M10.4 are built in Pass 2 and train every model after; S-M10a checks the derivations; the optional M10.5 and M10.6 are C1 diagnostics and an A/B experiment.
Concepts & Techniques
Section titled “Concepts & Techniques”Core Insight
Section titled “Core Insight”Training is minimization of an average loss with noisy gradients. Everything in an optimizer answers one of three questions: which direction (the gradient, smoothed by momentum), how far in each coordinate (the learning rate, scaled per parameter by Adam), and how that step size changes over time (the schedule). Convex theory gives the vocabulary and the rates; neural networks break its assumptions but keep its intuitions.
1. Convexity, smoothness, and gradient descent
Section titled “1. Convexity, smoothness, and gradient descent”Key ideas:
- Definitions: is -smooth if its gradient is -Lipschitz and -strongly convex if .
- Armijo line search backtracks until the decrease is at least (
M10.1).
2. The Optimizer protocol: SGD, momentum, AdamW
Section titled “2. The Optimizer protocol: SGD, momentum, AdamW”Key ideas:
- Protocol:
step,zero_grad,state_dict,load_state_dict, shared by every optimizer the course trains with (M10.2,M10.3). - AdamW: , , step .
3. Schedules and clipping
Section titled “3. Schedules and clipping”Key ideas:
- Cosine with warmup, warmup-stable-decay (the C1 schedule), and Noam (the 2017 transformer), plus global-norm gradient clipping (
M10.4).
4. Curvature and beyond (optional)
Section titled “4. Curvature and beyond (optional)”Key ideas:
- Top Hessian eigenvalue by power iteration on Hessian-vector products; gradient descent is stable only while (
M10.5). - Muon orthogonalizes the momentum of 2D weights with Newton-Schulz iterations (
M10.6). - Lagrangians and KKT conditions are solve-only (
M10.7); they derive the closed form behind DPO inL12.2.
Course modules
Section titled “Course modules”| Module | Topic | Kind | Pass |
|---|---|---|---|
S-M10a | Descent methods and optimizers problem set | solve | 2 |
S-M10b | Lagrangians and KL-regularized objectives problem set | solve | 10, optional |
M10.1 | Convexity, L-smoothness, gradient descent, Armijo line search. Beat 2 defines the condition number κ = L/μ (M09.3 later generalizes it to matrices) | build | 2 |
M10.2 | Optimizer protocol, SGD, momentum, Nesterov, weight decay | build | 2 |
M10.3 | Adam/AdamW, bias correction, decoupled decay | build | 2 |
M10.4 | Schedules (cosine, WSD, Noam) and gradient clipping | build | 2 |
M10.5 | Curvature: Newton step, top Hessian eigenvalue, edge of stability | build | 9, optional |
M10.6 | Muon: Newton-Schulz orthogonalized momentum | build | 9, optional |
M10.7 | Lagrangians, KKT, duality, KL-regularized objectives | solve | solve set |
Chapters
Section titled “Chapters”Connections to Other Tracks
Section titled “Connections to Other Tracks”| Track | Connection |
|---|---|
| Calculus 3 | gradients, Hessians, and Taylor models of the loss |
| Matrix Calculus and Autodiff | where the gradients come from; Hessian-vector products |
| tinyllm Part 0 | L0.5 trains with these optimizers |
| Training and Post-Training | optimizer memory and large-scale training practice |