Skip to content

Optimization

  • Gradient descent converges at a rate set by the condition number κ=L/μ\kappa = L/\mu: the gap shrinks by about (1−1/κ)(1 - 1/\kappa) per step on a quadratic.
  • Momentum and Nesterov average past gradients to move faster along shallow directions; Adam adds a per-parameter scale from the second moment, with bias correction for early steps.
  • Weight decay is not L2 regularization under Adam: AdamW decouples it from the adaptive scale.
  • The schedule is part of the optimizer. Warmup, cosine or warmup-stable-decay, and gradient clipping decide whether a run trains at all.
  • Optimizers are stateful: resuming a run bitwise needs the moments, the step count, and the schedule position in the checkpoint.

Read Boyd and Vandenberghe chapters 2, 3, and 9 for convexity and descent methods, then the Adam and AdamW papers in full. Plot every optimizer on a 2D ill-conditioned quadratic before you trust it on a network. In the course, M10.1 to M10.4 are built in Pass 2 and train every model after; S-M10a checks the derivations; the optional M10.5 and M10.6 are C1 diagnostics and an A/B experiment.


Training is minimization of an average loss with noisy gradients. Everything in an optimizer answers one of three questions: which direction (the gradient, smoothed by momentum), how far in each coordinate (the learning rate, scaled per parameter by Adam), and how that step size changes over time (the schedule). Convex theory gives the vocabulary and the rates; neural networks break its assumptions but keep its intuitions.

1. Convexity, smoothness, and gradient descent

Section titled “1. Convexity, smoothness, and gradient descent”

Key ideas:

  • Definitions: ff is LL-smooth if its gradient is LL-Lipschitz and μ\mu-strongly convex if f(y)≥f(x)+∇f(x)⊤(y−x)+μ2∥y−x∥2f(y) \ge f(x) + \nabla f(x)^{\top}(y - x) + \tfrac{\mu}{2}\|y - x\|^2.
  • Armijo line search backtracks until the decrease is at least c α ∇f⊤dc\,\alpha\,\nabla f^{\top} d (M10.1).

2. The Optimizer protocol: SGD, momentum, AdamW

Section titled “2. The Optimizer protocol: SGD, momentum, AdamW”

Key ideas:

  • Protocol: step, zero_grad, state_dict, load_state_dict, shared by every optimizer the course trains with (M10.2, M10.3).
  • AdamW: mt=β1mt−1+(1−β1)gtm_t = \beta_1 m_{t-1} + (1 - \beta_1) g_t, vt=β2vt−1+(1−β2)gt2v_t = \beta_2 v_{t-1} + (1 - \beta_2) g_t^2, step η m^t/(v^t+ϵ)+ηλθ\eta\,\hat{m}_t / (\sqrt{\hat{v}_t} + \epsilon) + \eta\lambda\theta.

Key ideas:

  • Cosine with warmup, warmup-stable-decay (the C1 schedule), and Noam (the 2017 transformer), plus global-norm gradient clipping (M10.4).

Key ideas:

  • Top Hessian eigenvalue by power iteration on Hessian-vector products; gradient descent is stable only while η<2/λmax⁡\eta < 2/\lambda_{\max} (M10.5).
  • Muon orthogonalizes the momentum of 2D weights with Newton-Schulz iterations (M10.6).
  • Lagrangians and KKT conditions are solve-only (M10.7); they derive the closed form behind DPO in L12.2.
ModuleTopicKindPass
S-M10aDescent methods and optimizers problem setsolve2
S-M10bLagrangians and KL-regularized objectives problem setsolve10, optional
M10.1Convexity, L-smoothness, gradient descent, Armijo line search. Beat 2 defines the condition number κ = L/μ (M09.3 later generalizes it to matrices)build2
M10.2Optimizer protocol, SGD, momentum, Nesterov, weight decaybuild2
M10.3Adam/AdamW, bias correction, decoupled decaybuild2
M10.4Schedules (cosine, WSD, Noam) and gradient clippingbuild2
M10.5Curvature: Newton step, top Hessian eigenvalue, edge of stabilitybuild9, optional
M10.6Muon: Newton-Schulz orthogonalized momentumbuild9, optional
M10.7Lagrangians, KKT, duality, KL-regularized objectivessolvesolve set
#ModuleChapterKindPass
1M10.1Convexity, smoothness, gradient descent, and Armijo line searchbuild2
2M10.2The Optimizer protocol: SGD, momentum, Nesterov, weight decaybuild2
3M10.3Adam and AdamW: bias correction and decoupled weight decaybuild2
4M10.4Learning-rate schedules (cosine, WSD, Noam) and gradient clippingbuild2
5S-M10aOptimization problem set, part a: convexity, GD rates, momentum, Adamsolve2
6S-M10bOptimization problem set, part b: constrained optimization, duality, and DPOsolve10
TrackConnection
Calculus 3gradients, Hessians, and Taylor models of the loss
Matrix Calculus and Autodiffwhere the gradients come from; Hessian-vector products
tinyllm Part 0L0.5 trains with these optimizers
Training and Post-Trainingoptimizer memory and large-scale training practice