Calculus 3 problem set: vectors, gradients, the chain rule, Hessians, multiple integrals, Lagrange
Overview
Section titled “Overview”| Module | S-M04 · solve · none · Pass 2 · 6 to 8 h |
| You build | answers in solve/S-M04.toml (56 checked by SymPy) and 2 proofs in solve/S-M04/q34.md and q58.md (self-graded against their rubrics) |
| Contract | none: a pen and paper set |
| Tests | course/solve/S-M04/key.toml (hidden): typed answers plus reject canaries; the problems are in course/solve/S-M04/problems.md and in section 4 |
| Needs | no module. Reading: S-M01 (one-variable derivatives and integrals) and the Calculus 3 topic |
| Used by | no call site (a solve set). Take it after M04.1 (gradcheck) and M04.2 (Jacobians, VJP, JVP) in Pass 2. It is also the whole of the solve-only M04.3 (Hessian), M04.4 (multiple integrals), and M04.5 (Lagrange), which M10.1, M07.3, and M11.5 read |
| Milestone | MS-P2 (the Pass 2 gate runs ol check on every solve part of the pass) |
| Optional depth | OpenStax, Calculus Volume 3 (free), ch. 2, 4, and 5; Boyd and Vandenberghe, Convex Optimization, section 5.1, for Lagrange duality |
Key Takeaways
Section titled “Key Takeaways”- The gradient collects the partial derivatives; it points uphill fastest, and its length is the steepest slope (q11, q30, q34).
- The multivariable chain rule multiplies Jacobians, and reverse mode multiplies them from the loss backward, one vector-Jacobian product at a time (q23, q27, q28).
- At a critical point the Hessian’s eigenvalues decide minimum, maximum, or saddle, and their ratio is the condition number that slows gradient descent (q37, q42).
- A change of variables multiplies the area element by : the in polar coordinates is why (q26, q46, q50).
- At a constrained optimum the gradient is parallel to the constraint’s gradient; maximum entropy under an energy constraint is a softmax (q51 to q58).
How to work this chapter
Section titled “How to work this chapter”ol start S-M04 # writes solve/S-M04.toml and the two proof filesol check S-M04 # SymPy checks the answers, then asks each proof rubric (y/n)ol check S-M04 --regrade # ask the rubrics again after you change a proof1. Why now
Section titled “1. Why now”Your model has many parameters, and training moves all of them at once along the gradient. M04.1 builds gradcheck, which compares your analytic gradients with central differences coordinate by coordinate, and M04.2 computes Jacobians and the vector-Jacobian products that reverse-mode autodiff (M08.2, L0.1) chains together. M10.1 then descends along gradients, and its convergence rate is set by the Hessian’s condition number. M07.3 sizes initial weights with Gaussian integrals in several variables, and M11.5 reads the softmax as the solution of a constrained maximization. This set checks the hand computations behind all of them.
2. Principles
Section titled “2. Principles”| Symbol | Meaning | Type / shape |
|---|---|---|
| , | dot product; cross product (in ) | scalar; vector |
| Euclidean length, | real | |
| partial derivative: vary , hold the others fixed | function | |
| gradient, the vector of partials | vector | |
| directional derivative along a unit vector , | real | |
| Jacobian of , | ||
| Hessian, | , symmetric | |
| condition number of a positive definite , | real | |
| double integral over a region | real | |
| a Lagrange multiplier | real |
2.1 Vectors and geometry
Section titled “2.1 Vectors and geometry”, so exactly when . The cross product is perpendicular to both, with length the area of the parallelogram they span. A plane is with normal (a cross product of two edge vectors), and the distance from a point to it is .
2.2 Partial derivatives and the gradient
Section titled “2.2 Partial derivatives and the gradient”A partial derivative differentiates in one variable with the rest held fixed, by the one-variable rules. Mixed second partials agree for smooth functions: . Two gradients recur in machine learning: , and the gradient of log-sum-exp, , is the softmax of .
2.3 The multivariable chain rule
Section titled “2.3 The multivariable chain rule”If , then : one term per path from to . In general the Jacobian of a composite is the product of Jacobians, . A loss is a scalar at the end of a long chain, and reverse mode evaluates the product from the left: start with and multiply by one Jacobian at a time, . Each step is a vector-Jacobian product (VJP), and backpropagation is nothing more.
2.4 Directional derivatives
Section titled “2.4 Directional derivatives”The rate of change of at along a unit vector is , by the chain rule. By Cauchy-Schwarz it is at most , reached at : the gradient is the direction of steepest ascent, of steepest descent, and directions perpendicular to follow a level curve.
2.5 The Hessian and the second-derivative test
Section titled “2.5 The Hessian and the second-derivative test”At a critical point (), the second-order Taylor expansion is . If every eigenvalue of is positive (positive definite), the point is a local minimum; all negative, a maximum; mixed signs, a saddle. For two variables: and is a minimum, is a saddle. For the quadratic , gradient descent’s error shrinks by about per step with (M10.1).
2.6 Multiple integrals and change of variables
Section titled “2.6 Multiple integrals and change of variables”A double integral over a rectangle is an iterated integral in either order; over other regions, the limits of the inner integral depend on the outer variable. Changing variables multiplies the area element by the absolute Jacobian determinant: . Polar coordinates have , so .
2.7 Lagrange multipliers
Section titled “2.7 Lagrange multipliers”To optimize subject to , look for points where and : at a constrained optimum, moving along the constraint surface cannot change to first order, so has no component along the surface. With several constraints, . Solve the system, then compare the candidates’ values.
3. Worked example by hand
Section titled “3. Worked example by hand”This is a sibling of q27 and q52, not one of the graded problems.
A backward pass by hand. , , (the cross-entropy of one positive example). At : forward, , , . Backward, one factor at a time: ; ; . So . The middle two factors combine to , the familiar “prediction minus label” of a logistic output. In solve/ the answer would be answer = "-1".
A constrained minimum. Minimize subject to . Lagrange: , so , , and gives , , value . Check by substitution: , at .
4. The problem set
Section titled “4. The problem set”Write each answer in solve/S-M04.toml:
[q5]answer = "x + y + z = 1"[q11]answer = "[8, 7]"[q35]answer = "[[2, 3], [3, 4]]"[q37]answer = "saddle"[q34]proof = "S-M04/q34.md"Vectors are flat lists, matrices lists of rows, and entries are exact (-1/sqrt(5)). For q33 either sign of the direction passes.
Vectors and geometry
Section titled “Vectors and geometry”q1. . [number]
q2. . [vector]
q3. Give the angle between and , in radians. [number]
q4. . [number]
q5. Give the equation of the plane through , , and . [equation in x, y, z]
q6. Give the distance from the point to the plane . [number]
q7. Give the area of the triangle with vertices , , . [number]
q8. Are and orthogonal? [bool]
Partial derivatives and the gradient
Section titled “Partial derivatives and the gradient”q9. . Give . [expr in x, y]
q10. Give for the of q9. [expr in x, y]
q11. . Give . [vector]
q12. . Give . [vector]
q13. . Give . [vector]
q14. . [expr in x, y]
q15. (log-sum-exp). Give . [vector]
q16. with and . Give . [vector]
q17. . Give at . [number]
q18. The squared error of a one-weight model is . Give . [expr in w, x, y]
The multivariable chain rule
Section titled “The multivariable chain rule”q19. with , . Give . [number]
q20. with , . Give . [expr in t]
q21. with , . Give . [expr in x, y]
q22. Give for the of q21. [expr in x, y]
q23. with the sigmoid . Give at , . [number]
q24. with , . Give . [expr in r]
q25. Give the Jacobian matrix of at . [matrix]
q26. Give the determinant of the Jacobian of . [expr in r]
q27. , , . Give at (backpropagate one factor at a time). [number]
q28. with . Give the vector-Jacobian product for , as a column. [vector]
Directional derivatives
Section titled “Directional derivatives”q29. . Give the directional derivative at in the direction . [number]
q30. . Give the largest directional derivative at over all unit directions. [number]
q31. . Give the unit direction of steepest descent at . [vector]
q32. . Give the directional derivative at in the direction . [number]
q33. . Give a unit direction at along which the directional derivative is 0 (either sign). [vector]
q34. Prove that at a point where , the unit direction that maximizes the directional derivative is , and that the maximum is . [proof]
The Hessian and the second-derivative test
Section titled “The Hessian and the second-derivative test”q35. Give the Hessian of . [matrix]
q36. Give the Hessian of at . [matrix]
q37. Classify the critical point of : minimum, maximum, or saddle. [choice]
q38. Classify the critical point of : minimum, maximum, or saddle. [choice]
q39. Give the critical point of . [vector]
q40. Classify the critical point of : minimum, maximum, or saddle. [choice]
q41. Is the Hessian of q35 positive definite? [bool]
q42. with . Give the condition number of its Hessian. [number]
Multiple integrals and change of variables
Section titled “Multiple integrals and change of variables”q43. . [number]
q44. Give the area of the triangle as a double integral. [number]
q45. . [number]
q46. . [number]
q47. . [number]
q48. Give the volume under and above the unit disk . [number]
q49. , the total mass of a 2-D standard normal. [number]
q50. With , , give the area of the region , . [number]
Lagrange multipliers
Section titled “Lagrange multipliers”q51. Give the maximum of subject to . [number]
q52. Give the minimum of subject to . [number]
q53. Give the maximum of subject to . [number]
q54. Maximize the entropy subject to . Give the optimal . [number]
q55. Give the maximum of subject to with . [number]
q56. Give the minimum of subject to with . [number]
q57. Give the point of the plane closest to the origin. [vector]
q58. With Lagrange multipliers, show that the distribution over states with energies that maximizes the entropy subject to and has the form (a softmax of ). [proof]
5. Pitfalls
Section titled “5. Pitfalls”| Pitfall | Symptom | Caught by |
|---|---|---|
| Dropping an inner derivative in a partial | given as ; gradcheck fails on every product inside a nonlinearity | q9, q10, q18 (canaries) |
| Forgetting the in a least-squares gradient | descent along the wrong direction for any non-identity | q16 (canary) |
| Multiplying the derivatives of the parts instead of summing over paths | for given as | q20 (canary 6*t^3) |
| Skipping a factor in a chain of local derivatives | a backward pass off by the sigmoid’s slope or a weight | q23, q27 (canaries) |
| Computing when the backward pass needs | the gradient of a non-symmetric layer comes out transposed | q28 (canary) |
| Using an unnormalized direction | a directional derivative scaled by the direction’s length | q29 (canary 4), q31 (canary) |
| Halving the mixed term of a Hessian | a saddle misread as a minimum | q35 (canary) |
| Forgetting the Jacobian factor in a change of variables | polar integrals off by the missing | q45 (canary 2*pi), q50 (canary 4) |
| Reporting the optimizer instead of the optimum | the argmax given as the max | q52, q53 (canaries) |
6. Where it’s used next
Section titled “6. Where it’s used next”| Direction | Module | How it uses this |
|---|---|---|
| Back | S-M01 | one-variable derivatives, the chain rule, and integrals |
| Forward | M04.1 | gradcheck compares your partials (q9 to q18) with central differences |
| Forward | M04.2 | jacobian, vjp_numeric, jvp_numeric: the products of q25 and q28 |
| Forward | M08.2 | reverse-mode autodiff is the chain of VJPs of q27 |
| Forward | M10.1 | gradient descent, Armijo steps, and the condition number of q42 |
| Forward | M07.3 | Gaussian integrals in two variables (q46, q49) behind initialization variances |
| Forward | M11.5 | softmax as maximum entropy under an energy constraint (q58) |