Skip to content

Calculus 3 problem set: vectors, gradients, the chain rule, Hessians, multiple integrals, Lagrange

ModuleS-M04 · solve · none · Pass 2 · 6 to 8 h
You buildanswers in solve/S-M04.toml (56 checked by SymPy) and 2 proofs in solve/S-M04/q34.md and q58.md (self-graded against their rubrics)
Contractnone: a pen and paper set
Testscourse/solve/S-M04/key.toml (hidden): typed answers plus reject canaries; the problems are in course/solve/S-M04/problems.md and in section 4
Needsno module. Reading: S-M01 (one-variable derivatives and integrals) and the Calculus 3 topic
Used byno call site (a solve set). Take it after M04.1 (gradcheck) and M04.2 (Jacobians, VJP, JVP) in Pass 2. It is also the whole of the solve-only M04.3 (Hessian), M04.4 (multiple integrals), and M04.5 (Lagrange), which M10.1, M07.3, and M11.5 read
MilestoneMS-P2 (the Pass 2 gate runs ol check on every solve part of the pass)
Optional depthOpenStax, Calculus Volume 3 (free), ch. 2, 4, and 5; Boyd and Vandenberghe, Convex Optimization, section 5.1, for Lagrange duality
  • The gradient collects the partial derivatives; it points uphill fastest, and its length is the steepest slope (q11, q30, q34).
  • The multivariable chain rule multiplies Jacobians, and reverse mode multiplies them from the loss backward, one vector-Jacobian product at a time (q23, q27, q28).
  • At a critical point the Hessian’s eigenvalues decide minimum, maximum, or saddle, and their ratio is the condition number that slows gradient descent (q37, q42).
  • A change of variables multiplies the area element by ∣det⁡J∣|\det J|: the rr in polar coordinates is why ∬e−(x2+y2)=π\iint e^{-(x^2 + y^2)} = \pi (q26, q46, q50).
  • At a constrained optimum the gradient is parallel to the constraint’s gradient; maximum entropy under an energy constraint is a softmax (q51 to q58).
Terminal window
ol start S-M04 # writes solve/S-M04.toml and the two proof files
ol check S-M04 # SymPy checks the answers, then asks each proof rubric (y/n)
ol check S-M04 --regrade # ask the rubrics again after you change a proof

Your model has many parameters, and training moves all of them at once along the gradient. M04.1 builds gradcheck, which compares your analytic gradients with central differences coordinate by coordinate, and M04.2 computes Jacobians and the vector-Jacobian products that reverse-mode autodiff (M08.2, L0.1) chains together. M10.1 then descends along gradients, and its convergence rate is set by the Hessian’s condition number. M07.3 sizes initial weights with Gaussian integrals in several variables, and M11.5 reads the softmax as the solution of a constrained maximization. This set checks the hand computations behind all of them.

SymbolMeaningType / shape
u⋅vu \cdot v, u×vu \times vdot product; cross product (in R3\mathbb{R}^3)scalar; vector
∥v∥\lVert v \rVertEuclidean length, v⋅v\sqrt{v \cdot v}real
∂f/∂xi\partial f / \partial x_ipartial derivative: vary xix_i, hold the others fixedfunction
∇f\nabla fgradient, the vector of partialsvector
DufD_u fdirectional derivative along a unit vector uu, ∇f⋅u\nabla f \cdot ureal
JgJ_gJacobian of g:Rn→Rmg: \mathbb{R}^n \to \mathbb{R}^m, (Jg)ij=∂gi/∂xj(J_g)_{ij} = \partial g_i / \partial x_jm×nm \times n
HfH_fHessian, (Hf)ij=∂2f/∂xi∂xj(H_f)_{ij} = \partial^2 f / \partial x_i \partial x_jn×nn \times n, symmetric
κ\kappacondition number of a positive definite HH, λmax⁡/λmin⁡\lambda_{\max}/\lambda_{\min}real
∬Df dA\iint_D f\,dAdouble integral over a region DDreal
λ\lambdaa Lagrange multiplierreal

u⋅v=∑iuivi=∥u∥∥v∥cos⁡θu \cdot v = \sum_i u_i v_i = \lVert u \rVert \lVert v \rVert \cos\theta, so u⊥vu \perp v exactly when u⋅v=0u \cdot v = 0. The cross product u×vu \times v is perpendicular to both, with length the area of the parallelogram they span. A plane is n⋅x=dn \cdot x = d with normal nn (a cross product of two edge vectors), and the distance from a point pp to it is ∣n⋅p−d∣/∥n∥|n \cdot p - d| / \lVert n \rVert.

A partial derivative differentiates in one variable with the rest held fixed, by the one-variable rules. Mixed second partials agree for smooth functions: ∂2f/∂x ∂y=∂2f/∂y ∂x\partial^2 f/\partial x\,\partial y = \partial^2 f/\partial y\,\partial x. Two gradients recur in machine learning: ∇x12∥Ax−b∥2=A⊤(Ax−b)\nabla_x \frac12 \lVert Ax - b \rVert^2 = A^\top(Ax - b), and the gradient of log-sum-exp, ∇ln⁡∑jexj\nabla \ln \sum_j e^{x_j}, is the softmax of xx.

If z=f(x(t),y(t))z = f(x(t), y(t)), then dzdt=∂f∂xdxdt+∂f∂ydydt\frac{dz}{dt} = \frac{\partial f}{\partial x}\frac{dx}{dt} + \frac{\partial f}{\partial y}\frac{dy}{dt}: one term per path from tt to zz. In general the Jacobian of a composite is the product of Jacobians, Jf∘g=Jf JgJ_{f \circ g} = J_f\, J_g. A loss is a scalar at the end of a long chain, and reverse mode evaluates the product from the left: start with u⊤=∂L/∂(output)u^\top = \partial L / \partial(\text{output}) and multiply by one Jacobian at a time, u⊤←u⊤Ju^\top \leftarrow u^\top J. Each step is a vector-Jacobian product (VJP), and backpropagation is nothing more.

The rate of change of ff at xx along a unit vector uu is Duf=ddsf(x+su)∣s=0=∇f⋅uD_u f = \frac{d}{ds} f(x + s u)\big|_{s=0} = \nabla f \cdot u, by the chain rule. By Cauchy-Schwarz it is at most ∥∇f∥\lVert \nabla f \rVert, reached at u=∇f/∥∇f∥u = \nabla f / \lVert \nabla f \rVert: the gradient is the direction of steepest ascent, −∇f-\nabla f of steepest descent, and directions perpendicular to ∇f\nabla f follow a level curve.

2.5 The Hessian and the second-derivative test

Section titled “2.5 The Hessian and the second-derivative test”

At a critical point (∇f=0\nabla f = 0), the second-order Taylor expansion is f(x+h)≈f(x)+12h⊤Hhf(x + h) \approx f(x) + \frac12 h^\top H h. If every eigenvalue of HH is positive (positive definite), the point is a local minimum; all negative, a maximum; mixed signs, a saddle. For two variables: det⁡H>0\det H > 0 and fxx>0f_{xx} > 0 is a minimum, det⁡H<0\det H < 0 is a saddle. For the quadratic 12x⊤Hx\frac12 x^\top H x, gradient descent’s error shrinks by about 1−1/κ1 - 1/\kappa per step with κ=λmax⁡/λmin⁡\kappa = \lambda_{\max}/\lambda_{\min} (M10.1).

2.6 Multiple integrals and change of variables

Section titled “2.6 Multiple integrals and change of variables”

A double integral over a rectangle is an iterated integral in either order; over other regions, the limits of the inner integral depend on the outer variable. Changing variables (u,v)↦(x,y)(u, v) \mapsto (x, y) multiplies the area element by the absolute Jacobian determinant: dx dy=∣det⁡J∣ du dvdx\,dy = |\det J|\,du\,dv. Polar coordinates have ∣det⁡J∣=r|\det J| = r, so dA=r dr dθdA = r\,dr\,d\theta.

To optimize ff subject to g=0g = 0, look for points where ∇f=λ∇g\nabla f = \lambda \nabla g and g=0g = 0: at a constrained optimum, moving along the constraint surface cannot change ff to first order, so ∇f\nabla f has no component along the surface. With several constraints, ∇f=∑kλk∇gk\nabla f = \sum_k \lambda_k \nabla g_k. Solve the system, then compare the candidates’ values.

This is a sibling of q27 and q52, not one of the graded problems.

A backward pass by hand. h=2x−1h = 2x - 1, y=σ(h)y = \sigma(h), L=−ln⁡yL = -\ln y (the cross-entropy of one positive example). At x=1/2x = 1/2: forward, h=0h = 0, y=σ(0)=1/2y = \sigma(0) = 1/2, L=ln⁡2L = \ln 2. Backward, one factor at a time: dLdy=−1y=−2\frac{dL}{dy} = -\frac{1}{y} = -2; dydh=σ(h)(1−σ(h))=14\frac{dy}{dh} = \sigma(h)(1 - \sigma(h)) = \frac14; dhdx=2\frac{dh}{dx} = 2. So dLdx=(−2)⋅14⋅2=−1\frac{dL}{dx} = (-2)\cdot\frac14\cdot 2 = -1. The middle two factors combine to dLdh=−(1−y)=y−1=−12\frac{dL}{dh} = -(1 - y) = y - 1 = -\frac12, the familiar “prediction minus label” of a logistic output. In solve/ the answer would be answer = "-1".

A constrained minimum. Minimize x2+2y2x^2 + 2y^2 subject to x+y=3x + y = 3. Lagrange: ∇f=(2x,4y)=λ(1,1)\nabla f = (2x, 4y) = \lambda (1, 1), so x=λ/2x = \lambda/2, y=λ/4y = \lambda/4, and x+y=3λ/4=3x + y = 3\lambda/4 = 3 gives λ=4\lambda = 4, (x,y)=(2,1)(x, y) = (2, 1), value 4+2=64 + 2 = 6. Check by substitution: f=x2+2(3−x)2f = x^2 + 2(3 - x)^2, f′=2x−4(3−x)=6x−12=0f' = 2x - 4(3 - x) = 6x - 12 = 0 at x=2x = 2.

Write each answer in solve/S-M04.toml:

[q5]
answer = "x + y + z = 1"
[q11]
answer = "[8, 7]"
[q35]
answer = "[[2, 3], [3, 4]]"
[q37]
answer = "saddle"
[q34]
proof = "S-M04/q34.md"

Vectors are flat lists, matrices lists of rows, and entries are exact (-1/sqrt(5)). For q33 either sign of the direction passes.

q1. (1,2,3)⋅(4,−5,6)(1, 2, 3) \cdot (4, -5, 6). [number]

q2. (1,2,3)×(4,5,6)(1, 2, 3) \times (4, 5, 6). [vector]

q3. Give the angle between (1,0)(1, 0) and (1,1)(1, 1), in radians. [number]

q4. ∥(2,3,6)∥\lVert (2, 3, 6) \rVert. [number]

q5. Give the equation of the plane through (1,0,0)(1, 0, 0), (0,1,0)(0, 1, 0), and (0,0,1)(0, 0, 1). [equation in x, y, z]

q6. Give the distance from the point (1,2,3)(1, 2, 3) to the plane x+y+z=0x + y + z = 0. [number]

q7. Give the area of the triangle with vertices (1,0,0)(1, 0, 0), (0,1,0)(0, 1, 0), (0,0,1)(0, 0, 1). [number]

q8. Are (1,2,−1)(1, 2, -1) and (3,−1,1)(3, -1, 1) orthogonal? [bool]

q9. f(x,y)=x2y+sin⁡(xy)f(x, y) = x^2 y + \sin(xy). Give ∂f/∂x\partial f / \partial x. [expr in x, y]

q10. Give ∂f/∂y\partial f / \partial y for the ff of q9. [expr in x, y]

q11. f(x,y)=x2+3xy+y2f(x, y) = x^2 + 3xy + y^2. Give ∇f(1,2)\nabla f(1, 2). [vector]

q12. f(a,b,c)=a2+b2+c2f(a, b, c) = a^2 + b^2 + c^2. Give ∇f\nabla f. [vector]

q13. f(x,y)=3x−2y+7f(x, y) = 3x - 2y + 7. Give ∇f\nabla f. [vector]

q14. ∂2∂x ∂y x3y2\dfrac{\partial^2}{\partial x\, \partial y}\, x^3 y^2. [expr in x, y]

q15. f(x,y)=ln⁡(ex+ey)f(x, y) = \ln(e^x + e^y) (log-sum-exp). Give ∇f\nabla f. [vector]

q16. f(x)=12∥Ax−b∥2f(x) = \frac{1}{2}\lVert A x - b \rVert^2 with A=(1002)A = \begin{pmatrix} 1 & 0 \\ 0 & 2 \end{pmatrix} and b=(1,1)b = (1, 1). Give ∇f(0,0)\nabla f(0, 0). [vector]

q17. f(x,y)=xeyf(x, y) = x e^y. Give ∂f/∂y\partial f / \partial y at (2,0)(2, 0). [number]

q18. The squared error of a one-weight model is L(w)=(wx−y)2L(w) = (w x - y)^2. Give ∂L/∂w\partial L / \partial w. [expr in w, x, y]

q19. z=x2+y2z = x^2 + y^2 with x=cos⁡tx = \cos t, y=sin⁡ty = \sin t. Give dz/dtdz/dt. [number]

q20. z=xyz = xy with x=t2x = t^2, y=t3y = t^3. Give dz/dtdz/dt. [expr in t]

q21. f(u,v)=uvf(u, v) = uv with u=x+yu = x + y, v=x−yv = x - y. Give ∂f/∂x\partial f / \partial x. [expr in x, y]

q22. Give ∂f/∂y\partial f / \partial y for the ff of q21. [expr in x, y]

q23. L(w)=(σ(wx)−1)2L(w) = (\sigma(w x) - 1)^2 with the sigmoid σ\sigma. Give dL/dwdL/dw at w=0w = 0, x=2x = 2. [number]

q24. f(x,y)=x2+y2f(x, y) = x^2 + y^2 with x=rcos⁡θx = r\cos\theta, y=rsin⁡θy = r\sin\theta. Give ∂f/∂r\partial f / \partial r. [expr in r]

q25. Give the Jacobian matrix of (x,y)↦(x+y, xy)(x, y) \mapsto (x + y,\ xy) at (1,2)(1, 2). [matrix]

q26. Give the determinant of the Jacobian of (r,θ)↦(rcos⁡θ, rsin⁡θ)(r, \theta) \mapsto (r\cos\theta,\ r\sin\theta). [expr in r]

q27. h=3x+1h = 3x + 1, y=h2y = h^2, L=(y−4)2L = (y - 4)^2. Give dL/dxdL/dx at x=0x = 0 (backpropagate one factor at a time). [number]

q28. f(x)=Axf(x) = A x with A=(1234)A = \begin{pmatrix} 1 & 2 \\ 3 & 4 \end{pmatrix}. Give the vector-Jacobian product u⊤Jfu^\top J_f for u=(1,1)u = (1, 1), as a column. [vector]

q29. f(x,y)=x2+y2f(x, y) = x^2 + y^2. Give the directional derivative at (1,1)(1, 1) in the direction u=(3/5,4/5)u = (3/5, 4/5). [number]

q30. f(x,y)=xyf(x, y) = xy. Give the largest directional derivative at (2,3)(2, 3) over all unit directions. [number]

q31. f(x,y)=x2+2y2f(x, y) = x^2 + 2y^2. Give the unit direction of steepest descent at (1,1)(1, 1). [vector]

q32. f(x,y)=exsin⁡yf(x, y) = e^x \sin y. Give the directional derivative at (0,π/2)(0, \pi/2) in the direction (1,0)(1, 0). [number]

q33. f(x,y)=x2−y2f(x, y) = x^2 - y^2. Give a unit direction at (1,1)(1, 1) along which the directional derivative is 0 (either sign). [vector]

q34. Prove that at a point where ∇f≠0\nabla f \ne 0, the unit direction uu that maximizes the directional derivative Duf=∇f⋅uD_u f = \nabla f \cdot u is u=∇f/∥∇f∥u = \nabla f / \lVert \nabla f \rVert, and that the maximum is ∥∇f∥\lVert \nabla f \rVert. [proof]

The Hessian and the second-derivative test

Section titled “The Hessian and the second-derivative test”

q35. Give the Hessian of f(x,y)=x2+3xy+2y2f(x, y) = x^2 + 3xy + 2y^2. [matrix]

q36. Give the Hessian of f(x,y)=x3+y3f(x, y) = x^3 + y^3 at (1,1)(1, 1). [matrix]

q37. Classify the critical point (0,0)(0, 0) of f(x,y)=x2−y2f(x, y) = x^2 - y^2: minimum, maximum, or saddle. [choice]

q38. Classify the critical point (0,0)(0, 0) of f(x,y)=x2+xy+y2f(x, y) = x^2 + xy + y^2: minimum, maximum, or saddle. [choice]

q39. Give the critical point of f(x,y)=x2+y2−2x+4yf(x, y) = x^2 + y^2 - 2x + 4y. [vector]

q40. Classify the critical point (1,0)(1, 0) of f(x,y)=x3−3x+y2f(x, y) = x^3 - 3x + y^2: minimum, maximum, or saddle. [choice]

q41. Is the Hessian of q35 positive definite? [bool]

q42. f(x)=12x⊤Hxf(x) = \frac{1}{2} x^\top H x with H=diag⁡(1,100)H = \operatorname{diag}(1, 100). Give the condition number λmax⁡/λmin⁡\lambda_{\max} / \lambda_{\min} of its Hessian. [number]

Multiple integrals and change of variables

Section titled “Multiple integrals and change of variables”

q43. ∫01∫02xy dy dx\displaystyle\int_0^1 \int_0^2 x y\, dy\, dx. [number]

q44. Give the area of the triangle 0≤y≤x≤10 \le y \le x \le 1 as a double integral. [number]

q45. ∬x2+y2≤11 dA\displaystyle\iint_{x^2 + y^2 \le 1} 1\, dA. [number]

q46. ∬R2e−(x2+y2) dA\displaystyle\iint_{\mathbb{R}^2} e^{-(x^2 + y^2)}\, dA. [number]

q47. ∫01∫0x2y dy dx\displaystyle\int_0^1 \int_0^x 2y\, dy\, dx. [number]

q48. Give the volume under z=1−x2−y2z = 1 - x^2 - y^2 and above the unit disk x2+y2≤1x^2 + y^2 \le 1. [number]

q49. ∬R212πe−(x2+y2)/2 dA\displaystyle\iint_{\mathbb{R}^2} \frac{1}{2\pi} e^{-(x^2 + y^2)/2}\, dA, the total mass of a 2-D standard normal. [number]

q50. With u=x+yu = x + y, v=x−yv = x - y, give the area of the region ∣x+y∣≤1|x + y| \le 1, ∣x−y∣≤1|x - y| \le 1. [number]

q51. Give the maximum of x+yx + y subject to x2+y2=1x^2 + y^2 = 1. [number]

q52. Give the minimum of x2+y2x^2 + y^2 subject to x+y=1x + y = 1. [number]

q53. Give the maximum of xyxy subject to x+y=10x + y = 10. [number]

q54. Maximize the entropy −∑i=13piln⁡pi-\sum_{i=1}^{3} p_i \ln p_i subject to p1+p2+p3=1p_1 + p_2 + p_3 = 1. Give the optimal p1p_1. [number]

q55. Give the maximum of xyzxyz subject to x+y+z=3x + y + z = 3 with x,y,z>0x, y, z > 0. [number]

q56. Give the minimum of 2x+3y2x + 3y subject to xy=6xy = 6 with x,y>0x, y > 0. [number]

q57. Give the point of the plane x+2y+2z=9x + 2y + 2z = 9 closest to the origin. [vector]

q58. With Lagrange multipliers, show that the distribution pp over states 1,…,n1, \dots, n with energies EiE_i that maximizes the entropy −∑ipiln⁡pi-\sum_i p_i \ln p_i subject to ∑ipi=1\sum_i p_i = 1 and ∑ipiEi=U\sum_i p_i E_i = U has the form pi=e−βEi/∑je−βEjp_i = e^{-\beta E_i} / \sum_j e^{-\beta E_j} (a softmax of −βE-\beta E). [proof]

PitfallSymptomCaught by
Dropping an inner derivative in a partial∂xsin⁡(xy)\partial_x \sin(xy) given as cos⁡(xy)\cos(xy); gradcheck fails on every product inside a nonlinearityq9, q10, q18 (canaries)
Forgetting the A⊤A^\top in a least-squares gradientdescent along the wrong direction for any non-identity AAq16 (canary)
Multiplying the derivatives of the parts instead of summing over pathsdz/dtdz/dt for z=xyz = xy given as x′(t) y′(t)x'(t)\,y'(t)q20 (canary 6*t^3)
Skipping a factor in a chain of local derivativesa backward pass off by the sigmoid’s slope or a weightq23, q27 (canaries)
Computing JuJ u when the backward pass needs u⊤Ju^\top Jthe gradient of a non-symmetric layer comes out transposedq28 (canary)
Using an unnormalized directiona directional derivative scaled by the direction’s lengthq29 (canary 4), q31 (canary)
Halving the mixed term of a Hessiana saddle misread as a minimumq35 (canary)
Forgetting the Jacobian factor in a change of variablespolar integrals off by the missing rrq45 (canary 2*pi), q50 (canary 4)
Reporting the optimizer instead of the optimumthe argmax given as the maxq52, q53 (canaries)
DirectionModuleHow it uses this
BackS-M01one-variable derivatives, the chain rule, and integrals
ForwardM04.1gradcheck compares your partials (q9 to q18) with central differences
ForwardM04.2jacobian, vjp_numeric, jvp_numeric: the products of q25 and q28
ForwardM08.2reverse-mode autodiff is the chain of VJPs of q27
ForwardM10.1gradient descent, Armijo steps, and the condition number of q42
ForwardM07.3Gaussian integrals in two variables (q46, q49) behind initialization variances
ForwardM11.5softmax as maximum entropy under an energy constraint (q58)