Book 2B

© 2026 NeckPinch · www.neckpinch.com · All rights reserved.

Course 2Book 2B: Spaces, Functions and ChangeChapter 8

Calculus in Several Variables

The derivative as a linear map, done rigorously, and the Hessian at a maximum.

37 min read · Updated Oct 2, 2026

Read with Tao, Analysis II, chapter "Several variable differential calculus", sections "Linear transformations" through "Double derivatives and Clairaut's theorem". (The section "The contraction mapping theorem" was read with [[2B.2]]; the last two sections go with [[2B.9]].) Munkres, Analysis on Manifolds, chapter "Differentiation", is a good second voice.

In this chapter · 9 sections
  1. 8.1Propagating uncertainty
  2. 8.2Linear maps and their size
  3. 8.3The derivative
  4. 8.3.1Partial and directional derivatives
  5. 8.4The chain rule and the mean value inequality
  6. 8.5Second derivatives
  7. 8.6Critical points and the Hessian
  8. 8.7The Laplacian at a maximum
  9. 8.8History
  10. 8.9Exercises

Course 1 met derivatives of several variables as computations: gradients, Jacobian matrices, Hessians (1A.7 The Derivative as a Linear Approximation, 1A.8 Gradient, Jacobian and Hessian). This chapter makes them rigorous. The central definition is the one stated informally in Course 1: the derivative of a map at a point is the best linear approximation to the map near that point. In one variable, a linear map is multiplication by a number, so the derivative could be treated as a number. In several variables it is a linear map, a matrix, and treating it as one is what makes everything work.

The second half of the chapter is about second derivatives. The Hessian, the matrix of second partial derivatives, is symmetric under a mild hypothesis, and so it is a quadratic form: the spectral theorem of 1A.6 Symmetric Matrices and the Spectral Theorem diagonalises it, and its eigenvalues decide whether a critical point is a peak, a pit or a pass. This is the first link of thread Q, which leads to the Ricci tensor. And at an interior maximum, the Hessian is negative semidefinite, so its trace, the Laplacian, is at most zero. That single observation is the seed of every maximum principle for the heat equation, and of Hamilton's maximum principle for Ricci flow.

By the end of this chapter you will be able to:

  • define the derivative of a map between Euclidean spaces as a linear map, and compute it as the Jacobian matrix;
  • give examples where partial derivatives exist but the function isn't differentiable, and prove that continuous partials are enough;
  • prove and use the chain rule and the mean value inequality;
  • prove Clairaut's theorem on mixed partials, and write the second-order Taylor expansion with the Hessian;
  • classify critical points by the eigenvalues of the Hessian, and prove that Δu≤0\Delta u \leq 0 at an interior maximum.

Propagating uncertainty

In the world In use The area of a room, with its uncertainty

You measure a rectangular room as x=4.20x = 4.20 m by y=3.60y = 3.60 m, each with a standard uncertainty of 0.010.01 m. The area is A=xy=15.12A = xy = 15.12 m². How uncertain is it?

The international Guide to the Expression of Uncertainty in Measurement (the GUM, published by the Joint Committee for Guides in Metrology as JCGM 100:2008) answers with its law of propagation of uncertainty (clause 5.1.2). For a quantity y=f(x1,…,xN)y = f(x_1, \dots, x_N) computed from independent inputs with standard uncertainties u(xi)u(x_i), the combined standard uncertainty is

uc(y)=∑i=1N(∂f∂xi)2u(xi)2.u_c(y) = \sqrt{\sum_{i=1}^N\Big(\frac{\partial f}{\partial x_i}\Big)^2u(x_i)^2}.

The partial derivatives are called sensitivity coefficients. For the room, ∂A/∂x=y=3.60\partial A/\partial x = y = 3.60 and ∂A/∂y=x=4.20\partial A/\partial y = x = 4.20, so

uc(A)=(3.60×0.01)2+(4.20×0.01)2≈0.055 m2.u_c(A) = \sqrt{(3.60 \times 0.01)^2 + (4.20 \times 0.01)^2} \approx 0.055 \text{ m}^2.

The formula comes from replacing ff near the measured values by its best linear approximation: small input errors Δxi\Delta x_i produce an output error Δy≈∑i∂f∂xiΔxi\Delta y \approx \sum_i\frac{\partial f}{\partial x_i}\Delta x_i, and independent errors add in quadrature. The GUM itself notes that when ff is significantly nonlinear over the range of the uncertainties, higher-order terms of the Taylor expansion must be included. Both halves of that advice are this chapter: the first-order term is the derivative, and the next one is the Hessian.

Linear maps and their size

A linear map L:Rn→RmL : \mathbb{R}^n \to \mathbb{R}^m is given by an m×nm \times n matrix. Its operator norm

∥L∥=sup⁡{∣Lv∣:∣v∣≤1}\|L\| = \sup\{|Lv| : |v| \leq 1\}

is the largest factor by which LL stretches a vector, so ∣Lv∣≤∥L∥ ∣v∣|Lv| \leq \|L\|\,|v| for every vv. It is finite: if LL has entries aija_{ij}, Cauchy–Schwarz gives ∣Lv∣≤(∑ijaij2)1/2∣v∣|Lv| \leq \big(\sum_{ij}a_{ij}^2\big)^{1/2}|v|. Hence every linear map between Euclidean spaces is Lipschitz, and so continuous. The operator norm is a norm on the space of matrices, and it satisfies ∥LM∥≤∥L∥ ∥M∥\|LM\| \leq \|L\|\,\|M\|, which was used to make the matrix exponential converge in 2B.6 Power Series, Exponentials and Bump Functions. (In infinite dimensions linear maps can be unbounded, and that is the starting point of functional analysis, 4A.1 Banach Spaces and Bounded Operators.)

The derivative

Throughout, UU is an open subset of Rn\mathbb{R}^n, and ∣⋅∣|\cdot| is the Euclidean norm.

Definition 8.1 Differentiability

A map f:U→Rmf : U \to \mathbb{R}^m is differentiable at x0∈Ux_0 \in U if there is a linear map L:Rn→RmL : \mathbb{R}^n \to \mathbb{R}^m such that

lim⁡h→0∣f(x0+h)−f(x0)−Lh∣∣h∣=0.\lim_{h \to 0}\frac{|f(x_0 + h) - f(x_0) - Lh|}{|h|} = 0.

Then LL is the derivative of ff at x0x_0, written Df(x0)Df(x_0) or f′(x0)f'(x_0).

The error f(x0+h)−f(x0)−Lhf(x_0 + h) - f(x_0) - Lh must be small compared with ∣h∣|h|, as h→0h \to 0 from every direction at once. The derivative is unique: if L1L_1 and L2L_2 both work, then ∣(L1−L2)h∣/∣h∣→0|(L_1 - L_2)h|/|h| \to 0; taking h=tvh = tv with t→0t \to 0 gives (L1−L2)v=0(L_1 - L_2)v = 0 for every vv. Differentiable maps are continuous at x0x_0, since ∣f(x0+h)−f(x0)∣≤∣Lh∣+o(∣h∣)→0|f(x_0 + h) - f(x_0)| \leq |Lh| + o(|h|) \to 0.

Example 8.2 A quadratic form

Let f(x)=xTAxf(x) = x^{\mathsf T}Ax for a symmetric n×nn \times n matrix AA. Then f(x0+h)=f(x0)+2x0TAh+hTAhf(x_0 + h) = f(x_0) + 2x_0^{\mathsf T}Ah + h^{\mathsf T}Ah, and ∣hTAh∣≤∥A∥ ∣h∣2|h^{\mathsf T}Ah| \leq \|A\|\,|h|^2, which is o(∣h∣)o(|h|). So Df(x0)h=2x0TAhDf(x_0)h = 2x_0^{\mathsf T}Ah; as a row vector, Df(x0)=2x0TADf(x_0) = 2x_0^{\mathsf T}A, and the gradient is ∇f(x0)=2Ax0\nabla f(x_0) = 2Ax_0.

Figure 8.1. The graph of f(x,y)=12x2−14y2+13xyf(x, y) = \tfrac12x^2 - \tfrac14y^2 + \tfrac13xy at the point (0.6,0.4)(0.6, 0.4), with its tangent plane there, the graph of the best linear approximation f(x0)+Df(x0)hf(x_0) + Df(x_0)h. The gap between them is of order ∣h∣2|h|^2: it is the second-order term, governed by the Hessian.

Partial and directional derivatives

The directional derivative of ff at x0x_0 in direction vv is Dvf(x0)=lim⁡t→0f(x0+tv)−f(x0)tD_vf(x_0) = \lim_{t\to 0}\frac{f(x_0 + tv) - f(x_0)}{t}, if it exists. The partial derivatives ∂jf=∂f∂xj\partial_jf = \frac{\partial f}{\partial x_j} are the directional derivatives along the coordinate directions eje_j. If ff is differentiable, then Dvf(x0)=Df(x0)vD_vf(x_0) = Df(x_0)v (take h=tvh = tv in the definition), so the matrix of Df(x0)Df(x_0) has the partial derivatives ∂jfi(x0)\partial_jf_i(x_0) as its entries: it is the Jacobian matrix.

The converse fails, and badly.

Example 8.3 Partial derivatives are not enough

Let f(x,y)=xyx2+y2f(x, y) = \frac{xy}{x^2 + y^2} for (x,y)≠0(x, y) \neq 0 and f(0,0)=0f(0, 0) = 0. On the axes f=0f = 0, so both partial derivatives exist at the origin and equal 00. But along the ray at angle θ\theta, f=12sin⁡2θf = \tfrac12\sin 2\theta is constant, so near the origin ff takes every value in [−12,12][-\tfrac12, \tfrac12] (Figure 8.2). It is not even continuous at 00, let alone differentiable.

Even having every directional derivative is not enough. For g(x,y)=x2yx4+y2g(x, y) = \frac{x^2y}{x^4 + y^2} (and g(0)=0g(0) = 0), every directional derivative at 00 exists (Exercise 8.13), but g=12g = \tfrac12 along the parabola y=x2y = x^2, so gg is not continuous at 00. Differentiability is a statement about all ways of approaching the point at once, not about straight lines.

Figure 8.2. f=xyx2+y2f = \frac{xy}{x^2 + y^2} is constant along rays from the origin, equal to 12sin⁡2θ\tfrac12\sin 2\theta. Its partial derivatives at 00 exist (it vanishes on the axes), but every value in [−12,12][-\tfrac12, \tfrac12] occurs in every neighbourhood of 00.

What rescues the situation is continuity of the partials.

Theorem 8.4 Continuous partials imply differentiability

If all partial derivatives ∂jf\partial_jf exist on UU and are continuous at x0x_0, then ff is differentiable at x0x_0.

Proof. It suffices to treat each component, so let ff be real-valued, and take n=2n = 2 to keep the notation light (the general case changes one coordinate at a time in the same way). Write h=(h1,h2)h = (h_1, h_2) and split the increment into a horizontal and a vertical step:

f(x0+h)−f(x0)=[f(a+h1,b+h2)−f(a,b+h2)]+[f(a,b+h2)−f(a,b)],f(x_0 + h) - f(x_0) = \big[f(a + h_1, b + h_2) - f(a, b + h_2)\big] + \big[f(a, b + h_2) - f(a, b)\big],

where x0=(a,b)x_0 = (a, b). By the one-variable mean value theorem (2A.10 Derivatives), the brackets equal ∂1f(a+θ1h1,b+h2) h1\partial_1f(a + \theta_1h_1, b + h_2)\,h_1 and ∂2f(a,b+θ2h2) h2\partial_2f(a, b + \theta_2h_2)\,h_2 for some θ1,θ2∈(0,1)\theta_1, \theta_2 \in (0, 1). The evaluation points tend to x0x_0 as h→0h \to 0, so by continuity of the partials these are ∂1f(x0)h1+∂2f(x0)h2+o(∣h∣)\partial_1f(x_0)h_1 + \partial_2f(x_0)h_2 + o(|h|).

A map whose partial derivatives exist and are continuous on UU is called continuously differentiable, or C1C^1. In practice nearly every map we differentiate is C1C^1, and the theorem says that for them the Jacobian matrix is the derivative.

The chain rule and the mean value inequality

Theorem 8.5 Chain rule

If ff is differentiable at x0x_0 and gg is differentiable at f(x0)f(x_0), then g∘fg \circ f is differentiable at x0x_0 and

D(g∘f)(x0)=Dg(f(x0)) Df(x0).D(g \circ f)(x_0) = Dg(f(x_0))\,Df(x_0).

Proof. Write f(x0+h)=f(x0)+Ah+ε1(h)∣h∣f(x_0 + h) = f(x_0) + Ah + \varepsilon_1(h)|h| and g(y0+k)=g(y0)+Bk+ε2(k)∣k∣g(y_0 + k) = g(y_0) + Bk + \varepsilon_2(k)|k|, where A=Df(x0)A = Df(x_0), y0=f(x0)y_0 = f(x_0), B=Dg(y0)B = Dg(y_0), and ε1(h)→0\varepsilon_1(h) \to 0, ε2(k)→0\varepsilon_2(k) \to 0 (set ε2(0)=0\varepsilon_2(0) = 0). With k=f(x0+h)−y0=Ah+ε1(h)∣h∣k = f(x_0 + h) - y_0 = Ah + \varepsilon_1(h)|h|, we have ∣k∣≤(∥A∥+∣ε1(h)∣)∣h∣|k| \leq (\|A\| + |\varepsilon_1(h)|)|h|, so k→0k \to 0 and ∣k∣/∣h∣|k|/|h| stays bounded. Then

g(f(x0+h))−g(y0)−BAh=Bε1(h)∣h∣+ε2(k)∣k∣,g(f(x_0 + h)) - g(y_0) - BAh = B\varepsilon_1(h)|h| + \varepsilon_2(k)|k|,

and dividing by ∣h∣|h| gives something that tends to 00.

Two consequences are used constantly.

The gradient is perpendicular to level sets. For real-valued ff, the derivative is a row vector, and its transpose is the gradient ∇f\nabla f, so that Df(x0)v=∇f(x0)⋅vDf(x_0)v = \nabla f(x_0)\cdot v. If γ(t)\gamma(t) is a curve in a level set {f=c}\{f = c\} with γ(0)=x0\gamma(0) = x_0, then f(γ(t))=cf(\gamma(t)) = c, and the chain rule gives ∇f(x0)⋅γ′(0)=0\nabla f(x_0)\cdot\gamma'(0) = 0. Among unit vectors vv, ∇f⋅v\nabla f\cdot v is largest for v=∇f/∣∇f∣v = \nabla f/|\nabla f| (Cauchy–Schwarz), so the gradient points in the direction of steepest ascent. 2B.9 The Inverse and Implicit Function Theorems shows that, when ∇f(x0)≠0\nabla f(x_0) \neq 0, the level set really is a smooth hypersurface near x0x_0.

Proposition 8.6 Mean value inequality

Let f:U→Rmf : U \to \mathbb{R}^m be C1C^1, and suppose the segment from aa to bb lies in UU. Then

∣f(b)−f(a)∣≤sup⁡0≤t≤1∥Df(a+t(b−a))∥ ∣b−a∣.|f(b) - f(a)| \leq \sup_{0 \leq t \leq 1}\|Df(a + t(b - a))\|\,|b - a|.

Proof. Let v=f(b)−f(a)v = f(b) - f(a) (if v=0v = 0 there is nothing to prove) and ϕ(t)=v⋅f(a+t(b−a))\phi(t) = v\cdot f(a + t(b - a)), a real function on [0,1][0, 1]. By the chain rule and the mean value theorem, ∣v∣2=ϕ(1)−ϕ(0)=ϕ′(c)=v⋅Df(a+c(b−a))(b−a)≤∣v∣ ∥Df(…)∥ ∣b−a∣|v|^2 = \phi(1) - \phi(0) = \phi'(c) = v\cdot Df(a + c(b - a))(b - a) \leq |v|\,\|Df(\ldots)\|\,|b - a|. Divide by ∣v∣|v|.

There is no mean value equality for vector-valued maps (the circle t↦(cos⁡t,sin⁡t)t \mapsto (\cos t, \sin t) returns to its start, but its derivative never vanishes), and the inequality is the correct replacement. It is what turns a bound on DfDf into a Lipschitz bound, and so into a contraction, in the proof of the inverse function theorem (2B.9 The Inverse and Implicit Function Theorems). It also proves the fact promised in 2B.4 Connectedness: a map with Df=0Df = 0 on a connected open set is constant. On each ball it is constant by the inequality, so the set where ff equals its value at a chosen point is open, and it is closed by continuity.

Second derivatives

If the partial derivatives ∂jf\partial_jf are themselves differentiable, we get second partials ∂i∂jf\partial_i\partial_jf. A function whose partials up to order kk exist and are continuous is CkC^k, and C∞C^\infty (smooth) if this holds for every kk.

Theorem 8.7 Clairaut's theorem (symmetry of mixed partials)

If ff is C2C^2 on UU, then ∂i∂jf=∂j∂if\partial_i\partial_jf = \partial_j\partial_if.

Proof. Only two variables are involved, so take n=2n = 2 and x0=(a,b)x_0 = (a, b). Consider the second difference

Δ(h)=f(a+h,b+h)−f(a+h,b)−f(a,b+h)+f(a,b).\Delta(h) = f(a + h, b + h) - f(a + h, b) - f(a, b + h) + f(a, b).

With ϕ(s)=f(s,b+h)−f(s,b)\phi(s) = f(s, b + h) - f(s, b), Δ(h)=ϕ(a+h)−ϕ(a)=h ϕ′(a+θh)\Delta(h) = \phi(a + h) - \phi(a) = h\,\phi'(a + \theta h) by the mean value theorem, and ϕ′(s)=∂1f(s,b+h)−∂1f(s,b)=h ∂2∂1f(s,b+θ′h)\phi'(s) = \partial_1f(s, b + h) - \partial_1f(s, b) = h\,\partial_2\partial_1f(s, b + \theta'h) by the mean value theorem again. So Δ(h)/h2=∂2∂1f(ph)\Delta(h)/h^2 = \partial_2\partial_1f(p_h) for some point php_h within 2∣h∣\sqrt2|h| of x0x_0. Doing the same with the roles of the variables swapped gives Δ(h)/h2=∂1∂2f(qh)\Delta(h)/h^2 = \partial_1\partial_2f(q_h) with qh→x0q_h \to x_0. Let h→0h \to 0 and use continuity of both mixed partials.

The second difference Δ(h)\Delta(h) treats the two variables symmetrically, which is why the order doesn't matter in the limit. Without continuity it can matter: for f(x,y)=xy(x2−y2)x2+y2f(x, y) = \frac{xy(x^2 - y^2)}{x^2 + y^2} (and f(0)=0f(0) = 0), ∂x∂yf(0,0)=1\partial_x\partial_yf(0, 0) = 1 but ∂y∂xf(0,0)=−1\partial_y\partial_xf(0, 0) = -1 (Exercise 8.14).

Definition 8.8 The Hessian

For a C2C^2 function f:U→Rf : U \to \mathbb{R}, the Hessian at xx is the symmetric matrix ∇2f(x)=(∂i∂jf(x))i,j\nabla^2f(x) = \big(\partial_i\partial_jf(x)\big)_{i,j}, or equivalently the symmetric bilinear form ∇2f(x)(v,w)=vT∇2f(x) w\nabla^2f(x)(v, w) = v^{\mathsf T}\nabla^2f(x)\,w. Its trace is the Laplacian, Δf=∑i∂i2f\Delta f = \sum_i\partial_i^2f.

Theorem 8.9 Second-order Taylor expansion

If ff is C2C^2 near x0x_0, then

f(x0+h)=f(x0)+∇f(x0)⋅h+12 hT∇2f(x0) h+o(∣h∣2).f(x_0 + h) = f(x_0) + \nabla f(x_0)\cdot h + \tfrac12\,h^{\mathsf T}\nabla^2f(x_0)\,h + o(|h|^2).

Proof. Let ϕ(t)=f(x0+th)\phi(t) = f(x_0 + th). By the chain rule ϕ′(t)=∇f(x0+th)⋅h\phi'(t) = \nabla f(x_0 + th)\cdot h and ϕ′′(t)=hT∇2f(x0+th) h\phi''(t) = h^{\mathsf T}\nabla^2f(x_0 + th)\,h. Taylor's theorem in one variable with Lagrange remainder (2A.10 Derivatives) gives ϕ(1)=ϕ(0)+ϕ′(0)+12ϕ′′(c)\phi(1) = \phi(0) + \phi'(0) + \tfrac12\phi''(c) for some c∈(0,1)c \in (0, 1). The difference 12hT(∇2f(x0+ch)−∇2f(x0))h\tfrac12 h^{\mathsf T}\big(\nabla^2f(x_0 + ch) - \nabla^2f(x_0)\big)h is at most 12∥∇2f(x0+ch)−∇2f(x0)∥ ∣h∣2\tfrac12\|\nabla^2f(x_0 + ch) - \nabla^2f(x_0)\|\,|h|^2, which is o(∣h∣2)o(|h|^2) by continuity of the second partials.

Critical points and the Hessian

A critical point of ff is a point where ∇f=0\nabla f = 0. At a local maximum or minimum in the interior of UU, every directional derivative vanishes (the one-variable test of 2A.10 Derivatives along each line), so local extrema are critical points. Which critical points are extrema is decided by the Hessian.

Because ∇2f(x0)\nabla^2f(x_0) is symmetric, the spectral theorem (1A.6 Symmetric Matrices and the Spectral Theorem) gives an orthonormal basis of eigenvectors v1,…,vnv_1, \dots, v_n with real eigenvalues λ1,…,λn\lambda_1, \dots, \lambda_n. In these coordinates, at a critical point, the Taylor expansion reads

f(x0+h)=f(x0)+12∑iλici2+o(∣h∣2),h=∑icivi.f(x_0 + h) = f(x_0) + \tfrac12\sum_i\lambda_ic_i^2 + o(|h|^2), \qquad h = \sum_ic_iv_i.
Theorem 8.10 Second-derivative test

Let ff be C2C^2 with a critical point at x0x_0, and let λmin⁡\lambda_{\min} and λmax⁡\lambda_{\max} be the extreme eigenvalues of ∇2f(x0)\nabla^2f(x_0).

  1. If all eigenvalues are positive (positive definite), x0x_0 is a strict local minimum.
  2. If all are negative (negative definite), it is a strict local maximum.
  3. If there are eigenvalues of both signs (indefinite), it is a saddle: neither a maximum nor a minimum.
  4. If some eigenvalue is 00 and the others share a sign, the test is inconclusive.

Proof. (1) ∑λici2≥λmin⁡∣h∣2\sum\lambda_ic_i^2 \geq \lambda_{\min}|h|^2, so f(x0+h)−f(x0)≥(12λmin⁡−o(∣h∣2)∣h∣2)∣h∣2>0f(x_0 + h) - f(x_0) \geq \big(\tfrac12\lambda_{\min} - \frac{o(|h|^2)}{|h|^2}\big)|h|^2 > 0 for small h≠0h \neq 0. (2) Apply (1) to −f-f. (3) Along an eigenvector with λ>0\lambda > 0, ff increases for small steps; along one with λ<0\lambda < 0, it decreases. (4) x4+y2x^4 + y^2, −x4+y2-x^4 + y^2 and x3+y2x^3 + y^2 all have Hessian diag(0,2)\mathrm{diag}(0, 2) at 00, and have a minimum, a saddle and neither, respectively.

The number of negative eigenvalues of the Hessian is the index of the critical point: 00 for a minimum, nn for a maximum, in between for saddles. The quantity is robust: a small perturbation of ff moves a non-degenerate critical point slightly but doesn't change its index. That robustness is the basis of Morse theory (7A.10 Morse Theory), which reads the topology of a space off the indices of the critical points of a function on it.

In the world Model Peaks, pits and passes

On a landscape with elevation h(x,y)h(x, y), the critical points are the summits, the bottoms of hollows and the passes (cols), the lowest points on the highest routes between valleys. At a summit the Hessian of hh is negative definite; at a hollow, positive definite; at a pass it is indefinite, curving down along the ridge and up along the route through the pass (Figure 8.3). Mountaineers and hydrologists use the same classification: a pass is where two drainage basins meet, and the routes over a range go through its saddles because along the route the pass is the highest point, but across it, it is the lowest.

On a smooth landscape there is a further relation, a first glimpse of topology constraining analysis. For a generic island (a landscape above sea level on a disc, sloping down to the shore), the counts satisfy

#summits−#passes+#pits=1,\#\text{summits} - \#\text{passes} + \#\text{pits} = 1,

the Euler characteristic of the disc (Exercise 8.18, 7A.7 Smooth Topology). Two summits force at least one pass between them.

Figure 8.3. Contours of a smooth two-peaked landscape (illustrative, not real terrain). The summits (dots) have negative definite Hessians. The pass (cross) has an indefinite Hessian: along the ridge joining the summits the surface curves down (eigenvalue <0< 0), across the ridge it curves up (eigenvalue >0> 0). The contour through the pass crosses itself there.
In the world In use Newton's method needs a positive definite Hessian

To minimise a smooth function ff of many variables, as in fitting a model to data, Newton's method jumps from xx to the critical point of the second-order Taylor approximation at xx:

xnew=x−∇2f(x)−1∇f(x).x_{\text{new}} = x - \nabla^2f(x)^{-1}\nabla f(x).

Near a non-degenerate minimum this converges quadratically, as in one variable (2B.2 Completeness and Contraction). But the step goes to the critical point of the quadratic model, whatever its type. If ∇2f(x)\nabla^2f(x) is indefinite, the model's critical point is a saddle and Newton's method happily walks towards it. This is why practical optimisation software checks whether the Hessian (or its approximation) is positive definite, and modifies the step if it isn't, for instance by adding a multiple of the identity to shift the eigenvalues positive.

The Laplacian at a maximum

Here is the observation that this chapter contributes to thread M.

Proposition 8.11 At an interior maximum, Δu≤0\Delta u \leq 0

Let uu be C2C^2 on an open set U⊆RnU \subseteq \mathbb{R}^n, with a local maximum at x0x_0. Then ∇u(x0)=0\nabla u(x_0) = 0, the Hessian ∇2u(x0)\nabla^2u(x_0) is negative semidefinite (all eigenvalues ≤0\leq 0), and so

Δu(x0)=tr⁡∇2u(x0)=∑iλi≤0.\Delta u(x_0) = \operatorname{tr}\nabla^2u(x_0) = \sum_i\lambda_i \leq 0.

Proof. For any unit vector vv, ϕ(t)=u(x0+tv)\phi(t) = u(x_0 + tv) has a local maximum at t=0t = 0, so ϕ′(0)=0\phi'(0) = 0 and ϕ′′(0)≤0\phi''(0) \leq 0 (2A.10 Derivatives). That is, ∇u(x0)⋅v=0\nabla u(x_0)\cdot v = 0 and vT∇2u(x0) v≤0v^{\mathsf T}\nabla^2u(x_0)\,v \leq 0 for every vv. Taking vv to be the eigenvectors shows that every eigenvalue is ≤0\leq 0, and the trace is their sum.

The statement is elementary, and its consequences are not. If uu solves the heat equation ∂tu=Δu\partial_tu = \Delta u, then at a point where u(⋅,t)u(\cdot, t) is largest, ∂tu=Δu≤0\partial_tu = \Delta u \leq 0: the maximum can't increase. Made rigorous with Hamilton's trick from 2A.10 Derivatives, this is the weak maximum principle for the heat equation (6A.4 Maximum Principles). On a ring, 2B.7 Fourier Series and the First Heat Equation proved it from a formula; this argument needs no formula, and works on any domain and, with the Laplacian of a Riemannian metric, on any closed manifold.

Where this goes From the Hessian to the Ricci tensor

Two extensions of Proposition 8.11 will be needed. First, for an operator ∑aij∂i∂ju\sum a_{ij}\partial_i\partial_ju with a positive semidefinite coefficient matrix (aij)(a_{ij}), the same conclusion holds at a maximum: tr⁡(A ∇2u)≤0\operatorname{tr}(A\,\nabla^2u) \leq 0 (Exercise 8.17). This is what makes the maximum principle work for every elliptic and parabolic equation in Course 6, and for the DeTurck–Ricci flow. Second, in 11A.4 Maximum Principles under Ricci Flow Hamilton applies the argument not to a function but to a symmetric 2-tensor, the Ricci tensor or the curvature operator, at a point and in a direction where its smallest eigenvalue is attained. The Hessian, as a quadratic form whose eigenvalues carry the geometry, is the prototype: in 9A.4 Curvature and What It Means the Riemann curvature appears as the failure of second covariant derivatives to commute (a Clairaut theorem that fails), and in 11B.1 Ricci Solitons a gradient Ricci soliton is defined by the equation Ric+∇2f=λg\mathrm{Ric} + \nabla^2f = \lambda g, a Hessian set equal to a curvature.

History

The symmetry of mixed partial derivatives was used by Euler and by Clairaut in the 1730s and 1740s; Hermann Amandus Schwarz gave a rigorous proof in 1873, and the counterexample xy(x2−y2)x2+y2\frac{xy(x^2-y^2)}{x^2+y^2} appeared in Giuseppe Peano's 1884 edition of Angelo Genocchi's calculus lectures. The modern definition of the derivative as a linear approximation appears in Otto Stolz's calculus text of 1893 and was championed by W. H. Young (1910) and Maurice Fréchet (1911), whose name it carries in infinite dimensions. Its advantage, that it treats all directions at once and makes the chain rule a product of matrices, is why every later book in this guidebook uses it.

Recall Where we stand

The derivative of a map f:Rn→Rmf : \mathbb{R}^n \to \mathbb{R}^m at a point is the linear map that approximates it to first order, represented by the Jacobian matrix. Partial derivatives alone don't guarantee it, but continuous partials do. The chain rule multiplies derivatives, and the mean value inequality turns derivative bounds into Lipschitz bounds. For C2C^2 functions, mixed partials commute, the Hessian is a symmetric matrix, and the second-order Taylor expansion lets its eigenvalues classify critical points. At an interior maximum the Hessian is negative semidefinite and Δu≤0\Delta u \leq 0, the seed of the maximum principle. 2B.9 The Inverse and Implicit Function Theorems uses the derivative and the contraction principle together to decide when a nonlinear equation f(x)=yf(x) = y can be solved.

Exercises

Exercise 8.12 Jacobians

Compute DfDf for (a) f(x,y)=(x2−y2,2xy)f(x, y) = (x^2 - y^2, 2xy); (b) polar coordinates f(r,θ)=(rcos⁡θ,rsin⁡θ)f(r, \theta) = (r\cos\theta, r\sin\theta); (c) f(x)=∣x∣2f(x) = |x|^2 on Rn\mathbb{R}^n; (d) f(x)=∣x∣f(x) = |x| on Rn∖{0}\mathbb{R}^n \setminus \{0\}. Where is (b) not invertible?

Solution

(a) (2x−2y2y2x)\begin{pmatrix}2x & -2y\\ 2y & 2x\end{pmatrix}. (b) (cos⁡θ−rsin⁡θsin⁡θrcos⁡θ)\begin{pmatrix}\cos\theta & -r\sin\theta\\ \sin\theta & r\cos\theta\end{pmatrix}, with determinant rr: singular at r=0r = 0. (c) 2xT2x^{\mathsf T}. (d) xT/∣x∣x^{\mathsf T}/|x|.

Exercise 8.13 Every direction, still not differentiable

For g(x,y)=x2yx4+y2g(x, y) = \frac{x^2y}{x^4 + y^2} (with g(0)=0g(0) = 0), show that Dvg(0)D_vg(0) exists for every vv, and compute it. Show that g(x,x2)=12g(x, x^2) = \tfrac12, so gg is not continuous at 00.

Solution

For v=(a,b)v = (a, b) with b≠0b \neq 0: g(ta,tb)/t=a2bt2a4+b2→a2bg(ta, tb)/t = \frac{a^2b}{t^2a^4 + b^2} \to \frac{a^2}{b}. For b=0b = 0, g(ta,0)=0g(ta, 0) = 0, so Dvg(0)=0D_vg(0) = 0. On y=x2y = x^2: x42x4=12\frac{x^4}{2x^4} = \frac12.

Exercise 8.14 When mixed partials disagree

For f(x,y)=xy(x2−y2)x2+y2f(x, y) = \frac{xy(x^2 - y^2)}{x^2 + y^2} (with f(0)=0f(0) = 0), show that ∂xf(0,y)=−y\partial_xf(0, y) = -y and ∂yf(x,0)=x\partial_yf(x, 0) = x. Deduce ∂y∂xf(0,0)=−1\partial_y\partial_xf(0, 0) = -1 and ∂x∂yf(0,0)=1\partial_x\partial_yf(0, 0) = 1. Which hypothesis of Clairaut's theorem fails?

Solution

For y≠0y \neq 0, ∂xf(0,y)=lim⁡x→0f(x,y)x=lim⁡y(x2−y2)x2+y2=−y\partial_xf(0, y) = \lim_{x\to0}\frac{f(x, y)}{x} = \lim\frac{y(x^2 - y^2)}{x^2 + y^2} = -y; also ∂xf(0,0)=0\partial_xf(0, 0) = 0. Similarly ∂yf(x,0)=x\partial_yf(x, 0) = x. So ∂y(∂xf)(0,0)=−1\partial_y(\partial_xf)(0, 0) = -1 and ∂x(∂yf)(0,0)=1\partial_x(\partial_yf)(0, 0) = 1. The second partials exist but are not continuous at 00.

Exercise 8.15 The monkey saddle

Let f(x,y)=x3−3xy2f(x, y) = x^3 - 3xy^2. Show that 00 is a critical point with Hessian 00, so the second-derivative test says nothing. Show that ff takes both signs in every neighbourhood of 00 (so 00 is not an extremum), and that ff changes sign six times on a small circle around 00: it has three "downhill valleys", one more than an ordinary saddle, room for a monkey's tail.

Solution

∇f=(3x2−3y2,−6xy)\nabla f = (3x^2 - 3y^2, -6xy) and ∇2f=(6x−6y−6y−6x)\nabla^2f = \begin{pmatrix}6x & -6y\\ -6y & -6x\end{pmatrix}, both 00 at the origin. In polar coordinates f=r3cos⁡3θf = r^3\cos 3\theta, which changes sign six times as θ\theta goes round.

Exercise 8.16 Propagating uncertainty

A pendulum's period TT gives the gravitational acceleration by g=4π2L/T2g = 4\pi^2L/T^2. With L=1.000±0.002L = 1.000 \pm 0.002 m and T=2.006±0.004T = 2.006 \pm 0.004 s (standard uncertainties), compute gg and its combined standard uncertainty by the law of propagation of uncertainty. Which input contributes more?

Solution

g=4π2/2.0062≈9.811g = 4\pi^2/2.006^2 \approx 9.811 m/s². The relative sensitivities are 11 for LL and −2-2 for TT, so u(g)g=(0.002)2+(2×0.004/2.006)2≈4×10−6+1.59×10−5≈0.0045\frac{u(g)}{g} = \sqrt{(0.002)^2 + (2 \times 0.004/2.006)^2} \approx \sqrt{4\times10^{-6} + 1.59\times10^{-5}} \approx 0.0045, so u(g)≈0.044u(g) \approx 0.044 m/s². The period contributes about four times as much variance, because it enters squared.

Exercise 8.17 Maximum principle with variable coefficients

Let AA be a symmetric positive semidefinite n×nn \times n matrix and BB a symmetric negative semidefinite one. Show that tr⁡(AB)≤0\operatorname{tr}(AB) \leq 0. (Diagonalise A=∑iμiviviTA = \sum_i\mu_iv_iv_i^{\mathsf T} with μi≥0\mu_i \geq 0, and compute tr⁡(AB)=∑iμi viTBvi\operatorname{tr}(AB) = \sum_i\mu_i\,v_i^{\mathsf T}Bv_i.) Deduce that if uu has an interior local maximum at x0x_0 and (aij(x0))(a_{ij}(x_0)) is positive semidefinite, then ∑i,jaij(x0) ∂i∂ju(x0)≤0\sum_{i,j}a_{ij}(x_0)\,\partial_i\partial_ju(x_0) \leq 0.

Exercise 8.18 Summits, passes and pits

On the closed unit disc, let hh be smooth with h=0h = 0 on the boundary circle, h>0h > 0 inside, and only non-degenerate critical points inside. (a) Check the formula #max−#saddles+#min=1\#\text{max} - \#\text{saddles} + \#\text{min} = 1 for h=1−∣x∣2h = 1 - |x|^2 and for a landscape with two summits joined by a ridge through one pass. (b) Explain, using the intermediate value theorem along a path from one summit to another, why two summits force the existence of a point on every such path that is the lowest point of the path, and why this suggests (but doesn't prove) the existence of a pass. A proof uses the mountain pass theorem, a min–max argument (10A.8 Min–Max and Width).

Exercise 8.19 Rehearsal: the Gaussian soliton's potential

Let f(x)=∣x∣24f(x) = \frac{|x|^2}{4} on Rn\mathbb{R}^n. (a) Show that ∇f=x2\nabla f = \frac x2, ∇2f=12I\nabla^2f = \tfrac12 I, and Δf=n2\Delta f = \frac n2. (b) Show that ∣∇f∣2=f|\nabla f|^2 = f, and hence 2Δf−∣∇f∣2+f−n=02\Delta f - |\nabla f|^2 + f - n = 0. (c) Check that ∫Rn(4π)−n/2e−f dx=1\int_{\mathbb{R}^n}(4\pi)^{-n/2}e^{-f}\,dx = 1 (take ∫Re−x2/4dx=4π\int_{\mathbb{R}}e^{-x^2/4}dx = \sqrt{4\pi} on trust until 3A.5 Product Measures and Change of Variables).

In 11B.1 Ricci Solitons the flat metric on Rn\mathbb{R}^n with this ff is the Gaussian shrinking soliton: its Ricci tensor is 00, so the soliton equation Ric+∇2f=12g\mathrm{Ric} + \nabla^2f = \frac{1}{2}g reduces to (a). The identities in (b) and (c) are exactly the normalisations in Perelman's W\mathcal{W}-entropy at τ=1\tau = 1 (12A.3 The 𝓦-Entropy), where the Gaussian is the case of equality: the flat Gaussian has entropy 00, and every other shrinker has less.

Solution

(a) ∂if=xi/2\partial_if = x_i/2, ∂i∂jf=12δij\partial_i\partial_jf = \tfrac12\delta_{ij}, trace n/2n/2. (b) ∣∇f∣2=∣x∣2/4=f|\nabla f|^2 = |x|^2/4 = f, and 2⋅n2−f+f−n=02\cdot\frac n2 - f + f - n = 0. (c) The integral factors into nn one-dimensional integrals, each (4π)−1/24π=1(4\pi)^{-1/2}\sqrt{4\pi} = 1.

© 2026 NeckPinch (www.neckpinch.com). All content in the guidebook (text, mathematics, figures and exercises) is protected by copyright. All rights reserved. No part may be copied, republished or redistributed without written permission.