© 2026 NeckPinch · www.neckpinch.com · All rights reserved.
Course 2Book 2B: Spaces, Functions and ChangeChapter 8
Calculus in Several Variables
The derivative as a linear map, done rigorously, and the Hessian at a maximum.
Read with Tao, Analysis II, chapter "Several variable differential calculus", sections "Linear transformations" through "Double derivatives and Clairaut's theorem". (The section "The contraction mapping theorem" was read with [[2B.2]]; the last two sections go with [[2B.9]].) Munkres, Analysis on Manifolds, chapter "Differentiation", is a good second voice.
Course 1 met derivatives of several variables as computations: gradients, Jacobian matrices, Hessians (1A.7 The Derivative as a Linear Approximation, 1A.8 Gradient, Jacobian and Hessian). This chapter makes them rigorous. The central definition is the one stated informally in Course 1: the derivative of a map at a point is the best linear approximation to the map near that point. In one variable, a linear map is multiplication by a number, so the derivative could be treated as a number. In several variables it is a linear map, a matrix, and treating it as one is what makes everything work.
The second half of the chapter is about second derivatives. The Hessian, the matrix of second partial derivatives, is symmetric under a mild hypothesis, and so it is a quadratic form: the spectral theorem of 1A.6 Symmetric Matrices and the Spectral Theorem diagonalises it, and its eigenvalues decide whether a critical point is a peak, a pit or a pass. This is the first link of thread Q, which leads to the Ricci tensor. And at an interior maximum, the Hessian is negative semidefinite, so its trace, the Laplacian, is at most zero. That single observation is the seed of every maximum principle for the heat equation, and of Hamilton's maximum principle for Ricci flow.
By the end of this chapter you will be able to:
- define the derivative of a map between Euclidean spaces as a linear map, and compute it as the Jacobian matrix;
- give examples where partial derivatives exist but the function isn't differentiable, and prove that continuous partials are enough;
- prove and use the chain rule and the mean value inequality;
- prove Clairaut's theorem on mixed partials, and write the second-order Taylor expansion with the Hessian;
- classify critical points by the eigenvalues of the Hessian, and prove that at an interior maximum.
Propagating uncertainty
You measure a rectangular room as m by m, each with a standard uncertainty of m. The area is m². How uncertain is it?
The international Guide to the Expression of Uncertainty in Measurement (the GUM, published by the Joint Committee for Guides in Metrology as JCGM 100:2008) answers with its law of propagation of uncertainty (clause 5.1.2). For a quantity computed from independent inputs with standard uncertainties , the combined standard uncertainty is
The partial derivatives are called sensitivity coefficients. For the room, and , so
The formula comes from replacing near the measured values by its best linear approximation: small input errors produce an output error , and independent errors add in quadrature. The GUM itself notes that when is significantly nonlinear over the range of the uncertainties, higher-order terms of the Taylor expansion must be included. Both halves of that advice are this chapter: the first-order term is the derivative, and the next one is the Hessian.
Linear maps and their size
A linear map is given by an matrix. Its operator norm
is the largest factor by which stretches a vector, so for every . It is finite: if has entries , Cauchy–Schwarz gives . Hence every linear map between Euclidean spaces is Lipschitz, and so continuous. The operator norm is a norm on the space of matrices, and it satisfies , which was used to make the matrix exponential converge in 2B.6 Power Series, Exponentials and Bump Functions. (In infinite dimensions linear maps can be unbounded, and that is the starting point of functional analysis, 4A.1 Banach Spaces and Bounded Operators.)
The derivative
Throughout, is an open subset of , and is the Euclidean norm.
A map is differentiable at if there is a linear map such that
Then is the derivative of at , written or .
The error must be small compared with , as from every direction at once. The derivative is unique: if and both work, then ; taking with gives for every . Differentiable maps are continuous at , since .
Let for a symmetric matrix . Then , and , which is . So ; as a row vector, , and the gradient is .
Partial and directional derivatives
The directional derivative of at in direction is , if it exists. The partial derivatives are the directional derivatives along the coordinate directions . If is differentiable, then (take in the definition), so the matrix of has the partial derivatives as its entries: it is the Jacobian matrix.
The converse fails, and badly.
Let for and . On the axes , so both partial derivatives exist at the origin and equal . But along the ray at angle , is constant, so near the origin takes every value in (Figure 8.2). It is not even continuous at , let alone differentiable.
Even having every directional derivative is not enough. For (and ), every directional derivative at exists (Exercise 8.13), but along the parabola , so is not continuous at . Differentiability is a statement about all ways of approaching the point at once, not about straight lines.
What rescues the situation is continuity of the partials.
If all partial derivatives exist on and are continuous at , then is differentiable at .
Proof. It suffices to treat each component, so let be real-valued, and take to keep the notation light (the general case changes one coordinate at a time in the same way). Write and split the increment into a horizontal and a vertical step:
where . By the one-variable mean value theorem (2A.10 Derivatives), the brackets equal and for some . The evaluation points tend to as , so by continuity of the partials these are .
A map whose partial derivatives exist and are continuous on is called continuously differentiable, or . In practice nearly every map we differentiate is , and the theorem says that for them the Jacobian matrix is the derivative.
The chain rule and the mean value inequality
If is differentiable at and is differentiable at , then is differentiable at and
Proof. Write and , where , , , and , (set ). With , we have , so and stays bounded. Then
and dividing by gives something that tends to .
Two consequences are used constantly.
The gradient is perpendicular to level sets. For real-valued , the derivative is a row vector, and its transpose is the gradient , so that . If is a curve in a level set with , then , and the chain rule gives . Among unit vectors , is largest for (Cauchy–Schwarz), so the gradient points in the direction of steepest ascent. 2B.9 The Inverse and Implicit Function Theorems shows that, when , the level set really is a smooth hypersurface near .
Let be , and suppose the segment from to lies in . Then
Proof. Let (if there is nothing to prove) and , a real function on . By the chain rule and the mean value theorem, . Divide by .
There is no mean value equality for vector-valued maps (the circle returns to its start, but its derivative never vanishes), and the inequality is the correct replacement. It is what turns a bound on into a Lipschitz bound, and so into a contraction, in the proof of the inverse function theorem (2B.9 The Inverse and Implicit Function Theorems). It also proves the fact promised in 2B.4 Connectedness: a map with on a connected open set is constant. On each ball it is constant by the inequality, so the set where equals its value at a chosen point is open, and it is closed by continuity.
Second derivatives
If the partial derivatives are themselves differentiable, we get second partials . A function whose partials up to order exist and are continuous is , and (smooth) if this holds for every .
If is on , then .
Proof. Only two variables are involved, so take and . Consider the second difference
With , by the mean value theorem, and by the mean value theorem again. So for some point within of . Doing the same with the roles of the variables swapped gives with . Let and use continuity of both mixed partials.
The second difference treats the two variables symmetrically, which is why the order doesn't matter in the limit. Without continuity it can matter: for (and ), but (Exercise 8.14).
For a function , the Hessian at is the symmetric matrix , or equivalently the symmetric bilinear form . Its trace is the Laplacian, .
If is near , then
Proof. Let . By the chain rule and . Taylor's theorem in one variable with Lagrange remainder (2A.10 Derivatives) gives for some . The difference is at most , which is by continuity of the second partials.
Critical points and the Hessian
A critical point of is a point where . At a local maximum or minimum in the interior of , every directional derivative vanishes (the one-variable test of 2A.10 Derivatives along each line), so local extrema are critical points. Which critical points are extrema is decided by the Hessian.
Because is symmetric, the spectral theorem (1A.6 Symmetric Matrices and the Spectral Theorem) gives an orthonormal basis of eigenvectors with real eigenvalues . In these coordinates, at a critical point, the Taylor expansion reads
Let be with a critical point at , and let and be the extreme eigenvalues of .
- If all eigenvalues are positive (positive definite), is a strict local minimum.
- If all are negative (negative definite), it is a strict local maximum.
- If there are eigenvalues of both signs (indefinite), it is a saddle: neither a maximum nor a minimum.
- If some eigenvalue is and the others share a sign, the test is inconclusive.
Proof. (1) , so for small . (2) Apply (1) to . (3) Along an eigenvector with , increases for small steps; along one with , it decreases. (4) , and all have Hessian at , and have a minimum, a saddle and neither, respectively.
The number of negative eigenvalues of the Hessian is the index of the critical point: for a minimum, for a maximum, in between for saddles. The quantity is robust: a small perturbation of moves a non-degenerate critical point slightly but doesn't change its index. That robustness is the basis of Morse theory (7A.10 Morse Theory), which reads the topology of a space off the indices of the critical points of a function on it.
On a landscape with elevation , the critical points are the summits, the bottoms of hollows and the passes (cols), the lowest points on the highest routes between valleys. At a summit the Hessian of is negative definite; at a hollow, positive definite; at a pass it is indefinite, curving down along the ridge and up along the route through the pass (Figure 8.3). Mountaineers and hydrologists use the same classification: a pass is where two drainage basins meet, and the routes over a range go through its saddles because along the route the pass is the highest point, but across it, it is the lowest.
On a smooth landscape there is a further relation, a first glimpse of topology constraining analysis. For a generic island (a landscape above sea level on a disc, sloping down to the shore), the counts satisfy
the Euler characteristic of the disc (Exercise 8.18, 7A.7 Smooth Topology). Two summits force at least one pass between them.
To minimise a smooth function of many variables, as in fitting a model to data, Newton's method jumps from to the critical point of the second-order Taylor approximation at :
Near a non-degenerate minimum this converges quadratically, as in one variable (2B.2 Completeness and Contraction). But the step goes to the critical point of the quadratic model, whatever its type. If is indefinite, the model's critical point is a saddle and Newton's method happily walks towards it. This is why practical optimisation software checks whether the Hessian (or its approximation) is positive definite, and modifies the step if it isn't, for instance by adding a multiple of the identity to shift the eigenvalues positive.
The Laplacian at a maximum
Here is the observation that this chapter contributes to thread M.
Let be on an open set , with a local maximum at . Then , the Hessian is negative semidefinite (all eigenvalues ), and so
Proof. For any unit vector , has a local maximum at , so and (2A.10 Derivatives). That is, and for every . Taking to be the eigenvectors shows that every eigenvalue is , and the trace is their sum.
The statement is elementary, and its consequences are not. If solves the heat equation , then at a point where is largest, : the maximum can't increase. Made rigorous with Hamilton's trick from 2A.10 Derivatives, this is the weak maximum principle for the heat equation (6A.4 Maximum Principles). On a ring, 2B.7 Fourier Series and the First Heat Equation proved it from a formula; this argument needs no formula, and works on any domain and, with the Laplacian of a Riemannian metric, on any closed manifold.
Two extensions of Proposition 8.11 will be needed. First, for an operator with a positive semidefinite coefficient matrix , the same conclusion holds at a maximum: (Exercise 8.17). This is what makes the maximum principle work for every elliptic and parabolic equation in Course 6, and for the DeTurck–Ricci flow. Second, in 11A.4 Maximum Principles under Ricci Flow Hamilton applies the argument not to a function but to a symmetric 2-tensor, the Ricci tensor or the curvature operator, at a point and in a direction where its smallest eigenvalue is attained. The Hessian, as a quadratic form whose eigenvalues carry the geometry, is the prototype: in 9A.4 Curvature and What It Means the Riemann curvature appears as the failure of second covariant derivatives to commute (a Clairaut theorem that fails), and in 11B.1 Ricci Solitons a gradient Ricci soliton is defined by the equation , a Hessian set equal to a curvature.
History
The symmetry of mixed partial derivatives was used by Euler and by Clairaut in the 1730s and 1740s; Hermann Amandus Schwarz gave a rigorous proof in 1873, and the counterexample appeared in Giuseppe Peano's 1884 edition of Angelo Genocchi's calculus lectures. The modern definition of the derivative as a linear approximation appears in Otto Stolz's calculus text of 1893 and was championed by W. H. Young (1910) and Maurice Fréchet (1911), whose name it carries in infinite dimensions. Its advantage, that it treats all directions at once and makes the chain rule a product of matrices, is why every later book in this guidebook uses it.
The derivative of a map at a point is the linear map that approximates it to first order, represented by the Jacobian matrix. Partial derivatives alone don't guarantee it, but continuous partials do. The chain rule multiplies derivatives, and the mean value inequality turns derivative bounds into Lipschitz bounds. For functions, mixed partials commute, the Hessian is a symmetric matrix, and the second-order Taylor expansion lets its eigenvalues classify critical points. At an interior maximum the Hessian is negative semidefinite and , the seed of the maximum principle. 2B.9 The Inverse and Implicit Function Theorems uses the derivative and the contraction principle together to decide when a nonlinear equation can be solved.
Exercises
Compute for (a) ; (b) polar coordinates ; (c) on ; (d) on . Where is (b) not invertible?
Solution
(a) . (b) , with determinant : singular at . (c) . (d) .
For (with ), show that exists for every , and compute it. Show that , so is not continuous at .
Solution
For with : . For , , so . On : .
For (with ), show that and . Deduce and . Which hypothesis of Clairaut's theorem fails?
Solution
For , ; also . Similarly . So and . The second partials exist but are not continuous at .
Let . Show that is a critical point with Hessian , so the second-derivative test says nothing. Show that takes both signs in every neighbourhood of (so is not an extremum), and that changes sign six times on a small circle around : it has three "downhill valleys", one more than an ordinary saddle, room for a monkey's tail.
Solution
and , both at the origin. In polar coordinates , which changes sign six times as goes round.
A pendulum's period gives the gravitational acceleration by . With m and s (standard uncertainties), compute and its combined standard uncertainty by the law of propagation of uncertainty. Which input contributes more?
Solution
m/s². The relative sensitivities are for and for , so , so m/s². The period contributes about four times as much variance, because it enters squared.
Let be a symmetric positive semidefinite matrix and a symmetric negative semidefinite one. Show that . (Diagonalise with , and compute .) Deduce that if has an interior local maximum at and is positive semidefinite, then .
On the closed unit disc, let be smooth with on the boundary circle, inside, and only non-degenerate critical points inside. (a) Check the formula for and for a landscape with two summits joined by a ridge through one pass. (b) Explain, using the intermediate value theorem along a path from one summit to another, why two summits force the existence of a point on every such path that is the lowest point of the path, and why this suggests (but doesn't prove) the existence of a pass. A proof uses the mountain pass theorem, a min–max argument (10A.8 Min–Max and Width).
Let on . (a) Show that , , and . (b) Show that , and hence . (c) Check that (take on trust until 3A.5 Product Measures and Change of Variables).
In 11B.1 Ricci Solitons the flat metric on with this is the Gaussian shrinking soliton: its Ricci tensor is , so the soliton equation reduces to (a). The identities in (b) and (c) are exactly the normalisations in Perelman's -entropy at (12A.3 The 𝓦-Entropy), where the Gaussian is the case of equality: the flat Gaussian has entropy , and every other shrinker has less.
Solution
(a) , , trace . (b) , and . (c) The integral factors into one-dimensional integrals, each .
© 2026 NeckPinch (www.neckpinch.com). All content in the guidebook (text, mathematics, figures and exercises) is protected by copyright. All rights reserved. No part may be copied, republished or redistributed without written permission.