Book 2B

© 2026 NeckPinch · www.neckpinch.com · All rights reserved.

Course 2Book 2B: Spaces, Functions and ChangeChapter 9

The Inverse and Implicit Function Theorems

When a nonlinear equation can be solved: linearise, invert, and use a contraction.

34 min read · Updated Oct 2, 2026

Read with Tao, Analysis II, chapter "Several variable differential calculus", sections "The inverse function theorem in several variable calculus" and "The implicit function theorem". Munkres, Analysis on Manifolds, chapter "Inverse Functions and the Implicit Function Theorem", has a good geometric account of the implicit function theorem and the rank of a map.

In this chapter · 6 sections
  1. 9.1When a robot arm gets stuck
  2. 9.2The inverse function theorem
  3. 9.2.1The proof
  4. 9.2.2Local, not global
  5. 9.3The implicit function theorem
  6. 9.3.1Level sets are smooth where the gradient isn't zero
  7. 9.4Lagrange multipliers
  8. 9.5History
  9. 9.6Exercises

When can a nonlinear equation be solved? For a linear system Ax=yAx = y the answer is clean: if AA is invertible, there is exactly one solution, x=A−1yx = A^{-1}y, and it depends continuously on yy. For a nonlinear system f(x)=yf(x) = y there is no general answer. But there is a local one, and it is the most important theorem of this book: if the linearisation is invertible, the nonlinear problem is solvable nearby. Precisely, if f(x0)=y0f(x_0) = y_0 and the derivative Df(x0)Df(x_0) is an invertible matrix, then for every yy close to y0y_0 the equation f(x)=yf(x) = y has exactly one solution xx close to x0x_0, and that solution depends smoothly on yy. This is the inverse function theorem.

Its proof combines the two engines built so far: the derivative (2B.8 Calculus in Several Variables) says ff is nearly linear near x0x_0, and the contraction mapping principle (2B.2 Completeness and Contraction) turns "nearly linear and invertible" into "invertible". The implicit function theorem follows, and with it the geometric fact that a level set {F=c}\{F = c\} is a smooth surface wherever ∇F≠0\nabla F \neq 0, and the method of Lagrange multipliers. The pattern of the proof, linearise, invert, contract, is thread L of the guidebook, and it is exactly how the existence of Ricci flow is proved (11A.3 Short-Time Existence and Uniqueness).

By the end of this chapter you will be able to:

  • state the inverse function theorem and prove it from the contraction mapping principle;
  • explain why the theorem is local, with examples of maps that are locally but not globally invertible;
  • state and use the implicit function theorem, including the formula for the derivative of an implicitly defined function;
  • show that a level set is locally a smooth graph where the gradient is non-zero, and find where it fails;
  • derive and use Lagrange multipliers.

When a robot arm gets stuck

In the world In use Singular configurations of a robot arm

A robot arm is controlled through its joint angles, but its job is defined by where its tool is. The map from joint angles to tool position is nonlinear, and to move the tool along a planned path the controller must invert it, continuously, many times a second.

Take the simplest case: a planar arm with two links of lengths ℓ1\ell_1 and ℓ2\ell_2, the first rotating about a fixed shoulder by angle θ1\theta_1, the second about the elbow by angle θ2\theta_2 relative to the first. The tool is at

f(θ1,θ2)=(ℓ1cos⁡θ1+ℓ2cos⁡(θ1+θ2),  ℓ1sin⁡θ1+ℓ2sin⁡(θ1+θ2)).f(\theta_1, \theta_2) = \big(\ell_1\cos\theta_1 + \ell_2\cos(\theta_1 + \theta_2),\ \ \ell_1\sin\theta_1 + \ell_2\sin(\theta_1 + \theta_2)\big).

Its Jacobian determinant is det⁡Df=ℓ1ℓ2sin⁡θ2\det Df = \ell_1\ell_2\sin\theta_2 (Exercise 9.6). It vanishes exactly when θ2=0\theta_2 = 0 (arm fully stretched) or θ2=π\theta_2 = \pi (arm folded back). At those singular configurations the arm can't move its tool in one direction at all: when fully stretched, it can't move the tool straight outwards. Near them, the joint speeds needed to move the tool at a given speed are proportional to 1/∣sin⁡θ2∣1/|\sin\theta_2| and blow up. Away from them, the inverse function theorem guarantees that small tool motions correspond to small, smoothly varying joint motions (Figure 9.1).

Industrial six-axis arms have the same problem in more dimensions. The best-known case is the wrist singularity: when the fifth joint passes through 0°0°, the axes of the fourth and sixth joints line up, the two joints do the same thing, and the arm loses a degree of freedom. Robot programmers plan paths around it, for instance by mounting the tool at a slight angle so that the two axes can't line up, because a controller that tries to move the tool in a straight line through the singularity demands unbounded joint speeds.

Figure 9.1. A two-link arm with ℓ1=1\ell_1 = 1, ℓ2=0.6\ell_2 = 0.6. Its tool can reach the annulus 0.4≤r≤1.60.4 \leq r \leq 1.6. At the singular configurations, fully stretched (outer circle) and fully folded (inner circle), det⁡Df=ℓ1ℓ2sin⁡θ2=0\det Df = \ell_1\ell_2\sin\theta_2 = 0, and the tool can't move radially outward (respectively inward). Everywhere else the joint angles can be recovered smoothly from the tool position.

The inverse function theorem

Theorem 9.1 Inverse function theorem

Let U⊆RnU \subseteq \mathbb{R}^n be open, f:U→Rnf : U \to \mathbb{R}^n continuously differentiable, and x0∈Ux_0 \in U a point where Df(x0)Df(x_0) is invertible. Then there are open sets V∋x0V \ni x_0 and W∋f(x0)W \ni f(x_0) such that ff is a bijection from VV onto WW, its inverse f−1:W→Vf^{-1} : W \to V is continuously differentiable, and

D(f−1)(y)=(Df(f−1(y)))−1.D(f^{-1})(y) = \big(Df(f^{-1}(y))\big)^{-1}.

If ff is CkC^k (or smooth), so is f−1f^{-1}.

The formula for D(f−1)D(f^{-1}) is forced by the chain rule: differentiating f(f−1(y))=yf(f^{-1}(y)) = y gives Df(x) D(f−1)(y)=IDf(x)\,D(f^{-1})(y) = I. The content of the theorem is that the inverse exists and is differentiable.

The proof

Proof. Step 0: reduce to Df(x0)=IDf(x_0) = I. Let A=Df(x0)A = Df(x_0) and replace ff by f~(x)=A−1(f(x0+x)−f(x0))\tilde f(x) = A^{-1}\big(f(x_0 + x) - f(x_0)\big). Then f~(0)=0\tilde f(0) = 0 and Df~(0)=ID\tilde f(0) = I, and ff is locally invertible exactly when f~\tilde f is. So assume x0=0x_0 = 0, f(0)=0f(0) = 0 and Df(0)=IDf(0) = I.

Step 1: a contraction. Let g(x)=x−f(x)g(x) = x - f(x), the nonlinear part of ff. Then Dg(0)=0Dg(0) = 0, and DgDg is continuous, so there is r>0r > 0 such that ∥Dg(x)∥≤12\|Dg(x)\| \leq \tfrac12 on the closed ball Bˉ=Bˉ(0,r)\bar B = \bar B(0, r); shrinking rr, we may also assume Df(x)Df(x) is invertible on Bˉ\bar B (the determinant is continuous and det⁡Df(0)=1\det Df(0) = 1). By the mean value inequality (2B.8 Calculus in Several Variables),

∣g(x)−g(x′)∣≤12∣x−x′∣for x,x′∈Bˉ.(1)|g(x) - g(x')| \leq \tfrac12|x - x'| \qquad \text{for } x, x' \in \bar B. \tag{1}

Now fix yy with ∣y∣<r/2|y| < r/2. Solving f(x)=yf(x) = y is the same as finding a fixed point of

Φy(x)=x−f(x)+y=g(x)+y.\Phi_y(x) = x - f(x) + y = g(x) + y.

For x∈Bˉx \in \bar B, ∣Φy(x)∣≤∣g(x)−g(0)∣+∣y∣<12r+12r=r|\Phi_y(x)| \leq |g(x) - g(0)| + |y| < \tfrac12 r + \tfrac12 r = r, so Φy\Phi_y maps the complete space Bˉ\bar B into itself (indeed into the open ball), and by (1) it is a contraction with constant 12\tfrac12. By the contraction mapping theorem (2B.2 Completeness and Contraction), there is exactly one x∈Bˉx \in \bar B with f(x)=yf(x) = y, and it lies in the open ball B(0,r)B(0, r).

Step 2: the sets VV and WW. Let W=B(0,r/2)W = B(0, r/2) and V=B(0,r)∩f−1(W)V = B(0, r) \cap f^{-1}(W), which is open because ff is continuous. By step 1, ff maps VV bijectively onto WW.

Step 3: the inverse is Lipschitz. From (1), for x,x′∈Bˉx, x' \in \bar B,

∣x−x′∣≤∣f(x)−f(x′)∣+∣g(x)−g(x′)∣≤∣f(x)−f(x′)∣+12∣x−x′∣,|x - x'| \leq |f(x) - f(x')| + |g(x) - g(x')| \leq |f(x) - f(x')| + \tfrac12|x - x'|,

so ∣x−x′∣≤2∣f(x)−f(x′)∣|x - x'| \leq 2|f(x) - f(x')|. With x=f−1(y)x = f^{-1}(y) and x′=f−1(y′)x' = f^{-1}(y'): ∣f−1(y)−f−1(y′)∣≤2∣y−y′∣|f^{-1}(y) - f^{-1}(y')| \leq 2|y - y'|. (This is 2B.2 Completeness and Contraction's continuous dependence of a fixed point on a parameter, here the parameter yy.)

Step 4: the inverse is differentiable. Fix y∈Wy \in W, x=f−1(y)x = f^{-1}(y), L=Df(x)L = Df(x) (invertible), and let y+k∈Wy + k \in W, x+h=f−1(y+k)x + h = f^{-1}(y + k). By step 3, ∣h∣≤2∣k∣|h| \leq 2|k|. Differentiability of ff at xx gives k=f(x+h)−f(x)=Lh+o(∣h∣)k = f(x + h) - f(x) = Lh + o(|h|), so

h−L−1k=−L−1 o(∣h∣)=o(∣k∣),h - L^{-1}k = -L^{-1}\,o(|h|) = o(|k|),

since ∣h∣≤2∣k∣|h| \leq 2|k|. That says f−1f^{-1} is differentiable at yy with derivative L−1L^{-1}. Since y↦L−1=Df(f−1(y))−1y \mapsto L^{-1} = Df(f^{-1}(y))^{-1} is a composition of continuous maps (inversion of matrices is continuous, by Cramer's rule), f−1f^{-1} is C1C^1. If ff is CkC^k, the formula shows inductively that f−1f^{-1} is CkC^k.

Note

The iteration in step 1, xn+1=xn−(f(xn)−y)x_{n+1} = x_n - (f(x_n) - y), which in the original coordinates is xn+1=xn−Df(x0)−1(f(xn)−y)x_{n+1} = x_n - Df(x_0)^{-1}\big(f(x_n) - y\big), is a version of Newton's method that keeps the derivative fixed at x0x_0 instead of updating it at each step (the "chord method"). It converges only linearly, but the proof needs only that it converges.

Local, not global

The theorem is local, and it has to be. The map f(x,y)=(excos⁡y,exsin⁡y)f(x, y) = (e^x\cos y, e^x\sin y) has det⁡Df=e2x≠0\det Df = e^{2x} \neq 0 everywhere, so it is locally invertible at every point. But it is not injective: f(x,y+2π)=f(x,y)f(x, y + 2\pi) = f(x, y). (It is the complex exponential z↦ezz \mapsto e^z written in real coordinates, and its local inverses are branches of the logarithm.) An invertible derivative tells you that nearby points have distinct images, not that distant ones do.

Nor can the hypothesis be dropped. f(x)=x3f(x) = x^3 is a bijection of R\mathbb{R}, but f′(0)=0f'(0) = 0, and the inverse y1/3y^{1/3} is not differentiable at 00. Where the derivative is singular, the inverse, if it exists at all, is not differentiable.

The implicit function theorem

Often the equation to solve has more unknowns than equations. A single equation F(x,y)=0F(x, y) = 0 in two unknowns usually defines a curve, and the question is whether that curve can be written as a graph y=g(x)y = g(x), and whether gg is differentiable.

Write points of Rn+m\mathbb{R}^{n + m} as (x,y)(x, y) with x∈Rnx \in \mathbb{R}^n, y∈Rmy \in \mathbb{R}^m. For F:Rn+m→RmF : \mathbb{R}^{n+m} \to \mathbb{R}^m, the derivative DF(x,y)DF(x, y) is an m×(n+m)m \times (n + m) matrix, which splits into an m×nm \times n block DxFD_xF (derivatives in the xx-variables) and an m×mm \times m block DyFD_yF.

Theorem 9.2 Implicit function theorem

Let FF be continuously differentiable near (a,b)∈Rn×Rm(a, b) \in \mathbb{R}^n \times \mathbb{R}^m, with F(a,b)=0F(a, b) = 0 and DyF(a,b)D_yF(a, b) invertible. Then there are open sets A∋aA \ni a and B∋bB \ni b and a continuously differentiable g:A→Bg : A \to B such that, for (x,y)∈A×B(x, y) \in A \times B,

F(x,y)=0  ⟺  y=g(x).F(x, y) = 0 \iff y = g(x).

Its derivative is Dg(x)=−(DyF(x,g(x)))−1DxF(x,g(x))Dg(x) = -\big(D_yF(x, g(x))\big)^{-1}D_xF(x, g(x)).

Proof. Apply the inverse function theorem to Φ(x,y)=(x,F(x,y))\Phi(x, y) = (x, F(x, y)), a map from Rn+m\mathbb{R}^{n+m} to itself. Its derivative at (a,b)(a, b) is the block matrix (I0DxFDyF)\begin{pmatrix} I & 0\\ D_xF & D_yF\end{pmatrix}, which is invertible because DyFD_yF is. So Φ\Phi has a C1C^1 local inverse near Φ(a,b)=(a,0)\Phi(a, b) = (a, 0), which must have the form Φ−1(x,z)=(x,h(x,z))\Phi^{-1}(x, z) = (x, h(x, z)) because Φ\Phi doesn't change the first coordinate. Then F(x,y)=0F(x, y) = 0 with (x,y)(x, y) near (a,b)(a, b) if and only if Φ(x,y)=(x,0)\Phi(x, y) = (x, 0), if and only if y=h(x,0)y = h(x, 0). Set g(x)=h(x,0)g(x) = h(x, 0). The derivative formula comes from differentiating F(x,g(x))=0F(x, g(x)) = 0 by the chain rule: DxF+DyF Dg=0D_xF + D_yF\,Dg = 0.

Level sets are smooth where the gradient isn't zero

The case m=1m = 1 is the geometric heart of the theorem. Let F:Rn+1→RF : \mathbb{R}^{n+1} \to \mathbb{R} be C1C^1 and pp a point of the level set {F=c}\{F = c\} where ∇F(p)≠0\nabla F(p) \neq 0. Some partial derivative is non-zero at pp; renumber the coordinates so it is the last one. Then the implicit function theorem says that near pp, the level set is the graph of a C1C^1 function of the other nn coordinates: a smooth hypersurface, with ∇F(p)\nabla F(p) as its normal (2B.8 Calculus in Several Variables). A value cc such that ∇F≠0\nabla F \neq 0 at every point of {F=c}\{F = c\} is called a regular value, and then the whole level set is a smooth hypersurface. This is the regular value theorem, the main way manifolds are produced in practice (8A.4 Submanifolds).

Example 9.3 The sphere is smooth

For F(x,y,z)=x2+y2+z2F(x, y, z) = x^2 + y^2 + z^2, ∇F=2(x,y,z)\nabla F = 2(x, y, z), which vanishes only at the origin, not on the sphere {F=1}\{F = 1\}. So 11 is a regular value, and the unit sphere is a smooth surface. Near the north pole it is the graph z=1−x2−y2z = \sqrt{1 - x^2 - y^2}; near a point on the equator such as (1,0,0)(1, 0, 0), ∂zF=0\partial_zF = 0 there, but ∂xF≠0\partial_xF \neq 0, and the sphere is the graph x=1−y2−z2x = \sqrt{1 - y^2 - z^2} instead.

Where the gradient vanishes, the level set can do anything. For the lemniscate of Bernoulli, F(x,y)=(x2+y2)2−2(x2−y2)=0F(x, y) = (x^2 + y^2)^2 - 2(x^2 - y^2) = 0, the gradient vanishes at the origin, and there the curve crosses itself: near the origin it is not the graph of any function of either variable (Figure 9.2). The crossing contour through the mountain pass in 2B.8 Calculus in Several Variables is the same phenomenon: a pass is a critical point of the elevation, and the contour through it is singular.

Figure 9.2. Left: near a point where ∇F≠0\nabla F \neq 0, the level set is a graph over the tangent line, with ∇F\nabla F normal to it. Right: the lemniscate (x2+y2)2=2(x2−y2)(x^2 + y^2)^2 = 2(x^2 - y^2) crosses itself at the origin, where ∇F=0\nabla F = 0. No box around the crossing contains a graph.
In the world In use Comparative statics: who pays a tax?

In economics, an equilibrium is the solution of a system of equations, and "comparative statics" asks how it moves when a parameter changes. That is the implicit function theorem, and its derivative formula is the answer.

Take a single market with demand D(p)D(p) at the price pp buyers pay, and supply S(p−τ)S(p - \tau) when a tax τ\tau per unit separates what buyers pay from what sellers receive. Equilibrium is F(p,τ)=D(p)−S(p−τ)=0F(p, \tau) = D(p) - S(p - \tau) = 0. With D′<0D' < 0 and S′>0S' > 0, ∂pF=D′−S′<0\partial_pF = D' - S' < 0, so near an equilibrium the price is a smooth function p(τ)p(\tau) of the tax, and

dpdτ=−∂τF∂pF=−S′D′−S′=S′S′−D′∈(0,1).\frac{dp}{d\tau} = -\frac{\partial_\tau F}{\partial_p F} = -\frac{S'}{D' - S'} = \frac{S'}{S' - D'} \in (0, 1).

Buyers bear the fraction S′S′−D′\frac{S'}{S' - D'} of a small tax and sellers the rest: whichever side responds less to price bears more of the tax. This is the standard result on tax incidence, and it holds only locally, near an equilibrium where ∂pF≠0\partial_pF \neq 0, exactly as the theorem says.

Lagrange multipliers

To find the maximum of ff on a constraint set {g=0}\{g = 0\}, we can't simply set ∇f=0\nabla f = 0: the maximum on the surface need not be a critical point in the whole space.

Theorem 9.4 Lagrange multipliers

Let f,g:U→Rf, g : U \to \mathbb{R} be C1C^1 on an open U⊆RnU \subseteq \mathbb{R}^n, and let x0x_0 be a local maximum or minimum of ff restricted to M={g=0}M = \{g = 0\}, with ∇g(x0)≠0\nabla g(x_0) \neq 0. Then there is a number λ\lambda (the Lagrange multiplier) with

∇f(x0)=λ ∇g(x0).\nabla f(x_0) = \lambda\,\nabla g(x_0).

Proof. By the implicit function theorem, near x0x_0 the set MM is a C1C^1 hypersurface, and for every vector vv with ∇g(x0)⋅v=0\nabla g(x_0)\cdot v = 0 there is a C1C^1 curve γ\gamma in MM with γ(0)=x0\gamma(0) = x_0 and γ′(0)=v\gamma'(0) = v (take γ\gamma to be the graph of the implicit function over the line through x0x_0 in direction vv, in coordinates where the last partial of gg is non-zero). Since f∘γf\circ\gamma has a local extremum at 00, ∇f(x0)⋅v=(f∘γ)′(0)=0\nabla f(x_0)\cdot v = (f\circ\gamma)'(0) = 0. So ∇f(x0)\nabla f(x_0) is perpendicular to every vector perpendicular to ∇g(x0)\nabla g(x_0), which means it is a multiple of ∇g(x0)\nabla g(x_0).

In words: at a constrained extremum, the level set of ff is tangent to the constraint surface, so their normals are parallel. With several constraints g1=⋯=gk=0g_1 = \cdots = g_k = 0 whose gradients are linearly independent at x0x_0, the same proof gives ∇f=∑iλi∇gi\nabla f = \sum_i\lambda_i\nabla g_i.

Example 9.5 The box of greatest volume

Among closed rectangular boxes with total surface area SS, which has the largest volume? Maximise f=xyzf = xyz subject to g=2(xy+yz+zx)−S=0g = 2(xy + yz + zx) - S = 0, with x,y,z>0x, y, z > 0. The multiplier condition ∇f=λ∇g\nabla f = \lambda\nabla g reads

yz=2λ(y+z),xz=2λ(x+z),xy=2λ(x+y).yz = 2\lambda(y + z), \quad xz = 2\lambda(x + z), \quad xy = 2\lambda(x + y).

Subtracting the first two gives z(y−x)=2λ(y−x)z(y - x) = 2\lambda(y - x), so x=yx = y or z=2λz = 2\lambda; checking cases shows the only solution with positive sides is x=y=zx = y = z. So the best box is a cube, of side S/6\sqrt{S/6}.

The multiplier rule only finds candidates. That a maximum exists at all needs a separate argument, by compactness (2B.3 Compactness): a box with a side close to 00 or a very long side has small volume for its surface area, so the maximum, if any, lies in a compact region of the constraint surface, where the extreme value theorem provides it. Then it must be the cube. This two-step structure, existence by compactness and identification by the first-order condition, is the direct method of the calculus of variations (4A.6 Weak Convergence and the Direct Method), and it is how Perelman's μ\mu-functional is shown to have a minimiser (12A.3 The 𝓦-Entropy).

In the world In use Satellite geometry and GPS accuracy

A GPS receiver finds its position x∈R3x \in \mathbb{R}^3 and its clock error bb from four or more measured pseudoranges ρi=∣si−x∣+cb\rho_i = |s_i - x| + cb, where sis_i is the position of satellite ii and cc the speed of light. This is a nonlinear system. Near the solution it is replaced by its linearisation, whose matrix HH has one row per satellite: the unit vector from the receiver towards the satellite, and a 11 for the clock. Small errors δρ\delta\rho in the ranges produce errors δx\delta x governed by HH, and if the range errors are independent with equal variance σ2\sigma^2, the position-and-clock errors have covariance σ2(HTH)−1\sigma^2(H^{\mathsf T}H)^{-1}. The dilution of precision is the factor by which geometry amplifies range errors: the geometric DOP is tr⁡(HTH)−1\sqrt{\operatorname{tr}(H^{\mathsf T}H)^{-1}}, and the position DOP (PDOP) uses only the three position entries of the trace.

When the satellites are spread across the sky, the rows of HH point in very different directions, HTHH^{\mathsf T}H is well conditioned, and the DOP is small. When they are bunched together, the rows are nearly parallel, HTHH^{\mathsf T}H is nearly singular, and errors along the poorly constrained direction are hugely amplified (Figure 9.3). This is the arm's singularity again: the inverse of a nearly singular linearisation has a large norm. Geometry matters enough that the U.S. government's 2001 GPS Standard Positioning Service Performance Standard stated its availability commitment in these terms: a global PDOP of 66 or less at least 98%98\% of the time.

Figure 9.3. A two-dimensional illustration (three beacons, no clock term). Left: well-spread beacons give DOP ≈1.15\approx 1.15 and a small, round error ellipse. Right: beacons within 30°30° of each other give DOP ≈2.8\approx 2.8, with the error stretched along the direction across the line of sight, which the bunched beacons barely constrain.
Where this goes Linearise, invert, contract

The inverse function theorem is the finite-dimensional model of how nonlinear PDE are solved. The steps are always the same: write the equation as F(u)=0\mathcal{F}(u) = 0, compute the linearisation DFD\mathcal{F} at an approximate solution, show it is invertible between suitable complete function spaces, and run the contraction of step 1.

  • Nonlinear heat equations (6A.7 Nonlinear Parabolic Equations): the linearisation is a linear heat equation, which is invertible on Hölder spaces by Schauder theory (6A.6 Parabolic Regularity), and the contraction gives a solution for a short time.
  • Ricci flow (11A.3 Short-Time Existence and Uniqueness): the linearisation of ∂tg=−2 Ric(g)\partial_tg = -2\,\mathrm{Ric}(g) is not invertible, because Ricci flow is unchanged by diffeomorphisms, so its linearisation has a large kernel. DeTurck's trick adds a term that breaks the symmetry, making the linearisation an invertible heat-type operator, and then the argument of this chapter applies.
  • Regular values and manifolds (8A.4 Submanifolds): the implicit function theorem shows that level sets of maps with surjective derivative are manifolds.
  • Stability of solitons: a soliton is a solution of a nonlinear equation, and if the linearised operator at it has no kernel beyond the symmetries, nearby solutions are controlled by it. Questions of this kind run through the study of singularity models (11B.1 Ricci Solitons).

History

Lagrange introduced his multipliers in mechanics, in the Méchanique analitique of 1788, to handle constrained motion. Cauchy proved an implicit function theorem for analytic functions in the 1830s, using power series. Ulisse Dini proved the real-variable implicit function theorem, with continuously differentiable functions, in his Pisa lectures of 1877–78. The proof through a contraction is from the 20th century, once Banach's principle (2B.2 Completeness and Contraction) was available, and it is the version that extends to infinite-dimensional spaces, where most of its modern uses lie.

Recall Where we stand

If ff is C1C^1 and Df(x0)Df(x_0) is invertible, ff has a C1C^1 inverse near x0x_0, proved by rewriting f(x)=yf(x) = y as a fixed-point problem for a contraction. The theorem is local, and fails without an invertible derivative. The implicit function theorem solves F(x,y)=0F(x, y) = 0 for yy when DyFD_yF is invertible, with Dg=−DyF−1DxFDg = -D_yF^{-1}D_xF; geometrically, level sets are smooth hypersurfaces where the gradient doesn't vanish. At a constrained extremum, ∇f\nabla f is a multiple of ∇g\nabla g. In applications, singular or nearly singular linearisations show up as robot singularities and as poor GPS geometry. 2B.10 Ordinary Differential Equations uses the contraction principle once more, in a space of functions, to solve ordinary differential equations.

Exercises

Exercise 9.6 The two-link arm

For the two-link arm, (a) compute Df(θ1,θ2)Df(\theta_1, \theta_2) and show det⁡Df=ℓ1ℓ2sin⁡θ2\det Df = \ell_1\ell_2\sin\theta_2; (b) show that the reachable set is the annulus ∣ℓ1−ℓ2∣≤r≤ℓ1+ℓ2|\ell_1 - \ell_2| \leq r \leq \ell_1 + \ell_2; (c) show that each point strictly inside the annulus is reached by exactly two configurations (elbow up and elbow down), and each point on its boundary by one. How do the two branches of the inverse meet?

Solution

(a) Df=(−ℓ1s1−ℓ2s12−ℓ2s12ℓ1c1+ℓ2c12ℓ2c12)Df = \begin{pmatrix}-\ell_1s_1 - \ell_2s_{12} & -\ell_2s_{12}\\ \ell_1c_1 + \ell_2c_{12} & \ell_2c_{12}\end{pmatrix} with s1=sin⁡θ1s_1 = \sin\theta_1, s12=sin⁡(θ1+θ2)s_{12} = \sin(\theta_1 + \theta_2), etc. Its determinant is ℓ1ℓ2(c1s12−s1c12)=ℓ1ℓ2sin⁡θ2\ell_1\ell_2(c_1s_{12} - s_1c_{12}) = \ell_1\ell_2\sin\theta_2. (b) By the law of cosines, r2=ℓ12+ℓ22+2ℓ1ℓ2cos⁡θ2r^2 = \ell_1^2 + \ell_2^2 + 2\ell_1\ell_2\cos\theta_2, which ranges over [(ℓ1−ℓ2)2,(ℓ1+ℓ2)2][(\ell_1 - \ell_2)^2, (\ell_1 + \ell_2)^2]. (c) For rr strictly inside, cos⁡θ2\cos\theta_2 is determined, giving two values ±θ2\pm\theta_2, and then θ1\theta_1 is determined by the direction. On the boundary θ2∈{0,π}\theta_2 \in \{0, \pi\} is unique. The two branches meet at the singular configurations, where the Jacobian is singular, as the inverse function theorem requires.

Exercise 9.7 Polar coordinates

Let f(r,θ)=(rcos⁡θ,rsin⁡θ)f(r, \theta) = (r\cos\theta, r\sin\theta) on {r>0}\{r > 0\}. Show that DfDf is invertible everywhere, compute D(f−1)D(f^{-1}) at (x,y)(x, y) using the formula of Theorem 9.1, and check it against r=x2+y2r = \sqrt{x^2 + y^2}, θ=arctan⁡(y/x)\theta = \arctan(y/x) (for x>0x > 0). Is ff globally injective?

Exercise 9.8 Locally but not globally invertible

(a) Verify that f(x,y)=(excos⁡y,exsin⁡y)f(x, y) = (e^x\cos y, e^x\sin y) has det⁡Df=e2x\det Df = e^{2x} and is not injective. (b) Find a C1C^1 map R→R\mathbb{R} \to \mathbb{R} with f′(x)≠0f'(x) \neq 0 everywhere that is not surjective. Can a C1C^1 map R→R\mathbb{R} \to \mathbb{R} with f′≠0f' \neq 0 everywhere fail to be injective?

Solution

(b) arctan⁡\arctan, or exe^x. No: by the intermediate value theorem for derivatives (Darboux) or, for C1C^1 maps, the intermediate value theorem for f′f', f′f' has constant sign, so ff is strictly monotone. In one dimension, local invertibility everywhere implies global injectivity; in two dimensions it doesn't.

Exercise 9.9 An implicit curve

The equation x3+y3−3xy=0x^3 + y^3 - 3xy = 0 defines the folium of Descartes. (a) Find the points where the implicit function theorem fails to give yy as a function of xx, and the points where it fails to give xx as a function of yy. (b) Near the point (32,32)(\tfrac32, \tfrac32), compute dy/dxdy/dx.

Solution

(a) ∂yF=3y2−3x=0\partial_yF = 3y^2 - 3x = 0 with F=0F = 0: x=y2x = y^2 and y6+y3−3y3=0y^6 + y^3 - 3y^3 = 0, so y=0y = 0 (the origin) or y3=2y^3 = 2, giving (22/3,21/3)(2^{2/3}, 2^{1/3}). Symmetrically ∂xF=0\partial_xF = 0 at the origin and (21/3,22/3)(2^{1/3}, 2^{2/3}). At the origin the curve crosses itself. (b) dy/dx=−3x2−3y3y2−3x=−9/4−3/29/4−3/2=−1dy/dx = -\frac{3x^2 - 3y}{3y^2 - 3x} = -\frac{9/4 - 3/2}{9/4 - 3/2} = -1.

Exercise 9.10 Lagrange multipliers and eigenvalues

Let AA be a symmetric n×nn \times n matrix. Use Lagrange multipliers to show that the maximum and minimum of f(x)=xTAxf(x) = x^{\mathsf T}Ax on the unit sphere {∣x∣2=1}\{|x|^2 = 1\} are attained at eigenvectors of AA, with values the largest and smallest eigenvalues. (The maximum exists because the sphere is compact.) This gives a proof of the existence of a real eigenvalue of a symmetric matrix, the first step of the spectral theorem (1A.6 Symmetric Matrices and the Spectral Theorem).

Solution

∇f=2Ax\nabla f = 2Ax and ∇g=2x\nabla g = 2x for g=∣x∣2−1g = |x|^2 - 1, so at an extremum Ax=λxAx = \lambda x: an eigenvector, with f(x)=xTλx=λf(x) = x^{\mathsf T}\lambda x = \lambda. The maximum of ff is therefore the largest eigenvalue, and the minimum the smallest.

Exercise 9.11 Dilution of precision in the plane

Three beacons at unit distance in directions u1,u2,u3u_1, u_2, u_3 (unit vectors) give the matrix HH with rows uiTu_i^{\mathsf T}. (a) For directions 90°,210°,330°90°, 210°, 330°, show HTH=32IH^{\mathsf T}H = \tfrac32I and DOP =tr⁡(HTH)−1=4/3≈1.15= \sqrt{\operatorname{tr}(H^{\mathsf T}H)^{-1}} = \sqrt{4/3} \approx 1.15. (b) For directions 75°,90°,105°75°, 90°, 105°, compute HTHH^{\mathsf T}H and the DOP. Which direction is poorly determined, and why?

Solution

(a) ∑iuiuiT=32I\sum_iu_iu_i^{\mathsf T} = \tfrac32I for three unit vectors 120°120° apart. (b) HTH=diag(2cos⁡275°, 1+2sin⁡275°)≈diag(0.134,2.866)H^{\mathsf T}H = \mathrm{diag}\big(2\cos^2 75°,\ 1 + 2\sin^2 75°\big) \approx \mathrm{diag}(0.134, 2.866), so the DOP is 1/0.134+1/2.866≈2.80\sqrt{1/0.134 + 1/2.866} \approx 2.80. The horizontal direction, across the line of sight, is poorly determined: all three beacons are nearly overhead, so moving sideways barely changes any range.

Exercise 9.12 Rehearsal: non-degenerate critical points persist

Let ff be C2C^2 on Rn\mathbb{R}^n with a critical point x0x_0 whose Hessian ∇2f(x0)\nabla^2f(x_0) is invertible (a non-degenerate critical point), and let hh be C2C^2. Apply the implicit function theorem to F(x,ε)=∇f(x)+ε∇h(x)F(x, \varepsilon) = \nabla f(x) + \varepsilon\nabla h(x) to show that for small ∣ε∣|\varepsilon|, f+εhf + \varepsilon h has a unique critical point x(ε)x(\varepsilon) near x0x_0, depending smoothly on ε\varepsilon, with x′(0)=−∇2f(x0)−1∇h(x0)x'(0) = -\nabla^2f(x_0)^{-1}\nabla h(x_0). Show that it has the same index (number of negative Hessian eigenvalues) as x0x_0. This is the simplest perturbation argument: a non-degenerate solution survives small changes of the equation. It is the reason Morse functions are stable (7A.10 Morse Theory), and its infinite-dimensional versions govern which singularity models of Ricci flow are stable under perturbation (11B.1 Ricci Solitons, 11B.4 Singularities).

Solution

DxF(x0,0)=∇2f(x0)D_xF(x_0, 0) = \nabla^2f(x_0) is invertible, so near (x0,0)(x_0, 0) the zeros of FF are a smooth curve x(ε)x(\varepsilon), with x′(0)=−DxF−1∂εF=−∇2f(x0)−1∇h(x0)x'(0) = -D_xF^{-1}\partial_\varepsilon F = -\nabla^2f(x_0)^{-1}\nabla h(x_0). The Hessian ∇2f(x(ε))+ε∇2h(x(ε))\nabla^2f(x(\varepsilon)) + \varepsilon\nabla^2h(x(\varepsilon)) depends continuously on ε\varepsilon and is invertible at ε=0\varepsilon = 0; its eigenvalues depend continuously on it, and none can cross 00 while it stays invertible (by continuity of the determinant, for small ε\varepsilon), so the number of negative ones is constant.

© 2026 NeckPinch (www.neckpinch.com). All content in the guidebook (text, mathematics, figures and exercises) is protected by copyright. All rights reserved. No part may be copied, republished or redistributed without written permission.