© 2026 NeckPinch · www.neckpinch.com · All rights reserved.
Course 2Book 2B: Spaces, Functions and ChangeChapter 9
The Inverse and Implicit Function Theorems
When a nonlinear equation can be solved: linearise, invert, and use a contraction.
Read with Tao, Analysis II, chapter "Several variable differential calculus", sections "The inverse function theorem in several variable calculus" and "The implicit function theorem". Munkres, Analysis on Manifolds, chapter "Inverse Functions and the Implicit Function Theorem", has a good geometric account of the implicit function theorem and the rank of a map.
When can a nonlinear equation be solved? For a linear system the answer is clean: if is invertible, there is exactly one solution, , and it depends continuously on . For a nonlinear system there is no general answer. But there is a local one, and it is the most important theorem of this book: if the linearisation is invertible, the nonlinear problem is solvable nearby. Precisely, if and the derivative is an invertible matrix, then for every close to the equation has exactly one solution close to , and that solution depends smoothly on . This is the inverse function theorem.
Its proof combines the two engines built so far: the derivative (2B.8 Calculus in Several Variables) says is nearly linear near , and the contraction mapping principle (2B.2 Completeness and Contraction) turns "nearly linear and invertible" into "invertible". The implicit function theorem follows, and with it the geometric fact that a level set is a smooth surface wherever , and the method of Lagrange multipliers. The pattern of the proof, linearise, invert, contract, is thread L of the guidebook, and it is exactly how the existence of Ricci flow is proved (11A.3 Short-Time Existence and Uniqueness).
By the end of this chapter you will be able to:
- state the inverse function theorem and prove it from the contraction mapping principle;
- explain why the theorem is local, with examples of maps that are locally but not globally invertible;
- state and use the implicit function theorem, including the formula for the derivative of an implicitly defined function;
- show that a level set is locally a smooth graph where the gradient is non-zero, and find where it fails;
- derive and use Lagrange multipliers.
When a robot arm gets stuck
A robot arm is controlled through its joint angles, but its job is defined by where its tool is. The map from joint angles to tool position is nonlinear, and to move the tool along a planned path the controller must invert it, continuously, many times a second.
Take the simplest case: a planar arm with two links of lengths and , the first rotating about a fixed shoulder by angle , the second about the elbow by angle relative to the first. The tool is at
Its Jacobian determinant is (Exercise 9.6). It vanishes exactly when (arm fully stretched) or (arm folded back). At those singular configurations the arm can't move its tool in one direction at all: when fully stretched, it can't move the tool straight outwards. Near them, the joint speeds needed to move the tool at a given speed are proportional to and blow up. Away from them, the inverse function theorem guarantees that small tool motions correspond to small, smoothly varying joint motions (Figure 9.1).
Industrial six-axis arms have the same problem in more dimensions. The best-known case is the wrist singularity: when the fifth joint passes through , the axes of the fourth and sixth joints line up, the two joints do the same thing, and the arm loses a degree of freedom. Robot programmers plan paths around it, for instance by mounting the tool at a slight angle so that the two axes can't line up, because a controller that tries to move the tool in a straight line through the singularity demands unbounded joint speeds.
The inverse function theorem
Let be open, continuously differentiable, and a point where is invertible. Then there are open sets and such that is a bijection from onto , its inverse is continuously differentiable, and
If is (or smooth), so is .
The formula for is forced by the chain rule: differentiating gives . The content of the theorem is that the inverse exists and is differentiable.
The proof
Proof. Step 0: reduce to . Let and replace by . Then and , and is locally invertible exactly when is. So assume , and .
Step 1: a contraction. Let , the nonlinear part of . Then , and is continuous, so there is such that on the closed ball ; shrinking , we may also assume is invertible on (the determinant is continuous and ). By the mean value inequality (2B.8 Calculus in Several Variables),
Now fix with . Solving is the same as finding a fixed point of
For , , so maps the complete space into itself (indeed into the open ball), and by (1) it is a contraction with constant . By the contraction mapping theorem (2B.2 Completeness and Contraction), there is exactly one with , and it lies in the open ball .
Step 2: the sets and . Let and , which is open because is continuous. By step 1, maps bijectively onto .
Step 3: the inverse is Lipschitz. From (1), for ,
so . With and : . (This is 2B.2 Completeness and Contraction's continuous dependence of a fixed point on a parameter, here the parameter .)
Step 4: the inverse is differentiable. Fix , , (invertible), and let , . By step 3, . Differentiability of at gives , so
since . That says is differentiable at with derivative . Since is a composition of continuous maps (inversion of matrices is continuous, by Cramer's rule), is . If is , the formula shows inductively that is .
The iteration in step 1, , which in the original coordinates is , is a version of Newton's method that keeps the derivative fixed at instead of updating it at each step (the "chord method"). It converges only linearly, but the proof needs only that it converges.
Local, not global
The theorem is local, and it has to be. The map has everywhere, so it is locally invertible at every point. But it is not injective: . (It is the complex exponential written in real coordinates, and its local inverses are branches of the logarithm.) An invertible derivative tells you that nearby points have distinct images, not that distant ones do.
Nor can the hypothesis be dropped. is a bijection of , but , and the inverse is not differentiable at . Where the derivative is singular, the inverse, if it exists at all, is not differentiable.
The implicit function theorem
Often the equation to solve has more unknowns than equations. A single equation in two unknowns usually defines a curve, and the question is whether that curve can be written as a graph , and whether is differentiable.
Write points of as with , . For , the derivative is an matrix, which splits into an block (derivatives in the -variables) and an block .
Let be continuously differentiable near , with and invertible. Then there are open sets and and a continuously differentiable such that, for ,
Its derivative is .
Proof. Apply the inverse function theorem to , a map from to itself. Its derivative at is the block matrix , which is invertible because is. So has a local inverse near , which must have the form because doesn't change the first coordinate. Then with near if and only if , if and only if . Set . The derivative formula comes from differentiating by the chain rule: .
Level sets are smooth where the gradient isn't zero
The case is the geometric heart of the theorem. Let be and a point of the level set where . Some partial derivative is non-zero at ; renumber the coordinates so it is the last one. Then the implicit function theorem says that near , the level set is the graph of a function of the other coordinates: a smooth hypersurface, with as its normal (2B.8 Calculus in Several Variables). A value such that at every point of is called a regular value, and then the whole level set is a smooth hypersurface. This is the regular value theorem, the main way manifolds are produced in practice (8A.4 Submanifolds).
For , , which vanishes only at the origin, not on the sphere . So is a regular value, and the unit sphere is a smooth surface. Near the north pole it is the graph ; near a point on the equator such as , there, but , and the sphere is the graph instead.
Where the gradient vanishes, the level set can do anything. For the lemniscate of Bernoulli, , the gradient vanishes at the origin, and there the curve crosses itself: near the origin it is not the graph of any function of either variable (Figure 9.2). The crossing contour through the mountain pass in 2B.8 Calculus in Several Variables is the same phenomenon: a pass is a critical point of the elevation, and the contour through it is singular.
In economics, an equilibrium is the solution of a system of equations, and "comparative statics" asks how it moves when a parameter changes. That is the implicit function theorem, and its derivative formula is the answer.
Take a single market with demand at the price buyers pay, and supply when a tax per unit separates what buyers pay from what sellers receive. Equilibrium is . With and , , so near an equilibrium the price is a smooth function of the tax, and
Buyers bear the fraction of a small tax and sellers the rest: whichever side responds less to price bears more of the tax. This is the standard result on tax incidence, and it holds only locally, near an equilibrium where , exactly as the theorem says.
Lagrange multipliers
To find the maximum of on a constraint set , we can't simply set : the maximum on the surface need not be a critical point in the whole space.
Let be on an open , and let be a local maximum or minimum of restricted to , with . Then there is a number (the Lagrange multiplier) with
Proof. By the implicit function theorem, near the set is a hypersurface, and for every vector with there is a curve in with and (take to be the graph of the implicit function over the line through in direction , in coordinates where the last partial of is non-zero). Since has a local extremum at , . So is perpendicular to every vector perpendicular to , which means it is a multiple of .
In words: at a constrained extremum, the level set of is tangent to the constraint surface, so their normals are parallel. With several constraints whose gradients are linearly independent at , the same proof gives .
Among closed rectangular boxes with total surface area , which has the largest volume? Maximise subject to , with . The multiplier condition reads
Subtracting the first two gives , so or ; checking cases shows the only solution with positive sides is . So the best box is a cube, of side .
The multiplier rule only finds candidates. That a maximum exists at all needs a separate argument, by compactness (2B.3 Compactness): a box with a side close to or a very long side has small volume for its surface area, so the maximum, if any, lies in a compact region of the constraint surface, where the extreme value theorem provides it. Then it must be the cube. This two-step structure, existence by compactness and identification by the first-order condition, is the direct method of the calculus of variations (4A.6 Weak Convergence and the Direct Method), and it is how Perelman's -functional is shown to have a minimiser (12A.3 The 𝓦-Entropy).
A GPS receiver finds its position and its clock error from four or more measured pseudoranges , where is the position of satellite and the speed of light. This is a nonlinear system. Near the solution it is replaced by its linearisation, whose matrix has one row per satellite: the unit vector from the receiver towards the satellite, and a for the clock. Small errors in the ranges produce errors governed by , and if the range errors are independent with equal variance , the position-and-clock errors have covariance . The dilution of precision is the factor by which geometry amplifies range errors: the geometric DOP is , and the position DOP (PDOP) uses only the three position entries of the trace.
When the satellites are spread across the sky, the rows of point in very different directions, is well conditioned, and the DOP is small. When they are bunched together, the rows are nearly parallel, is nearly singular, and errors along the poorly constrained direction are hugely amplified (Figure 9.3). This is the arm's singularity again: the inverse of a nearly singular linearisation has a large norm. Geometry matters enough that the U.S. government's 2001 GPS Standard Positioning Service Performance Standard stated its availability commitment in these terms: a global PDOP of or less at least of the time.
The inverse function theorem is the finite-dimensional model of how nonlinear PDE are solved. The steps are always the same: write the equation as , compute the linearisation at an approximate solution, show it is invertible between suitable complete function spaces, and run the contraction of step 1.
- Nonlinear heat equations (6A.7 Nonlinear Parabolic Equations): the linearisation is a linear heat equation, which is invertible on Hölder spaces by Schauder theory (6A.6 Parabolic Regularity), and the contraction gives a solution for a short time.
- Ricci flow (11A.3 Short-Time Existence and Uniqueness): the linearisation of is not invertible, because Ricci flow is unchanged by diffeomorphisms, so its linearisation has a large kernel. DeTurck's trick adds a term that breaks the symmetry, making the linearisation an invertible heat-type operator, and then the argument of this chapter applies.
- Regular values and manifolds (8A.4 Submanifolds): the implicit function theorem shows that level sets of maps with surjective derivative are manifolds.
- Stability of solitons: a soliton is a solution of a nonlinear equation, and if the linearised operator at it has no kernel beyond the symmetries, nearby solutions are controlled by it. Questions of this kind run through the study of singularity models (11B.1 Ricci Solitons).
History
Lagrange introduced his multipliers in mechanics, in the Méchanique analitique of 1788, to handle constrained motion. Cauchy proved an implicit function theorem for analytic functions in the 1830s, using power series. Ulisse Dini proved the real-variable implicit function theorem, with continuously differentiable functions, in his Pisa lectures of 1877–78. The proof through a contraction is from the 20th century, once Banach's principle (2B.2 Completeness and Contraction) was available, and it is the version that extends to infinite-dimensional spaces, where most of its modern uses lie.
If is and is invertible, has a inverse near , proved by rewriting as a fixed-point problem for a contraction. The theorem is local, and fails without an invertible derivative. The implicit function theorem solves for when is invertible, with ; geometrically, level sets are smooth hypersurfaces where the gradient doesn't vanish. At a constrained extremum, is a multiple of . In applications, singular or nearly singular linearisations show up as robot singularities and as poor GPS geometry. 2B.10 Ordinary Differential Equations uses the contraction principle once more, in a space of functions, to solve ordinary differential equations.
Exercises
For the two-link arm, (a) compute and show ; (b) show that the reachable set is the annulus ; (c) show that each point strictly inside the annulus is reached by exactly two configurations (elbow up and elbow down), and each point on its boundary by one. How do the two branches of the inverse meet?
Solution
(a) with , , etc. Its determinant is . (b) By the law of cosines, , which ranges over . (c) For strictly inside, is determined, giving two values , and then is determined by the direction. On the boundary is unique. The two branches meet at the singular configurations, where the Jacobian is singular, as the inverse function theorem requires.
Let on . Show that is invertible everywhere, compute at using the formula of Theorem 9.1, and check it against , (for ). Is globally injective?
(a) Verify that has and is not injective. (b) Find a map with everywhere that is not surjective. Can a map with everywhere fail to be injective?
Solution
(b) , or . No: by the intermediate value theorem for derivatives (Darboux) or, for maps, the intermediate value theorem for , has constant sign, so is strictly monotone. In one dimension, local invertibility everywhere implies global injectivity; in two dimensions it doesn't.
The equation defines the folium of Descartes. (a) Find the points where the implicit function theorem fails to give as a function of , and the points where it fails to give as a function of . (b) Near the point , compute .
Solution
(a) with : and , so (the origin) or , giving . Symmetrically at the origin and . At the origin the curve crosses itself. (b) .
Let be a symmetric matrix. Use Lagrange multipliers to show that the maximum and minimum of on the unit sphere are attained at eigenvectors of , with values the largest and smallest eigenvalues. (The maximum exists because the sphere is compact.) This gives a proof of the existence of a real eigenvalue of a symmetric matrix, the first step of the spectral theorem (1A.6 Symmetric Matrices and the Spectral Theorem).
Solution
and for , so at an extremum : an eigenvector, with . The maximum of is therefore the largest eigenvalue, and the minimum the smallest.
Three beacons at unit distance in directions (unit vectors) give the matrix with rows . (a) For directions , show and DOP . (b) For directions , compute and the DOP. Which direction is poorly determined, and why?
Solution
(a) for three unit vectors apart. (b) , so the DOP is . The horizontal direction, across the line of sight, is poorly determined: all three beacons are nearly overhead, so moving sideways barely changes any range.
Let be on with a critical point whose Hessian is invertible (a non-degenerate critical point), and let be . Apply the implicit function theorem to to show that for small , has a unique critical point near , depending smoothly on , with . Show that it has the same index (number of negative Hessian eigenvalues) as . This is the simplest perturbation argument: a non-degenerate solution survives small changes of the equation. It is the reason Morse functions are stable (7A.10 Morse Theory), and its infinite-dimensional versions govern which singularity models of Ricci flow are stable under perturbation (11B.1 Ricci Solitons, 11B.4 Singularities).
Solution
is invertible, so near the zeros of are a smooth curve , with . The Hessian depends continuously on and is invertible at ; its eigenvalues depend continuously on it, and none can cross while it stays invertible (by continuity of the determinant, for small ), so the number of negative ones is constant.
© 2026 NeckPinch (www.neckpinch.com). All content in the guidebook (text, mathematics, figures and exercises) is protected by copyright. All rights reserved. No part may be copied, republished or redistributed without written permission.