© 2026 NeckPinch · www.neckpinch.com · All rights reserved.
Course 4Book 4A: Function Spaces and Sobolev SpacesChapter 3
Hahn–Banach and Duality
Separating convex sets, dual spaces, and the dual of Lᵖ.
Read with Brezis, chapter 1 (the Hahn–Banach theorems, analytic and geometric forms; the section on conjugate convex functions can wait), and chapter 4 for the dual of ; or Kreyszig §4.1–4.6 (Zorn's lemma, Hahn–Banach, the dual of , adjoint operators, reflexive spaces).
A bounded linear functional is a measurement: a continuous linear way of turning a vector into a number. Evaluating a function at a point, integrating it against a weight, taking a Fourier coefficient: these are all linear functionals. The dual space is the space of all of them, and much of functional analysis consists of studying a space through its dual. Two questions come first. Are there enough functionals to tell vectors apart? And what does the dual of a familiar space look like?
The first question is answered by the Hahn–Banach theorem, in two forms. The analytic form extends functionals from a subspace without increasing their norm. The geometric form, easier to picture and just as important, says that two disjoint convex sets can be separated by a hyperplane. The second question has satisfying answers: the dual of is for conjugate exponents (with one exception), and the dual of is the space of measures. That last fact is why limits of probability densities, the objects of 3A.4 Measures, Probability and Weights, can be measures, and it reappears at the frontier of Ricci flow research.
By the end of this chapter you will be able to:
- state and use the analytic Hahn–Banach theorem, and find norming functionals;
- prove that a point outside a closed convex set in can be separated from it by a hyperplane, and state the general version;
- identify the duals of , and , and explain which spaces are reflexive;
- use separating hyperplanes to price assets by no arbitrage, and to classify data.
Prices from separation
An arbitrage is a trading strategy that costs nothing (or less than nothing) today and can never lose money, with a chance of making some. In a well-functioning market arbitrage opportunities are competed away, and the fundamental theorem of asset pricing (J. Michael Harrison and David Kreps, 1979; Harrison and Stanley Pliska, 1981) says what that implies: in a market with no arbitrage there is a positive linear pricing functional, equivalently a set of "risk-neutral" probabilities, under which every asset's price is its expected discounted payoff. In a finite model the proof is a separating hyperplane.
The simplest case. A stock costs today; tomorrow it will be worth either or . Money can be held at zero interest. Look for state prices that reproduce today's prices: (the price of for sure) and . The unique solution is , . Then a call option paying , which is in the up state and in the down state, must cost . Any other price creates an arbitrage, by combining the option with the stock and cash (Exercise 3.7).
Why separation? Consider the set of payoff vectors achievable at zero cost, a subspace, and the cone of non-negative, non-zero payoffs. "No arbitrage" says they don't meet. The separating hyperplane between them is a linear functional that vanishes on the costless payoffs and is positive on the cone, and that functional is the pricing rule. With more possible states than independent assets, the separating functional need not be unique, and the market is called incomplete.
The analytic Hahn–Banach theorem
A function on a real vector space is sublinear if and for . A norm is sublinear.
Let be a real vector space, sublinear on , a subspace, and linear with on . Then extends to a linear with on all of .
Proof architecture. One step: to extend from to , one must choose a value with for all and real . Dividing by , this says must lie between and , and sublinearity is exactly what makes the supremum at most the infimum (Exercise 3.8). All steps: extend one dimension at a time. In a space with a countable dense set (a separable space, such as for ) ordinary induction and continuity suffice; in general, Zorn's lemma (equivalent to the axiom of choice, 2A.8 Infinite Sets) provides a maximal extension, which must be defined everywhere.
The case used most is .
Let be a normed space.
- Every bounded linear functional on a subspace extends to with the same norm.
- For every there is with and .
- Consequently , and separates points: if for all , then .
Proof. (1) Apply the theorem with (for complex spaces, apply it to the real part and recover the imaginary part). (2) On the line define , of norm , and extend. (3) when , with equality for the functional of (2).
In concrete spaces the norming functional of (2) can usually be written down. In with , it is , the case of equality in Hölder's inequality (Exercise 3.9). The theorem is needed when there is no formula, and to know that enough functionals exist in general.
Separating convex sets
Let be non-empty, closed and convex, and . Then there is a vector and a number with
Proof. Let be the point of nearest to , which exists because closed bounded subsets of are compact and is unique by convexity (2B.10 Ordinary Differential Equations's projection exercise). Set . For , the projection property gives , so . And . Take .
The hyperplane touches at and separates it from . If is a boundary point of instead, a limiting argument gives a supporting hyperplane: one that passes through with on one side (Figure 3.1).
In a general normed space the nearest point need not exist, but the conclusion survives, with a continuous linear functional in place of .
Let and be disjoint non-empty convex subsets of a real normed space .
- If is open, there are and with for all , .
- If is closed and is compact, there are and with : strict separation.
Proof architecture (Brezis, Theorems 1.6 and 1.7). For (1), reduce to separating a point from an open convex set containing , and apply the analytic form with the Minkowski gauge , which is sublinear and comparable to the norm. For (2), thicken slightly into an open convex set that still misses , and apply (1).
Given data points of two kinds (say, measurements from healthy and faulty machines), a linear classifier looks for a hyperplane with one kind on each side. If the two convex hulls of the data are disjoint, the geometric Hahn–Banach theorem (in , Theorem 3.3) says such a hyperplane exists, and the support vector machine of Corinna Cortes and Vladimir Vapnik (1995) chooses the one with the largest margin, the widest empty strip between the classes (Figure 3.2). That choice is itself a convex optimisation problem, solved through its dual. When the data are not separable, the method penalises violations, or maps the data into a higher-dimensional space where they become separable.
Linear programming is a third use. Minimising a linear cost over a convex polyhedron of feasible plans has a dual problem whose variables are prices, the "shadow prices" of the constraints: the marginal value of relaxing each constraint by one unit. The duality theorem of linear programming, that the optimal costs of the two problems agree, is a separation theorem for polyhedra (Farkas's lemma, Exercise 3.10).
Dual spaces
The dual of a normed space is (or ), with the operator norm. It is always a Banach space (4A.1 Banach Spaces and Bounded Operators), even if is not complete. For the spaces of Book 3A the duals are concrete.
Let and the conjugate exponent, and let be σ-finite. Every defines a bounded functional on with , and every bounded linear functional on is of this form for a unique . So .
Proof architecture. Hölder gives , and the norming function of Exercise 3.9 gives equality. For the converse, given , define a set function on sets of finite measure. It is countably additive (continuity of plus dominated convergence) and absolutely continuous with respect to , so the Radon–Nikodym theorem (3A.4 Measures, Probability and Weights) gives a density with ; one then checks and extends from indicators to all of by density. Brezis gives a different proof for through uniform convexity.
The case is the exception. Every gives a functional on , but not every functional on comes from an function: Hahn–Banach produces, for example, a "limit at " functional on that extends from the functions that have such a limit, and no integrable can represent it (Exercise 3.12). For sequences, for , and the dual of (sequences tending to , with the sup norm) is .
Let be a compact metric space. Every bounded linear functional on has the form for a unique finite signed Borel measure , and , the total variation of . Positive functionals ( when ) correspond to positive measures.
(Stated, not proved here; Kreyszig §4.4 treats , which was Riesz's original 1909 theorem, with Stieltjes integrals.) The Dirac mass , evaluation at a point, is the functional , of norm : it lives naturally in even though it has no density (3A.4 Measures, Probability and Weights).
Reflexivity
Each defines a functional on , namely , and by Corollary 3.2 this embeds isometrically into its bidual . If the embedding is onto, is reflexive. Hilbert spaces are reflexive (4A.4 Hilbert Spaces and Lax–Milgram); so is for , since . But , and (for infinite ) are not.
Why care? Reflexivity is exactly the condition under which bounded sequences have weakly convergent subsequences (4A.6 Weak Convergence and the Direct Method), the replacement for the compactness that infinite dimensions take away. That is why existence theory in PDE prefers and Sobolev spaces with , and handles and separately and with care.
Because is a space of measures, a sequence of probability densities on a compact space always has a subsequence converging, against continuous test functions, to a probability measure (4A.6 Weak Convergence and the Direct Method, Banach–Alaoglu) — but the limit need not have a density. The Dirac mass as a limit of narrowing Gaussians (3A.4 Measures, Probability and Weights) is the simplest example. In Ricci flow the conjugate heat kernel is a probability density that can concentrate, and in recent work on the structure of limits of Ricci flows (Bamler's theory of "metric flows", 12C.5 After Perelman) the natural objects are families of probability measures obtained in exactly this way.
History
Eduard Helly proved an extension theorem for functionals on in 1912. Hans Hahn (1927) and Stefan Banach (1929) independently proved the general theorem now named after them. Hermann Minkowski developed separation and supporting hyperplanes for convex bodies around 1900, and Gyula Farkas published his lemma on linear inequalities in 1902. Frigyes Riesz identified the dual of in 1909 and the duals of in 1910; the representation for general compact spaces is due to Andrey Markov (1938) and Shizuo Kakutani (1941). Linear programming duality grew out of work by Leonid Kantorovich (1939), John von Neumann and George Dantzig (1947). Harrison, Kreps and Pliska's papers on arbitrage and martingales appeared in 1979 and 1981.
The analytic Hahn–Banach theorem extends functionals dominated by a sublinear function; it gives norming functionals and shows the dual separates points. The geometric form separates disjoint convex sets by hyperplanes, with the nearest-point construction in as the model; it underlies no-arbitrage pricing, linear classifiers and linear programming duality. The dual of is for , the dual of is the measures, and is reflexive for . 4A.4 Hilbert Spaces and Lax–Milgram specialises to Hilbert spaces, where the dual is the space itself and projections exist.
Exercises
In the one-period model of the anchor, show that if the option traded at instead of , one could sell the option, buy of a share and borrow the difference, and make a riskless profit. Find a replicating portfolio: holdings of stock and cash whose value tomorrow equals the option's payoff in both states, and check its cost is .
Solution
Replicate with shares and in cash: , , so , , at cost . If the option sells for , sell it and buy the replicating portfolio: profit now, and tomorrow the portfolio exactly pays what the option owes.
Show that for , , using . Deduce that a value can be chosen, and check that the extension satisfies on .
For , , , let . Show and . What is the norming functional for ? For ?
Let be an matrix and . Show that exactly one of the following holds: (i) for some (componentwise); (ii) there is with and . (Apply Theorem 3.3 to and the closed convex cone ; closedness of a finitely generated cone may be assumed.) Explain how the no-arbitrage theorem in a finite model is an instance.
Show and . Then show that is not reflexive, assuming the fact that a normed space whose dual is separable is itself separable. ( is not separable: the indicator sequences of subsets of are uncountably many and pairwise at distance . If were reflexive, would be isometric to , which is separable.)
On , let be the subspace of functions with a limit as (essentially), and . Extend to by Hahn–Banach. Show there is no with for all : test on and let .
Solution
for every , but by dominated convergence.
On , let with . Show and for every continuous . So the functionals converge, pointwise on , to , which is in but not of the form for any . In 4A.6 Weak Convergence and the Direct Method this pointwise convergence of functionals is called weak-* convergence, and Banach–Alaoglu guarantees that every bounded sequence of probability densities has a subsequence converging in this sense to a probability measure.
© 2026 NeckPinch (www.neckpinch.com). All content in the guidebook (text, mathematics, figures and exercises) is protected by copyright. All rights reserved. No part may be copied, republished or redistributed without written permission.