Book 4A

© 2026 NeckPinch · www.neckpinch.com · All rights reserved.

Course 4Book 4A: Function Spaces and Sobolev SpacesChapter 3

Hahn–Banach and Duality

Separating convex sets, dual spaces, and the dual of Lᵖ.

19 min read · Updated Oct 2, 2026

Read with Brezis, chapter 1 (the Hahn–Banach theorems, analytic and geometric forms; the section on conjugate convex functions can wait), and chapter 4 for the dual of LpL^p; or Kreyszig §4.1–4.6 (Zorn's lemma, Hahn–Banach, the dual of C[a,b]C[a, b], adjoint operators, reflexive spaces).

In this chapter · 6 sections
  1. 3.1Prices from separation
  2. 3.2The analytic Hahn–Banach theorem
  3. 3.3Separating convex sets
  4. 3.4Dual spaces
  5. 3.4.1Reflexivity
  6. 3.5History
  7. 3.6Exercises

A bounded linear functional is a measurement: a continuous linear way of turning a vector into a number. Evaluating a function at a point, integrating it against a weight, taking a Fourier coefficient: these are all linear functionals. The dual space X∗X^* is the space of all of them, and much of functional analysis consists of studying a space through its dual. Two questions come first. Are there enough functionals to tell vectors apart? And what does the dual of a familiar space look like?

The first question is answered by the Hahn–Banach theorem, in two forms. The analytic form extends functionals from a subspace without increasing their norm. The geometric form, easier to picture and just as important, says that two disjoint convex sets can be separated by a hyperplane. The second question has satisfying answers: the dual of LpL^p is LqL^q for conjugate exponents (with one exception), and the dual of C(K)C(K) is the space of measures. That last fact is why limits of probability densities, the objects of 3A.4 Measures, Probability and Weights, can be measures, and it reappears at the frontier of Ricci flow research.

By the end of this chapter you will be able to:

  • state and use the analytic Hahn–Banach theorem, and find norming functionals;
  • prove that a point outside a closed convex set in Rn\mathbb{R}^n can be separated from it by a hyperplane, and state the general version;
  • identify the duals of ℓp\ell^p, LpL^p and C(K)C(K), and explain which spaces are reflexive;
  • use separating hyperplanes to price assets by no arbitrage, and to classify data.

Prices from separation

In the world In use No-arbitrage pricing

An arbitrage is a trading strategy that costs nothing (or less than nothing) today and can never lose money, with a chance of making some. In a well-functioning market arbitrage opportunities are competed away, and the fundamental theorem of asset pricing (J. Michael Harrison and David Kreps, 1979; Harrison and Stanley Pliska, 1981) says what that implies: in a market with no arbitrage there is a positive linear pricing functional, equivalently a set of "risk-neutral" probabilities, under which every asset's price is its expected discounted payoff. In a finite model the proof is a separating hyperplane.

The simplest case. A stock costs 100100 today; tomorrow it will be worth either 120120 or 9090. Money can be held at zero interest. Look for state prices q↑,q↓≥0q_{\uparrow}, q_{\downarrow} \geq 0 that reproduce today's prices: q↑+q↓=1q_\uparrow + q_\downarrow = 1 (the price of 11 for sure) and 120q↑+90q↓=100120q_\uparrow + 90q_\downarrow = 100. The unique solution is q↑=13q_\uparrow = \tfrac13, q↓=23q_\downarrow = \tfrac23. Then a call option paying max⁡(S−100,0)\max(S - 100, 0), which is 2020 in the up state and 00 in the down state, must cost 20×13≈6.6720 \times \tfrac13 \approx 6.67. Any other price creates an arbitrage, by combining the option with the stock and cash (Exercise 3.7).

Why separation? Consider the set of payoff vectors achievable at zero cost, a subspace, and the cone of non-negative, non-zero payoffs. "No arbitrage" says they don't meet. The separating hyperplane between them is a linear functional that vanishes on the costless payoffs and is positive on the cone, and that functional is the pricing rule. With more possible states than independent assets, the separating functional need not be unique, and the market is called incomplete.

The analytic Hahn–Banach theorem

A function p:X→Rp : X \to \mathbb{R} on a real vector space is sublinear if p(x+y)≤p(x)+p(y)p(x + y) \leq p(x) + p(y) and p(λx)=λp(x)p(\lambda x) = \lambda p(x) for λ≥0\lambda \geq 0. A norm is sublinear.

Theorem 3.1 Hahn–Banach, analytic form

Let XX be a real vector space, pp sublinear on XX, MM a subspace, and f:M→Rf : M \to \mathbb{R} linear with f≤pf \leq p on MM. Then ff extends to a linear F:X→RF : X \to \mathbb{R} with F≤pF \leq p on all of XX.

Proof architecture. One step: to extend ff from MM to M+Rx0M + \mathbb{R}x_0, one must choose a value F(x0)=cF(x_0) = c with f(m)+tc≤p(m+tx0)f(m) + tc \leq p(m + tx_0) for all m∈Mm \in M and real tt. Dividing by ∣t∣|t|, this says cc must lie between sup⁡m(f(m)−p(m−x0))\sup_m(f(m) - p(m - x_0)) and inf⁡m(p(m+x0)−f(m))\inf_m(p(m + x_0) - f(m)), and sublinearity is exactly what makes the supremum at most the infimum (Exercise 3.8). All steps: extend one dimension at a time. In a space with a countable dense set (a separable space, such as LpL^p for p<∞p < \infty) ordinary induction and continuity suffice; in general, Zorn's lemma (equivalent to the axiom of choice, 2A.8 Infinite Sets) provides a maximal extension, which must be defined everywhere.

The case used most is p(x)=C∥x∥p(x) = C\|x\|.

Corollary 3.2 Functionals are plentiful

Let XX be a normed space.

  1. Every bounded linear functional on a subspace extends to XX with the same norm.
  2. For every x0∈Xx_0 \in X there is ϕ∈X∗\phi \in X^* with ∥ϕ∥=1\|\phi\| = 1 and ϕ(x0)=∥x0∥\phi(x_0) = \|x_0\|.
  3. Consequently ∥x∥=max⁡{∣ϕ(x)∣:ϕ∈X∗,∥ϕ∥≤1}\|x\| = \max\{|\phi(x)| : \phi \in X^*, \|\phi\| \leq 1\}, and X∗X^* separates points: if ϕ(x)=ϕ(y)\phi(x) = \phi(y) for all ϕ\phi, then x=yx = y.

Proof. (1) Apply the theorem with p=∥f∥ ∥⋅∥p = \|f\|\,\|\cdot\| (for complex spaces, apply it to the real part and recover the imaginary part). (2) On the line Rx0\mathbb{R}x_0 define f(tx0)=t∥x0∥f(tx_0) = t\|x_0\|, of norm 11, and extend. (3) ∣ϕ(x)∣≤∥x∥|\phi(x)| \leq \|x\| when ∥ϕ∥≤1\|\phi\| \leq 1, with equality for the functional of (2).

In concrete spaces the norming functional of (2) can usually be written down. In LpL^p with 1<p<∞1 < p < \infty, it is g↦∫g ∣f∣p−1sign⁡f‾∥f∥pp−1g \mapsto \int g\,\frac{|f|^{p-1}\overline{\operatorname{sign}f}}{\|f\|_p^{p-1}}, the case of equality in Hölder's inequality (Exercise 3.9). The theorem is needed when there is no formula, and to know that enough functionals exist in general.

Separating convex sets

Theorem 3.3 Separation from a closed convex set, in Rn\mathbb{R}^n

Let C⊆RnC \subseteq \mathbb{R}^n be non-empty, closed and convex, and x0∉Cx_0 \notin C. Then there is a vector vv and a number cc with

v⋅x≤c<v⋅x0for all x∈C.v\cdot x \leq c < v\cdot x_0 \quad \text{for all } x \in C.

Proof. Let pp be the point of CC nearest to x0x_0, which exists because closed bounded subsets of Rn\mathbb{R}^n are compact and is unique by convexity (2B.10 Ordinary Differential Equations's projection exercise). Set v=x0−p≠0v = x_0 - p \neq 0. For x∈Cx \in C, the projection property gives v⋅(x−p)≤0v\cdot(x - p) \leq 0, so v⋅x≤v⋅pv\cdot x \leq v\cdot p. And v⋅x0−v⋅p=∣v∣2>0v\cdot x_0 - v\cdot p = |v|^2 > 0. Take c=v⋅pc = v\cdot p.

The hyperplane {v⋅x=c}\{v\cdot x = c\} touches CC at pp and separates it from x0x_0. If x0x_0 is a boundary point of CC instead, a limiting argument gives a supporting hyperplane: one that passes through x0x_0 with CC on one side (Figure 3.1).

Figure 3.1. Left: separating a point from a closed convex set by the hyperplane through the nearest point pp, perpendicular to x0−px_0 - p. Right: a supporting hyperplane at a boundary point. Both are what Hahn–Banach provides in any normed space.

In a general normed space the nearest point need not exist, but the conclusion survives, with a continuous linear functional in place of v⋅v\cdot{}.

Theorem 3.4 Hahn–Banach, geometric form

Let AA and BB be disjoint non-empty convex subsets of a real normed space XX.

  1. If AA is open, there are ϕ∈X∗\phi \in X^* and cc with ϕ(a)<c≤ϕ(b)\phi(a) < c \leq \phi(b) for all a∈Aa \in A, b∈Bb \in B.
  2. If AA is closed and BB is compact, there are ϕ∈X∗\phi \in X^* and c1<c2c_1 < c_2 with ϕ(a)≤c1<c2≤ϕ(b)\phi(a) \leq c_1 < c_2 \leq \phi(b): strict separation.

Proof architecture (Brezis, Theorems 1.6 and 1.7). For (1), reduce to separating a point from an open convex set CC containing 00, and apply the analytic form with the Minkowski gauge p(x)=inf⁡{t>0:x/t∈C}p(x) = \inf\{t > 0 : x/t \in C\}, which is sublinear and comparable to the norm. For (2), thicken BB slightly into an open convex set that still misses AA, and apply (1).

In the world In use A separating hyperplane classifies data

Given data points of two kinds (say, measurements from healthy and faulty machines), a linear classifier looks for a hyperplane {w⋅x=c}\{w\cdot x = c\} with one kind on each side. If the two convex hulls of the data are disjoint, the geometric Hahn–Banach theorem (in Rn\mathbb{R}^n, Theorem 3.3) says such a hyperplane exists, and the support vector machine of Corinna Cortes and Vladimir Vapnik (1995) chooses the one with the largest margin, the widest empty strip between the classes (Figure 3.2). That choice is itself a convex optimisation problem, solved through its dual. When the data are not separable, the method penalises violations, or maps the data into a higher-dimensional space where they become separable.

Figure 3.2. A maximal-margin separating line for two classes of points (computed for an illustrative data set). The dashed lines touch the nearest points of each class, the support vectors; the classifier depends only on them.

Linear programming is a third use. Minimising a linear cost over a convex polyhedron of feasible plans has a dual problem whose variables are prices, the "shadow prices" of the constraints: the marginal value of relaxing each constraint by one unit. The duality theorem of linear programming, that the optimal costs of the two problems agree, is a separation theorem for polyhedra (Farkas's lemma, Exercise 3.10).

Dual spaces

The dual of a normed space XX is X∗=L(X,R)X^* = L(X, \mathbb{R}) (or C\mathbb{C}), with the operator norm. It is always a Banach space (4A.1 Banach Spaces and Bounded Operators), even if XX is not complete. For the spaces of Book 3A the duals are concrete.

Theorem 3.5 The dual of LpL^p

Let 1≤p<∞1 \leq p < \infty and qq the conjugate exponent, and let μ\mu be σ-finite. Every g∈Lqg \in L^q defines a bounded functional ϕg(f)=∫fg dμ\phi_g(f) = \int fg\,d\mu on LpL^p with ∥ϕg∥=∥g∥q\|\phi_g\| = \|g\|_q, and every bounded linear functional on LpL^p is of this form for a unique gg. So (Lp)∗=Lq(L^p)^* = L^q.

Proof architecture. Hölder gives ∥ϕg∥≤∥g∥q\|\phi_g\| \leq \|g\|_q, and the norming function of Exercise 3.9 gives equality. For the converse, given ϕ\phi, define a set function ν(E)=ϕ(1E)\nu(E) = \phi(1_E) on sets of finite measure. It is countably additive (continuity of ϕ\phi plus dominated convergence) and absolutely continuous with respect to μ\mu, so the Radon–Nikodym theorem (3A.4 Measures, Probability and Weights) gives a density gg with ϕ(1E)=∫Eg dμ\phi(1_E) = \int_Eg\,d\mu; one then checks g∈Lqg \in L^q and extends from indicators to all of LpL^p by density. Brezis gives a different proof for 1<p<∞1 < p < \infty through uniform convexity.

The case p=∞p = \infty is the exception. Every g∈L1g \in L^1 gives a functional on L∞L^\infty, but not every functional on L∞L^\infty comes from an L1L^1 function: Hahn–Banach produces, for example, a "limit at 00" functional on L∞(0,1)L^\infty(0, 1) that extends f↦lim⁡x→0f(x)f \mapsto \lim_{x\to0}f(x) from the functions that have such a limit, and no integrable gg can represent it (Exercise 3.12). For sequences, (ℓp)∗=ℓq(\ell^p)^* = \ell^q for 1≤p<∞1 \leq p < \infty, and the dual of c0c_0 (sequences tending to 00, with the sup norm) is ℓ1\ell^1.

Theorem 3.6 The dual of C(K)C(K) (Riesz–Markov)

Let KK be a compact metric space. Every bounded linear functional on C(K)C(K) has the form ϕ(f)=∫Kf dν\phi(f) = \int_Kf\,d\nu for a unique finite signed Borel measure ν\nu, and ∥ϕ∥=∣ν∣(K)\|\phi\| = |\nu|(K), the total variation of ν\nu. Positive functionals (ϕ(f)≥0\phi(f) \geq 0 when f≥0f \geq 0) correspond to positive measures.

(Stated, not proved here; Kreyszig §4.4 treats C[a,b]C[a, b], which was Riesz's original 1909 theorem, with Stieltjes integrals.) The Dirac mass δx0\delta_{x_0}, evaluation at a point, is the functional f↦f(x0)f \mapsto f(x_0), of norm 11: it lives naturally in C(K)∗C(K)^* even though it has no density (3A.4 Measures, Probability and Weights).

Reflexivity

Each x∈Xx \in X defines a functional on X∗X^*, namely ϕ↦ϕ(x)\phi \mapsto \phi(x), and by Corollary 3.2 this embeds XX isometrically into its bidual X∗∗X^{**}. If the embedding is onto, XX is reflexive. Hilbert spaces are reflexive (4A.4 Hilbert Spaces and Lax–Milgram); so is LpL^p for 1<p<∞1 < p < \infty, since (Lp)∗∗=(Lq)∗=Lp(L^p)^{**} = (L^q)^* = L^p. But L1L^1, L∞L^\infty and C(K)C(K) (for infinite KK) are not.

Why care? Reflexivity is exactly the condition under which bounded sequences have weakly convergent subsequences (4A.6 Weak Convergence and the Direct Method), the replacement for the compactness that infinite dimensions take away. That is why existence theory in PDE prefers LpL^p and Sobolev spaces with 1<p<∞1 < p < \infty, and handles p=1p = 1 and p=∞p = \infty separately and with care.

Where this goes Measures as limits

Because C(K)∗C(K)^* is a space of measures, a sequence of probability densities ρn\rho_n on a compact space always has a subsequence converging, against continuous test functions, to a probability measure (4A.6 Weak Convergence and the Direct Method, Banach–Alaoglu) — but the limit need not have a density. The Dirac mass as a limit of narrowing Gaussians (3A.4 Measures, Probability and Weights) is the simplest example. In Ricci flow the conjugate heat kernel is a probability density that can concentrate, and in recent work on the structure of limits of Ricci flows (Bamler's theory of "metric flows", 12C.5 After Perelman) the natural objects are families of probability measures obtained in exactly this way.

History

Eduard Helly proved an extension theorem for functionals on C[a,b]C[a, b] in 1912. Hans Hahn (1927) and Stefan Banach (1929) independently proved the general theorem now named after them. Hermann Minkowski developed separation and supporting hyperplanes for convex bodies around 1900, and Gyula Farkas published his lemma on linear inequalities in 1902. Frigyes Riesz identified the dual of C[a,b]C[a, b] in 1909 and the duals of LpL^p in 1910; the representation for general compact spaces is due to Andrey Markov (1938) and Shizuo Kakutani (1941). Linear programming duality grew out of work by Leonid Kantorovich (1939), John von Neumann and George Dantzig (1947). Harrison, Kreps and Pliska's papers on arbitrage and martingales appeared in 1979 and 1981.

Recall Where we stand

The analytic Hahn–Banach theorem extends functionals dominated by a sublinear function; it gives norming functionals and shows the dual separates points. The geometric form separates disjoint convex sets by hyperplanes, with the nearest-point construction in Rn\mathbb{R}^n as the model; it underlies no-arbitrage pricing, linear classifiers and linear programming duality. The dual of LpL^p is LqL^q for 1≤p<∞1 \leq p < \infty, the dual of C(K)C(K) is the measures, and LpL^p is reflexive for 1<p<∞1 < p < \infty. 4A.4 Hilbert Spaces and Lax–Milgram specialises to Hilbert spaces, where the dual is the space itself and projections exist.

Exercises

Exercise 3.7 Pricing the option

In the one-period model of the anchor, show that if the option traded at 88 instead of 203\frac{20}3, one could sell the option, buy 23\frac{2}{3} of a share and borrow the difference, and make a riskless profit. Find a replicating portfolio: holdings of stock and cash whose value tomorrow equals the option's payoff in both states, and check its cost is 203\frac{20}{3}.

Solution

Replicate with Δ\Delta shares and BB in cash: 120Δ+B=20120\Delta + B = 20, 90Δ+B=090\Delta + B = 0, so Δ=23\Delta = \frac23, B=−60B = -60, at cost 100⋅23−60=203100\cdot\frac23 - 60 = \frac{20}3. If the option sells for 88, sell it and buy the replicating portfolio: 8−203=438 - \frac{20}{3} = \frac43 profit now, and tomorrow the portfolio exactly pays what the option owes.

Exercise 3.8 The one-step extension

Show that for m1,m2∈Mm_1, m_2 \in M, f(m1)−p(m1−x0)≤p(m2+x0)−f(m2)f(m_1) - p(m_1 - x_0) \leq p(m_2 + x_0) - f(m_2), using f(m1)+f(m2)=f(m1+m2)≤p(m1+m2)≤p(m1−x0)+p(m2+x0)f(m_1) + f(m_2) = f(m_1 + m_2) \leq p(m_1 + m_2) \leq p(m_1 - x_0) + p(m_2 + x_0). Deduce that a value c=F(x0)c = F(x_0) can be chosen, and check that the extension satisfies F≤pF \leq p on M+Rx0M + \mathbb{R}x_0.

Exercise 3.9 Norming functionals in LpL^p

For f∈Lpf \in L^p, 1<p<∞1 < p < \infty, f≠0f \neq 0, let g=∣f∣p−1sign⁡f‾/∥f∥pp−1g = |f|^{p-1}\overline{\operatorname{sign}f}/\|f\|_p^{p-1}. Show ∥g∥q=1\|g\|_q = 1 and ∫fg=∥f∥p\int fg = \|f\|_p. What is the norming functional for p=1p = 1? For p=2p = 2?

Exercise 3.10 Farkas's lemma

Let AA be an m×nm \times n matrix and b∈Rmb \in \mathbb{R}^m. Show that exactly one of the following holds: (i) Ax=bAx = b for some x≥0x \geq 0 (componentwise); (ii) there is yy with ATy≥0A^{\mathsf T}y \geq 0 and y⋅b<0y\cdot b < 0. (Apply Theorem 3.3 to bb and the closed convex cone {Ax:x≥0}\{Ax : x \geq 0\}; closedness of a finitely generated cone may be assumed.) Explain how the no-arbitrage theorem in a finite model is an instance.

Exercise 3.11 ℓ1\ell^1 is not reflexive

Show (c0)∗=ℓ1(c_0)^* = \ell^1 and (ℓ1)∗=ℓ∞(\ell^1)^* = \ell^\infty. Then show that ℓ1\ell^1 is not reflexive, assuming the fact that a normed space whose dual is separable is itself separable. (ℓ∞\ell^\infty is not separable: the indicator sequences of subsets of N\mathbb{N} are uncountably many and pairwise at distance 11. If ℓ1\ell^1 were reflexive, (ℓ∞)∗=(ℓ1)∗∗(\ell^\infty)^* = (\ell^1)^{**} would be isometric to ℓ1\ell^1, which is separable.)

Exercise 3.12 A functional on L∞L^\infty with no density

On L∞(0,1)L^\infty(0, 1), let MM be the subspace of functions with a limit as x→0+x \to 0^+ (essentially), and f(u)=lim⁡x→0+u(x)f(u) = \lim_{x\to0^+}u(x). Extend ff to ϕ∈(L∞)∗\phi \in (L^\infty)^* by Hahn–Banach. Show there is no g∈L1g \in L^1 with ϕ(u)=∫ug\phi(u) = \int ug for all uu: test on u=1(0,ε)u = 1_{(0, \varepsilon)} and let ε→0\varepsilon \to 0.

Solution

ϕ(1(0,ε))=1\phi(1_{(0,\varepsilon)}) = 1 for every ε\varepsilon, but ∫0εg→0\int_0^\varepsilon g \to 0 by dominated convergence.

Exercise 3.13 Rehearsal: measures as limits of densities

On C([−1,1])C([-1, 1]), let ϕn(f)=∫fρn\phi_n(f) = \int f\rho_n with ρn(x)=n2 1[−1/n,1/n](x)\rho_n(x) = \frac n2\,1_{[-1/n, 1/n]}(x). Show ∥ϕn∥=1\|\phi_n\| = 1 and ϕn(f)→f(0)\phi_n(f) \to f(0) for every continuous ff. So the functionals converge, pointwise on C([−1,1])C([-1, 1]), to δ0\delta_0, which is in C([−1,1])∗C([-1, 1])^* but not of the form ∫fρ\int f\rho for any ρ∈L1\rho \in L^1. In 4A.6 Weak Convergence and the Direct Method this pointwise convergence of functionals is called weak-* convergence, and Banach–Alaoglu guarantees that every bounded sequence of probability densities has a subsequence converging in this sense to a probability measure.

© 2026 NeckPinch (www.neckpinch.com). All content in the guidebook (text, mathematics, figures and exercises) is protected by copyright. All rights reserved. No part may be copied, republished or redistributed without written permission.