© 2026 NeckPinch · www.neckpinch.com · All rights reserved.
Course 3Book 3A: Measure, Integration and LᵖChapter 7
Lᵖ Spaces and Jensen’s Inequality
Hölder, Minkowski, completeness, and convexity as the seed of entropy.
Tao's measure book stops before spaces; read Tao, An Epsilon of Room I, §1.3 " spaces" (Hölder, Minkowski, completeness, density, duality, interpolation). Stein and Shakarchi, Real Analysis, is an optional second voice for and . Jensen's inequality is developed here.
2B.1 Metric Spaces noticed that one space of functions can carry several natural distances: the largest difference, the area between the graphs. Measure theory now gives a whole family of them, one for each exponent :
Small cares about averages and ignores rare large values; large cares about the largest values and ignores averages. The spaces of functions with are the settings for nearly all of the PDE in Courses 4 and 6, and the inequalities proved in this chapter, Hölder's and Minkowski's, are used on almost every page from here on.
The chapter ends with Jensen's inequality: for a convex function, the function of the average is at most the average of the function. It is the source of the inequalities between means, and, applied to the logarithm, it proves that relative entropy is never negative. That is the first link of the entropy thread E, which leads through the log-Sobolev inequality (6A.10 Entropy, Information and Diffusion) to Perelman's -entropy (12A.3 The 𝓦-Entropy).
By the end of this chapter you will be able to:
- compute norms, and decide which spaces a given function belongs to;
- prove Young's, Hölder's and Minkowski's inequalities, and use Hölder to compare norms;
- prove that is complete, and say which functions are dense in it;
- prove Jensen's inequality and derive from it the AM–GM inequality and Gibbs' inequality;
- choose the right norm for a given practical question.
Which norm is your specification?
European mains electricity is specified as volts. That number is not the largest voltage: it is the root-mean-square (RMS) value of the alternating voltage over a cycle,
an average over one period, normalised by the period. The peak voltage is therefore volts, an quantity (Figure 7.1).
Both numbers matter, for different decisions. The heat a resistive appliance produces is proportional to averaged over time, so it depends on , which is why RMS is the specification. A kettle on V RMS AC heats exactly as on V DC. Insulation, on the other hand, must withstand the largest instantaneous voltage, so it is rated against the peak. And the mean value, the -type average of itself, is , which tells you nothing useful at all.
The same choice appears in forecasting and machine learning, where a model's errors are summarised by the mean absolute error (an norm), the root-mean-square error (), or the worst case (). They rank models differently. RMS error punishes occasional large errors more than mean absolute error does, and the worst case cares about nothing else. Which is right depends on the cost of an error, not on the mathematics.
spaces
Let be a measure space and . For measurable , . The essential supremum is the least with almost everywhere: the supremum, ignoring null sets.
For , is the set of measurable with , with functions equal almost everywhere identified. On with Lebesgue measure we write ; for counting measure on we write .
The identification is what makes imply (3A.3 The Lebesgue Integral). The triangle inequality is Minkowski's inequality below; with it, is a metric for . (For the triangle inequality fails, which is why .) The unit balls of the norms on , the norms on two points, show the family interpolating between the diamond of and the square of (Figure 7.2, extending 2B.1 Metric Spaces).
Which contains which? On , neither inclusion holds in general: is in near when (small forgives singularities) and in near infinity when (large forgives slow decay). So is in but not , and is in but not . On a space of finite measure, large is more demanding: for (Exercise 7.9). On it is the other way round: for , since small terms raised to higher powers get smaller.
Young, Hölder and Minkowski
Call conjugate exponents if : , , .
For and conjugate , , with equality if and only if .
Proof. If or is there is nothing to prove. Otherwise, since is concave,
with equality exactly when (strict concavity). Exponentiate.
For conjugate and measurable ,
Proof. The cases or are immediate ( almost everywhere). For , if or is or the inequality is trivial. Otherwise normalise: replace by and by , so that both norms are . Then apply Young pointwise and integrate:
Normalising to reduce to the case of norm , then using a pointwise inequality, is a standard move: the inequality is homogeneous, so the normalised case is the general case. For Hölder is the Cauchy–Schwarz inequality .
For , .
Proof. For and it follows from . For , first note when are (since ). Then write and apply Hölder to each term, with exponents and :
Divide by (if it is there is nothing to prove).
Two consequences of Hölder are used constantly:
- Finite measure spaces. If and , then (Exercise 7.9). On a probability space, is increasing in .
- Interpolation. If and , then (Exercise 7.10). Control at two exponents gives control at every exponent in between. This log-convexity is the simplest case of interpolation, a theme that returns in the Sobolev inequalities (4A.10 Sobolev Embeddings and Critical Exponents) and in Gagliardo–Nirenberg-type estimates for PDE.
Completeness and density
For , is complete.
Proof. For . Let be Cauchy in . Choose with , a "fast Cauchy" subsequence (3A.6 Modes of Convergence and Differentiation). Let . By monotone convergence and Minkowski (applied to the partial sums), , so is finite almost everywhere. Where is finite, the series converges absolutely, so the subsequence converges almost everywhere, to some with , so .
It remains to show in . Given , choose with for . For , Fatou's lemma (3A.3 The Lebesgue Integral) applied to gives . (The case is simpler: a Cauchy sequence in is uniformly Cauchy off a null set.)
Two ideas from earlier chapters made this short: a fast subsequence converges almost everywhere, and Fatou passes the Cauchy bound to the limit. Completeness is the property the Riemann integral lacked (2B.2 Completeness and Contraction, 3A.3 The Lebesgue Integral); now every has it. A complete normed vector space is a Banach space, the subject of Book 4A.
Density. For , the following are dense in : simple functions with finite-measure support; finite combinations of indicators of boxes; continuous functions with compact support; and (after 3A.8 Convolution and Mollifiers) smooth functions with compact support. So any statement that is continuous in the norm can be proved for nice functions and extended by approximation. This fails for : the uniform limit of continuous functions is continuous, so is at distance at least from every continuous function in .
The Hilbert space
alone among the spaces has an inner product, , with . A complete inner product space is a Hilbert space (4A.4 Hilbert Spaces and Lax–Milgram), and is the model one. Fourier series find their natural home here: the characters form an orthonormal basis of , and the map is a bijection between and preserving inner products. In 2B.7 Fourier Series and the First Heat Equation only continuous were allowed and the map was not onto ; with Lebesgue's integral, every square-summable sequence of coefficients is the Fourier series of some function. This was the content of the Riesz–Fischer theorem as first proved in 1907.
Duality is stated now and proved in 4A.3 Hahn–Banach and Duality: for and conjugate , every continuous linear functional on is for a unique , and Hölder's inequality is sharp: .
Jensen's inequality
Let be a probability space, real-valued, and convex. Then
If is strictly convex, equality holds only when is constant almost everywhere.
Proof. Let . A convex function has a supporting line at : a line with everywhere (take between the left and right derivatives of at , which exist by convexity, Exercise 7.12). Then for every . Integrate against the probability measure : the right side integrates to . For strict convexity, equality forces almost everywhere, and a strictly convex function meets a line in at most one point, so almost everywhere.
The pattern "function of average versus average of function" is everywhere.
- AM–GM. With and taking values with equal probability: .
- Power means. With for : on a probability space.
- Variance. With : , that is, variance is non-negative.
In 2008, Richard Larrick and Jack Soll showed in Science that people systematically misjudge fuel savings when economy is quoted in miles per gallon. Improving a car from to mpg saves gallons every miles; improving another from to mpg, a much bigger-looking jump, saves only . Fuel used is proportional to , a convex function, so equal steps in mpg are worth less and less. The same convexity spoils averages: a car that does half its mileage at mpg and half at mpg averages not mpg but the harmonic mean, mpg, and by Jensen (with ) the harmonic mean never exceeds the arithmetic one. The U.S. fuel-economy label introduced for the 2013 model year shows gallons per 100 miles alongside mpg; the regulators' final rule cited the Larrick–Soll paper in explaining why.
A call option pays if the asset's price at expiry exceeds the strike , and nothing otherwise. The payoff is a convex function of . By Jensen, its expected value is at least the payoff at the expected price: , and the gap grows as becomes more spread out. So, other things equal, a more volatile asset makes the option more valuable to its holder, which is one of the basic facts of option pricing. Convexity of a payoff in this sense is called its "gamma" by traders.
Entropy is never negative
Let and be probability densities on (or any measure space), with where . Their relative entropy (Kullback–Leibler divergence) is
, with equality if and only if almost everywhere.
Proof. Integrate against the probability measure and use Jensen with the strictly convex function , applied to :
Equality forces to be constant -almost everywhere and , hence .
Relative entropy measures how far one probability distribution is from another, in units of information. It is not a metric (it is not symmetric), but it is non-negative and vanishes only for equal distributions, which is all that is needed to use it as a measure of distance from equilibrium.
The entropy of a probability density relative to the Gaussian measure, , is the quantity controlled by the logarithmic Sobolev inequality (6A.10 Entropy, Information and Diffusion): Gross's theorem bounds it by the Fisher information of 3A.4 Measures, Probability and Weights, with a constant that doesn't depend on the dimension. Under the heat flow, entropy decreases, and the log-Sobolev inequality makes it decrease exponentially fast, as Wirtinger's inequality did for the energy in 2B.7 Fourier Series and the First Heat Equation. Perelman's -entropy (12A.3 The 𝓦-Entropy) is built so that, for flat space and the Gaussian, it reduces to exactly this log-Sobolev functional; its derivative along Ricci flow is a sum of squares, and its lower bound, a log-Sobolev inequality on the manifold, is what prevents collapsing (12A.4 κ-Noncollapsing). Gibbs' inequality is the first rung of that ladder.
History
Leonard James Rogers (1888) and Otto Hölder (1889) proved the inequality now named after Hölder; Hermann Minkowski's appeared in his Geometrie der Zahlen (1896). W. H. Young's inequality for products dates from 1912. Frigyes Riesz introduced the spaces in 1910. Riesz and Ernst Fischer independently proved the completeness of in 1907. Johan Jensen published his inequality for convex functions in 1906. J. Willard Gibbs stated the inequality now named after him in his work on statistical mechanics (1902), and Solomon Kullback and Richard Leibler introduced relative entropy as a measure of information in 1951.
The norms measure functions with different emphases, from averages () to worst cases (). Young's inequality gives Hölder's by normalisation, and Hölder gives Minkowski's triangle inequality, comparisons of norms on finite measure spaces, and interpolation. Every is complete (Riesz–Fischer), and for nice functions are dense; is a Hilbert space in which Fourier series are a perfect dictionary. Jensen's inequality, for convex functions and probability measures, gives AM–GM, power-mean inequalities and Gibbs' inequality: relative entropy is non-negative. 3A.8 Convolution and Mollifiers studies convolution, proves that smooth functions are dense in , and shows that Gaussian blur is heat flow.
Exercises
For which does each function belong to ? (a) ; (b) ; (c) ; (d) .
Solution
(a) , including . (b) . (c) Near : is integrable iff , or (the log helps: ). Near : integrable iff , or . So exactly . (d) All .
(a) On a measure space with and , apply Hölder to with exponents and its conjugate to show . (b) Show for . (c) Show as when .
Let with . Apply Hölder to with exponents and to show .
For , show that equality holds in Hölder's inequality (with finite non-zero norms) if and only if and are proportional almost everywhere.
Let be convex on . (a) Show that the slope is increasing in (for ). (b) Deduce that the one-sided derivatives exist, and that any between them gives a supporting line. (c) Deduce that convex functions on open intervals are continuous.
(a) Show that has RMS value over a period. (b) A square wave alternating between has RMS value , and a symmetric triangle wave of amplitude has RMS value . (c) By Plancherel (2B.7 Fourier Series and the First Heat Equation), the mean square of a periodic signal is the sum of the mean squares of its harmonics. Check this for the square wave using its Fourier coefficients.
Solution
(a) . (b) Square: always. Triangle: by symmetry, the mean of for uniform on , which is . (c) For the square wave, the -th odd harmonic has amplitude and mean square ; summing over odd gives .
Let be the Gaussian density on with mean and variance , and the standard one (). Show that
and check directly that it is non-negative with equality only at . (Use , 3A.5 Product Measures and Change of Variables.) In 6A.10 Entropy, Information and Diffusion and 12A.3 The 𝓦-Entropy, entropies are computed for densities that are Gaussian to leading order, and a computation of exactly this kind identifies which scale the Gaussian should be measured at: the minimum over is attained at the matching scale.
Solution
. Integrating against : . The function (with ) is convex with minimum at .
© 2026 NeckPinch (www.neckpinch.com). All content in the guidebook (text, mathematics, figures and exercises) is protected by copyright. All rights reserved. No part may be copied, republished or redistributed without written permission.