Book 3A

© 2026 NeckPinch · www.neckpinch.com · All rights reserved.

Course 3Book 3A: Measure, Integration and LᵖChapter 7

Lᵖ Spaces and Jensen’s Inequality

Hölder, Minkowski, completeness, and convexity as the seed of entropy.

27 min read · Updated Oct 2, 2026

Tao's measure book stops before LpL^p spaces; read Tao, An Epsilon of Room I, §1.3 "LpL^p spaces" (Hölder, Minkowski, completeness, density, duality, interpolation). Stein and Shakarchi, Real Analysis, is an optional second voice for L1L^1 and L2L^2. Jensen's inequality is developed here.

In this chapter · 7 sections
  1. 7.1Which norm is your specification?
  2. 7.2LpL^pLp spaces
  3. 7.3Young, Hölder and Minkowski
  4. 7.4Completeness and density
  5. 7.4.1The Hilbert space L2L^2L2
  6. 7.5Jensen's inequality
  7. 7.5.1Entropy is never negative
  8. 7.6History
  9. 7.7Exercises

2B.1 Metric Spaces noticed that one space of functions can carry several natural distances: the largest difference, the area between the graphs. Measure theory now gives a whole family of them, one for each exponent p≥1p \geq 1:

∥f∥p=(∫∣f∣p dμ)1/p,∥f∥∞=the essential supremum of ∣f∣.\|f\|_p = \Big(\int|f|^p\,d\mu\Big)^{1/p}, \qquad \|f\|_\infty = \text{the essential supremum of }|f|.

Small pp cares about averages and ignores rare large values; large pp cares about the largest values and ignores averages. The spaces LpL^p of functions with ∥f∥p<∞\|f\|_p < \infty are the settings for nearly all of the PDE in Courses 4 and 6, and the inequalities proved in this chapter, Hölder's and Minkowski's, are used on almost every page from here on.

The chapter ends with Jensen's inequality: for a convex function, the function of the average is at most the average of the function. It is the source of the inequalities between means, and, applied to the logarithm, it proves that relative entropy is never negative. That is the first link of the entropy thread E, which leads through the log-Sobolev inequality (6A.10 Entropy, Information and Diffusion) to Perelman's W\mathcal{W}-entropy (12A.3 The 𝓦-Entropy).

By the end of this chapter you will be able to:

  • compute LpL^p norms, and decide which LpL^p spaces a given function belongs to;
  • prove Young's, Hölder's and Minkowski's inequalities, and use Hölder to compare LpL^p norms;
  • prove that LpL^p is complete, and say which functions are dense in it;
  • prove Jensen's inequality and derive from it the AM–GM inequality and Gibbs' inequality;
  • choose the right norm for a given practical question.

Which norm is your specification?

In the world Data Mains voltage: RMS and peak

European mains electricity is specified as 230230 volts. That number is not the largest voltage: it is the root-mean-square (RMS) value of the alternating voltage V(t)=V0sin⁡(2πft)V(t) = V_0\sin(2\pi ft) over a cycle,

Vrms=(1T∫0TV(t)2 dt)1/2=V02,V_{\text{rms}} = \Big(\frac1T\int_0^TV(t)^2\,dt\Big)^{1/2} = \frac{V_0}{\sqrt2},

an L2L^2 average over one period, normalised by the period. The peak voltage is therefore V0=2302≈325V_0 = 230\sqrt2 \approx 325 volts, an L∞L^\infty quantity (Figure 7.1).

Both numbers matter, for different decisions. The heat a resistive appliance produces is proportional to V2/RV^2/R averaged over time, so it depends on VrmsV_{\text{rms}}, which is why RMS is the specification. A kettle on 230230 V RMS AC heats exactly as on 230230 V DC. Insulation, on the other hand, must withstand the largest instantaneous voltage, so it is rated against the peak. And the mean value, the L1L^1-type average of VV itself, is 00, which tells you nothing useful at all.

The same choice appears in forecasting and machine learning, where a model's errors e1,…,eNe_1, \ldots, e_N are summarised by the mean absolute error 1N∑∣ei∣\frac1N\sum|e_i| (an L1L^1 norm), the root-mean-square error (1N∑ei2)1/2\big(\frac1N\sum e_i^2\big)^{1/2} (L2L^2), or the worst case max⁡∣ei∣\max|e_i| (L∞L^\infty). They rank models differently. RMS error punishes occasional large errors more than mean absolute error does, and the worst case cares about nothing else. Which is right depends on the cost of an error, not on the mathematics.

Figure 7.1. One cycle of 230230 V mains. The peak, 325325 V, is the L∞L^\infty size; the RMS value, 230230 V, is the L2L^2 average, the square root of the mean of V2V^2 (lower panel). Power depends on the second; insulation on the first.

LpL^p spaces

Let (X,μ)(X, \mu) be a measure space and 0<p<∞0 < p < \infty. For measurable ff, ∥f∥p=(∫∣f∣p dμ)1/p\|f\|_p = \big(\int|f|^p\,d\mu\big)^{1/p}. The essential supremum ∥f∥∞\|f\|_\infty is the least MM with ∣f∣≤M|f| \leq M almost everywhere: the supremum, ignoring null sets.

Definition 7.1 LpL^p spaces

For 1≤p≤∞1 \leq p \leq \infty, Lp(X,μ)L^p(X, \mu) is the set of measurable ff with ∥f∥p<∞\|f\|_p < \infty, with functions equal almost everywhere identified. On Rd\mathbb{R}^d with Lebesgue measure we write Lp(Rd)L^p(\mathbb{R}^d); for counting measure on N\mathbb{N} we write ℓp\ell^p.

The identification is what makes ∥f∥p=0\|f\|_p = 0 imply f=0f = 0 (3A.3 The Lebesgue Integral). The triangle inequality is Minkowski's inequality below; with it, d(f,g)=∥f−g∥pd(f, g) = \|f - g\|_p is a metric for 1≤p≤∞1 \leq p \leq \infty. (For p<1p < 1 the triangle inequality fails, which is why p≥1p \geq 1.) The unit balls of the norms (∣x1∣p+∣x2∣p)1/p\big(|x_1|^p + |x_2|^p\big)^{1/p} on R2\mathbb{R}^2, the ℓp\ell^p norms on two points, show the family interpolating between the diamond of ℓ1\ell^1 and the square of ℓ∞\ell^\infty (Figure 7.2, extending 2B.1 Metric Spaces).

Figure 7.2. Unit balls of the ℓp\ell^p norm on R2\mathbb{R}^2 for p=1,1.5,2,4,∞p = 1, 1.5, 2, 4, \infty. All are convex for p≥1p \geq 1, which is the triangle inequality. As pp grows, the ball fills out towards the square.

Which LpL^p contains which? On R\mathbb{R}, neither inclusion holds in general: ∣x∣−a|x|^{-a} is in LpL^p near 00 when ap<1ap < 1 (small pp forgives singularities) and in LpL^p near infinity when ap>1ap > 1 (large pp forgives slow decay). So ∣x∣−1/21(0,1)|x|^{-1/2}1_{(0,1)} is in L1L^1 but not L2L^2, and ∣x∣−11(1,∞)|x|^{-1}1_{(1,\infty)} is in L2L^2 but not L1L^1. On a space of finite measure, large pp is more demanding: Lp⊆LqL^p \subseteq L^q for q≤pq \leq p (Exercise 7.9). On ℓp\ell^p it is the other way round: ℓq⊆ℓp\ell^q \subseteq \ell^p for q≤pq \leq p, since small terms raised to higher powers get smaller.

Young, Hölder and Minkowski

Call p,q∈[1,∞]p, q \in [1, \infty] conjugate exponents if 1p+1q=1\frac1p + \frac1q = 1: (2,2)(2, 2), (1,∞)(1, \infty), (3,32)(3, \frac32).

Lemma 7.2 Young's inequality

For a,b≥0a, b \geq 0 and conjugate 1<p,q<∞1 < p, q < \infty,   ab≤app+bqq\;ab \leq \dfrac{a^p}{p} + \dfrac{b^q}{q}, with equality if and only if ap=bqa^p = b^q.

Proof. If aa or bb is 00 there is nothing to prove. Otherwise, since log⁡\log is concave,

log⁡(app+bqq)≥1plog⁡ap+1qlog⁡bq=log⁡(ab),\log\Big(\frac{a^p}{p} + \frac{b^q}{q}\Big) \geq \frac1p\log a^p + \frac1q\log b^q = \log(ab),

with equality exactly when ap=bqa^p = b^q (strict concavity). Exponentiate.

Theorem 7.3 Hölder's inequality

For conjugate p,q∈[1,∞]p, q \in [1, \infty] and measurable f,gf, g,

∫∣fg∣ dμ≤∥f∥p ∥g∥q.\int|fg|\,d\mu \leq \|f\|_p\,\|g\|_q.

Proof. The cases p=1p = 1 or p=∞p = \infty are immediate (∣fg∣≤∣f∣ ∥g∥∞|fg| \leq |f|\,\|g\|_\infty almost everywhere). For 1<p<∞1 < p < \infty, if ∥f∥p\|f\|_p or ∥g∥q\|g\|_q is 00 or ∞\infty the inequality is trivial. Otherwise normalise: replace ff by f/∥f∥pf/\|f\|_p and gg by g/∥g∥qg/\|g\|_q, so that both norms are 11. Then apply Young pointwise and integrate:

∫∣fg∣≤∫∣f∣pp+∫∣g∣qq=1p+1q=1.\int|fg| \leq \int\frac{|f|^p}{p} + \int\frac{|g|^q}{q} = \frac1p + \frac1q = 1.

Normalising to reduce to the case of norm 11, then using a pointwise inequality, is a standard move: the inequality is homogeneous, so the normalised case is the general case. For p=q=2p = q = 2 Hölder is the Cauchy–Schwarz inequality ∣⟨f,g⟩∣≤∥f∥2∥g∥2|\langle f, g\rangle| \leq \|f\|_2\|g\|_2.

Theorem 7.4 Minkowski's inequality

For 1≤p≤∞1 \leq p \leq \infty,   ∥f+g∥p≤∥f∥p+∥g∥p\;\|f + g\|_p \leq \|f\|_p + \|g\|_p.

Proof. For p=1p = 1 and p=∞p = \infty it follows from ∣f+g∣≤∣f∣+∣g∣|f + g| \leq |f| + |g|. For 1<p<∞1 < p < \infty, first note f+g∈Lpf + g \in L^p when f,gf, g are (since ∣f+g∣p≤2pmax⁡(∣f∣,∣g∣)p≤2p(∣f∣p+∣g∣p)|f + g|^p \leq 2^p\max(|f|, |g|)^p \leq 2^p(|f|^p + |g|^p)). Then write ∣f+g∣p≤∣f∣ ∣f+g∣p−1+∣g∣ ∣f+g∣p−1|f + g|^p \leq |f|\,|f + g|^{p-1} + |g|\,|f + g|^{p-1} and apply Hölder to each term, with exponents pp and q=pp−1q = \frac{p}{p-1}:

∥f+g∥pp≤(∥f∥p+∥g∥p) ∥∣f+g∣p−1∥q=(∥f∥p+∥g∥p) ∥f+g∥pp−1.\|f + g\|_p^p \leq \big(\|f\|_p + \|g\|_p\big)\,\big\||f + g|^{p-1}\big\|_q = \big(\|f\|_p + \|g\|_p\big)\,\|f + g\|_p^{p-1}.

Divide by ∥f+g∥pp−1\|f + g\|_p^{p-1} (if it is 00 there is nothing to prove).

Two consequences of Hölder are used constantly:

  • Finite measure spaces. If μ(X)<∞\mu(X) < \infty and q≤pq \leq p, then ∥f∥q≤μ(X)1q−1p∥f∥p\|f\|_q \leq \mu(X)^{\frac1q - \frac1p}\|f\|_p (Exercise 7.9). On a probability space, ∥f∥p\|f\|_p is increasing in pp.
  • Interpolation. If p<r<qp < r < q and 1r=θp+1−θq\frac1r = \frac\theta p + \frac{1 - \theta}q, then ∥f∥r≤∥f∥pθ ∥f∥q1−θ\|f\|_r \leq \|f\|_p^\theta\,\|f\|_q^{1-\theta} (Exercise 7.10). Control at two exponents gives control at every exponent in between. This log-convexity is the simplest case of interpolation, a theme that returns in the Sobolev inequalities (4A.10 Sobolev Embeddings and Critical Exponents) and in Gagliardo–Nirenberg-type estimates for PDE.

Completeness and density

Theorem 7.5 Riesz–Fischer

For 1≤p≤∞1 \leq p \leq \infty, Lp(X,μ)L^p(X, \mu) is complete.

Proof. For p<∞p < \infty. Let (fn)(f_n) be Cauchy in LpL^p. Choose n1<n2<⋯n_1 < n_2 < \cdots with ∥fnk+1−fnk∥p≤2−k\|f_{n_{k+1}} - f_{n_k}\|_p \leq 2^{-k}, a "fast Cauchy" subsequence (3A.6 Modes of Convergence and Differentiation). Let G=∣fn1∣+∑k∣fnk+1−fnk∣G = |f_{n_1}| + \sum_k|f_{n_{k+1}} - f_{n_k}|. By monotone convergence and Minkowski (applied to the partial sums), ∥G∥p≤∥fn1∥p+∑k2−k<∞\|G\|_p \leq \|f_{n_1}\|_p + \sum_k2^{-k} < \infty, so GG is finite almost everywhere. Where GG is finite, the series fn1+∑k(fnk+1−fnk)f_{n_1} + \sum_k(f_{n_{k+1}} - f_{n_k}) converges absolutely, so the subsequence fnkf_{n_k} converges almost everywhere, to some ff with ∣f∣≤G|f| \leq G, so f∈Lpf \in L^p.

It remains to show fn→ff_n \to f in LpL^p. Given ε\varepsilon, choose NN with ∥fn−fm∥p≤ε\|f_n - f_m\|_p \leq \varepsilon for n,m≥Nn, m \geq N. For n≥Nn \geq N, Fatou's lemma (3A.3 The Lebesgue Integral) applied to ∣fn−fnk∣p→∣fn−f∣p|f_n - f_{n_k}|^p \to |f_n - f|^p gives ∫∣fn−f∣p≤lim inf⁡k∫∣fn−fnk∣p≤εp\int|f_n - f|^p \leq \liminf_k\int|f_n - f_{n_k}|^p \leq \varepsilon^p. (The case p=∞p = \infty is simpler: a Cauchy sequence in L∞L^\infty is uniformly Cauchy off a null set.)

Two ideas from earlier chapters made this short: a fast subsequence converges almost everywhere, and Fatou passes the Cauchy bound to the limit. Completeness is the property the Riemann integral lacked (2B.2 Completeness and Contraction, 3A.3 The Lebesgue Integral); now every LpL^p has it. A complete normed vector space is a Banach space, the subject of Book 4A.

Density. For 1≤p<∞1 \leq p < \infty, the following are dense in Lp(Rd)L^p(\mathbb{R}^d): simple functions with finite-measure support; finite combinations of indicators of boxes; continuous functions with compact support; and (after 3A.8 Convolution and Mollifiers) smooth functions with compact support. So any statement that is continuous in the LpL^p norm can be proved for nice functions and extended by approximation. This fails for p=∞p = \infty: the uniform limit of continuous functions is continuous, so 1[0,∞)1_{[0,\infty)} is at distance at least 12\tfrac12 from every continuous function in L∞L^\infty.

The Hilbert space L2L^2

L2L^2 alone among the LpL^p spaces has an inner product, ⟨f,g⟩=∫fgˉ dμ\langle f, g\rangle = \int f\bar g\,d\mu, with ∥f∥22=⟨f,f⟩\|f\|_2^2 = \langle f, f\rangle. A complete inner product space is a Hilbert space (4A.4 Hilbert Spaces and Lax–Milgram), and L2L^2 is the model one. Fourier series find their natural home here: the characters ek(x)=e2πikxe_k(x) = e^{2\pi ikx} form an orthonormal basis of L2(R/Z)L^2(\mathbb{R}/\mathbb{Z}), and the map f↦(f^(k))k∈Zf \mapsto (\hat f(k))_{k\in\mathbb{Z}} is a bijection between L2(R/Z)L^2(\mathbb{R}/\mathbb{Z}) and ℓ2(Z)\ell^2(\mathbb{Z}) preserving inner products. In 2B.7 Fourier Series and the First Heat Equation only continuous ff were allowed and the map was not onto ℓ2\ell^2; with Lebesgue's integral, every square-summable sequence of coefficients is the Fourier series of some function. This was the content of the Riesz–Fischer theorem as first proved in 1907.

Duality is stated now and proved in 4A.3 Hahn–Banach and Duality: for 1≤p<∞1 \leq p < \infty and conjugate qq, every continuous linear functional on LpL^p is f↦∫fgf \mapsto \int fg for a unique g∈Lqg \in L^q, and Hölder's inequality is sharp: ∥f∥p=sup⁡{∣∫fg∣:∥g∥q≤1}\|f\|_p = \sup\big\{\big|\int fg\big| : \|g\|_q \leq 1\big\}.

Jensen's inequality

Theorem 7.6 Jensen's inequality

Let (X,μ)(X, \mu) be a probability space, f∈L1(μ)f \in L^1(\mu) real-valued, and ϕ:R→R\phi : \mathbb{R} \to \mathbb{R} convex. Then

ϕ(∫f dμ)≤∫ϕ(f) dμ.\phi\Big(\int f\,d\mu\Big) \leq \int\phi(f)\,d\mu.

If ϕ\phi is strictly convex, equality holds only when ff is constant almost everywhere.

Proof. Let m=∫f dμm = \int f\,d\mu. A convex function has a supporting line at mm: a line ℓ(t)=ϕ(m)+c(t−m)\ell(t) = \phi(m) + c(t - m) with ℓ≤ϕ\ell \leq \phi everywhere (take cc between the left and right derivatives of ϕ\phi at mm, which exist by convexity, Exercise 7.12). Then ϕ(f(x))≥ϕ(m)+c(f(x)−m)\phi(f(x)) \geq \phi(m) + c(f(x) - m) for every xx. Integrate against the probability measure μ\mu: the right side integrates to ϕ(m)+c(m−m)=ϕ(m)\phi(m) + c(m - m) = \phi(m). For strict convexity, equality forces ϕ(f(x))=ℓ(f(x))\phi(f(x)) = \ell(f(x)) almost everywhere, and a strictly convex function meets a line in at most one point, so f=mf = m almost everywhere.

Figure 7.3. Jensen's inequality. The average of ϕ\phi at two points (midpoint of the chord) lies above ϕ\phi at the average point; the supporting line at the average (dashed) lies below the graph, which is the proof.

The pattern "function of average versus average of function" is everywhere.

  • AM–GM. With ϕ=exp⁡\phi = \exp and ff taking values log⁡a1,…,log⁡an\log a_1, \ldots, \log a_n with equal probability: (a1⋯an)1/n≤a1+⋯+ann(a_1\cdots a_n)^{1/n} \leq \frac{a_1 + \cdots + a_n}{n}.
  • Power means. With ϕ(t)=tq/p\phi(t) = t^{q/p} for q≥pq \geq p: ∥f∥p≤∥f∥q\|f\|_p \leq \|f\|_q on a probability space.
  • Variance. With ϕ(t)=t2\phi(t) = t^2: (∫f)2≤∫f2\big(\int f\big)^2 \leq \int f^2, that is, variance is non-negative.
In the world Data The MPG illusion

In 2008, Richard Larrick and Jack Soll showed in Science that people systematically misjudge fuel savings when economy is quoted in miles per gallon. Improving a car from 1010 to 2020 mpg saves 55 gallons every 100100 miles; improving another from 2525 to 5050 mpg, a much bigger-looking jump, saves only 22. Fuel used is proportional to 1/mpg1/\text{mpg}, a convex function, so equal steps in mpg are worth less and less. The same convexity spoils averages: a car that does half its mileage at 2020 mpg and half at 4040 mpg averages not 3030 mpg but the harmonic mean, 21/20+1/40≈26.7\frac{2}{1/20 + 1/40} \approx 26.7 mpg, and by Jensen (with ϕ(t)=1/t\phi(t) = 1/t) the harmonic mean never exceeds the arithmetic one. The U.S. fuel-economy label introduced for the 2013 model year shows gallons per 100 miles alongside mpg; the regulators' final rule cited the Larrick–Soll paper in explaining why.

In the world In use Why volatility has value

A call option pays max⁡(S−K,0)\max(S - K, 0) if the asset's price SS at expiry exceeds the strike KK, and nothing otherwise. The payoff is a convex function of SS. By Jensen, its expected value is at least the payoff at the expected price: Emax⁡(S−K,0)≥max⁡(ES−K,0)\mathbb{E}\max(S - K, 0) \geq \max(\mathbb{E}S - K, 0), and the gap grows as SS becomes more spread out. So, other things equal, a more volatile asset makes the option more valuable to its holder, which is one of the basic facts of option pricing. Convexity of a payoff in this sense is called its "gamma" by traders.

Entropy is never negative

Let pp and qq be probability densities on Rd\mathbb{R}^d (or any measure space), with q>0q > 0 where p>0p > 0. Their relative entropy (Kullback–Leibler divergence) is

D(p ∥ q)=∫plog⁡pq dx.D(p\,\|\,q) = \int p\log\frac pq\,dx.
Corollary 7.7 Gibbs' inequality

D(p ∥ q)≥0D(p\,\|\,q) \geq 0, with equality if and only if p=qp = q almost everywhere.

Proof. Integrate against the probability measure p dxp\,dx and use Jensen with the strictly convex function ϕ(t)=−log⁡t\phi(t) = -\log t, applied to f=q/pf = q/p:

D(p ∥ q)=∫(−log⁡qp)p dx≥−log⁡∫qp p dx=−log⁡∫{p>0}q dx≥−log⁡1=0.D(p\,\|\,q) = \int\Big(-\log\frac qp\Big)p\,dx \geq -\log\int\frac qp\,p\,dx = -\log\int_{\{p > 0\}}q\,dx \geq -\log 1 = 0.

Equality forces q/pq/p to be constant pp-almost everywhere and ∫{p>0}q=1\int_{\{p>0\}}q = 1, hence p=qp = q.

Relative entropy measures how far one probability distribution is from another, in units of information. It is not a metric (it is not symmetric), but it is non-negative and vanishes only for equal distributions, which is all that is needed to use it as a measure of distance from equilibrium.

Where this goes Thread E: entropy, from Jensen to Perelman

The entropy of a probability density uu relative to the Gaussian measure, ∫ulog⁡u dγ\int u\log u\,d\gamma, is the quantity controlled by the logarithmic Sobolev inequality (6A.10 Entropy, Information and Diffusion): Gross's theorem bounds it by the Fisher information ∫∣∇u∣2udγ\int\frac{|\nabla u|^2}{u}d\gamma of 3A.4 Measures, Probability and Weights, with a constant that doesn't depend on the dimension. Under the heat flow, entropy decreases, and the log-Sobolev inequality makes it decrease exponentially fast, as Wirtinger's inequality did for the energy in 2B.7 Fourier Series and the First Heat Equation. Perelman's W\mathcal{W}-entropy (12A.3 The 𝓦-Entropy) is built so that, for flat space and the Gaussian, it reduces to exactly this log-Sobolev functional; its derivative along Ricci flow is a sum of squares, and its lower bound, a log-Sobolev inequality on the manifold, is what prevents collapsing (12A.4 κ-Noncollapsing). Gibbs' inequality is the first rung of that ladder.

History

Leonard James Rogers (1888) and Otto Hölder (1889) proved the inequality now named after Hölder; Hermann Minkowski's appeared in his Geometrie der Zahlen (1896). W. H. Young's inequality for products dates from 1912. Frigyes Riesz introduced the LpL^p spaces in 1910. Riesz and Ernst Fischer independently proved the completeness of L2L^2 in 1907. Johan Jensen published his inequality for convex functions in 1906. J. Willard Gibbs stated the inequality now named after him in his work on statistical mechanics (1902), and Solomon Kullback and Richard Leibler introduced relative entropy as a measure of information in 1951.

Recall Where we stand

The LpL^p norms measure functions with different emphases, from averages (p=1p = 1) to worst cases (p=∞p = \infty). Young's inequality gives Hölder's by normalisation, and Hölder gives Minkowski's triangle inequality, comparisons of norms on finite measure spaces, and interpolation. Every LpL^p is complete (Riesz–Fischer), and for p<∞p < \infty nice functions are dense; L2L^2 is a Hilbert space in which Fourier series are a perfect dictionary. Jensen's inequality, for convex functions and probability measures, gives AM–GM, power-mean inequalities and Gibbs' inequality: relative entropy is non-negative. 3A.8 Convolution and Mollifiers studies convolution, proves that smooth functions are dense in LpL^p, and shows that Gaussian blur is heat flow.

Exercises

Exercise 7.8 Which LpL^p?

For which p∈[1,∞]p \in [1, \infty] does each function belong to Lp(R)L^p(\mathbb{R})? (a) 11+∣x∣\frac{1}{1 + |x|}; (b) ∣x∣−1/31(0,1)|x|^{-1/3}1_{(0,1)}; (c) 1∣x∣(1+∣log⁡∣x∣∣)\frac{1}{\sqrt{|x|}(1 + |\log|x||)}; (d) e−x2e^{-x^2}.

Solution

(a) p>1p > 1, including ∞\infty. (b) p<3p < 3. (c) Near 00: ∣x∣−p/2∣log⁡∣x∣∣−p|x|^{-p/2}|\log|x||^{-p} is integrable iff p<2p < 2, or p=2p = 2 (the log helps: ∫dxxlog⁡2x<∞\int\frac{dx}{x\log^2x} < \infty). Near ∞\infty: integrable iff p>2p > 2, or p=2p = 2. So exactly p=2p = 2. (d) All pp.

Exercise 7.9 Comparing norms

(a) On a measure space with μ(X)<∞\mu(X) < \infty and 1≤q≤p≤∞1 \leq q \leq p \leq \infty, apply Hölder to ∣f∣q⋅1|f|^q\cdot1 with exponents pq\frac pq and its conjugate to show ∥f∥q≤μ(X)1q−1p∥f∥p\|f\|_q \leq \mu(X)^{\frac1q - \frac1p}\|f\|_p. (b) Show ∥a∥ℓp≤∥a∥ℓq\|a\|_{\ell^p} \leq \|a\|_{\ell^q} for q≤pq \leq p. (c) Show ∥f∥p→∥f∥∞\|f\|_p \to \|f\|_\infty as p→∞p \to \infty when μ(X)<∞\mu(X) < \infty.

Exercise 7.10 Interpolation

Let p<r<qp < r < q with 1r=θp+1−θq\frac1r = \frac\theta p + \frac{1-\theta}q. Apply Hölder to ∣f∣r=∣f∣θr∣f∣(1−θ)r|f|^r = |f|^{\theta r}|f|^{(1-\theta)r} with exponents pθr\frac{p}{\theta r} and q(1−θ)r\frac{q}{(1-\theta)r} to show ∥f∥r≤∥f∥pθ∥f∥q1−θ\|f\|_r \leq \|f\|_p^\theta\|f\|_q^{1-\theta}.

Exercise 7.11 When is Hölder an equality?

For 1<p<∞1 < p < \infty, show that equality holds in Hölder's inequality (with finite non-zero norms) if and only if ∣f∣p|f|^p and ∣g∣q|g|^q are proportional almost everywhere.

Exercise 7.12 Convex functions have supporting lines

Let ϕ\phi be convex on R\mathbb{R}. (a) Show that the slope ϕ(t)−ϕ(m)t−m\frac{\phi(t) - \phi(m)}{t - m} is increasing in tt (for t≠mt \neq m). (b) Deduce that the one-sided derivatives ϕ−′(m)≤ϕ+′(m)\phi'_-(m) \leq \phi'_+(m) exist, and that any cc between them gives a supporting line. (c) Deduce that convex functions on open intervals are continuous.

Exercise 7.13 RMS of periodic signals

(a) Show that V0sin⁡(2πft)V_0\sin(2\pi ft) has RMS value V0/2V_0/\sqrt2 over a period. (b) A square wave alternating between ±V0\pm V_0 has RMS value V0V_0, and a symmetric triangle wave of amplitude V0V_0 has RMS value V0/3V_0/\sqrt3. (c) By Plancherel (2B.7 Fourier Series and the First Heat Equation), the mean square of a periodic signal is the sum of the mean squares of its harmonics. Check this for the square wave using its Fourier coefficients.

Solution

(a) 1T∫0Tsin⁡2=12\frac1T\int_0^T\sin^2 = \frac12. (b) Square: V2=V02V^2 = V_0^2 always. Triangle: by symmetry, the mean of (V0s)2(V_0 s)^2 for ss uniform on [−1,1][-1, 1], which is V02/3V_0^2/3. (c) For the square wave, the kk-th odd harmonic has amplitude 4V0πk\frac{4V_0}{\pi k} and mean square 8V02π2k2\frac{8V_0^2}{\pi^2k^2}; summing over odd kk gives 8V02π2⋅π28=V02\frac{8V_0^2}{\pi^2}\cdot\frac{\pi^2}{8} = V_0^2.

Exercise 7.14 Rehearsal: the relative entropy of two Gaussians

Let pp be the Gaussian density on R\mathbb{R} with mean 00 and variance σ2\sigma^2, and qq the standard one (σ=1\sigma = 1). Show that

D(p ∥ q)=12(σ2−1−log⁡σ2),D(p\,\|\,q) = \tfrac12\big(\sigma^2 - 1 - \log\sigma^2\big),

and check directly that it is non-negative with equality only at σ=1\sigma = 1. (Use ∫x2p=σ2\int x^2p = \sigma^2, 3A.5 Product Measures and Change of Variables.) In 6A.10 Entropy, Information and Diffusion and 12A.3 The 𝓦-Entropy, entropies are computed for densities that are Gaussian to leading order, and a computation of exactly this kind identifies which scale τ\tau the Gaussian should be measured at: the minimum over σ\sigma is attained at the matching scale.

Solution

log⁡pq=−log⁡σ−x22σ2+x22\log\frac pq = -\log\sigma - \frac{x^2}{2\sigma^2} + \frac{x^2}{2}. Integrating against pp: −log⁡σ−12+σ22-\log\sigma - \frac12 + \frac{\sigma^2}2. The function s−1−log⁡ss - 1 - \log s (with s=σ2s = \sigma^2) is convex with minimum 00 at s=1s = 1.

© 2026 NeckPinch (www.neckpinch.com). All content in the guidebook (text, mathematics, figures and exercises) is protected by copyright. All rights reserved. No part may be copied, republished or redistributed without written permission.