Book 3A

© 2026 NeckPinch · www.neckpinch.com · All rights reserved.

Course 3Book 3A: Measure, Integration and LᵖChapter 8

Convolution and Mollifiers

Smoothing by convolution, approximate identities, and Gaussian blur as heat flow.

29 min read · Updated Oct 2, 2026

Not in Tao's measure book. Read Tao, An Epsilon of Room I, §1.3 for density in LpL^p, and either Stein and Shakarchi, Real Analysis, chapter 3 (approximations to the identity), or Evans, Partial Differential Equations, Appendix C.5 (mollifiers). This chapter is the canonical home for both; later repeats can be skipped.

In this chapter · 6 sections
  1. 8.1Gaussian blur is heat flow
  2. 8.2Convolution
  3. 8.2.1Convolution smooths
  4. 8.3Mollifiers
  5. 8.3.1Approximate identities
  6. 8.4Convolution with the Gaussian solves the heat equation
  7. 8.5History
  8. 8.6Exercises

Convolution is averaging with a weight. Given a function ff and a weight gg, the convolution f∗gf * g replaces each value of ff by a weighted average of the values of ff nearby. The idea has appeared three times already, without the name: in the polynomial kernels of the Weierstrass theorem (2B.5 Uniform Convergence and Arzelà–Ascoli), in Fejér's kernel and the heat kernel on the ring (2B.7 Fourier Series and the First Heat Equation), and in the concentrating densities that converge to a Dirac mass (3A.4 Measures, Probability and Weights). This chapter gives it a home.

Three facts make convolution indispensable. It smooths: the convolution is as smooth as the smoother of the two factors, because derivatives can be moved onto the weight. It approximates: convolving with a narrow weight of total mass 11 barely moves a function, in every LpL^p norm with p<∞p < \infty. And with the Gaussian weight it is the heat equation: blurring by a Gaussian of variance 2t2t is the same as letting heat flow for time tt. The first two facts make smooth functions dense in LpL^p, which is how every theorem in Books 4A and 6A is first proved for nice functions and then extended. The third is the first quantitative statement of thread H, which becomes Ricci flow's description as a heat equation for the metric.

By the end of this chapter you will be able to:

  • compute convolutions, and prove Young's inequality ∥f∗g∥p≤∥f∥1∥g∥p\|f * g\|_p \leq \|f\|_1\|g\|_p;
  • show that convolution with a smooth compactly supported function is smooth, and differentiate it;
  • build mollifiers and prove that f∗φε→ff * \varphi_\varepsilon \to f in LpL^p, so smooth compactly supported functions are dense;
  • recognise approximate identities, and prove their convergence;
  • show that convolution with the Gaussian solves the heat equation, with the semigroup law.

Gaussian blur is heat flow

In the world Model Blurring an image is running the heat equation

Image editors offer a Gaussian blur: each pixel is replaced by a weighted average of its neighbours, with weights proportional to e−∣x∣2/2σ2e^{-|x|^2/2\sigma^2} for a chosen standard deviation σ\sigma. That operation is exactly convolution with the Gaussian density of variance σ2\sigma^2. And convolution with the Gaussian (4πt)−1e−∣x∣2/4t(4\pi t)^{-1}e^{-|x|^2/4t}, of variance 2t2t in each direction, is the solution at time tt of the heat equation ∂tu=Δu\partial_tu = \Delta u in the plane, started from the image (proved below). So a Gaussian blur with standard deviation σ\sigma is heat flow for time t=σ2/2t = \sigma^2/2, with the image's brightness playing the role of temperature (Figure 8.1).

This is not a loose analogy; it is an identity, and it has consequences you can check in any editor. Blurring twice, with standard deviations σ1\sigma_1 and σ2\sigma_2, is the same as blurring once with σ12+σ22\sqrt{\sigma_1^2 + \sigma_2^2}, because running heat flow for time t1t_1 and then t2t_2 is running it for t1+t2t_1 + t_2. In computer vision this observation is the basis of scale-space theory (Witkin, 1983; Koenderink, 1984), which represents an image by the whole family of its blurs and argues that blurring governed by the heat equation is the natural way to pass from fine detail to coarse structure without inventing new features.

Figure 8.1. A synthetic image (a square, a disc and a thin bar on a 32×3232 \times 32 grid) and its Gaussian blurs at times t=0.5,2,8t = 0.5, 2, 8 (standard deviations σ=2t=1,2,4\sigma = \sqrt{2t} = 1, 2, 4 pixels), computed by convolution. Fine features vanish first: the heat equation damps high frequencies fastest, as on the ring of 2B.7 Fourier Series and the First Heat Equation.

Convolution

Definition 8.1 Convolution

For measurable f,gf, g on Rd\mathbb{R}^d, the convolution is

(f∗g)(x)=∫Rdf(x−y) g(y) dy,(f * g)(x) = \int_{\mathbb{R}^d}f(x - y)\,g(y)\,dy,

wherever the integral converges absolutely.

Read gg as a weight and f∗g(x)f * g(x) as an average of the values f(x−y)f(x - y), the values of ff around xx, weighted by g(y)g(y). Substituting y↦x−yy \mapsto x - y shows f∗g=g∗ff * g = g * f, and Fubini gives associativity, (f∗g)∗h=f∗(g∗h)(f * g) * h = f * (g * h). The support of f∗gf * g lies in the closure of {a+b:a∈supp⁡f, b∈supp⁡g}\{a + b : a \in \operatorname{supp}f,\ b \in \operatorname{supp}g\}: averaging spreads a function out by at most the width of the weight.

Example 8.2 Convolving two boxes

Let g=1[−1/2,1/2]g = 1_{[-1/2, 1/2]}. Then g∗g(x)g * g(x) is the length of the overlap of [−12,12][-\tfrac12, \tfrac12] with its translate by xx, which is max⁡(1−∣x∣,0)\max(1 - |x|, 0): a triangle. Convolving again gives a piecewise quadratic, then a cubic, and so on, each smoother than the last; suitably rescaled, they approach a Gaussian. (That is the central limit theorem, for sums of uniform random numbers: the density of a sum of independent random variables is the convolution of their densities.)

Theorem 8.3 Young's inequality

For f∈L1(Rd)f \in L^1(\mathbb{R}^d) and g∈Lp(Rd)g \in L^p(\mathbb{R}^d), 1≤p≤∞1 \leq p \leq \infty, the convolution f∗gf * g is defined almost everywhere, and

∥f∗g∥p≤∥f∥1 ∥g∥p.\|f * g\|_p \leq \|f\|_1\,\|g\|_p.

Proof. For p=∞p = \infty, ∣f∗g(x)∣≤∫∣f(x−y)∣ ∥g∥∞ dy=∥f∥1∥g∥∞|f * g(x)| \leq \int|f(x - y)|\,\|g\|_\infty\,dy = \|f\|_1\|g\|_\infty. For 1≤p<∞1 \leq p < \infty, write ∣f(x−y)g(y)∣=∣f(x−y)∣1/q⋅∣f(x−y)∣1/p∣g(y)∣|f(x - y)g(y)| = |f(x - y)|^{1/q}\cdot|f(x - y)|^{1/p}|g(y)| with 1p+1q=1\frac1p + \frac1q = 1, and apply Hölder (3A.7 Lᵖ Spaces and Jensen’s Inequality) in yy:

∣f∗g(x)∣p≤∥f∥1p/q∫∣f(x−y)∣ ∣g(y)∣p dy.|f * g(x)|^p \leq \|f\|_1^{p/q}\int|f(x - y)|\,|g(y)|^p\,dy.

Integrate in xx and use Tonelli (3A.5 Product Measures and Change of Variables) on the right: ∫∫∣f(x−y)∣∣g(y)∣p dy dx=∥f∥1∥g∥pp\int\int|f(x - y)||g(y)|^p\,dy\,dx = \|f\|_1\|g\|_p^p. So ∥f∗g∥pp≤∥f∥1p/q+1∥g∥pp=∥f∥1p∥g∥pp\|f * g\|_p^p \leq \|f\|_1^{p/q + 1}\|g\|_p^p = \|f\|_1^p\|g\|_p^p. (Finiteness of the right side shows the integral defining f∗g(x)f * g(x) converges absolutely for almost every xx.)

Young's inequality says that averaging against a weight of total mass at most 11 cannot increase any LpL^p norm. The general version, ∥f∗g∥r≤∥f∥p∥g∥q\|f * g\|_r \leq \|f\|_p\|g\|_q when 1+1r=1p+1q1 + \frac1r = \frac1p + \frac1q, is in Tao's Epsilon of Room I and is needed in 4A.10 Sobolev Embeddings and Critical Exponents.

Convolution smooths

Proposition 8.4 Derivatives fall on the smooth factor

Let f∈Lloc1(Rd)f \in L^1_{\mathrm{loc}}(\mathbb{R}^d) (integrable on every bounded set) and g∈Cck(Rd)g \in C_c^k(\mathbb{R}^d), kk times continuously differentiable with compact support. Then f∗g∈Ckf * g \in C^k, and ∂α(f∗g)=f∗∂αg\partial^\alpha(f * g) = f * \partial^\alpha g for every derivative of order ∣α∣≤k|\alpha| \leq k.

Proof. Write f∗g(x)=∫f(y) g(x−y) dyf * g(x) = \int f(y)\,g(x - y)\,dy. For xx in a bounded set BB, only yy in a bounded set B′B' contribute (since gg has compact support), and ∣∂xi[f(y)g(x−y)]∣≤∣f(y)∣ sup⁡∣∂ig∣ 1B′(y)|\partial_{x_i}[f(y)g(x - y)]| \leq |f(y)|\,\sup|\partial_ig|\,1_{B'}(y), which is integrable in yy and independent of xx. Differentiation under the integral sign (3A.3 The Lebesgue Integral) gives ∂i(f∗g)=f∗∂ig\partial_i(f * g) = f * \partial_ig, and continuity of f∗∂igf * \partial_ig follows from dominated convergence the same way. Repeat kk times.

So the convolution of any locally integrable function, however rough, with a smooth compactly supported weight is smooth. The roughness of ff is invisible after averaging, because only the weight is ever differentiated.

Mollifiers

Take the smooth bump of 2B.6 Power Series, Exponentials and Bump Functions, ϕ(x)=c exp⁡(−11−∣x∣2)\phi(x) = c\,\exp\big(-\frac{1}{1 - |x|^2}\big) for ∣x∣<1|x| < 1 and 00 otherwise, with cc chosen so that ∫ϕ=1\int\phi = 1. It is smooth, non-negative, supported in the unit ball, and has total mass 11. For ε>0\varepsilon > 0, the mollifier at scale ε\varepsilon is

ϕε(x)=ε−dϕ(x/ε),\phi_\varepsilon(x) = \varepsilon^{-d}\phi(x/\varepsilon),

supported in the ball of radius ε\varepsilon, still of mass 11 (3A.5 Product Measures and Change of Variables's scaling). The mollification of ff is fε=f∗ϕεf_\varepsilon = f * \phi_\varepsilon: the average of ff over the ball of radius ε\varepsilon around each point, weighted towards the centre.

Figure 8.2. A unit step and its mollifications f∗ϕεf * \phi_\varepsilon for ε=0.4,0.2,0.05\varepsilon = 0.4, 0.2, 0.05. Each is infinitely differentiable, equals the step outside an ε\varepsilon-neighbourhood of the jump, and converges to it in every LpL^p with p<∞p < \infty, but not uniformly.

The key lemma is that translation is continuous in LpL^p.

Lemma 8.5 Continuity of translation

For f∈Lp(Rd)f \in L^p(\mathbb{R}^d) with 1≤p<∞1 \leq p < \infty,   ∥f(⋅−h)−f∥p→0\;\|f(\cdot - h) - f\|_p \to 0 as h→0h \to 0.

Proof. For ff continuous with compact support it follows from uniform continuity: ∣f(x−h)−f(x)∣|f(x - h) - f(x)| tends to 00 uniformly and vanishes outside a fixed bounded set. For general ff, pick such a gg with ∥f−g∥p≤ε\|f - g\|_p \leq \varepsilon (3A.7 Lᵖ Spaces and Jensen’s Inequality); then by Minkowski and translation invariance,

∥f(⋅−h)−f∥p≤∥f(⋅−h)−g(⋅−h)∥p+∥g(⋅−h)−g∥p+∥g−f∥p≤2ε+∥g(⋅−h)−g∥p,\|f(\cdot - h) - f\|_p \leq \|f(\cdot - h) - g(\cdot - h)\|_p + \|g(\cdot - h) - g\|_p + \|g - f\|_p \leq 2\varepsilon + \|g(\cdot - h) - g\|_p,

and the last term tends to 00.

This fails for p=∞p = \infty: translating the indicator of a half-line by any h≠0h \neq 0 moves it a distance 11 in L∞L^\infty. It is another face of the fact that continuous functions are not dense in L∞L^\infty.

Theorem 8.6 Mollification converges

Let f∈Lp(Rd)f \in L^p(\mathbb{R}^d) with 1≤p<∞1 \leq p < \infty. Then fε=f∗ϕεf_\varepsilon = f * \phi_\varepsilon is C∞C^\infty, ∥fε∥p≤∥f∥p\|f_\varepsilon\|_p \leq \|f\|_p, and ∥fε−f∥p→0\|f_\varepsilon - f\|_p \to 0 as ε→0\varepsilon \to 0. If ff is continuous, fε→ff_\varepsilon \to f uniformly on compact sets; and fε(x)→f(x)f_\varepsilon(x) \to f(x) at every Lebesgue point of ff.

Proof. Smoothness is Proposition 8.4 and the norm bound is Young. For convergence, since ∫ϕε=1\int\phi_\varepsilon = 1,

fε(x)−f(x)=∫(f(x−y)−f(x))ϕε(y) dy.f_\varepsilon(x) - f(x) = \int\big(f(x - y) - f(x)\big)\phi_\varepsilon(y)\,dy.

Treat the right side as an average of the functions x↦f(x−y)−f(x)x \mapsto f(x - y) - f(x) over yy, weighted by ϕε\phi_\varepsilon, and apply Minkowski's inequality for integrals (the LpL^p norm of an average is at most the average of the LpL^p norms, Exercise 8.10):

∥fε−f∥p≤∫∥f(⋅−y)−f∥p ϕε(y) dy≤sup⁡∣y∣≤ε∥f(⋅−y)−f∥p→0,\|f_\varepsilon - f\|_p \leq \int\|f(\cdot - y) - f\|_p\,\phi_\varepsilon(y)\,dy \leq \sup_{|y| \leq \varepsilon}\|f(\cdot - y) - f\|_p \to 0,

by Lemma 8.5. The statements about continuous ff and Lebesgue points follow from the same identity: ∣fε(x)−f(x)∣≤sup⁡ϕ⋅ε−d∫B(0,ε)∣f(x−y)−f(x)∣ dy|f_\varepsilon(x) - f(x)| \leq \sup\phi\cdot\varepsilon^{-d}\int_{B(0,\varepsilon)}|f(x - y) - f(x)|\,dy, which is a constant times the average of ∣f−f(x)∣|f - f(x)| over B(x,ε)B(x, \varepsilon) (3A.6 Modes of Convergence and Differentiation).

Corollary 8.7 Smooth functions are dense

For 1≤p<∞1 \leq p < \infty, smooth compactly supported functions Cc∞(Rd)C_c^\infty(\mathbb{R}^d) are dense in Lp(Rd)L^p(\mathbb{R}^d). The same holds on any open set U⊆RdU \subseteq \mathbb{R}^d.

Proof. Given f∈Lpf \in L^p and ε>0\varepsilon > 0, first approximate ff by g=f 1B(0,R)g = f\,1_{B(0, R)} for large RR (dominated convergence), then mollify: g∗ϕδg * \phi_\delta is smooth, supported in B(0,R+δ)B(0, R + \delta), and close to gg for small δ\delta. On an open set UU, cut off to a compact subset of UU first, so that mollifying doesn't spill outside UU.

Approximate identities

The proof used only three properties of ϕε\phi_\varepsilon, which define the general notion. A family KεK_\varepsilon of integrable functions is an approximate identity if (i) ∫Kε=1\int K_\varepsilon = 1; (ii) ∫∣Kε∣≤C\int|K_\varepsilon| \leq C for all ε\varepsilon; (iii) for every δ>0\delta > 0, ∫∣y∣>δ∣Kε∣→0\int_{|y| > \delta}|K_\varepsilon| \to 0 as ε→0\varepsilon \to 0. For every such family, f∗Kε→ff * K_\varepsilon \to f in LpL^p (p<∞p < \infty) and uniformly for uniformly continuous bounded ff (Exercise 8.11). The family has many members already met: the polynomial kernels of Weierstrass (2B.5 Uniform Convergence and Arzelà–Ascoli), Fejér's kernel (2B.7 Fourier Series and the First Heat Equation, on the circle), the Gaussian, and, later, the Poisson kernel for harmonic functions (6A.2 Harmonic Functions). The Dirichlet kernel of 2B.7 Fourier Series and the First Heat Equation fails condition (ii), since ∫∣DN∣\int|D_N| grows like log⁡N\log N, which is exactly why Fourier partial sums can fail to converge.

Convolution with the Gaussian solves the heat equation

Let

H(x,t)=(4πt)−d/2 e−∣x∣2/4t(x∈Rd, t>0),H(x, t) = (4\pi t)^{-d/2}\,e^{-|x|^2/4t} \qquad (x \in \mathbb{R}^d,\ t > 0),

the Gaussian of total mass 11 (3A.5 Product Measures and Change of Variables) and variance 2t2t in each coordinate.

Theorem 8.8 The Gaussian solves the heat equation
  1. ∂tH=ΔH\partial_tH = \Delta H for t>0t > 0.
  2. {H(⋅,t)}t>0\{H(\cdot, t)\}_{t > 0} is an approximate identity as t→0t \to 0.
  3. (Semigroup law) H(⋅,s)∗H(⋅,t)=H(⋅,s+t)H(\cdot, s) * H(\cdot, t) = H(\cdot, s + t).
  4. For f∈Lp(Rd)f \in L^p(\mathbb{R}^d), 1≤p<∞1 \leq p < \infty, the function u(x,t)=(f∗H(⋅,t))(x)u(x, t) = (f * H(\cdot, t))(x) is smooth for t>0t > 0, solves ∂tu=Δu\partial_tu = \Delta u, and u(⋅,t)→fu(\cdot, t) \to f in LpL^p as t→0t \to 0.

Proof. (1) A direct computation (Exercise 8.12): with r=∣x∣r = |x|, ∂tH=(r24t2−d2t)H\partial_tH = \big(\frac{r^2}{4t^2} - \frac{d}{2t}\big)H, and ΔH=(r24t2−d2t)H\Delta H = \big(\frac{r^2}{4t^2} - \frac{d}{2t}\big)H as well. (2) With ε=t\varepsilon = \sqrt t, H(x,t)=ε−dH(x/ε,1)H(x, t) = \varepsilon^{-d}H(x/\varepsilon, 1): a rescaling of a single positive Gaussian of mass 11, which concentrates at 00 as t→0t \to 0 by the argument of 3A.4 Measures, Probability and Weights. (3) Exercise 8.13. (4) Smoothness and the equation come from differentiating under the integral sign, with domination by derivatives of HH on t≥t0>0t \geq t_0 > 0 (a polynomial times a Gaussian); convergence from (2) and the approximate-identity theorem.

So the explicit formula

u(x,t)=∫Rd(4πt)−d/2e−∣x−y∣2/4tf(y) dyu(x, t) = \int_{\mathbb{R}^d}(4\pi t)^{-d/2}e^{-|x - y|^2/4t}f(y)\,dy

solves the heat equation on all of Rd\mathbb{R}^d, for any LpL^p initial temperature. It is the analogue of Fourier's series solution on the ring (2B.7 Fourier Series and the First Heat Equation), where the periodic heat kernel played the role of HH, and the properties seen there all carry over: instant smoothing, the maximum principle (an average against a positive kernel lies between the extremes), mass conservation (∫u=∫f\int u = \int f for f∈L1f \in L^1), and the semigroup law, which says the solution at time s+ts + t is the solution at time tt started from the solution at time ss.

In the world Model Optical blur and why it can't simply be undone

For an ideal camera whose blur is the same everywhere in the frame, the recorded image is the scene convolved with the camera's point-spread function, the image of a single point of light: a disc for a defocused lens, a Gaussian-like spot for atmospheric blur in astronomy. Undoing the blur, deconvolution, means solving g=f∗kg = f * k for ff. In Fourier terms convolution is multiplication, g^=f^ k^\hat g = \hat f\,\hat k, so in principle f^=g^/k^\hat f = \hat g/\hat k. But for a Gaussian kk, k^\hat k decays like e−4π2∣ξ∣2te^{-4\pi^2|\xi|^2t}, and dividing by it multiplies the high frequencies of the noise by enormous factors: the backward heat equation of 2B.7 Fourier Series and the First Heat Equation, which can't be run. Practical deblurring methods therefore add prior assumptions about the image (regularisation), and recover only a limited range of frequencies.

In the world In use Convolution reverb

The sound of a room is captured by its impulse response: record what a microphone hears after a single sharp click, including all the echoes from walls and ceiling. For a room that behaves linearly and doesn't change over time, the sound of any source played in it is the source's signal convolved with the impulse response. Audio software uses exactly this to make a recording made in a dry studio sound as if it were played in a cathedral or a concert hall, by convolving with that hall's measured impulse response. It is Young's inequality's "averaging with a weight", with the weight a few seconds of echoes.

Where this goes Mollifiers and heat kernels in the rest of the guidebook
  • Sobolev spaces and distributions (4A.8 Distributions and Weak Derivatives, 4A.9 Sobolev Spaces): weak derivatives are defined by moving derivatives onto smooth test functions, exactly as in Proposition 8.4, and mollification shows that smooth functions are dense in Sobolev spaces, so inequalities proved for smooth functions hold for all.
  • The heat equation (6A.3 The Heat Equation on ℝⁿ): the formula above, with uniqueness, the maximum principle and Gaussian bounds; on a Riemannian manifold the heat kernel is no longer explicit, but is still an approximate identity (9B.7 The Heat Equation on a Manifold).
  • Ricci flow (11A.1 The Equation and Its First Solutions): in suitable coordinates, ∂tg=−2 Ric(g)\partial_tg = -2\,\mathrm{Ric}(g) is a heat equation for the metric, ∂tgij≈Δgij+lower order\partial_tg_{ij} \approx \Delta g_{ij} + \text{lower order}, and it smooths a metric the way a Gaussian blur smooths an image. That intuition, made precise by Shi's estimates (11A.3 Short-Time Existence and Uniqueness), is behind Hamilton's original picture of Ricci flow as a way of making a metric "rounder".
  • Perelman (12A.6 Pseudolocality): the conjugate heat kernel, the analogue of HH for the backward heat equation coupled to Ricci flow, starts as a Dirac mass at a point; estimates for it play the role that the explicit Gaussian plays here, and they are central to pseudolocality.

History

Integrals of convolution type appear in the 19th-century solutions of the heat equation on the line, by Fourier and Poisson, as averages of the initial temperature against a Gaussian. Weierstrass's 1885 proof of his approximation theorem convolved with the Gaussian. Kurt Friedrichs introduced mollifiers under that name in 1944, in work on weak and strong solutions of differential equations; Sergei Sobolev had used similar averaging in the 1930s. W. H. Young proved his convolution inequality in 1912. The identification of Gaussian blur with diffusion as the basis of multi-scale image analysis is due to Andrew Witkin (1983) and Jan Koenderink (1984).

Recall Book 3A in one paragraph

Jordan measure, behind the Riemann integral, fails for countable unions; Lebesgue measure, built from countable covers, is countably additive on a σ-algebra containing every set met in practice, though not on all sets. The Lebesgue integral obeys monotone convergence, Fatou and dominated convergence, and the three escapes to infinity show what can go wrong without domination. Abstract measures, densities and pushforwards give probability and the weighted measures e−fdVe^{-f}dV; Fubini–Tonelli and change of variables give the Gaussian integral and volumes of balls; maximal functions give the Lebesgue differentiation theorem. The LpL^p spaces are complete, Hölder and Minkowski hold, and Jensen's inequality makes relative entropy non-negative. Convolution smooths and approximates, smooth functions are dense in LpL^p, and convolution with the Gaussian is heat flow.

Where this goes Into Book 4A

The LpL^p spaces are complete normed vector spaces of functions: the first Banach spaces. The next questions are about such spaces in general. Which linear maps between them are continuous? When does a bounded sequence of functions have a convergent subsequence, now that in infinite dimensions the unit ball is not compact (2B.3 Compactness)? How should one differentiate a function that is merely in LpL^p, as solutions of PDE often are? Book 4A, starting with 4A.1 Banach Spaces and Bounded Operators, answers all of these, and ends with the Sobolev spaces in which the PDE of Course 6 are solved.

Exercises

Exercise 8.9 Convolving indicators

Compute 1[0,1]∗1[0,1]1_{[0,1]} * 1_{[0,1]} and 1[0,1]∗1[0,2]1_{[0,1]} * 1_{[0,2]}, and check ∫(f∗g)=∫f∫g\int(f * g) = \int f\int g in each case. Prove that identity in general for f,g∈L1f, g \in L^1, using Fubini.

Solution

The first is the triangle max⁡(0,min⁡(x,2−x))\max(0, \min(x, 2 - x)) on [0,2][0, 2], area 11. The second is a trapezoid rising on [0,1][0, 1], flat at height 11 on [1,2][1, 2], falling on [2,3][2, 3], area 22. In general ∫∫f(x−y)g(y) dy dx=∫g(y)∫f(x−y) dx dy=∫f∫g\int\int f(x - y)g(y)\,dy\,dx = \int g(y)\int f(x - y)\,dx\,dy = \int f\int g, with the exchange justified by Tonelli applied to ∣f∣∣g∣|f||g|.

Exercise 8.10 Minkowski's inequality for integrals

For F(x,y)≥0F(x, y) \geq 0 measurable and 1≤p<∞1 \leq p < \infty, show ∥∫F(⋅,y) dy∥p≤∫∥F(⋅,y)∥p dy\Big\|\int F(\cdot, y)\,dy\Big\|_p \leq \int\|F(\cdot, y)\|_p\,dy. (For p=1p = 1 it is Tonelli. In general, write the left side to the power pp as ∫(∫F(x,y)dy)(∫F(x,y′)dy′)p−1dx\int\big(\int F(x, y)dy\big)\big(\int F(x, y')dy'\big)^{p-1}dx, exchange integrals, and apply Hölder.) This is the "norm of an average is at most the average of norms" used in Theorem 8.6.

Exercise 8.11 Approximate identities converge

Let KεK_\varepsilon satisfy the three conditions of an approximate identity. Show that f∗Kε→ff * K_\varepsilon \to f uniformly for every bounded uniformly continuous ff, and in LpL^p for f∈Lpf \in L^p, p<∞p < \infty. (Split the integral into ∣y∣≤δ|y| \leq \delta and ∣y∣>δ|y| > \delta as in 2B.5 Uniform Convergence and Arzelà–Ascoli's proof of the Weierstrass theorem.)

Exercise 8.12 The heat kernel solves the heat equation

Verify ∂tH=ΔH\partial_tH = \Delta H for H(x,t)=(4πt)−d/2e−∣x∣2/4tH(x, t) = (4\pi t)^{-d/2}e^{-|x|^2/4t} by computing ∂tH\partial_tH, ∂iH=−xi2tH\partial_iH = -\frac{x_i}{2t}H and ∂i2H=(xi24t2−12t)H\partial_i^2H = \big(\frac{x_i^2}{4t^2} - \frac{1}{2t}\big)H.

Solution

∂tlog⁡H=−d2t+∣x∣24t2\partial_t\log H = -\frac{d}{2t} + \frac{|x|^2}{4t^2}. Summing ∂i2H\partial_i^2H over ii gives (∣x∣24t2−d2t)H\big(\frac{|x|^2}{4t^2} - \frac{d}{2t}\big)H, the same.

Exercise 8.13 Rehearsal: the semigroup law

Show that H(⋅,s)∗H(⋅,t)=H(⋅,s+t)H(\cdot, s) * H(\cdot, t) = H(\cdot, s + t) on R\mathbb{R} (the dd-dimensional case then follows by Fubini, since HH is a product over coordinates). Complete the square in

(x−y)24s+y24t=x24(s+t)+s+t4st(y−ts+tx)2,\frac{(x - y)^2}{4s} + \frac{y^2}{4t} = \frac{x^2}{4(s + t)} + \frac{s + t}{4st}\Big(y - \frac{t}{s + t}x\Big)^2,

and use the Gaussian integral. In probability, this says the sum of independent normal variables with variances 2s2s and 2t2t is normal with variance 2(s+t)2(s + t). For Ricci flow, the same structure appears in 6A.3 The Heat Equation on ℝⁿ and 12A.6 Pseudolocality: solving forward for time ss and then tt is solving for s+ts + t, which is what makes heat-kernel estimates composable.

Solution

The integral of (4πs)−1/2(4πt)−1/2e−x2/4(s+t)e−s+t4st(y−…)2(4\pi s)^{-1/2}(4\pi t)^{-1/2}e^{-x^2/4(s+t)}e^{-\frac{s+t}{4st}(y - \ldots)^2} over yy is (4πs)−1/2(4πt)−1/2e−x2/4(s+t)4πsts+t=(4π(s+t))−1/2e−x2/4(s+t)(4\pi s)^{-1/2}(4\pi t)^{-1/2}e^{-x^2/4(s+t)}\sqrt{\frac{4\pi st}{s + t}} = (4\pi(s + t))^{-1/2}e^{-x^2/4(s+t)}.

Exercise 8.14 Blur in Fourier terms

On the circle R/Z\mathbb{R}/\mathbb{Z}, convolution with the periodic heat kernel at time tt multiplies the kk-th Fourier coefficient by e−4π2k2te^{-4\pi^2k^2t} (2B.7 Fourier Series and the First Heat Equation). (a) If a blurred signal is measured with an error of size η\eta in each coefficient, how large is the error in the kk-th coefficient of the naive deconvolution? (b) For t=0.01t = 0.01 and η=10−3\eta = 10^{-3}, find the largest kk for which the recovered coefficient has error below 11. This cut-off is why deblurring recovers only a band of frequencies.

Solution

(a) η e4π2k2t\eta\,e^{4\pi^2k^2t}. (b) e0.3948k2<1000e^{0.3948k^2} < 1000 requires k2<ln⁡1000/0.3948≈17.5k^2 < \ln 1000/0.3948 \approx 17.5, so k≤4k \leq 4.

Exercise 8.15 Smooth approximation on a domain

Let U⊆RdU \subseteq \mathbb{R}^d be open and f∈Lp(U)f \in L^p(U), 1≤p<∞1 \leq p < \infty. Show that there are fn∈Cc∞(U)f_n \in C_c^\infty(U) with ∥fn−f∥Lp(U)→0\|f_n - f\|_{L^p(U)} \to 0. (Use the compact sets Kn={x∈U:∣x∣≤n, dist(x,∂U)≥1n}K_n = \{x \in U : |x| \leq n,\ \mathrm{dist}(x, \partial U) \geq \frac1n\}, approximate ff by f1Knf1_{K_n}, and mollify at a scale smaller than 12n\frac1{2n}.)

© 2026 NeckPinch (www.neckpinch.com). All content in the guidebook (text, mathematics, figures and exercises) is protected by copyright. All rights reserved. No part may be copied, republished or redistributed without written permission.