© 2026 NeckPinch · www.neckpinch.com · All rights reserved.
Course 3Book 3A: Measure, Integration and LᵖChapter 8
Convolution and Mollifiers
Smoothing by convolution, approximate identities, and Gaussian blur as heat flow.
Not in Tao's measure book. Read Tao, An Epsilon of Room I, §1.3 for density in , and either Stein and Shakarchi, Real Analysis, chapter 3 (approximations to the identity), or Evans, Partial Differential Equations, Appendix C.5 (mollifiers). This chapter is the canonical home for both; later repeats can be skipped.
Convolution is averaging with a weight. Given a function and a weight , the convolution replaces each value of by a weighted average of the values of nearby. The idea has appeared three times already, without the name: in the polynomial kernels of the Weierstrass theorem (2B.5 Uniform Convergence and Arzelà–Ascoli), in Fejér's kernel and the heat kernel on the ring (2B.7 Fourier Series and the First Heat Equation), and in the concentrating densities that converge to a Dirac mass (3A.4 Measures, Probability and Weights). This chapter gives it a home.
Three facts make convolution indispensable. It smooths: the convolution is as smooth as the smoother of the two factors, because derivatives can be moved onto the weight. It approximates: convolving with a narrow weight of total mass barely moves a function, in every norm with . And with the Gaussian weight it is the heat equation: blurring by a Gaussian of variance is the same as letting heat flow for time . The first two facts make smooth functions dense in , which is how every theorem in Books 4A and 6A is first proved for nice functions and then extended. The third is the first quantitative statement of thread H, which becomes Ricci flow's description as a heat equation for the metric.
By the end of this chapter you will be able to:
- compute convolutions, and prove Young's inequality ;
- show that convolution with a smooth compactly supported function is smooth, and differentiate it;
- build mollifiers and prove that in , so smooth compactly supported functions are dense;
- recognise approximate identities, and prove their convergence;
- show that convolution with the Gaussian solves the heat equation, with the semigroup law.
Gaussian blur is heat flow
Image editors offer a Gaussian blur: each pixel is replaced by a weighted average of its neighbours, with weights proportional to for a chosen standard deviation . That operation is exactly convolution with the Gaussian density of variance . And convolution with the Gaussian , of variance in each direction, is the solution at time of the heat equation in the plane, started from the image (proved below). So a Gaussian blur with standard deviation is heat flow for time , with the image's brightness playing the role of temperature (Figure 8.1).
This is not a loose analogy; it is an identity, and it has consequences you can check in any editor. Blurring twice, with standard deviations and , is the same as blurring once with , because running heat flow for time and then is running it for . In computer vision this observation is the basis of scale-space theory (Witkin, 1983; Koenderink, 1984), which represents an image by the whole family of its blurs and argues that blurring governed by the heat equation is the natural way to pass from fine detail to coarse structure without inventing new features.
Convolution
For measurable on , the convolution is
wherever the integral converges absolutely.
Read as a weight and as an average of the values , the values of around , weighted by . Substituting shows , and Fubini gives associativity, . The support of lies in the closure of : averaging spreads a function out by at most the width of the weight.
Let . Then is the length of the overlap of with its translate by , which is : a triangle. Convolving again gives a piecewise quadratic, then a cubic, and so on, each smoother than the last; suitably rescaled, they approach a Gaussian. (That is the central limit theorem, for sums of uniform random numbers: the density of a sum of independent random variables is the convolution of their densities.)
For and , , the convolution is defined almost everywhere, and
Proof. For , . For , write with , and apply Hölder (3A.7 Lᵖ Spaces and Jensen’s Inequality) in :
Integrate in and use Tonelli (3A.5 Product Measures and Change of Variables) on the right: . So . (Finiteness of the right side shows the integral defining converges absolutely for almost every .)
Young's inequality says that averaging against a weight of total mass at most cannot increase any norm. The general version, when , is in Tao's Epsilon of Room I and is needed in 4A.10 Sobolev Embeddings and Critical Exponents.
Convolution smooths
Let (integrable on every bounded set) and , times continuously differentiable with compact support. Then , and for every derivative of order .
Proof. Write . For in a bounded set , only in a bounded set contribute (since has compact support), and , which is integrable in and independent of . Differentiation under the integral sign (3A.3 The Lebesgue Integral) gives , and continuity of follows from dominated convergence the same way. Repeat times.
So the convolution of any locally integrable function, however rough, with a smooth compactly supported weight is smooth. The roughness of is invisible after averaging, because only the weight is ever differentiated.
Mollifiers
Take the smooth bump of 2B.6 Power Series, Exponentials and Bump Functions, for and otherwise, with chosen so that . It is smooth, non-negative, supported in the unit ball, and has total mass . For , the mollifier at scale is
supported in the ball of radius , still of mass (3A.5 Product Measures and Change of Variables's scaling). The mollification of is : the average of over the ball of radius around each point, weighted towards the centre.
The key lemma is that translation is continuous in .
For with , as .
Proof. For continuous with compact support it follows from uniform continuity: tends to uniformly and vanishes outside a fixed bounded set. For general , pick such a with (3A.7 Lᵖ Spaces and Jensen’s Inequality); then by Minkowski and translation invariance,
and the last term tends to .
This fails for : translating the indicator of a half-line by any moves it a distance in . It is another face of the fact that continuous functions are not dense in .
Let with . Then is , , and as . If is continuous, uniformly on compact sets; and at every Lebesgue point of .
Proof. Smoothness is Proposition 8.4 and the norm bound is Young. For convergence, since ,
Treat the right side as an average of the functions over , weighted by , and apply Minkowski's inequality for integrals (the norm of an average is at most the average of the norms, Exercise 8.10):
by Lemma 8.5. The statements about continuous and Lebesgue points follow from the same identity: , which is a constant times the average of over (3A.6 Modes of Convergence and Differentiation).
For , smooth compactly supported functions are dense in . The same holds on any open set .
Proof. Given and , first approximate by for large (dominated convergence), then mollify: is smooth, supported in , and close to for small . On an open set , cut off to a compact subset of first, so that mollifying doesn't spill outside .
Approximate identities
The proof used only three properties of , which define the general notion. A family of integrable functions is an approximate identity if (i) ; (ii) for all ; (iii) for every , as . For every such family, in () and uniformly for uniformly continuous bounded (Exercise 8.11). The family has many members already met: the polynomial kernels of Weierstrass (2B.5 Uniform Convergence and Arzelà–Ascoli), Fejér's kernel (2B.7 Fourier Series and the First Heat Equation, on the circle), the Gaussian, and, later, the Poisson kernel for harmonic functions (6A.2 Harmonic Functions). The Dirichlet kernel of 2B.7 Fourier Series and the First Heat Equation fails condition (ii), since grows like , which is exactly why Fourier partial sums can fail to converge.
Convolution with the Gaussian solves the heat equation
Let
the Gaussian of total mass (3A.5 Product Measures and Change of Variables) and variance in each coordinate.
- for .
- is an approximate identity as .
- (Semigroup law) .
- For , , the function is smooth for , solves , and in as .
Proof. (1) A direct computation (Exercise 8.12): with , , and as well. (2) With , : a rescaling of a single positive Gaussian of mass , which concentrates at as by the argument of 3A.4 Measures, Probability and Weights. (3) Exercise 8.13. (4) Smoothness and the equation come from differentiating under the integral sign, with domination by derivatives of on (a polynomial times a Gaussian); convergence from (2) and the approximate-identity theorem.
So the explicit formula
solves the heat equation on all of , for any initial temperature. It is the analogue of Fourier's series solution on the ring (2B.7 Fourier Series and the First Heat Equation), where the periodic heat kernel played the role of , and the properties seen there all carry over: instant smoothing, the maximum principle (an average against a positive kernel lies between the extremes), mass conservation ( for ), and the semigroup law, which says the solution at time is the solution at time started from the solution at time .
For an ideal camera whose blur is the same everywhere in the frame, the recorded image is the scene convolved with the camera's point-spread function, the image of a single point of light: a disc for a defocused lens, a Gaussian-like spot for atmospheric blur in astronomy. Undoing the blur, deconvolution, means solving for . In Fourier terms convolution is multiplication, , so in principle . But for a Gaussian , decays like , and dividing by it multiplies the high frequencies of the noise by enormous factors: the backward heat equation of 2B.7 Fourier Series and the First Heat Equation, which can't be run. Practical deblurring methods therefore add prior assumptions about the image (regularisation), and recover only a limited range of frequencies.
The sound of a room is captured by its impulse response: record what a microphone hears after a single sharp click, including all the echoes from walls and ceiling. For a room that behaves linearly and doesn't change over time, the sound of any source played in it is the source's signal convolved with the impulse response. Audio software uses exactly this to make a recording made in a dry studio sound as if it were played in a cathedral or a concert hall, by convolving with that hall's measured impulse response. It is Young's inequality's "averaging with a weight", with the weight a few seconds of echoes.
- Sobolev spaces and distributions (4A.8 Distributions and Weak Derivatives, 4A.9 Sobolev Spaces): weak derivatives are defined by moving derivatives onto smooth test functions, exactly as in Proposition 8.4, and mollification shows that smooth functions are dense in Sobolev spaces, so inequalities proved for smooth functions hold for all.
- The heat equation (6A.3 The Heat Equation on ℝⁿ): the formula above, with uniqueness, the maximum principle and Gaussian bounds; on a Riemannian manifold the heat kernel is no longer explicit, but is still an approximate identity (9B.7 The Heat Equation on a Manifold).
- Ricci flow (11A.1 The Equation and Its First Solutions): in suitable coordinates, is a heat equation for the metric, , and it smooths a metric the way a Gaussian blur smooths an image. That intuition, made precise by Shi's estimates (11A.3 Short-Time Existence and Uniqueness), is behind Hamilton's original picture of Ricci flow as a way of making a metric "rounder".
- Perelman (12A.6 Pseudolocality): the conjugate heat kernel, the analogue of for the backward heat equation coupled to Ricci flow, starts as a Dirac mass at a point; estimates for it play the role that the explicit Gaussian plays here, and they are central to pseudolocality.
History
Integrals of convolution type appear in the 19th-century solutions of the heat equation on the line, by Fourier and Poisson, as averages of the initial temperature against a Gaussian. Weierstrass's 1885 proof of his approximation theorem convolved with the Gaussian. Kurt Friedrichs introduced mollifiers under that name in 1944, in work on weak and strong solutions of differential equations; Sergei Sobolev had used similar averaging in the 1930s. W. H. Young proved his convolution inequality in 1912. The identification of Gaussian blur with diffusion as the basis of multi-scale image analysis is due to Andrew Witkin (1983) and Jan Koenderink (1984).
Jordan measure, behind the Riemann integral, fails for countable unions; Lebesgue measure, built from countable covers, is countably additive on a σ-algebra containing every set met in practice, though not on all sets. The Lebesgue integral obeys monotone convergence, Fatou and dominated convergence, and the three escapes to infinity show what can go wrong without domination. Abstract measures, densities and pushforwards give probability and the weighted measures ; Fubini–Tonelli and change of variables give the Gaussian integral and volumes of balls; maximal functions give the Lebesgue differentiation theorem. The spaces are complete, Hölder and Minkowski hold, and Jensen's inequality makes relative entropy non-negative. Convolution smooths and approximates, smooth functions are dense in , and convolution with the Gaussian is heat flow.
The spaces are complete normed vector spaces of functions: the first Banach spaces. The next questions are about such spaces in general. Which linear maps between them are continuous? When does a bounded sequence of functions have a convergent subsequence, now that in infinite dimensions the unit ball is not compact (2B.3 Compactness)? How should one differentiate a function that is merely in , as solutions of PDE often are? Book 4A, starting with 4A.1 Banach Spaces and Bounded Operators, answers all of these, and ends with the Sobolev spaces in which the PDE of Course 6 are solved.
Exercises
Compute and , and check in each case. Prove that identity in general for , using Fubini.
Solution
The first is the triangle on , area . The second is a trapezoid rising on , flat at height on , falling on , area . In general , with the exchange justified by Tonelli applied to .
For measurable and , show . (For it is Tonelli. In general, write the left side to the power as , exchange integrals, and apply Hölder.) This is the "norm of an average is at most the average of norms" used in Theorem 8.6.
Let satisfy the three conditions of an approximate identity. Show that uniformly for every bounded uniformly continuous , and in for , . (Split the integral into and as in 2B.5 Uniform Convergence and Arzelà–Ascoli's proof of the Weierstrass theorem.)
Verify for by computing , and .
Solution
. Summing over gives , the same.
Show that on (the -dimensional case then follows by Fubini, since is a product over coordinates). Complete the square in
and use the Gaussian integral. In probability, this says the sum of independent normal variables with variances and is normal with variance . For Ricci flow, the same structure appears in 6A.3 The Heat Equation on ℝⁿ and 12A.6 Pseudolocality: solving forward for time and then is solving for , which is what makes heat-kernel estimates composable.
Solution
The integral of over is .
On the circle , convolution with the periodic heat kernel at time multiplies the -th Fourier coefficient by (2B.7 Fourier Series and the First Heat Equation). (a) If a blurred signal is measured with an error of size in each coefficient, how large is the error in the -th coefficient of the naive deconvolution? (b) For and , find the largest for which the recovered coefficient has error below . This cut-off is why deblurring recovers only a band of frequencies.
Solution
(a) . (b) requires , so .
Let be open and , . Show that there are with . (Use the compact sets , approximate by , and mollify at a scale smaller than .)
© 2026 NeckPinch (www.neckpinch.com). All content in the guidebook (text, mathematics, figures and exercises) is protected by copyright. All rights reserved. No part may be copied, republished or redistributed without written permission.