© 2026 NeckPinch · www.neckpinch.com · All rights reserved.
Course 3Book 3A: Measure, Integration and LᵖChapter 6
Modes of Convergence and Differentiation
Covering lemmas, the maximal function and Lebesgue’s differentiation theorem.
Read with Tao, An Introduction to Measure Theory, §1.5 "Modes of convergence" (especially the typewriter example and §1.5.4 on fast convergence) and §1.6 "Differentiation theorems", §1.6.1–1.6.2 (Lebesgue differentiation in one and several dimensions, with the Hardy–Littlewood maximal inequality and the Vitali-type covering lemma). Egorov's and Lusin's theorems are in §1.3.5. The rest of §1.6 (monotone and absolutely continuous functions in detail) can be read lightly.
For sequences of numbers there is only one way to converge. For sequences of functions there are many, and they are genuinely different: a sequence can converge in the sense that its integrals of shrink to zero, without converging at any single point. This chapter sorts out the main modes of convergence and how they relate, a piece of bookkeeping that every later analysis chapter takes for granted.
It then proves the measure-theoretic fundamental theorem of calculus: the average of an integrable function over a small ball around tends to , for almost every . The proof introduces two tools that outlive the theorem. The Vitali covering lemma selects disjoint balls from an overlapping family without losing much; the Hardy–Littlewood maximal function controls all averages of a function at once. Covering and packing arguments of exactly this kind are how Riemannian geometry turns volume bounds into compactness (9B.2 Volume Comparison, 9B.4 Convergence of Manifolds).
By the end of this chapter you will be able to:
- define convergence almost everywhere, in , in measure and almost uniformly, and decide which imply which;
- give the typewriter sequence and the other standard counterexamples, and extract an a.e. convergent subsequence from an convergent one;
- state Egorov's and Lusin's theorems and say what they mean;
- prove the Vitali covering lemma and the Hardy–Littlewood maximal inequality;
- prove and use the Lebesgue differentiation theorem.
Averages over shrinking windows
A moving average replaces a signal by its average over a window around each time,
It is the simplest smoothing filter in signal processing, and the standard way of reading trends in noisy data such as daily temperatures or prices. Shrinking the window should give back the signal, and where is continuous it obviously does. But a recorded signal need not be continuous: it may have jumps, or be merely integrable.
The Lebesgue differentiation theorem says that as for almost every , for every locally integrable , however rough. Where it fails is instructive: at a jump from to , a symmetric window converges to the midpoint (Figure 6.1), not to either value. But jumps (of a function of bounded variation, say) occur at only countably many times, a null set. Averaging loses nothing that integration can see.
Modes of convergence
Let be measurable functions on a measure space .
- Uniformly: .
- Almost everywhere: for all outside a null set.
- In (in mean): .
- In measure: for every , .
- Almost uniformly: for every there is a set with outside which uniformly.
Some implications are immediate. Uniform convergence implies the others (on a space of finite measure, for ). Convergence in implies convergence in measure, by Markov's inequality: (3A.3 The Lebesgue Integral). Almost uniform implies almost everywhere (intersect over ).
The failures are given by the escapes of 3A.3 The Lebesgue Integral, plus one new example.
On , list the indicator functions of the dyadic intervals, row by row: ; then ; then ; and so on (Figure 6.2). Like a typewriter carriage, each row sweeps across the whole interval with a narrower block. Then , since the blocks shrink: the sequence converges to in and in measure. But every point is covered by one block in every row, so infinitely often and infinitely often: the sequence converges at no point at all.
Conversely, the escapes of 3A.3 The Lebesgue Integral converge to almost everywhere but not in : horizontal escape on an infinite measure space; vertical escape, , even on . The tall spike converges to in measure but not in .
The typewriter's problem is that its blocks keep returning. Passing to a subsequence removes that.
- If , then almost everywhere.
- If in (or merely in measure), then a subsequence converges to almost everywhere.
Proof. (1) By monotone convergence, , so for almost every (an integrable function is finite a.e.), and then the terms tend to . (2) In measure, choose increasing with . By Borel–Cantelli (3A.2 Lebesgue Measure), almost every lies in only finitely many of these sets, and so for all large .
Part 2 is used constantly: whenever a limit is known to exist in an integral sense, some subsequence also converges pointwise almost everywhere, and then pointwise arguments (Fatou, continuity) can be applied to it.
Littlewood's second and third principles
Two theorems quantify how close the measure-theoretic world is to the continuous one. Both are in Tao, §1.3.5.
On a space of finite measure, convergence almost everywhere implies almost uniform convergence.
Let be measurable and finite almost everywhere. For every there is a set with such that the restriction of to is continuous.
Egorov is Littlewood's third principle (every convergent sequence of measurable functions is nearly uniformly convergent), Lusin his second (every measurable function is nearly continuous). Note what "nearly" means in Lusin: the restriction is continuous, not itself. Dirichlet's function is discontinuous everywhere, but restricted to the irrationals (removing a null set) it is the constant .
Covering and the maximal function
To prove that averages converge almost everywhere, we need to control the worst average around each point, uniformly over all radii. Define, for , the average of over the ball and its supremum:
is the Hardy–Littlewood maximal function. It can be large, but not often: the set where it exceeds has measure at most , which is what Markov's inequality would give for itself. The proof needs a way of choosing, from many overlapping balls, a disjoint family that still accounts for most of the space they cover.
Let be finitely many balls in . There is a subcollection of pairwise disjoint balls such that
where is the ball with the same centre and three times the radius. Consequently .
Proof. The greedy algorithm. Choose the largest ball. Discard every ball that meets it. From what remains, choose the largest, discard every ball that meets it, and repeat until nothing remains. The chosen balls are disjoint by construction. Every discarded ball met a chosen ball that was chosen while was still available, so is at least as large as : radius . Two balls that meet, the first no larger than the second, satisfy : any point of is within of a point of , hence within of the centre of . The volume bound follows from .
For and ,
Proof. Let (it is open, so measurable), and let be compact. Each has a ball centred at with . The open balls cover , so finitely many do (2B.3 Compactness). Apply Vitali to those: disjoint balls among them with
the last step because the are disjoint. By inner regularity (3A.2 Lebesgue Measure), is the supremum of over compact .
An estimate of the form is called weak type (1, 1). The maximal function is not bounded in (for , far away, which is not integrable), so this weaker estimate is the right one.
Lebesgue's differentiation theorem
Let be locally integrable on . Then for almost every ,
Points where the first limit holds are called Lebesgue points of .
Proof. The question is local, so we may assume (multiply by the indicator of a large ball). For let be the set of where . It suffices to show for every .
Split into a nice part and a small part. Given , choose a continuous compactly supported with (3A.3 The Lebesgue Integral), and write . For the continuous part the limit is at every point. So for ,
and therefore either or .
Both happen on small sets. By the maximal inequality and Markov's inequality,
Since is arbitrary, .
The proof is a template worth naming: a maximal inequality plus a dense class gives almost-everywhere convergence. The dense class (continuous functions) is where the convergence is easy; the maximal inequality controls the error uniformly in , which is what lets the approximation pass to the limit. The same template proves almost-everywhere convergence of heat-kernel averages as (6A.3 The Heat Equation on ℝⁿ), and of many other approximate identities.
Two consequences:
- Density points. Applying the theorem to : almost every point of a measurable set is a point of density 1, , and almost every point outside has density . A measurable set looks, under a microscope at almost every point, either completely full or completely empty. (A fat Cantor set, 3A.2 Lebesgue Measure, has density at almost all of its points even though it contains no interval.)
- The fundamental theorem of calculus for Lebesgue integrals. In one dimension, if and , then for almost every : the difference quotient is a one-sided average of , and one-sided averages converge at Lebesgue points too.
The converse direction of the fundamental theorem needs care. The Cantor function (the "devil's staircase") is continuous and increasing, constant on each interval removed in building the Cantor set (3A.2 Lebesgue Measure), and climbs from to entirely on . So at every point outside , which is almost every point, and yet . Functions for which always holds are exactly the absolutely continuous ones (Tao, §1.6.4); Lipschitz functions are among them, which is the case used most in geometry.
The greedy selection in Vitali's lemma has a twin that runs through Riemannian geometry: choose a maximal collection of disjoint balls of radius ; then the balls of radius with the same centres cover everything (Exercise 6.16). Counting how many disjoint small balls fit inside a big one is a volume comparison. On a manifold with Ricci curvature bounded below, Bishop–Gromov volume comparison (9B.2 Volume Comparison) gives exactly that count, independently of the manifold, and the resulting uniform bound on covering numbers is the total boundedness that Gromov's compactness theorem needs (2B.3 Compactness, 9B.4 Convergence of Manifolds). The same packing argument, with volumes controlled by noncollapsing, appears when Perelman's κ-solutions are shown to form a compact family (12B.2 The Structure of κ-Solutions).
History
Lebesgue proved the one-dimensional differentiation theorem in 1904 and the version for averages in several dimensions in 1910. Giuseppe Vitali's covering theorem appeared in 1908. Dmitri Egorov proved his theorem in 1911, and his student Nikolai Lusin his in 1912. G. H. Hardy and J. E. Littlewood introduced the maximal function in 1930, in a paper motivated by questions about averages of sequences, which they illustrated with a cricketer's batting averages. Cantor described his function in 1884, and Lebesgue used it as a counterexample to the fundamental theorem of calculus for continuous functions.
Functions can converge uniformly, almost everywhere, in , in measure or almost uniformly. convergence implies convergence in measure, and a subsequence then converges almost everywhere; the typewriter shows that the full sequence need not converge anywhere. Egorov and Lusin make "nearly uniform" and "nearly continuous" precise. The Vitali covering lemma gives the Hardy–Littlewood maximal inequality, and a maximal inequality plus a dense class gives the Lebesgue differentiation theorem: averages over shrinking balls recover an integrable function almost everywhere. 3A.7 Lᵖ Spaces and Jensen’s Inequality introduces the spaces, proves the inequalities of Hölder, Minkowski and Jensen, and shows the spaces are complete.
Exercises
For each sequence on , decide whether it converges to uniformly, almost everywhere, in , and in measure: (a) ; (b) ; (c) ; (d) the typewriter sequence; (e) .
Solution
(a) a.e., , in measure; not uniformly. (b) a.e. (except at ), in measure; not (), not uniform. (c) a.e., in measure, and in (); not uniform. (d) and in measure only. (e) a.e., , in measure; not uniform.
Find a subsequence of the typewriter sequence that converges to almost everywhere, and identify the exceptional set.
Solution
Take the first function in each row, . It converges to at every ; the exceptional set is .
Let on . Compute for and show it is comparable to . Deduce , and check the weak-type bound directly.
Solution
For and the interval of radius centred at : if it misses ; if the average is , increasing in ; if it is , decreasing. So the best radius is and (similarly for ), which is not integrable. For , is the interval , of length less than , within the bound .
Let be measurable with for every interval . Show that . (Use density points.) Can a measurable set meet every interval in exactly half its length?
Solution
If , almost every point of has density , so for small intervals around such a point . Contradiction. So no: such a set would have and also violate the density theorem at its density points.
(a) Define on by writing in ternary with digits , halving each digit, and reading the result in binary; extend to by making it constant on each removed interval. Show is continuous and increasing. (b) Show that maps , a null set, onto . So a continuous function can map a set of measure zero onto a set of positive measure, which a Lipschitz function can't (Exercise 6.15).
Let be Lipschitz with constant . Show that if then . (A ball of radius maps into a ball of radius ; cover by small cubes.) This is why the change-of-variables formula (3A.5 Product Measures and Change of Variables) and Sard's theorem (7A.7 Smooth Topology) can ignore null sets.
Let be a subset of , and let be a maximal set of points of with for (no more points of can be added). (a) Show that the balls cover , so the points form an -net (2B.3 Compactness). (b) Show that the balls are disjoint. (c) If , deduce by comparing volumes.
In 9B.2 Volume Comparison the volume comparison in (c) is replaced by Bishop–Gromov: on a manifold with , the ratio is bounded by a constant depending only on , and . That gives a bound on independent of the manifold, the uniform total boundedness behind Gromov's compactness theorem (9B.4 Convergence of Manifolds).
Solution
(a) If some were at distance from all , it could be added, contradicting maximality. (b) If and met, . (c) The disjoint balls of radius lie in , so .
© 2026 NeckPinch (www.neckpinch.com). All content in the guidebook (text, mathematics, figures and exercises) is protected by copyright. All rights reserved. No part may be copied, republished or redistributed without written permission.