Book 3A

© 2026 NeckPinch · www.neckpinch.com · All rights reserved.

Course 3Book 3A: Measure, Integration and LᵖChapter 6

Modes of Convergence and Differentiation

Covering lemmas, the maximal function and Lebesgue’s differentiation theorem.

24 min read · Updated Oct 2, 2026

Read with Tao, An Introduction to Measure Theory, §1.5 "Modes of convergence" (especially the typewriter example and §1.5.4 on fast convergence) and §1.6 "Differentiation theorems", §1.6.1–1.6.2 (Lebesgue differentiation in one and several dimensions, with the Hardy–Littlewood maximal inequality and the Vitali-type covering lemma). Egorov's and Lusin's theorems are in §1.3.5. The rest of §1.6 (monotone and absolutely continuous functions in detail) can be read lightly.

In this chapter · 6 sections
  1. 6.1Averages over shrinking windows
  2. 6.2Modes of convergence
  3. 6.2.1Littlewood's second and third principles
  4. 6.3Covering and the maximal function
  5. 6.4Lebesgue's differentiation theorem
  6. 6.5History
  7. 6.6Exercises

For sequences of numbers there is only one way to converge. For sequences of functions there are many, and they are genuinely different: a sequence can converge in the sense that its integrals of ∣fn−f∣|f_n - f| shrink to zero, without converging at any single point. This chapter sorts out the main modes of convergence and how they relate, a piece of bookkeeping that every later analysis chapter takes for granted.

It then proves the measure-theoretic fundamental theorem of calculus: the average of an integrable function over a small ball around xx tends to f(x)f(x), for almost every xx. The proof introduces two tools that outlive the theorem. The Vitali covering lemma selects disjoint balls from an overlapping family without losing much; the Hardy–Littlewood maximal function controls all averages of a function at once. Covering and packing arguments of exactly this kind are how Riemannian geometry turns volume bounds into compactness (9B.2 Volume Comparison, 9B.4 Convergence of Manifolds).

By the end of this chapter you will be able to:

  • define convergence almost everywhere, in L1L^1, in measure and almost uniformly, and decide which imply which;
  • give the typewriter sequence and the other standard counterexamples, and extract an a.e. convergent subsequence from an L1L^1 convergent one;
  • state Egorov's and Lusin's theorems and say what they mean;
  • prove the Vitali covering lemma and the Hardy–Littlewood maximal inequality;
  • prove and use the Lebesgue differentiation theorem.

Averages over shrinking windows

In the world In use Moving averages recover a signal almost everywhere

A moving average replaces a signal f(t)f(t) by its average over a window around each time,

Ahf(t)=12h∫t−ht+hf(s) ds.A_hf(t) = \frac{1}{2h}\int_{t-h}^{t+h}f(s)\,ds.

It is the simplest smoothing filter in signal processing, and the standard way of reading trends in noisy data such as daily temperatures or prices. Shrinking the window should give back the signal, and where ff is continuous it obviously does. But a recorded signal need not be continuous: it may have jumps, or be merely integrable.

The Lebesgue differentiation theorem says that Ahf(t)→f(t)A_hf(t) \to f(t) as h→0h \to 0 for almost every tt, for every locally integrable ff, however rough. Where it fails is instructive: at a jump from aa to bb, a symmetric window converges to the midpoint a+b2\frac{a+b}2 (Figure 6.1), not to either value. But jumps (of a function of bounded variation, say) occur at only countably many times, a null set. Averaging loses nothing that integration can see.

Figure 6.1. Moving averages of a unit step for window half-widths h=0.4,0.2,0.05h = 0.4, 0.2, 0.05. Away from the jump they agree with the step once hh is small; at the jump they all equal 12\tfrac12, the midpoint.

Modes of convergence

Let fn,ff_n, f be measurable functions on a measure space (X,μ)(X, \mu).

Definition 6.1 Modes of convergence
  1. Uniformly: sup⁡x∣fn(x)−f(x)∣→0\sup_x|f_n(x) - f(x)| \to 0.
  2. Almost everywhere: fn(x)→f(x)f_n(x) \to f(x) for all xx outside a null set.
  3. In L1L^1 (in mean): ∫∣fn−f∣ dμ→0\int|f_n - f|\,d\mu \to 0.
  4. In measure: for every ε>0\varepsilon > 0, μ({∣fn−f∣≥ε})→0\mu(\{|f_n - f| \geq \varepsilon\}) \to 0.
  5. Almost uniformly: for every δ>0\delta > 0 there is a set EE with μ(E)<δ\mu(E) < \delta outside which fn→ff_n \to f uniformly.

Some implications are immediate. Uniform convergence implies the others (on a space of finite measure, for L1L^1). Convergence in L1L^1 implies convergence in measure, by Markov's inequality: μ({∣fn−f∣≥ε})≤1ε∫∣fn−f∣\mu(\{|f_n - f| \geq \varepsilon\}) \leq \frac1\varepsilon\int|f_n - f| (3A.3 The Lebesgue Integral). Almost uniform implies almost everywhere (intersect over δ=1k\delta = \frac1k).

The failures are given by the escapes of 3A.3 The Lebesgue Integral, plus one new example.

Example 6.2 The typewriter sequence

On [0,1][0, 1], list the indicator functions of the dyadic intervals, row by row: 1[0,1]1_{[0,1]}; then 1[0,1/2],1[1/2,1]1_{[0, 1/2]}, 1_{[1/2, 1]}; then 1[0,1/4],1[1/4,1/2],1[1/2,3/4],1[3/4,1]1_{[0, 1/4]}, 1_{[1/4, 1/2]}, 1_{[1/2, 3/4]}, 1_{[3/4, 1]}; and so on (Figure 6.2). Like a typewriter carriage, each row sweeps across the whole interval with a narrower block. Then ∫fn→0\int f_n \to 0, since the blocks shrink: the sequence converges to 00 in L1L^1 and in measure. But every point xx is covered by one block in every row, so fn(x)=1f_n(x) = 1 infinitely often and fn(x)=0f_n(x) = 0 infinitely often: the sequence converges at no point at all.

Figure 6.2. The typewriter sequence: rows of narrower and narrower blocks, each row sweeping across [0,1][0, 1]. The integrals tend to 00, but any vertical line (dashed) meets a block in every row, so the functions are 11 at that point infinitely often.

Conversely, the escapes of 3A.3 The Lebesgue Integral converge to 00 almost everywhere but not in L1L^1: horizontal escape on an infinite measure space; vertical escape, n 1[1/n,2/n]n\,1_{[1/n, 2/n]}, even on [0,1][0, 1]. The tall spike converges to 00 in measure but not in L1L^1.

The typewriter's problem is that its blocks keep returning. Passing to a subsequence removes that.

Proposition 6.3 Fast convergence and subsequences
  1. If ∑n∫∣fn−f∣ dμ<∞\sum_n\int|f_n - f|\,d\mu < \infty, then fn→ff_n \to f almost everywhere.
  2. If fn→ff_n \to f in L1L^1 (or merely in measure), then a subsequence converges to ff almost everywhere.

Proof. (1) By monotone convergence, ∫∑n∣fn−f∣=∑n∫∣fn−f∣<∞\int\sum_n|f_n - f| = \sum_n\int|f_n - f| < \infty, so ∑n∣fn(x)−f(x)∣<∞\sum_n|f_n(x) - f(x)| < \infty for almost every xx (an integrable function is finite a.e.), and then the terms tend to 00. (2) In measure, choose nkn_k increasing with μ({∣fnk−f∣≥2−k})≤2−k\mu(\{|f_{n_k} - f| \geq 2^{-k}\}) \leq 2^{-k}. By Borel–Cantelli (3A.2 Lebesgue Measure), almost every xx lies in only finitely many of these sets, and so ∣fnk(x)−f(x)∣<2−k|f_{n_k}(x) - f(x)| < 2^{-k} for all large kk.

Part 2 is used constantly: whenever a limit is known to exist in an integral sense, some subsequence also converges pointwise almost everywhere, and then pointwise arguments (Fatou, continuity) can be applied to it.

Figure 6.3. How the modes of convergence relate. Solid arrows always hold (on a finite measure space for the arrows out of "uniform"). Dashed arrows need the stated hypothesis: Egorov's theorem on finite measure spaces, a subsequence, or domination. Each missing arrow has a counterexample: the typewriter (no a.e. convergence) or one of the three escapes (no L1L^1 convergence).

Littlewood's second and third principles

Two theorems quantify how close the measure-theoretic world is to the continuous one. Both are in Tao, §1.3.5.

Theorem 6.4 Egorov's theorem

On a space of finite measure, convergence almost everywhere implies almost uniform convergence.

Theorem 6.5 Lusin's theorem

Let f:Rd→Cf : \mathbb{R}^d \to \mathbb{C} be measurable and finite almost everywhere. For every ε>0\varepsilon > 0 there is a set EE with m(E)<εm(E) < \varepsilon such that the restriction of ff to Rd∖E\mathbb{R}^d \setminus E is continuous.

Egorov is Littlewood's third principle (every convergent sequence of measurable functions is nearly uniformly convergent), Lusin his second (every measurable function is nearly continuous). Note what "nearly" means in Lusin: the restriction is continuous, not ff itself. Dirichlet's function is discontinuous everywhere, but restricted to the irrationals (removing a null set) it is the constant 00.

Covering and the maximal function

To prove that averages converge almost everywhere, we need to control the worst average around each point, uniformly over all radii. Define, for f∈L1(Rd)f \in L^1(\mathbb{R}^d), the average of ∣f∣|f| over the ball B(x,r)B(x, r) and its supremum:

⟨∣f∣⟩B(x,r)=1m(B(x,r))∫B(x,r)∣f∣ dm,Mf(x)=sup⁡r>0⟨∣f∣⟩B(x,r).\langle|f|\rangle_{B(x,r)} = \frac{1}{m(B(x, r))}\int_{B(x,r)}|f|\,dm, \qquad Mf(x) = \sup_{r > 0}\langle|f|\rangle_{B(x,r)}.

MfMf is the Hardy–Littlewood maximal function. It can be large, but not often: the set where it exceeds λ\lambda has measure at most C∥f∥1/λC\|f\|_1/\lambda, which is what Markov's inequality would give for ff itself. The proof needs a way of choosing, from many overlapping balls, a disjoint family that still accounts for most of the space they cover.

Lemma 6.6 Vitali covering lemma

Let B1,…,BnB_1, \ldots, B_n be finitely many balls in Rd\mathbb{R}^d. There is a subcollection Bi1,…,BikB_{i_1}, \ldots, B_{i_k} of pairwise disjoint balls such that

⋃j=1nBj⊆⋃l=1k3Bil,\bigcup_{j=1}^nB_j \subseteq \bigcup_{l=1}^k3B_{i_l},

where 3B3B is the ball with the same centre and three times the radius. Consequently m(⋃jBj)≤3d∑lm(Bil)m\big(\bigcup_jB_j\big) \leq 3^d\sum_lm(B_{i_l}).

Proof. The greedy algorithm. Choose the largest ball. Discard every ball that meets it. From what remains, choose the largest, discard every ball that meets it, and repeat until nothing remains. The chosen balls are disjoint by construction. Every discarded ball BB met a chosen ball B′B' that was chosen while BB was still available, so B′B' is at least as large as BB: radius r(B)≤r(B′)r(B) \leq r(B'). Two balls that meet, the first no larger than the second, satisfy B⊆3B′B \subseteq 3B': any point of BB is within 2r(B)≤2r(B′)2r(B) \leq 2r(B') of a point of B′B', hence within 3r(B′)3r(B') of the centre of B′B'. The volume bound follows from m(3B)=3dm(B)m(3B) = 3^dm(B).

Figure 6.4. The Vitali covering lemma. From overlapping balls, the greedy algorithm picks disjoint ones (shaded), largest first. Tripling their radii (dashed) covers every original ball. So the union has volume at most 3d3^d times the total volume of the chosen disjoint balls.
Theorem 6.7 Hardy–Littlewood maximal inequality

For f∈L1(Rd)f \in L^1(\mathbb{R}^d) and λ>0\lambda > 0,

m({x:Mf(x)>λ})≤3dλ∫Rd∣f∣ dm.m(\{x : Mf(x) > \lambda\}) \leq \frac{3^d}{\lambda}\int_{\mathbb{R}^d}|f|\,dm.

Proof. Let E={Mf>λ}E = \{Mf > \lambda\} (it is open, so measurable), and let K⊆EK \subseteq E be compact. Each x∈Kx \in K has a ball BxB_x centred at xx with ∫Bx∣f∣>λ m(Bx)\int_{B_x}|f| > \lambda\,m(B_x). The open balls BxB_x cover KK, so finitely many do (2B.3 Compactness). Apply Vitali to those: disjoint balls B1,…,BkB_1, \ldots, B_k among them with

m(K)≤3d∑lm(Bl)<3dλ∑l∫Bl∣f∣≤3dλ∫∣f∣,m(K) \leq 3^d\sum_lm(B_l) < \frac{3^d}{\lambda}\sum_l\int_{B_l}|f| \leq \frac{3^d}{\lambda}\int|f|,

the last step because the BlB_l are disjoint. By inner regularity (3A.2 Lebesgue Measure), m(E)m(E) is the supremum of m(K)m(K) over compact K⊆EK \subseteq E.

An estimate of the form m({∣g∣>λ})≤C∥f∥1/λm(\{|g| > \lambda\}) \leq C\|f\|_1/\lambda is called weak type (1, 1). The maximal function is not bounded in L1L^1 (for f=1B(0,1)f = 1_{B(0,1)}, Mf(x)≈∣x∣−dMf(x) \approx |x|^{-d} far away, which is not integrable), so this weaker estimate is the right one.

Lebesgue's differentiation theorem

Theorem 6.8 Lebesgue differentiation theorem

Let ff be locally integrable on Rd\mathbb{R}^d. Then for almost every xx,

lim⁡r→01m(B(x,r))∫B(x,r)∣f(y)−f(x)∣ dy=0,in particularlim⁡r→0⟨f⟩B(x,r)=f(x).\lim_{r\to0}\frac{1}{m(B(x, r))}\int_{B(x,r)}|f(y) - f(x)|\,dy = 0, \qquad\text{in particular}\qquad \lim_{r\to0}\langle f\rangle_{B(x,r)} = f(x).

Points where the first limit holds are called Lebesgue points of ff.

Proof. The question is local, so we may assume f∈L1f \in L^1 (multiply by the indicator of a large ball). For λ>0\lambda > 0 let EλE_\lambda be the set of xx where lim sup⁡r→0⟨∣f−f(x)∣⟩B(x,r)>2λ\limsup_{r\to0}\langle|f - f(x)|\rangle_{B(x,r)} > 2\lambda. It suffices to show m∗(Eλ)=0m^*(E_\lambda) = 0 for every λ\lambda.

Split ff into a nice part and a small part. Given ε>0\varepsilon > 0, choose a continuous compactly supported gg with ∥f−g∥1≤ε\|f - g\|_1 \leq \varepsilon (3A.3 The Lebesgue Integral), and write h=f−gh = f - g. For the continuous part the limit is 00 at every point. So for x∈Eλx \in E_\lambda,

2λ<lim sup⁡r→0⟨∣h−h(x)∣⟩B(x,r)≤Mh(x)+∣h(x)∣,2\lambda < \limsup_{r\to0}\langle|h - h(x)|\rangle_{B(x,r)} \leq Mh(x) + |h(x)|,

and therefore either Mh(x)>λMh(x) > \lambda or ∣h(x)∣>λ|h(x)| > \lambda.

Both happen on small sets. By the maximal inequality and Markov's inequality,

m∗(Eλ)≤m({Mh>λ})+m({∣h∣>λ})≤3dλε+1λε.m^*(E_\lambda) \leq m(\{Mh > \lambda\}) + m(\{|h| > \lambda\}) \leq \frac{3^d}{\lambda}\varepsilon + \frac1\lambda\varepsilon.

Since ε\varepsilon is arbitrary, m∗(Eλ)=0m^*(E_\lambda) = 0.

The proof is a template worth naming: a maximal inequality plus a dense class gives almost-everywhere convergence. The dense class (continuous functions) is where the convergence is easy; the maximal inequality controls the error uniformly in rr, which is what lets the approximation pass to the limit. The same template proves almost-everywhere convergence of heat-kernel averages f∗Ht→ff * H_t \to f as t→0t \to 0 (6A.3 The Heat Equation on ℝⁿ), and of many other approximate identities.

Two consequences:

  • Density points. Applying the theorem to f=1Ef = 1_E: almost every point of a measurable set EE is a point of density 1, m(E∩B(x,r))/m(B(x,r))→1m(E \cap B(x, r))/m(B(x, r)) \to 1, and almost every point outside has density 00. A measurable set looks, under a microscope at almost every point, either completely full or completely empty. (A fat Cantor set, 3A.2 Lebesgue Measure, has density 11 at almost all of its points even though it contains no interval.)
  • The fundamental theorem of calculus for Lebesgue integrals. In one dimension, if f∈L1([a,b])f \in L^1([a, b]) and F(x)=∫axfF(x) = \int_a^xf, then F′(x)=f(x)F'(x) = f(x) for almost every xx: the difference quotient F(x+h)−F(x)h\frac{F(x + h) - F(x)}{h} is a one-sided average of ff, and one-sided averages converge at Lebesgue points too.
Example 6.9 The Cantor function

The converse direction of the fundamental theorem needs care. The Cantor function (the "devil's staircase") c:[0,1]→[0,1]c : [0, 1] \to [0, 1] is continuous and increasing, constant on each interval removed in building the Cantor set CC (3A.2 Lebesgue Measure), and climbs from 00 to 11 entirely on CC. So c′=0c' = 0 at every point outside CC, which is almost every point, and yet c(1)−c(0)=1≠0=∫01c′c(1) - c(0) = 1 \neq 0 = \int_0^1c'. Functions for which F(b)−F(a)=∫abF′F(b) - F(a) = \int_a^bF' always holds are exactly the absolutely continuous ones (Tao, §1.6.4); Lipschitz functions are among them, which is the case used most in geometry.

Figure 6.5. The Cantor function, computed to six levels. It is flat on every removed interval (total length 11), so its derivative is 00 almost everywhere, yet it rises from 00 to 11. Integrating the derivative does not recover it: it is continuous but not absolutely continuous.
Where this goes Covering and packing in geometry

The greedy selection in Vitali's lemma has a twin that runs through Riemannian geometry: choose a maximal collection of disjoint balls of radius ε\varepsilon; then the balls of radius 2ε2\varepsilon with the same centres cover everything (Exercise 6.16). Counting how many disjoint small balls fit inside a big one is a volume comparison. On a manifold with Ricci curvature bounded below, Bishop–Gromov volume comparison (9B.2 Volume Comparison) gives exactly that count, independently of the manifold, and the resulting uniform bound on covering numbers is the total boundedness that Gromov's compactness theorem needs (2B.3 Compactness, 9B.4 Convergence of Manifolds). The same packing argument, with volumes controlled by noncollapsing, appears when Perelman's κ-solutions are shown to form a compact family (12B.2 The Structure of κ-Solutions).

History

Lebesgue proved the one-dimensional differentiation theorem in 1904 and the version for averages in several dimensions in 1910. Giuseppe Vitali's covering theorem appeared in 1908. Dmitri Egorov proved his theorem in 1911, and his student Nikolai Lusin his in 1912. G. H. Hardy and J. E. Littlewood introduced the maximal function in 1930, in a paper motivated by questions about averages of sequences, which they illustrated with a cricketer's batting averages. Cantor described his function in 1884, and Lebesgue used it as a counterexample to the fundamental theorem of calculus for continuous functions.

Recall Where we stand

Functions can converge uniformly, almost everywhere, in L1L^1, in measure or almost uniformly. L1L^1 convergence implies convergence in measure, and a subsequence then converges almost everywhere; the typewriter shows that the full sequence need not converge anywhere. Egorov and Lusin make "nearly uniform" and "nearly continuous" precise. The Vitali covering lemma gives the Hardy–Littlewood maximal inequality, and a maximal inequality plus a dense class gives the Lebesgue differentiation theorem: averages over shrinking balls recover an integrable function almost everywhere. 3A.7 Lᵖ Spaces and Jensen’s Inequality introduces the LpL^p spaces, proves the inequalities of Hölder, Minkowski and Jensen, and shows the spaces are complete.

Exercises

Exercise 6.10 Sorting the examples

For each sequence on [0,1][0, 1], decide whether it converges to 00 uniformly, almost everywhere, in L1L^1, and in measure: (a) xnx^n; (b) n 1[0,1/n]n\,1_{[0, 1/n]}; (c) n 1[0,1/n]\sqrt n\,1_{[0, 1/n]}; (d) the typewriter sequence; (e) 1[0,1/n]1_{[0, 1/n]}.

Solution

(a) a.e., L1L^1, in measure; not uniformly. (b) a.e. (except at 00), in measure; not L1L^1 (∫=1\int = 1), not uniform. (c) a.e., in measure, and in L1L^1 (∫=n−1/2\int = n^{-1/2}); not uniform. (d) L1L^1 and in measure only. (e) a.e., L1L^1, in measure; not uniform.

Exercise 6.11 A good subsequence of the typewriter

Find a subsequence of the typewriter sequence that converges to 00 almost everywhere, and identify the exceptional set.

Solution

Take the first function in each row, 1[0,2−k]1_{[0, 2^{-k}]}. It converges to 00 at every x>0x > 0; the exceptional set is {0}\{0\}.

Exercise 6.12 The maximal function isn't integrable

Let f=1[−1,1]f = 1_{[-1, 1]} on R\mathbb{R}. Compute Mf(x)Mf(x) for ∣x∣>1|x| > 1 and show it is comparable to 1∣x∣\frac{1}{|x|}. Deduce Mf∉L1Mf \notin L^1, and check the weak-type bound m({Mf>λ})≤3⋅2λm(\{Mf > \lambda\}) \leq \frac{3\cdot2}{\lambda} directly.

Solution

For x>1x > 1 and the interval of radius rr centred at xx: if r≤x−1r \leq x - 1 it misses [−1,1][-1, 1]; if x−1<r≤x+1x - 1 < r \leq x + 1 the average is r−(x−1)2r\frac{r - (x - 1)}{2r}, increasing in rr; if r≥x+1r \geq x + 1 it is 22r\frac{2}{2r}, decreasing. So the best radius is r=x+1r = x + 1 and Mf(x)=1x+1Mf(x) = \frac{1}{x + 1} (similarly for x<−1x < -1), which is not integrable. For λ<1\lambda < 1, {Mf>λ}\{Mf > \lambda\} is the interval ∣x∣<1λ−1|x| < \frac1\lambda - 1, of length less than 2λ\frac2\lambda, within the bound 6λ\frac6\lambda.

Exercise 6.13 Density points

Let E⊆RE \subseteq \mathbb{R} be measurable with m(E∩I)≤12m(I)m(E \cap I) \leq \frac12 m(I) for every interval II. Show that m(E)=0m(E) = 0. (Use density points.) Can a measurable set meet every interval in exactly half its length?

Solution

If m(E)>0m(E) > 0, almost every point of EE has density 11, so for small intervals around such a point m(E∩I)>12m(I)m(E \cap I) > \frac12m(I). Contradiction. So no: such a set would have m(E)>0m(E) > 0 and also violate the density theorem at its density points.

Exercise 6.14 The Cantor function

(a) Define cc on CC by writing x∈Cx \in C in ternary with digits 0,20, 2, halving each digit, and reading the result in binary; extend to [0,1][0, 1] by making it constant on each removed interval. Show cc is continuous and increasing. (b) Show that cc maps CC, a null set, onto [0,1][0, 1]. So a continuous function can map a set of measure zero onto a set of positive measure, which a Lipschitz function can't (Exercise 6.15).

Exercise 6.15 Lipschitz maps preserve null sets

Let ϕ:Rd→Rd\phi : \mathbb{R}^d \to \mathbb{R}^d be Lipschitz with constant LL. Show that if m(N)=0m(N) = 0 then m(ϕ(N))=0m(\phi(N)) = 0. (A ball of radius rr maps into a ball of radius LrLr; cover NN by small cubes.) This is why the change-of-variables formula (3A.5 Product Measures and Change of Variables) and Sard's theorem (7A.7 Smooth Topology) can ignore null sets.

Exercise 6.16 Rehearsal: packing and covering numbers

Let XX be a subset of Rd\mathbb{R}^d, and let x1,…,xNx_1, \ldots, x_N be a maximal set of points of XX with ∣xi−xj∣≥ε|x_i - x_j| \geq \varepsilon for i≠ji \neq j (no more points of XX can be added). (a) Show that the balls B(xi,ε)B(x_i, \varepsilon) cover XX, so the points form an ε\varepsilon-net (2B.3 Compactness). (b) Show that the balls B(xi,ε/2)B(x_i, \varepsilon/2) are disjoint. (c) If X⊆B(0,R)X \subseteq B(0, R), deduce N≤(2R+εε)dN \leq \big(\frac{2R + \varepsilon}{\varepsilon}\big)^d by comparing volumes.

In 9B.2 Volume Comparison the volume comparison in (c) is replaced by Bishop–Gromov: on a manifold with Ric≥−(n−1)\mathrm{Ric} \geq -(n-1), the ratio vol B(x,R)/vol B(xi,ε/2)\mathrm{vol}\,B(x, R)/\mathrm{vol}\,B(x_i, \varepsilon/2) is bounded by a constant depending only on nn, RR and ε\varepsilon. That gives a bound on NN independent of the manifold, the uniform total boundedness behind Gromov's compactness theorem (9B.4 Convergence of Manifolds).

Solution

(a) If some x∈Xx \in X were at distance ≥ε\geq \varepsilon from all xix_i, it could be added, contradicting maximality. (b) If B(xi,ε/2)B(x_i, \varepsilon/2) and B(xj,ε/2)B(x_j, \varepsilon/2) met, ∣xi−xj∣<ε|x_i - x_j| < \varepsilon. (c) The NN disjoint balls of radius ε/2\varepsilon/2 lie in B(0,R+ε/2)B(0, R + \varepsilon/2), so Nωd(ε/2)d≤ωd(R+ε/2)dN\omega_d(\varepsilon/2)^d \leq \omega_d(R + \varepsilon/2)^d.

© 2026 NeckPinch (www.neckpinch.com). All content in the guidebook (text, mathematics, figures and exercises) is protected by copyright. All rights reserved. No part may be copied, republished or redistributed without written permission.