Book 2B

© 2026 NeckPinch · www.neckpinch.com · All rights reserved.

Course 2Book 2B: Spaces, Functions and ChangeChapter 5

Uniform Convergence and Arzelà–Ascoli

Spaces of functions, Weierstrass approximation, and the compactness theorem behind every later compactness theorem.

38 min read · Updated Oct 2, 2026

Read with Tao, Analysis II, chapter "Uniform convergence", all sections, from "Limiting values of functions" to "Uniform approximation by polynomials". Tao does not cover the Arzelà–Ascoli theorem; this chapter proves it in full.

In this chapter · 8 sections
  1. 5.1How a computer evaluates sin⁡x\sin xsinx
  2. 5.2Pointwise and uniform convergence
  3. 5.3Uniform limits are continuous
  4. 5.4Limits, integrals and derivatives
  5. 5.4.1Series of functions
  6. 5.5The Weierstrass approximation theorem
  7. 5.6Equicontinuity and the Arzelà–Ascoli theorem
  8. 5.7History
  9. 5.8Exercises

So far the points of our metric spaces have mostly been points: numbers, places, strings. This chapter is about spaces whose points are functions, and about what it means for a sequence of functions to converge. There are two natural answers. Pointwise convergence asks that fn(x)→f(x)f_n(x) \to f(x) at each xx separately. Uniform convergence asks that the whole graph of fnf_n come within ε\varepsilon of the graph of ff, at every point at once. The difference matters, because only the second preserves continuity, integrals and (with care) derivatives.

The chapter has two summits. The first is the Weierstrass approximation theorem: every continuous function on a closed interval is a uniform limit of polynomials. Its proof introduces approximate identities, which become mollifiers in Course 3 and the heat kernel in Course 6. The second is the Arzelà–Ascoli theorem, which says when a family of functions is compact: when it is bounded and equicontinuous. This is the compactness theorem behind every later compactness theorem in geometry, from Cheeger–Gromov to Hamilton's compactness theorem for Ricci flows. Its slogan is worth memorising now: derivative bounds give equicontinuity, and equicontinuity gives convergence.

By the end of this chapter you will be able to:

  • tell pointwise from uniform convergence, and prove or disprove each for a given sequence;
  • prove that uniform limits of continuous functions are continuous, and that C(K)C(K) with the sup metric is complete;
  • say when limits can be exchanged with integrals and with derivatives, and use the Weierstrass M-test;
  • prove the Weierstrass approximation theorem by convolution with a polynomial approximate identity;
  • prove the Arzelà–Ascoli theorem, and use it to extract convergent subsequences from families with derivative bounds.

How a computer evaluates sin⁡x\sin x

In the world In use A polynomial that is never wrong by more than 2−582^{-58}

A computer can only add and multiply, so every value of sin⁡x\sin x it returns is computed from a polynomial, or a ratio of polynomials. The widely copied fdlibm library, written at Sun Microsystems in the early 1990s (and the reference for Java's StrictMath class), works in two steps. First it uses the periodicity and symmetries of sin⁡\sin to reduce any argument to a small interval, essentially [−π4,π4][-\tfrac\pi4, \tfrac\pi4]. Then it evaluates a polynomial. Its source code states the guarantee: on [0,π4][0, \tfrac\pi4], sin⁡x\sin x is approximated by an odd polynomial of degree 1313, x+S1x3+⋯+S6x13x + S_1x^3 + \cdots + S_6x^{13}, with

∣sin⁡xx−(1+S1x2+⋯+S6x12)∣≤2−58for every x in the interval.\left|\frac{\sin x}{x} - \big(1 + S_1x^2 + \cdots + S_6x^{12}\big)\right| \leq 2^{-58} \quad \text{for every } x \text{ in the interval.}

The important words are "for every xx". A library is called with arguments nobody can predict, so a small average error is useless: it must be accurate at the worst point. That is a bound in the sup metric of 2B.1 Metric Spaces: a guarantee about the worst case, not the average. Three facts from this chapter are behind this. Polynomials can approximate any continuous function uniformly (the Weierstrass theorem below). How fast they do so depends on how smooth the function is. And for a function as smooth as sin⁡\sin, degree 1313 is enough to reach the limits of double-precision arithmetic.

Pointwise and uniform convergence

Let XX be a set (usually a metric space) and fn,f:X→Rf_n, f : X \to \mathbb{R}.

Definition 5.1 Pointwise and uniform convergence

fn→ff_n \to f pointwise if fn(x)→f(x)f_n(x) \to f(x) for each x∈Xx \in X: for every xx and every ε>0\varepsilon > 0 there is NN with ∣fn(x)−f(x)∣≤ε|f_n(x) - f(x)| \leq \varepsilon for n≥Nn \geq N.

fn→ff_n \to f uniformly if for every ε>0\varepsilon > 0 there is NN with ∣fn(x)−f(x)∣≤ε|f_n(x) - f(x)| \leq \varepsilon for all n≥Nn \geq N and all x∈Xx \in X. Equivalently, sup⁡x∈X∣fn(x)−f(x)∣→0\sup_{x \in X}|f_n(x) - f(x)| \to 0.

The two definitions differ only in the order of quantifiers (2A.5 Quantifiers and the Shape of a Proof): in pointwise convergence, NN may depend on xx; in uniform convergence, one NN serves every xx. Geometrically, uniform convergence says that for n≥Nn \geq N the graph of fnf_n lies inside the ε-tube around the graph of ff (Figure 5.1). For bounded functions, it is exactly convergence in the sup metric d∞d_\infty.

Figure 5.1. Uniform convergence: from some NN on, every graph fnf_n lies inside the tube of half-width ε\varepsilon around ff, over the whole interval.
In the world Model Smooth signals with a jumpy limit

On [0,1][0, 1], let fn(x)=xnf_n(x) = x^n. For 0≤x<10 \leq x < 1, xn→0x^n \to 0; at x=1x = 1, xn=1x^n = 1. So fnf_n converges pointwise to the function that is 00 on [0,1)[0, 1) and 11 at 11, which is discontinuous (Figure 5.2). The convergence is not uniform: for every nn, xn=12x^n = \tfrac12 at x=2−1/nx = 2^{-1/n}, which is a point where fnf_n is 12\tfrac12 away from the limit. The closer xx is to 11, the longer we must wait for xnx^n to be small, and no single NN works for all xx.

The practical moral is that a sequence of perfectly smooth signals, each of which a sensor could record, can converge at every instant to a signal with a jump. If an algorithm relies on properties of the limit (continuity, a bound on its slope), pointwise convergence of the approximations is not enough to guarantee them.

Figure 5.2. xnx^n for n=1,2,5,20,100n = 1, 2, 5, 20, 100. Each is continuous; the pointwise limit (heavy line, with the dot at x=1x = 1) is not. The convergence fails to be uniform near x=1x = 1.

Uniform limits are continuous

Theorem 5.2 Uniform limits of continuous functions

Let XX be a metric space. If each fn:X→Rf_n : X \to \mathbb{R} is continuous at x0x_0 and fn→ff_n \to f uniformly on XX, then ff is continuous at x0x_0. In particular a uniform limit of continuous functions is continuous.

Proof. The ε/3 argument. Given ε>0\varepsilon > 0, choose nn with ∣fn(x)−f(x)∣≤ε/3|f_n(x) - f(x)| \leq \varepsilon/3 for all xx. Then choose δ>0\delta > 0 with ∣fn(x)−fn(x0)∣≤ε/3|f_n(x) - f_n(x_0)| \leq \varepsilon/3 for d(x,x0)<δd(x, x_0) < \delta. For such xx,

∣f(x)−f(x0)∣≤∣f(x)−fn(x)∣+∣fn(x)−fn(x0)∣+∣fn(x0)−f(x0)∣≤ε.|f(x) - f(x_0)| \leq |f(x) - f_n(x)| + |f_n(x) - f_n(x_0)| + |f_n(x_0) - f(x_0)| \leq \varepsilon.

The proof needs uniformity in the first step: the same nn must make ∣fn−f∣|f_n - f| small at xx and at x0x_0, and xx is not known when nn is chosen. With only pointwise convergence the argument collapses, and xnx^n shows the conclusion can fail.

Now the completeness promised in 2B.2 Completeness and Contraction.

Corollary 5.3 C(K)C(K) is complete

For a metric space XX, the space B(X)B(X) of bounded functions and the space Cb(X)C_b(X) of bounded continuous functions are complete with the sup metric. In particular, for compact KK, the space C(K)C(K) of continuous functions with d∞d_\infty is complete.

Proof. Let (fn)(f_n) be Cauchy in d∞d_\infty. For each xx, ∣fm(x)−fn(x)∣≤d∞(fm,fn)|f_m(x) - f_n(x)| \leq d_\infty(f_m, f_n), so (fn(x))(f_n(x)) is Cauchy in R\mathbb{R} and converges to some f(x)f(x). Given ε\varepsilon, pick NN with d∞(fm,fn)≤εd_\infty(f_m, f_n) \leq \varepsilon for m,n≥Nm, n \geq N. Letting m→∞m \to \infty in ∣fm(x)−fn(x)∣≤ε|f_m(x) - f_n(x)| \leq \varepsilon gives ∣f(x)−fn(x)∣≤ε|f(x) - f_n(x)| \leq \varepsilon for all xx and n≥Nn \geq N. So ff is bounded and fn→ff_n \to f uniformly. If the fnf_n are continuous, so is ff, by the theorem. On a compact KK, every continuous function is bounded, so Cb(K)=C(K)C_b(K) = C(K).

This is the space in which ODEs are solved by contraction (2B.10 Ordinary Differential Equations), and its completeness is used every time a solution is built as a uniform limit of approximations.

Limits, integrals and derivatives

Integrals. Uniform convergence on a bounded interval lets limits pass through integrals.

Proposition 5.4 Uniform convergence and integration

If fnf_n are Riemann integrable on [a,b][a, b] and fn→ff_n \to f uniformly, then ff is Riemann integrable and ∫abfn→∫abf\int_a^b f_n \to \int_a^b f.

Proof. If ∣fn−f∣≤ε|f_n - f| \leq \varepsilon everywhere, then fn−ε≤f≤fn+εf_n - \varepsilon \leq f \leq f_n + \varepsilon, so the upper and lower integrals of ff lie within ε(b−a)\varepsilon(b - a) of ∫fn\int f_n, and of each other. Since ε\varepsilon is arbitrary, ff is integrable, and ∣∫fn−∫f∣≤ε(b−a)\big|\int f_n - \int f\big| \leq \varepsilon(b - a).

The escaping spikes of 2A.11 The Riemann Integral converge pointwise but not uniformly, and their integrals don't converge to the integral of the limit. Uniform convergence is far stronger than necessary, though: it fails for almost every sequence that matters in PDE. The convergence theorems of 3A.3 The Lebesgue Integral replace it with much weaker hypotheses.

Derivatives. Here uniform convergence of the functions is not enough. The functions fn(x)=sin⁡(n2x)nf_n(x) = \frac{\sin(n^2 x)}{n} converge uniformly to 00 (since ∣fn∣≤1n|f_n| \leq \tfrac1n), but fn′(x)=ncos⁡(n2x)f_n'(x) = n\cos(n^2 x) doesn't converge at all. Small wiggles can have large slopes. What is needed is uniform convergence of the derivatives.

Theorem 5.5 Uniform convergence and differentiation

Let fnf_n be continuously differentiable on [a,b][a, b]. Suppose fn′→gf_n' \to g uniformly on [a,b][a, b], and fn(x0)f_n(x_0) converges for some x0∈[a,b]x_0 \in [a, b]. Then fnf_n converges uniformly to a continuously differentiable function ff, and f′=gf' = g.

Proof. By the fundamental theorem of calculus (2A.11 The Riemann Integral), fn(x)=fn(x0)+∫x0xfn′f_n(x) = f_n(x_0) + \int_{x_0}^x f_n'. Let L=lim⁡fn(x0)L = \lim f_n(x_0) and define f(x)=L+∫x0xgf(x) = L + \int_{x_0}^x g; gg is continuous as a uniform limit of continuous functions. Then

∣fn(x)−f(x)∣≤∣fn(x0)−L∣+∣∫x0x(fn′−g)∣≤∣fn(x0)−L∣+(b−a)sup⁡∣fn′−g∣→0,|f_n(x) - f(x)| \leq |f_n(x_0) - L| + \Big|\int_{x_0}^x (f_n' - g)\Big| \leq |f_n(x_0) - L| + (b - a)\sup|f_n' - g| \to 0,

uniformly in xx. By the fundamental theorem again, f′=gf' = g.

The pattern "control the derivatives to control the limit" is the theme of this chapter, and it reaches its full form in the Arzelà–Ascoli theorem.

Series of functions

A series ∑fk\sum f_k of functions converges uniformly if its partial sums do. The standard test is a comparison with a series of numbers.

Proposition 5.6 Weierstrass M-test

Let fk:X→Rf_k : X \to \mathbb{R} satisfy sup⁡X∣fk∣≤Mk\sup_X|f_k| \leq M_k, where ∑Mk<∞\sum M_k < \infty. Then ∑fk\sum f_k converges uniformly (and absolutely at each point). If the fkf_k are continuous, so is the sum.

Proof. For m>nm > n, ∣∑k=n+1mfk(x)∣≤∑k=n+1mMk\big|\sum_{k=n+1}^m f_k(x)\big| \leq \sum_{k=n+1}^m M_k, which is small for large nn, uniformly in xx. So the partial sums are Cauchy in d∞d_\infty, and they converge by completeness (Corollary 5.3).

The M-test makes continuity of power series (2B.6 Power Series, Exponentials and Bump Functions) and of Fourier series with summable coefficients (2B.7 Fourier Series and the First Heat Equation) immediate. It also proves that a famous monster is continuous. Weierstrass's function W(x)=∑k=0∞akcos⁡(bkπx)W(x) = \sum_{k=0}^\infty a^k\cos(b^k\pi x), with 0<a<10 < a < 1, is continuous by the M-test with Mk=akM_k = a^k. Weierstrass showed in 1872 that for suitable aa and bb (for example a=12a = \tfrac12, b=13b = 13) it is differentiable nowhere. A uniform limit of smooth functions can be continuous and as rough as possible, another reminder that uniform convergence of functions says nothing about their derivatives.

The Weierstrass approximation theorem

Theorem 5.7 Weierstrass approximation theorem

Let ff be continuous on [a,b][a, b]. For every ε>0\varepsilon > 0 there is a polynomial PP with ∣f(x)−P(x)∣≤ε|f(x) - P(x)| \leq \varepsilon for all x∈[a,b]x \in [a, b].

In metric language: the polynomials are dense in C([a,b])C([a, b]) with the sup metric. The proof below is by convolution with an approximate identity, and its idea is more important than the theorem.

The idea. Replace f(x)f(x) by a weighted average of the values of ff near xx:

P(x)=∫f(x+t) Q(t) dt,P(x) = \int f(x + t)\,Q(t)\,dt,

where the weight Q≥0Q \geq 0 has total integral 11 and is concentrated near t=0t = 0. If QQ is concentrated enough, the average is close to f(x)f(x), because ff is (uniformly) continuous. And if QQ is a polynomial, the average is a polynomial in xx.

Proof. Reductions. By the change of variables x↦a+(b−a)xx \mapsto a + (b - a)x, we may take [a,b]=[0,1][a, b] = [0, 1]. Subtracting the linear function f(0)+(f(1)−f(0))xf(0) + (f(1) - f(0))x, which is a polynomial, we may assume f(0)=f(1)=0f(0) = f(1) = 0. Extend ff by 00 outside [0,1][0, 1]. The extension is continuous on R\mathbb{R}, and uniformly continuous since it is uniformly continuous on [−1,2][-1, 2] (2B.3 Compactness) and constant outside [0,1][0, 1]. Let M=sup⁡∣f∣M = \sup|f|.

The kernels. For n≥1n \geq 1 let Qn(t)=cn(1−t2)nQ_n(t) = c_n(1 - t^2)^n for ∣t∣≤1|t| \leq 1, and Qn(t)=0Q_n(t) = 0 otherwise, with cnc_n chosen so that ∫−11Qn=1\int_{-1}^1 Q_n = 1 (Figure 5.3). Bernoulli's inequality (1−t2)n≥1−nt2(1 - t^2)^n \geq 1 - nt^2 gives

∫−11(1−t2)n dt≥2∫01/n(1−nt2) dt=43n>1n,\int_{-1}^1 (1 - t^2)^n\,dt \geq 2\int_0^{1/\sqrt n}(1 - nt^2)\,dt = \frac{4}{3\sqrt n} > \frac{1}{\sqrt n},

so cn<nc_n < \sqrt n. Hence, for any δ∈(0,1)\delta \in (0, 1), on δ≤∣t∣≤1\delta \leq |t| \leq 1,

Qn(t)≤n (1−δ2)n→0(n→∞).Q_n(t) \leq \sqrt n\,(1 - \delta^2)^n \to 0 \quad (n \to \infty).

This is the sense in which the kernels concentrate at 00: their total mass is 11, and the mass outside any fixed neighbourhood of 00 tends to 00.

The approximants are polynomials. For x∈[0,1]x \in [0, 1] let Pn(x)=∫−11f(x+t) Qn(t) dtP_n(x) = \int_{-1}^1 f(x + t)\,Q_n(t)\,dt. Since ff vanishes outside [0,1][0, 1], substituting s=x+ts = x + t gives Pn(x)=∫01f(s) Qn(s−x) dsP_n(x) = \int_0^1 f(s)\,Q_n(s - x)\,ds. Expanding Qn(s−x)=cn(1−(s−x)2)nQ_n(s - x) = c_n(1 - (s - x)^2)^n (valid since ∣s−x∣≤1|s - x| \leq 1) shows that Pn(x)P_n(x) is a polynomial in xx whose coefficients are integrals of ff against powers of ss.

The approximants are close. Given ε>0\varepsilon > 0, choose δ\delta by uniform continuity so that ∣f(x+t)−f(x)∣≤ε/2|f(x + t) - f(x)| \leq \varepsilon/2 for ∣t∣<δ|t| < \delta. Since ∫Qn=1\int Q_n = 1,

∣Pn(x)−f(x)∣=∣∫−11(f(x+t)−f(x))Qn(t) dt∣≤ε2∫∣t∣<δQn+2M∫δ≤∣t∣≤1Qn≤ε2+4Mn(1−δ2)n.|P_n(x) - f(x)| = \Big|\int_{-1}^1 \big(f(x + t) - f(x)\big)Q_n(t)\,dt\Big| \leq \frac\varepsilon2\int_{|t| < \delta} Q_n + 2M\int_{\delta \leq |t| \leq 1} Q_n \leq \frac\varepsilon2 + 4M\sqrt n(1 - \delta^2)^n.

The last term is less than ε/2\varepsilon/2 for large nn, independently of xx.

Figure 5.3. The polynomial kernels Qn(t)=cn(1−t2)nQ_n(t) = c_n(1 - t^2)^n for n=2,10,50n = 2, 10, 50. Each has area 11; as nn grows they concentrate at 00. Averaging ff against QnQ_n blurs ff less and less, and the blurred function is a polynomial.
Where this goes Approximate identities

A family of non-negative kernels with integral 11 that concentrate at a point is called an approximate identity, because convolving with it approximately does nothing. The proof above has three parts that will recur every time: mass 11, concentration, and uniform continuity of ff. The same three parts prove that mollifiers smooth functions without moving them far (3A.8 Convolution and Mollifiers); that Fejér's kernel recovers a continuous function from its Fourier series (2B.7 Fourier Series and the First Heat Equation); and that the heat kernel 14πte−x2/4t\frac{1}{\sqrt{4\pi t}}e^{-x^2/4t}, which is an approximate identity as t→0t \to 0, makes the heat equation attain its initial data (6A.3 The Heat Equation on ℝⁿ). On a Riemannian manifold the heat kernel plays the same role (9B.7 The Heat Equation on a Manifold), and Perelman's reduced volume (12A.5 Reduced Distance and Reduced Volume) is built from a heat-kernel-like weight that concentrates as time runs backwards.

In the world In use Bernstein polynomials and Bézier curves

In 1912 Sergei Bernstein gave a different, constructive proof of the Weierstrass theorem, with explicit polynomials:

Bnf(x)=∑k=0nf ⁣(kn)(nk)xk(1−x)n−k.B_n f(x) = \sum_{k=0}^n f\!\left(\tfrac kn\right)\binom nk x^k(1 - x)^{n-k}.

The weights (nk)xk(1−x)n−k\binom nk x^k(1-x)^{n-k} are the probabilities of kk successes in nn trials with success probability xx, and they concentrate near k/n≈xk/n \approx x as nn grows: another approximate identity, this time a discrete one (Exercise 5.16).

The same polynomials draw curves. Given control points p0,…,pnp_0, \dots, p_n in the plane, the curve t↦∑k(nk)tk(1−t)n−k pkt \mapsto \sum_k \binom nk t^k(1-t)^{n-k}\,p_k starts at p0p_0, ends at pnp_n, and is pulled towards the points in between. Paul de Casteljau developed this construction at Citroën in 1959, with a stable recursive algorithm for evaluating it, but the company kept it confidential; Pierre Bézier developed the same curves independently at Renault and published them in the 1960s, so they carry his name. Bézier curves are now the outlines of letters in digital fonts (quadratic ones in TrueType, cubic ones in PostScript Type 1 fonts) and the paths of vector-graphics programs.

Bernstein approximation is reliable but slow. For f(x)=∣x−12∣f(x) = |x - \tfrac12|, the worst error of BnfB_n f is about 0.190.19, 0.0980.098, 0.0500.050 and 0.0250.025 for n=4,16,64,256n = 4, 16, 64, 256 (Figure 5.4): quadrupling the degree only halves the error. The degree-1313 polynomial of the opening example does incomparably better because sin⁡\sin is smooth and its coefficients are optimised for the worst case.

Figure 5.4. Bernstein approximations of ∣x−12∣|x - \tfrac12| of degrees 44, 1616 and 6464. They converge uniformly, with the largest error at the corner, x=12x = \tfrac12, where it is about 1/2πn1/\sqrt{2\pi n}.

Equicontinuity and the Arzelà–Ascoli theorem

When does a sequence of functions have a uniformly convergent subsequence? In C(K)C(K), by 2B.3 Compactness, the question is when a set of functions is compact. Boundedness is not enough: xnx^n (2B.3 Compactness) and sin⁡(nx)\sin(nx) (Figure 5.5) are bounded by 11 and have no uniformly convergent subsequences. What goes wrong in both cases is that the functions get steeper and steeper. The condition that rules this out is the following.

Definition 5.8 Equicontinuity

A family F\mathcal{F} of functions from a metric space XX to R\mathbb{R} is equicontinuous if for every ε>0\varepsilon > 0 there is a δ>0\delta > 0 such that

d(x,y)<δ  ⟹  ∣f(x)−f(y)∣≤εfor all x,y∈X and all f∈F.d(x, y) < \delta \implies |f(x) - f(y)| \leq \varepsilon \quad \text{for all } x, y \in X \text{ and all } f \in \mathcal{F}.

The point is that one δ\delta works for all the functions in the family, as well as for all points. (Strictly, this is uniform equicontinuity; on a compact space the pointwise version, with δ\delta allowed to depend on xx, implies it, by the argument of 2B.3 Compactness.) The most important source of equicontinuity is a uniform derivative bound.

Example 5.9 Derivative bounds give equicontinuity

If every f∈Ff \in \mathcal{F} is differentiable on an interval with ∣f′∣≤L|f'| \leq L, then by the mean value theorem ∣f(x)−f(y)∣≤L∣x−y∣|f(x) - f(y)| \leq L|x - y| for every ff, and δ=ε/L\delta = \varepsilon / L works for the whole family. More generally, any family of functions with a common Lipschitz constant, or a common Hölder bound ∣f(x)−f(y)∣≤C∣x−y∣α|f(x) - f(y)| \leq C|x - y|^\alpha, is equicontinuous.

Theorem 5.10 Arzelà–Ascoli

Let KK be a compact metric space and (fn)(f_n) a sequence in C(K)C(K) that is uniformly bounded (∣fn(x)∣≤M|f_n(x)| \leq M for all nn and xx) and equicontinuous. Then (fn)(f_n) has a subsequence that converges uniformly on KK.

Proof. Step 1: a countable dense set. For each kk, KK has a finite 1k\tfrac1k-net (2B.3 Compactness). The union of these nets is a countable set D={q1,q2,q3,… }D = \{q_1, q_2, q_3, \dots\} that comes within 1k\tfrac1k of every point of KK, for every kk.

Step 2: convergence on DD, by the diagonal argument. The numbers fn(q1)f_n(q_1) lie in [−M,M][-M, M], so by Bolzano–Weierstrass some subsequence S1S_1 of (fn)(f_n) converges at q1q_1. The numbers f(q2)f(q_2), for ff in S1S_1, are bounded, so a further subsequence S2⊆S1S_2 \subseteq S_1 converges at q2q_2 as well. Continue: SjS_j converges at q1,…,qjq_1, \dots, q_j. Let gjg_j be the jj-th term of SjS_j. For each ii, the sequence (gj)j≥i(g_j)_{j \geq i} is a subsequence of SiS_i, so gj(qi)g_j(q_i) converges as j→∞j \to \infty. The diagonal subsequence (gj)(g_j) converges at every point of DD.

Step 3: equicontinuity spreads convergence from DD to all of KK, uniformly. Let ε>0\varepsilon > 0. Choose δ\delta from equicontinuity for ε/3\varepsilon/3, and kk with 1k<δ\tfrac1k < \delta. The 1k\tfrac1k-net from step 1 is a finite subset {p1,…,pm}\{p_1, \dots, p_m\} of DD. Since gjg_j converges at each of these finitely many points, there is NN with ∣gi(pl)−gj(pl)∣≤ε/3|g_i(p_l) - g_j(p_l)| \leq \varepsilon/3 for all i,j≥Ni, j \geq N and all ll. Now take any x∈Kx \in K, and a net point plp_l with d(x,pl)<1k<δd(x, p_l) < \tfrac1k < \delta. For i,j≥Ni, j \geq N,

∣gi(x)−gj(x)∣≤∣gi(x)−gi(pl)∣+∣gi(pl)−gj(pl)∣+∣gj(pl)−gj(x)∣≤ε.|g_i(x) - g_j(x)| \leq |g_i(x) - g_i(p_l)| + |g_i(p_l) - g_j(p_l)| + |g_j(p_l) - g_j(x)| \leq \varepsilon.

So (gj)(g_j) is Cauchy in the sup metric. By completeness of C(K)C(K) (Corollary 5.3), it converges uniformly.

The three steps are worth separating, because the same three steps prove every later compactness theorem. Step 1: the domain is compact, so finitely many points see everything at each resolution. Step 2: bounds at each point give convergence at countably many points, by a diagonal argument. Step 3: equicontinuity, which comes from a derivative bound, upgrades convergence at a dense set to uniform convergence.

The converse also holds: a subset of C(K)C(K) whose every sequence has a uniformly convergent subsequence is uniformly bounded and equicontinuous (Exercise 5.15). So, for subsets of C(K)C(K): precompact ⇔ bounded and equicontinuous.

Figure 5.5. Left: sin⁡(nx)\sin(nx) is bounded but not equicontinuous: its slopes grow like nn, and no subsequence converges uniformly. Right: functions with ∣f∣≤1|f| \leq 1 and ∣f′∣≤1|f'| \leq 1 form an equicontinuous family, and by Arzelà–Ascoli every sequence of them has a uniformly convergent subsequence.
Where this goes Derivative bounds, then compactness, then a limit

The slogan "derivative bounds give equicontinuity, and equicontinuity gives convergence" organises a remarkable amount of the route.

  • Existence for ODEs without uniqueness (Peano's theorem, 2B.10 Ordinary Differential Equations): approximate solutions have bounded derivatives, so a subsequence converges to a solution.
  • Compact embeddings (4A.10 Sobolev Embeddings and Critical Exponents): a bound on the derivative in an integral sense gives compactness in a weaker norm. This is Arzelà–Ascoli with integrals.
  • Cheeger–Gromov compactness (9B.4 Convergence of Manifolds): a sequence of Riemannian manifolds with bounded curvature (and all its derivatives) and injectivity radius bounded below has a subsequence converging, after choosing coordinates, to a smooth limit. The metrics are written as functions in coordinate charts, and the convergence comes from Arzelà–Ascoli applied to them and to all their derivatives, as in Exercise 5.17.
  • Hamilton's compactness theorem (11B.3 Compactness of Ricci Flows) does the same for sequences of Ricci flows, using Shi's estimates to turn a curvature bound into bounds on all derivatives of curvature.
  • κ-solutions (12B.2 The Structure of κ-Solutions): Perelman's classification of the possible blow-up limits of Ricci flow rests on the compactness of the space of κ-solutions, which again comes down to derivative bounds plus Arzelà–Ascoli.

History

In 1821 Augustin-Louis Cauchy stated, in his Cours d'analyse, that a convergent series of continuous functions has a continuous sum. In 1826 Niels Henrik Abel pointed out a counterexample, a Fourier series (sin⁡x−12sin⁡2x+13sin⁡3x−⋯\sin x - \tfrac12\sin 2x + \tfrac13\sin 3x - \cdots) converging to a function with jumps. The missing ingredient, uniform convergence, was identified in the 1840s, independently by Philipp Ludwig von Seidel and George Gabriel Stokes in 1847, and made central by Karl Weierstrass in his Berlin lectures. Weierstrass proved his approximation theorem in 1885, by convolution with the Gaussian heat kernel: the proof above is a polynomial version of his idea. Giulio Ascoli introduced equicontinuity in 1883–84 and proved that it suffices for compactness; Cesare Arzelà proved the converse and gave the first clear statement of the theorem in 1895. Bernstein's probabilistic proof appeared in 1912.

Recall Where we stand

Uniform convergence is convergence in the sup metric. It preserves continuity, so C(K)C(K) is complete; it passes through integrals; and it passes through derivatives only when the derivatives converge uniformly. The M-test gives uniform convergence of series. Polynomials are dense in C([a,b])C([a, b]), by convolution with an approximate identity, an idea that returns as mollifiers, Fejér's kernel and the heat kernel. A bounded, equicontinuous sequence on a compact space has a uniformly convergent subsequence (Arzelà–Ascoli), and equicontinuity comes from derivative bounds. 2B.6 Power Series, Exponentials and Bump Functions applies these tools to power series, and constructs the exponential function and the smooth bump functions used throughout geometry.

Exercises

Exercise 5.11 Pointwise or uniform?

Find the pointwise limit, and decide whether the convergence is uniform: (a) x1+nx2\frac{x}{1 + nx^2} on R\mathbb{R}; (b) nx(1−x)nnx(1 - x)^n on [0,1][0, 1]; (c) xn1+xn\frac{x^n}{1 + x^n} on [0,2][0, 2]; (d) sin⁡(nx)n\frac{\sin(nx)}{n} on R\mathbb{R}; (e) xn(1−x)x^n(1 - x) on [0,1][0, 1].

Solution

(a) 00, uniformly: ∣x∣1+nx2≤12n\frac{|x|}{1 + nx^2} \leq \frac{1}{2\sqrt n} (by AM–GM, 1+nx2≥2n∣x∣1 + nx^2 \geq 2\sqrt n|x|). (b) 00 pointwise, not uniformly: at x=1nx = \frac1{n} the value is (1−1n)n→e−1(1 - \tfrac1n)^n \to e^{-1}. (c) 00 on [0,1)[0, 1), 12\tfrac12 at 11, 11 on (1,2](1, 2]: discontinuous limit, so not uniform. (d) 00, uniformly (≤1n\leq \tfrac1n). (e) 00, uniformly: the maximum, at x=nn+1x = \frac{n}{n+1}, is 1n+1(1−1n+1)n≤1n+1\frac{1}{n+1}(1 - \tfrac1{n+1})^n \leq \frac1{n+1}.

Exercise 5.12 Using the M-test

Show that ∑k=1∞cos⁡(kx)k2\sum_{k=1}^\infty \frac{\cos(kx)}{k^2} converges uniformly on R\mathbb{R} to a continuous function, and that the series of derivatives, −∑sin⁡(kx)k-\sum \frac{\sin(kx)}{k}, can't be handled by the M-test. Then show that ∑k=1∞cos⁡(kx)k3\sum_{k=1}^\infty \frac{\cos(kx)}{k^3} is continuously differentiable, with derivative −∑sin⁡(kx)k2-\sum \frac{\sin(kx)}{k^2}.

Solution

∣cos⁡(kx)/k2∣≤1/k2|\cos(kx)/k^2| \leq 1/k^2, which is summable. The derivative series has terms bounded only by 1/k1/k, not summable. For the last part, apply Theorem 5.5 to the partial sums: their derivatives converge uniformly by the M-test with Mk=1/k2M_k = 1/k^2, and the partial sums converge at x=0x = 0.

Exercise 5.13 A compact set of functions

Let F={f∈C([0,1]):∣f(x)∣≤1, ∣f(x)−f(y)∣≤∣x−y∣ for all x,y}\mathcal{F} = \{f \in C([0, 1]) : |f(x)| \leq 1,\ |f(x) - f(y)| \leq |x - y| \text{ for all } x, y\}. Show that every sequence in F\mathcal{F} has a uniformly convergent subsequence whose limit is in F\mathcal{F}, so F\mathcal{F} is a compact subset of C([0,1])C([0, 1]). Then find, by the maximum principle on this compact set, a function in F\mathcal{F} that maximises ∫01f(x) sin⁡(2πx) dx\int_0^1 f(x)\,\sin(2\pi x)\,dx. (You don't need to identify it.)

Solution

F\mathcal{F} is uniformly bounded and equicontinuous (δ=ε\delta = \varepsilon), so Arzelà–Ascoli gives a uniformly convergent subsequence. The conditions ∣f∣≤1|f| \leq 1 and ∣f(x)−f(y)∣≤∣x−y∣|f(x) - f(y)| \leq |x - y| pass to pointwise limits, so the limit lies in F\mathcal{F}. The functional f↦∫fsin⁡(2πx)f \mapsto \int f\sin(2\pi x) is continuous on C([0,1])C([0, 1]) (it changes by at most d∞(f,g)d_\infty(f, g)), so it attains a maximum on the compact set F\mathcal{F} (2B.3 Compactness). This is the direct method of the calculus of variations in miniature (4A.6 Weak Convergence and the Direct Method).

Exercise 5.14 No convergent subsequence

Show that ∫02π(sin⁡mx−sin⁡nx)2 dx=2π\int_0^{2\pi}(\sin mx - \sin nx)^2\,dx = 2\pi for distinct positive integers m,nm, n. Deduce that sup⁡∣sin⁡mx−sin⁡nx∣≥1\sup|\sin mx - \sin nx| \geq 1 on [0,2π][0, 2\pi], so no subsequence of (sin⁡nx)(\sin nx) is uniformly Cauchy.

Hint

Expand the square and use ∫02πsin⁡mxsin⁡nx dx=0\int_0^{2\pi}\sin mx\sin nx\,dx = 0 for m≠nm \neq n and ∫02πsin⁡2nx dx=π\int_0^{2\pi}\sin^2 nx\,dx = \pi. If ∣sin⁡mx−sin⁡nx∣<1|\sin mx - \sin nx| < 1 everywhere, the integral would be less than 2π2\pi.

Exercise 5.15 The converse of Arzelà–Ascoli

Let KK be compact and F⊆C(K)\mathcal{F} \subseteq C(K) a set such that every sequence in F\mathcal{F} has a uniformly convergent subsequence. Show that F\mathcal{F} is uniformly bounded and equicontinuous. (For equicontinuity, take a finite ε\varepsilon-net f1,…,fmf_1, \dots, f_m of F\mathcal{F} in the sup metric, which exists by 2B.3 Compactness, and use the uniform continuity of each fif_i.)

Solution

F\mathcal{F} is totally bounded in d∞d_\infty (2B.3 Compactness's argument applies to its closure), so it is bounded. Given ε\varepsilon, take an ε/3\varepsilon/3-net f1,…,fmf_1, \dots, f_m and a δ\delta that works for each fif_i with ε/3\varepsilon/3. For f∈Ff \in \mathcal{F} pick fif_i within ε/3\varepsilon/3; then for d(x,y)<δd(x, y) < \delta, ∣f(x)−f(y)∣≤∣f(x)−fi(x)∣+∣fi(x)−fi(y)∣+∣fi(y)−f(y)∣≤ε|f(x) - f(y)| \leq |f(x) - f_i(x)| + |f_i(x) - f_i(y)| + |f_i(y) - f(y)| \leq \varepsilon.

Exercise 5.16 Bernstein's proof

Let ff be continuous on [0,1][0, 1], and wk(x)=(nk)xk(1−x)n−kw_k(x) = \binom nk x^k(1-x)^{n-k}. (a) Show that ∑kwk=1\sum_k w_k = 1, ∑kknwk=x\sum_k \tfrac kn w_k = x and ∑k(kn−x)2wk=x(1−x)n≤14n\sum_k \big(\tfrac kn - x\big)^2 w_k = \frac{x(1-x)}{n} \leq \frac{1}{4n}. (b) Split ∣Bnf(x)−f(x)∣≤∑k∣f(kn)−f(x)∣ wk|B_nf(x) - f(x)| \leq \sum_k |f(\tfrac kn) - f(x)|\,w_k into the terms with ∣kn−x∣<δ|\tfrac kn - x| < \delta and the rest, and bound the rest using (a) and (kn−x)2/δ2≥1\big(\tfrac kn - x\big)^2/\delta^2 \geq 1. Conclude that Bnf→fB_nf \to f uniformly.

Solution

(a) These are the total probability, the mean and the variance of a binomial distribution with nn trials and success probability xx, divided by nn and n2n^2; or differentiate (x+y)n=∑(nk)xkyn−k(x + y)^n = \sum\binom nk x^ky^{n-k} once and twice in xx and set y=1−xy = 1 - x. (b) Choose δ\delta with ∣f(s)−f(x)∣≤ε|f(s) - f(x)| \leq \varepsilon for ∣s−x∣<δ|s - x| < \delta. The near terms contribute at most ε\varepsilon. The far terms contribute at most 2M∑farwk≤2M∑k(k/n−x)2δ2wk≤2M4nδ22M\sum_{\text{far}} w_k \leq 2M\sum_k \frac{(k/n - x)^2}{\delta^2}w_k \leq \frac{2M}{4n\delta^2}, which is less than ε\varepsilon for nn large, uniformly in xx.

Exercise 5.17 Rehearsal: compactness with all derivatives

Let fn:[0,1]→Rf_n : [0, 1] \to \mathbb{R} be smooth, with bounds sup⁡∣fn(k)∣≤Ck\sup|f_n^{(k)}| \leq C_k for every k≥0k \geq 0, the constants CkC_k not depending on nn. Show that some subsequence converges uniformly, together with all its derivatives, to a smooth function ff: that is, fnj(k)→f(k)f_{n_j}^{(k)} \to f^{(k)} uniformly for every kk.

(Apply Arzelà–Ascoli to the kk-th derivatives, which are uniformly bounded and, by the bound on the (k+1)(k{+}1)-th derivatives, equicontinuous. Extract a subsequence for k=0k = 0, a further one for k=1k = 1, and so on, and take the diagonal. Use Theorem 5.5 to identify the limits as derivatives of ff.) This is the skeleton of the Cheeger–Gromov and Hamilton compactness theorems (9B.4 Convergence of Manifolds, 11B.3 Compactness of Ricci Flows): bounds on all derivatives of curvature give a subsequence converging smoothly. Most of the work in those theorems goes into producing such bounds and choosing good coordinates.

Solution

For each kk, the family {fn(k)}\{f_n^{(k)}\} is bounded by CkC_k and Lipschitz with constant Ck+1C_{k+1}, so it is equicontinuous. Choose a subsequence S0S_0 along which fnf_n converges uniformly; a further subsequence S1⊆S0S_1 \subseteq S_0 along which fn′f_n' also converges uniformly; and so on. The diagonal sequence gjg_j (jj-th term of SjS_j) is eventually a subsequence of every SkS_k, so gj(k)g_j^{(k)} converges uniformly to some hkh_k, for every kk. By Theorem 5.5 applied to gj(k)g_j^{(k)}, hkh_k is differentiable with hk′=hk+1h_k' = h_{k+1}. So f=h0f = h_0 is smooth with f(k)=hkf^{(k)} = h_k.

© 2026 NeckPinch (www.neckpinch.com). All content in the guidebook (text, mathematics, figures and exercises) is protected by copyright. All rights reserved. No part may be copied, republished or redistributed without written permission.