Book 4A

© 2026 NeckPinch · www.neckpinch.com · All rights reserved.

Course 4Book 4A: Function Spaces and Sobolev SpacesChapter 11

Hölder Spaces

Measuring roughness, and why Hölder spaces are the right setting for regularity.

21 min read · Updated Oct 2, 2026

Read with Evans, Partial Differential Equations, §5.1 (Hölder spaces), which is short; Gilbarg and Trudinger, Elliptic Partial Differential Equations of Second Order, chapter 4, is the reference for Hölder estimates when Book 6A needs them.

In this chapter · 6 sections
  1. 11.1How rough is a random path?
  2. 11.2Hölder spaces
  3. 11.3Compactness
  4. 11.4Why Hölder and not CkC^kCk
  5. 11.5History
  6. 11.6Exercises

Continuity says a function changes little over small distances. Hölder continuity says how little: ∣u(x)−u(y)∣≤C∣x−y∣α|u(x) - u(y)| \leq C|x - y|^\alpha for some exponent α∈(0,1]\alpha \in (0, 1]. The exponent measures roughness on a continuous scale between merely continuous (α\alpha near 00) and Lipschitz (α=1\alpha = 1), and the corresponding spaces Ck,αC^{k,\alpha} sit between CkC^k and Ck+1C^{k+1}. Morrey's inequality (4A.10 Sobolev Embeddings and Critical Exponents) already showed that Sobolev functions with enough integrability are Hölder continuous.

The reason for a whole chapter is a fact about elliptic and parabolic equations. If Δu\Delta u is continuous, uu need not be twice differentiable: the classical CkC^k spaces don't fit the Laplacian. But if Δu\Delta u is Hölder continuous, then uu is twice differentiable with Hölder second derivatives (Schauder's estimates, 6A.6 Parabolic Regularity). Hölder spaces are the spaces in which the Laplacian, and the heat operator, gain exactly two derivatives. That is why the existence theory for nonlinear parabolic equations, and in particular the short-time existence of Ricci flow, is set in Hölder spaces (6A.7 Nonlinear Parabolic Equations, 11A.3 Short-Time Existence and Uniqueness). This short chapter defines them, proves their compactness properties, and explains with an explicit example why C2C^2 would not do.

By the end of this chapter you will be able to:

  • compute Hölder seminorms and decide which Ck,αC^{k,\alpha} a given function belongs to;
  • prove that Hölder spaces are complete and that C0,α↪C0,βC^{0,\alpha} \hookrightarrow C^{0,\beta} is compact for β<α\beta < \alpha on compact sets;
  • give a function with continuous Laplacian that is not C2C^2, and state Schauder's estimate that rules this out in Hölder spaces;
  • relate Hölder exponents to the roughness of real signals and random paths.

How rough is a random path?

In the world Model Brownian motion and the Hurst exponent

The path of a particle in Brownian motion, the jittering of pollen grains in water that Robert Brown observed in 1827, is modelled mathematically by a random continuous function B(t)B(t) whose increments B(t)−B(s)B(t) - B(s) are independent, with mean 00 and variance ∣t−s∣|t - s|. Its typical displacement over a time hh is therefore about h\sqrt h, much larger than hh for small hh, so the path can't be differentiable. Norbert Wiener constructed this process rigorously in 1923, and it is now known that, with probability 11, the path is Hölder continuous with every exponent α<12\alpha < \tfrac12, and with no exponent α>12\alpha > \tfrac12, and that it is differentiable at no point (Paley, Wiener and Zygmund, 1933). The exponent 12\tfrac12 is the roughness of Brownian motion: zooming in by a factor λ\lambda in time and λ\sqrt\lambda in space gives a path with the same statistics (Figure 11.1).

Natural records are rougher or smoother than this. The hydrologist Harold Edwin Hurst, studying centuries of records of the Nile's floods to size the reservoirs planned at Aswan, found in 1951 that the range of cumulative departures from the mean grew like THT^H with H≈0.7H \approx 0.7 over time spans TT, not like T1/2T^{1/2} as independent random fluctuations would give: wet and dry years clustered. Benoit Mandelbrot and John van Ness introduced fractional Brownian motion in 1968 as a model with this scaling: its paths are Hölder continuous with every exponent below its Hurst exponent H∈(0,1)H \in (0, 1). Estimating Hölder or Hurst exponents of measured signals (river flows, financial prices, rough surfaces) is now a standard way of quantifying how rough they are, though the estimates need care.

Figure 11.1. A simulated Brownian path (a random walk with 2142^{14} steps), and two successive zooms by a factor of 44 in time and 2=42 = \sqrt4 in value. The rescaled pieces look alike: the path is equally rough at every scale, with Hölder exponent 12\tfrac12.

Hölder spaces

Let Ω⊆Rn\Omega \subseteq \mathbb{R}^n and 0<α≤10 < \alpha \leq 1. The Hölder seminorm of u:Ω→Ru : \Omega \to \mathbb{R} is

[u]α=sup⁡x≠y∣u(x)−u(y)∣∣x−y∣α.[u]_{\alpha} = \sup_{x \neq y}\frac{|u(x) - u(y)|}{|x - y|^\alpha}.
Definition 11.1 Hölder spaces

C0,α(Ωˉ)C^{0,\alpha}(\bar\Omega) is the space of bounded continuous uu with [u]α<∞[u]_\alpha < \infty, normed by ∥u∥C0,α=sup⁡∣u∣+[u]α\|u\|_{C^{0,\alpha}} = \sup|u| + [u]_\alpha. For an integer k≥0k \geq 0, Ck,α(Ωˉ)C^{k,\alpha}(\bar\Omega) consists of CkC^k functions whose derivatives up to order kk are bounded and whose kk-th derivatives are in C0,αC^{0,\alpha}, with

∥u∥Ck,α=∑∣β∣≤ksup⁡∣∂βu∣+∑∣β∣=k[∂βu]α.\|u\|_{C^{k,\alpha}} = \sum_{|\beta|\leq k}\sup|\partial^\beta u| + \sum_{|\beta| = k}[\partial^\beta u]_\alpha.

For α=1\alpha = 1 the seminorm is the Lipschitz constant. For α>1\alpha > 1 only constants qualify (on a connected set the difference quotient would tend to 00, forcing the derivative to vanish), which is why α≤1\alpha \leq 1. The standard example is a power of the distance.

Example 11.2 ∣x∣α|x|^\alpha

On [−1,1][-1, 1], u(x)=∣x∣αu(x) = |x|^\alpha with 0<α≤10 < \alpha \leq 1 has [u]α=1[u]_\alpha = 1: by the concavity of t↦tαt \mapsto t^\alpha, ∣∣x∣α−∣y∣α∣≤∣x−y∣α\big||x|^\alpha - |y|^\alpha\big| \leq |x - y|^\alpha, with equality when y=0y = 0. It is not in C0,βC^{0,\beta} for any β>α\beta > \alpha, since ∣x∣α−0∣x∣β→∞\frac{|x|^\alpha - 0}{|x|^\beta} \to \infty as x→0x \to 0 (Figure 11.2). So the exponent of a power singularity is exactly its Hölder exponent.

Figure 11.2. ∣x∣α|x|^\alpha for α=0.2,0.5,0.9\alpha = 0.2, 0.5, 0.9: Hölder continuous with exponent exactly α\alpha. Smaller α\alpha means a sharper cusp at 00.
Proposition 11.3 Completeness

Ck,α(Ωˉ)C^{k,\alpha}(\bar\Omega) is a Banach space.

Proof. For k=0k = 0: a Cauchy sequence converges uniformly to some uu (2B.5 Uniform Convergence and Arzelà–Ascoli). For fixed x≠yx \neq y, ∣(um−ul)(x)−(um−ul)(y)∣∣x−y∣α≤[um−ul]α≤ε\frac{|(u_m - u_l)(x) - (u_m - u_l)(y)|}{|x - y|^\alpha} \leq [u_m - u_l]_\alpha \leq \varepsilon for m,l≥Nm, l \geq N; letting l→∞l \to \infty gives [um−u]α≤ε[u_m - u]_\alpha \leq \varepsilon. For k≥1k \geq 1, apply this to the derivatives, using 2B.5 Uniform Convergence and Arzelà–Ascoli's theorem on uniform limits of derivatives.

Hölder spaces have one awkward feature: smooth functions are not dense in them. The function ∣x∣α|x|^\alpha can't be approximated in C0,αC^{0,\alpha} norm by smooth functions, because near 00 any smooth function has ∣v(x)−v(0)∣∣x∣α→0\frac{|v(x) - v(0)|}{|x|^\alpha} \to 0, while for ∣x∣α|x|^\alpha the ratio is 11. (The closure of the smooth functions is the "little Hölder space".) Hölder spaces are also not separable. Neither fact matters for the existence theory, which uses completeness and the compactness below.

Compactness

Proposition 11.4 Compact embeddings

Let KK be compact and 0<β<α≤10 < \beta < \alpha \leq 1. Then bounded sets in C0,α(K)C^{0,\alpha}(K) are precompact in C0,β(K)C^{0,\beta}(K). Likewise Ck,α(K)↪Ck,β(K)C^{k,\alpha}(K) \hookrightarrow C^{k,\beta}(K) is compact.

Proof. A bounded set in C0,αC^{0,\alpha} is uniformly bounded and equicontinuous (∣u(x)−u(y)∣≤M∣x−y∣α|u(x) - u(y)| \leq M|x - y|^\alpha for all its members), so by Arzelà–Ascoli (2B.5 Uniform Convergence and Arzelà–Ascoli) every sequence has a uniformly convergent subsequence. To upgrade to C0,βC^{0,\beta} convergence, use the interpolation inequality

[v]β≤[v]αβ/α (2sup⁡∣v∣)1−β/α,[v]_\beta \leq [v]_\alpha^{\beta/\alpha}\,\big(2\sup|v|\big)^{1 - \beta/\alpha},

which follows from ∣v(x)−v(y)∣=∣v(x)−v(y)∣β/α∣v(x)−v(y)∣1−β/α≤([v]α∣x−y∣α)β/α(2sup⁡∣v∣)1−β/α|v(x) - v(y)| = |v(x) - v(y)|^{\beta/\alpha}|v(x) - v(y)|^{1 - \beta/\alpha} \leq \big([v]_\alpha|x - y|^\alpha\big)^{\beta/\alpha}(2\sup|v|)^{1-\beta/\alpha}. Applied to v=um−ulv = u_m - u_l, whose α\alpha-seminorm is bounded and whose sup tends to 00, it shows (um)(u_m) is Cauchy in C0,βC^{0,\beta}.

The pattern is the one of the whole book: a stronger norm bounded, a weaker norm convergent, and an interpolation inequality between them (3A.7 Lᵖ Spaces and Jensen’s Inequality). It is how limits are taken in the existence theory for nonlinear PDE, and in the compactness theorems for Ricci flows, where bounds on curvature in a Hölder norm give convergence in a slightly weaker one (11B.3 Compactness of Ricci Flows).

Why Hölder and not CkC^k

For the Laplacian in the plane, C0C^0 data do not give C2C^2 solutions.

Example 11.5 Continuous Laplacian, unbounded second derivatives

On the disc ∣x∣<12|x| < \frac12 in R2\mathbb{R}^2, let

u(x,y)=(x2−y2) log⁡(−log⁡r),r=x2+y2,u(x, y) = (x^2 - y^2)\,\log\big(-\log r\big), \qquad r = \sqrt{x^2 + y^2},

with u(0)=0u(0) = 0. A computation (Exercise 11.10) gives, for r>0r > 0,

Δu=cos⁡(2θ) 4log⁡r−1(log⁡r)2,\Delta u = \cos(2\theta)\,\frac{4\log r - 1}{(\log r)^2},

which tends to 00 as r→0r \to 0: so Δu\Delta u extends continuously to the whole disc. But ∂x2u(x,0)=2log⁡(−log⁡∣x∣)+(bounded terms)→∞\partial_x^2u(x, 0) = 2\log(-\log|x|) + (\text{bounded terms}) \to \infty as x→0x \to 0. So uu is a solution of Δu=f\Delta u = f with ff continuous, and u∉C2u \notin C^2.

The failure is a slow, logarithmic one, and it disappears as soon as the data are a little better than continuous.

Theorem 11.6 Schauder estimate (statement)

Let 0<α<10 < \alpha < 1. If Δu=f\Delta u = f in a ball B2B_2 and f∈C0,α(B2)f \in C^{0,\alpha}(B_2), then u∈C2,α(B1)u \in C^{2,\alpha}(B_1) and

∥u∥C2,α(B1)≤C(sup⁡B2∣u∣+∥f∥C0,α(B2)),\|u\|_{C^{2,\alpha}(B_1)} \leq C\big(\sup_{B_2}|u| + \|f\|_{C^{0,\alpha}(B_2)}\big),

with CC depending only on nn and α\alpha. The same holds for uniformly elliptic operators with Hölder continuous coefficients, and for the heat equation in parabolic Hölder spaces.

The proof, by comparison with the fundamental solution (4A.8 Distributions and Weak Derivatives) and careful estimates of singular integrals, or by approximation with harmonic functions, is in 6A.6 Parabolic Regularity (Gilbarg–Trudinger, chapter 4). The point here is the form: the operator gains exactly two derivatives, measured in the same Hölder exponent. That makes the Laplacian an isomorphism between suitable C2,αC^{2,\alpha} and C0,αC^{0,\alpha} spaces, which is what the perturbation and contraction arguments of 4A.1 Banach Spaces and Bounded Operators and 2B.9 The Inverse and Implicit Function Theorems need. In C2→C0C^2 \to C^0 it is not an isomorphism (by the example), and neither is it from Ck+2C^{k+2} to CkC^k for any kk. Hölder spaces are the classical spaces in which elliptic and parabolic equations are well posed.

In the world Model The stress at a crack tip

In linear elastic fracture mechanics, the displacement of the material near the tip of a crack behaves like r\sqrt r times a function of angle, where rr is the distance to the tip, so the stresses (derivatives of the displacement) grow like 1r\frac{1}{\sqrt r}. The displacement is Hölder continuous with exponent 12\tfrac12 at the tip and no better. M. L. Williams (1957) derived this behaviour from the equations of elasticity, and George Irwin (1957) built fracture mechanics around the coefficient of the 1r\frac1{\sqrt r} singularity, the stress intensity factor: a crack grows when that coefficient reaches a critical value characteristic of the material. Engineers designing against fracture thus work directly with a Hölder-12\tfrac12 singularity, a reminder that the regularity theory of elliptic equations has corners (literally) where it stops: at boundary points that are not smooth, solutions are only Hölder continuous, with an exponent set by the angle.

Where this goes Hölder spaces in the rest of the guidebook

Schauder theory, interior and up to the boundary, for elliptic and parabolic equations (6A.6 Parabolic Regularity); short-time existence for quasilinear parabolic equations by contraction in parabolic Hölder spaces C2+α,1+α/2C^{2+\alpha, 1+\alpha/2} (6A.7 Nonlinear Parabolic Equations); and the short-time existence of Ricci flow, where after DeTurck's modification the flow is a strictly parabolic system for the metric, solved in exactly these spaces (11A.3 Short-Time Existence and Uniqueness). In compactness theorems, Ck,αC^{k,\alpha} bounds on metrics in harmonic coordinates give Ck,βC^{k,\beta} convergence for β<α\beta < \alpha, which is the form in which Cheeger–Gromov convergence is often stated (9B.4 Convergence of Manifolds).

History

Rudolf Lipschitz introduced his condition in 1864 (in work on Fourier series) and used it for differential equations in 1876; Otto Hölder used the fractional version in his 1882 dissertation on potential theory. Juliusz Schauder proved his estimates in 1934, building on work of Hölder, Korn and Lichtenstein. Robert Brown described the motion of pollen particles in 1827, Albert Einstein explained it in 1905, and Wiener constructed the mathematical process in 1923; Paley, Wiener and Zygmund proved nowhere differentiability in 1933. Hurst's Nile study appeared in 1951 and Mandelbrot and van Ness's fractional Brownian motion in 1968.

Recall Book 4A in one paragraph

Banach spaces carry analysis into infinite dimensions, where the unit ball is not compact (Riesz). Completeness alone gives uniform boundedness, open mapping and closed graph (Baire). Hahn–Banach supplies enough functionals, and identifies duals: LqL^q for LpL^p, measures for C(K)C(K). Hilbert spaces add projections, Riesz representation and orthonormal bases, and Lax–Milgram gives weak solutions. The Fourier transform diagonalises derivatives and is an isometry of L2L^2. Weak convergence restores compactness in reflexive spaces, and with convexity the direct method gives minimisers. Compact self-adjoint operators have eigenbases, and the Laplacian on a bounded domain has a discrete spectrum. Distributions differentiate everything; Sobolev spaces are the functions with weak derivatives in LpL^p, with Poincaré, Sobolev and Morrey inequalities, compact embeddings below the critical exponent, and concentration at it. Hölder spaces are where elliptic equations gain exactly two derivatives.

Where this goes Into Books 5A and 6A

Everything needed to state and solve linear PDE is now in place: spaces (LpL^p, HkH^k, Ck,αC^{k,\alpha}), compactness (Rellich, Arzelà–Ascoli), existence (Lax–Milgram, the direct method) and spectra. Book 5A is optional: it revisits harmonic functions and conformal geometry from the complex-analytic side, ending at Ricci flow on surfaces. Book 6A, starting with 6A.1 What a PDE Is, is the main line: the heat equation, in full.

Exercises

Exercise 11.7 Which Hölder space?

Find the best Hölder exponent on [0,1][0, 1] of (a) x\sqrt x; (b) xlog⁡xx\log x (with value 00 at 00); (c) 1log⁡(1/x)\frac{1}{\log(1/x)} on [0,12][0, \frac12]; (d) x3/2x^{3/2}, as an element of C1,αC^{1,\alpha}.

Solution

(a) 12\tfrac12. (b) Every α<1\alpha < 1, not 11 (the derivative log⁡x+1\log x + 1 is unbounded). (c) Continuous but no positive exponent: 1/log⁡(1/x)xα→∞\frac{1/\log(1/x)}{x^\alpha} \to \infty. (d) u′=32x1/2u' = \tfrac32x^{1/2}, so u∈C1,1/2u \in C^{1,1/2}.

Exercise 11.8 Products and compositions

Show that [uv]α≤sup⁡∣u∣ [v]α+sup⁡∣v∣ [u]α[uv]_\alpha \leq \sup|u|\,[v]_\alpha + \sup|v|\,[u]_\alpha, so C0,αC^{0,\alpha} is closed under products, and that if ϕ\phi is Lipschitz then [ϕ∘u]α≤Lip(ϕ)[u]α[\phi\circ u]_\alpha \leq \mathrm{Lip}(\phi)[u]_\alpha. Show also that composing two Hölder functions with exponents α\alpha and β\beta gives exponent αβ\alpha\beta.

Exercise 11.9 The interpolation inequality

Prove [v]β≤[v]αβ/α(2sup⁡∣v∣)1−β/α[v]_\beta \leq [v]_\alpha^{\beta/\alpha}(2\sup|v|)^{1-\beta/\alpha} for 0<β<α≤10 < \beta < \alpha \leq 1, and use it to show that C0,α([0,1])↪C0,β([0,1])C^{0,\alpha}([0, 1]) \hookrightarrow C^{0,\beta}([0, 1]) is compact. Why is the embedding C0,α↪C0,αC^{0,\alpha} \hookrightarrow C^{0,\alpha} (the identity) not compact? (Use un(x)=max⁡(0,n−α−∣x−12∣α)u_n(x) = \max(0, n^{-\alpha} - |x - \tfrac12|^\alpha)-type bumps, or 4A.1 Banach Spaces and Bounded Operators.)

Exercise 11.10 The classical example

For u=h(x,y) g(r)u = h(x, y)\,g(r) with h=x2−y2=r2cos⁡2θh = x^2 - y^2 = r^2\cos2\theta (a harmonic polynomial of degree 22) and gg radial, show Δu=h (g′′+5rg′)\Delta u = h\,(g'' + \frac5rg') in the plane. With g=log⁡(−log⁡r)g = \log(-\log r), compute g′=1rlog⁡rg' = \frac{1}{r\log r} and g′′=−1+log⁡rr2(log⁡r)2g'' = -\frac{1 + \log r}{r^2(\log r)^2}, and deduce the formula for Δu\Delta u in Example 11.5. Then show ∂x2u\partial_x^2u is unbounded near 00.

Solution

Δ(hg)=gΔh+2∇h⋅∇g+hΔg=0+2⋅2hrg′+h(g′′+1rg′)\Delta(hg) = g\Delta h + 2\nabla h\cdot\nabla g + h\Delta g = 0 + 2\cdot\frac{2h}{r}g' + h(g'' + \frac1rg'), using x⋅∇h=2hx\cdot\nabla h = 2h for a homogeneous polynomial of degree 22. So Δu=h(g′′+5rg′)\Delta u = h(g'' + \frac5rg'). With the given gg: g′′+5rg′=−(1+log⁡r)+5log⁡rr2(log⁡r)2=4log⁡r−1r2(log⁡r)2g'' + \frac5rg' = \frac{-(1 + \log r) + 5\log r}{r^2(\log r)^2} = \frac{4\log r - 1}{r^2(\log r)^2}, and h=r2cos⁡2θh = r^2\cos2\theta. On the xx-axis, u=x2g(∣x∣)u = x^2g(|x|) and u′′=2g+4xg′+x2g′′u'' = 2g + 4xg' + x^2g''; the last two terms are bounded and 2g=2log⁡(−log⁡∣x∣)→∞2g = 2\log(-\log|x|) \to \infty.

Exercise 11.11 Brownian scaling

If BB is a Brownian motion, show (assuming the definition) that B~(t)=λ−1/2B(λt)\tilde B(t) = \lambda^{-1/2}B(\lambda t) has independent increments with variance ∣t−s∣|t - s|, so it is again a Brownian motion. Explain why this self-similarity is consistent with Hölder exponent exactly 12\tfrac12 and with no larger one.

Exercise 11.12 Rehearsal: scaling the Schauder estimate

Let Δu=f\Delta u = f on B2rB_{2r} and set ur(x)=u(rx)u_r(x) = u(rx), fr(x)=r2f(rx)f_r(x) = r^2f(rx) on B2B_2, so Δur=fr\Delta u_r = f_r. Using the Schauder estimate on B2B_2 for uru_r and translating back, show

sup⁡Br∣∇2u∣+rα[∇2u]α;Br≤C(r−2sup⁡B2r∣u∣+sup⁡B2r∣f∣+rα[f]α;B2r).\sup_{B_r}|\nabla^2u| + r^\alpha[\nabla^2u]_{\alpha;B_r} \leq C\Big(r^{-2}\sup_{B_{2r}}|u| + \sup_{B_{2r}}|f| + r^\alpha[f]_{\alpha;B_{2r}}\Big).

Every term has the same "units" (those of u/r2u/r^2): this is the scale-invariant form of the estimate. In 11B.4 Singularities and 12B.3 The Canonical Neighbourhood Theorem, estimates for Ricci flow are always used in scale-invariant form, so that they survive the rescalings at singularities; checking the powers of rr as here is how one confirms an estimate is stated correctly.

Solution

∇2ur(x)=r2(∇2u)(rx)\nabla^2u_r(x) = r^2(\nabla^2u)(rx) and [∇2ur]α;B1=r2+α[∇2u]α;Br[\nabla^2u_r]_{\alpha;B_1} = r^{2+\alpha}[\nabla^2u]_{\alpha;B_r}; similarly sup⁡∣fr∣=r2sup⁡∣f∣\sup|f_r| = r^2\sup|f| and [fr]α=r2+α[f]α[f_r]_\alpha = r^{2+\alpha}[f]_\alpha. Substituting into the estimate for uru_r and dividing by r2r^2 gives the stated form.

© 2026 NeckPinch (www.neckpinch.com). All content in the guidebook (text, mathematics, figures and exercises) is protected by copyright. All rights reserved. No part may be copied, republished or redistributed without written permission.