Book 6A

© 2026 NeckPinch · www.neckpinch.com · All rights reserved.

Course 6Book 6A: The Heat Equation and Its RelativesChapter 3

The Heat Equation on ℝⁿ

The heat kernel, smoothing, uniqueness, Brownian motion, and a Gaussian that runs to Perelman.

32 min read · Updated Oct 2, 2026

Read with Evans, Partial Differential Equations, section 2.3 (the heat equation: fundamental solution, mean value formula, properties of solutions, energy methods), and section 4.3.1 for the Fourier transform derivation. The Fourier side was done in [[4A.5]].

In this chapter · 8 sections
  1. 3.1Brownian motion
  2. 3.2The heat kernel
  3. 3.3Solving the Cauchy problem
  4. 3.4Maximum principle and uniqueness
  5. 3.5Energy methods
  6. 3.6The heat kernel is a probability law
  7. 3.7History
  8. 3.8Exercises

This chapter solves the heat equation on all of space. The solution is a convolution with one function, the heat kernel

Φ(x,t)=1(4πt)n/2e−∣x∣2/4t,x∈Rn, t>0,\Phi(x, t) = \frac{1}{(4\pi t)^{n/2}}e^{-|x|^2/4t}, \qquad x \in \mathbb{R}^n,\ t > 0,

and almost everything about the equation can be read off from it. It is positive, it has total mass 11, it is a Gaussian whose width grows like t\sqrt t, and it is smooth. So solutions are averages of their initial data with Gaussian weights: they are instantly smooth, they spread at infinite speed, and their maxima can only decrease.

The heat kernel is also a probability law: the distribution of a Brownian particle. That is how Einstein used it in 1905 and how Perrin confirmed it, and it is the sense in which Perelman's and Bamler's heat kernels on Ricci flows are probability measures. And the Gaussian is the function Perelman puts at the centre of his entropy: this chapter's rehearsal computes it.

By the end of this chapter you will be able to:

  • derive the heat kernel by self-similarity and by the Fourier transform, and list its properties;
  • solve the Cauchy problem by convolution and prove that solutions are smooth for t>0t > 0;
  • prove a maximum principle and uniqueness on bounded domains, and explain why uniqueness on Rn\mathbb{R}^n needs a growth condition;
  • use energy methods for uniqueness and decay;
  • interpret the heat kernel as the law of Brownian motion, and estimate diffusion from physical data.

Brownian motion

In the world Data Einstein, Perrin and the size of atoms

In 1827 the botanist Robert Brown saw small particles released from pollen grains jiggling irregularly in water. In 1905 Albert Einstein explained the motion as the result of collisions with water molecules and derived two formulas. A particle performing such a random walk has, in each coordinate, a mean square displacement

⟨x2⟩=2Dt,\langle x^2\rangle = 2Dt,

because its position at time tt is distributed according to the heat kernel of ut=DΔuu_t = D\Delta u (Proposition 3.4). And for a sphere of radius aa in a fluid of viscosity η\eta at absolute temperature TT, the diffusion coefficient is

D=RTNA⋅16πηa,D = \frac{RT}{N_A}\cdot\frac{1}{6\pi\eta a},

where RR is the gas constant and NAN_A is Avogadro's number. Measuring DD by watching particles therefore measures the number of molecules in a mole.

Jean Perrin did this in 1908–09. He prepared suspensions of nearly uniform resin spheres, recorded their positions through a microscope at regular intervals, and computed mean square displacements. His values of Avogadro's number from these and related measurements were around 7×10237 \times 10^{23} (his published series range from about 6.96.9 to 7.15×10237.15 \times 10^{23}), compared with the modern exact value 6.022 140 76×10236.022\,140\,76 \times 10^{23}. The agreement, from several independent methods, settled the long dispute over whether atoms are real; Perrin received the 1926 Nobel Prize in Physics for this work.

To see the scale: a sphere of radius 0.50.5 µm in water at 2020 °C (viscosity about 1.0×10−31.0 \times 10^{-3} Pa·s) has, by Einstein's formula with the modern NAN_A, D≈4.3×10−13D \approx 4.3 \times 10^{-13} m²/s. In 3030 seconds it wanders a root-mean-square distance 2Dt≈5\sqrt{2Dt} \approx 5 µm in each direction: a few of its own diameters, easy to follow in a microscope (Exercise 3.8, Figure 3.3).

The heat kernel

There are two derivations, both already half done.

By scaling (6A.1 What a PDE Is). A solution of the form t−n/2F(∣x∣/t)t^{-n/2}F(|x|/\sqrt t) conserves ∫u dx\int u\,dx, and substituting into the heat equation forces F(r)=Ce−r2/4F(r) = Ce^{-r^2/4}. Choosing CC to make the total integral 11 gives Φ\Phi (Exercise 3.5).

By the Fourier transform (4A.5 The Fourier Transform). With u^(ξ,t)=∫u(x,t)e−2πix⋅ξ dx\hat u(\xi, t) = \int u(x, t)e^{-2\pi ix\cdot\xi}\,dx, the heat equation becomes ∂tu^=−4π2∣ξ∣2u^\partial_t\hat u = -4\pi^2|\xi|^2\hat u, so u^(ξ,t)=e−4π2t∣ξ∣2g^(ξ)\hat u(\xi, t) = e^{-4\pi^2t|\xi|^2}\hat g(\xi). A product of transforms is the transform of a convolution, and the inverse transform of the Gaussian e−4π2t∣ξ∣2e^{-4\pi^2t|\xi|^2} is Φ(⋅,t)\Phi(\cdot, t) (computed by contour shifting in the last exercise of 5A.3 Residues and Fourier Transforms). So u=Φ(⋅,t)∗gu = \Phi(\cdot, t) * g.

Proposition 3.1 Properties of the heat kernel

For t>0t > 0:

  1. Φ>0\Phi > 0, and ∫RnΦ(x,t) dx=1\int_{\mathbb{R}^n}\Phi(x, t)\,dx = 1;
  2. Φ\Phi is smooth and solves Φt=ΔΦ\Phi_t = \Delta\Phi on Rn×(0,∞)\mathbb{R}^n\times(0, \infty);
  3. (semigroup) Φ(⋅,t)∗Φ(⋅,s)=Φ(⋅,t+s)\Phi(\cdot, t) * \Phi(\cdot, s) = \Phi(\cdot, t + s);
  4. (approximate identity) for every δ>0\delta > 0, ∫∣x∣>δΦ(x,t) dx→0\int_{|x|>\delta}\Phi(x, t)\,dx \to 0 as t→0t \to 0;
  5. (scaling) Φ(x,t)=t−n/2Φ(x/t,1)\Phi(x, t) = t^{-n/2}\Phi(x/\sqrt t, 1); in each coordinate the variance of Φ(⋅,t)\Phi(\cdot, t) is 2t2t.

Proof. (1) The Gaussian integral (3A.5 Product Measures and Change of Variables): ∫e−∣x∣2/4tdx=(4πt)n/2\int e^{-|x|^2/4t}dx = (4\pi t)^{n/2}. (2) A direct computation (Exercise 3.6). (3) On the Fourier side, e−4π2t∣ξ∣2e−4π2s∣ξ∣2=e−4π2(t+s)∣ξ∣2e^{-4\pi^2t|\xi|^2}e^{-4\pi^2s|\xi|^2} = e^{-4\pi^2(t+s)|\xi|^2}. (4) Substituting x=t yx = \sqrt t\,y, the integral is ∫∣y∣>δ/tΦ(y,1) dy→0\int_{|y|>\delta/\sqrt t}\Phi(y, 1)\,dy \to 0 by dominated convergence (3A.3 The Lebesgue Integral). (5) Substitution, and ∫x12Φ(x,t) dx=2t\int x_1^2\Phi(x, t)\,dx = 2t (Exercise 3.5).

The semigroup property says that running the heat equation for time ss and then for time tt is the same as running it for time t+st + s. The approximate-identity property says that as t→0t \to 0 the kernel concentrates all its mass at the origin, so Φ(⋅,t)→δ0\Phi(\cdot, t) \to \delta_0: the heat kernel is the temperature after a unit of heat is released at a point at time 00 (Figure 3.1).

Figure 3.1. The one-dimensional heat kernel Φ(x,t)=(4πt)−1/2e−x2/4t\Phi(x, t) = (4\pi t)^{-1/2}e^{-x^2/4t} at t=0.01t = 0.01, 0.10.1 and 11 (computed). Each has area 11. Ten times the time means 10\sqrt{10} times the width and 1/101/\sqrt{10} of the height.

Solving the Cauchy problem

Theorem 3.2 Solution of the Cauchy problem

Let gg be bounded and continuous on Rn\mathbb{R}^n, and define

u(x,t)=∫RnΦ(x−y,t)g(y) dy(t>0).u(x, t) = \int_{\mathbb{R}^n}\Phi(x - y, t)g(y)\,dy \qquad (t > 0).

Then uu is C∞C^\infty on Rn×(0,∞)\mathbb{R}^n\times(0, \infty), solves ut=Δuu_t = \Delta u there, satisfies inf⁡g≤u≤sup⁡g\inf g \leq u \leq \sup g, and u(x,t)→g(x0)u(x, t) \to g(x_0) as (x,t)→(x0,0)(x, t) \to (x_0, 0), for every x0x_0.

Proof. Smoothness and the equation. Φ\Phi is smooth for t>0t > 0, and all its derivatives decay like a polynomial times e−∣x∣2/4te^{-|x|^2/4t}, so they can be taken under the integral sign (3A.3 The Lebesgue Integral); each derivative of uu is the integral of gg against the corresponding derivative of Φ\Phi, and ut−Δu=∫(Φt−ΔΦ)(x−y,t)g(y) dy=0u_t - \Delta u = \int(\Phi_t - \Delta\Phi)(x - y, t)g(y)\,dy = 0.

Bounds. uu is an average of gg with the positive weights Φ(x−⋅,t)\Phi(x - \cdot, t), of total mass 11.

Initial values. Given ε>0\varepsilon > 0, pick δ\delta with ∣g(y)−g(x0)∣<ε|g(y) - g(x_0)| < \varepsilon for ∣y−x0∣<2δ|y - x_0| < 2\delta. For ∣x−x0∣<δ|x - x_0| < \delta,

∣u(x,t)−g(x0)∣≤∫∣y−x0∣<2δΦ(x−y,t)∣g(y)−g(x0)∣ dy+2sup⁡∣g∣∫∣y−x∣>δΦ(x−y,t) dy,|u(x, t) - g(x_0)| \leq \int_{|y - x_0|<2\delta}\Phi(x - y, t)|g(y) - g(x_0)|\,dy + 2\sup|g|\int_{|y - x|>\delta}\Phi(x - y, t)\,dy,

and the first term is at most ε\varepsilon, the second tends to 00 as t→0t \to 0 by property 4.

Three features distinguish the heat equation from the wave equation and from ODEs.

Instant smoothing. The initial data gg need only be bounded and continuous (in fact bounded and measurable will do, with convergence almost everywhere), yet uu is C∞C^\infty for every t>0t > 0. Moreover, differentiating the kernel gives explicit bounds:

sup⁡x∣∇ku(x,t)∣≤Cn,ktk/2sup⁡∣g∣.\sup_x|\nabla^ku(x, t)| \leq \frac{C_{n,k}}{t^{k/2}}\sup|g|.

Each derivative costs a factor t−1/2t^{-1/2}, exactly as parabolic scaling predicts (Exercise 3.7). These smoothing estimates are the model for Shi's derivative estimates for the Ricci flow (11A.3 Short-Time Existence and Uniqueness).

Infinite speed of propagation. If g≥0g \geq 0 is not identically zero, then u(x,t)>0u(x, t) > 0 for every xx and every t>0t > 0, because Φ>0\Phi > 0 everywhere. A disturbance is felt everywhere immediately, though only with weight e−∣x∣2/4te^{-|x|^2/4t} at distance ∣x∣|x|. The heat equation is a model, accurate over the scales where the random-walk picture applies, not a law that violates relativity.

Inhomogeneous equations. The solution of ut−Δu=fu_t - \Delta u = f with u(⋅,0)=0u(\cdot, 0) = 0 is given by Duhamel's principle: superpose the solutions started at each earlier time ss with data f(⋅,s)f(\cdot, s),

u(x,t)=∫0t∫RnΦ(x−y,t−s)f(y,s) dy ds,u(x, t) = \int_0^t\int_{\mathbb{R}^n}\Phi(x - y, t - s)f(y, s)\,dy\,ds,

the same idea as variation of constants for linear ODE (2B.10 Ordinary Differential Equations). It is the starting point of the fixed-point arguments for nonlinear equations in 6A.7 Nonlinear Parabolic Equations.

In the world In use Black–Scholes is the heat equation

In 1973 Fischer Black and Myron Scholes, and independently Robert Merton, derived a PDE for the price V(S,t)V(S, t) of a European option on a stock with price SS, volatility σ\sigma and riskless interest rate rr:

Vt+12σ2S2VSS+rSVS−rV=0,V_t + \tfrac12\sigma^2S^2V_{SS} + rSV_S - rV = 0,

solved backward from the payoff at the expiry time TT. The substitutions S=KexS = Ke^x, t=T−2τσ2t = T - \frac{2\tau}{\sigma^2} and V=Keαx+βτw(x,τ)V = Ke^{\alpha x + \beta\tau}w(x, \tau), with k=2r/σ2k = 2r/\sigma^2, α=−k−12\alpha = -\frac{k - 1}{2} and β=−(k+1)24\beta = -\frac{(k + 1)^2}{4}, turn it into wτ=wxxw_\tau = w_{xx}, an ordinary heat equation run forward in τ\tau, the time remaining to expiry. The Black–Scholes formula for the price of a call option is the heat-kernel convolution of the transformed payoff, which is why it is written with the Gaussian distribution function. Scholes and Merton received the 1997 Nobel Memorial Prize in Economic Sciences for this work (Black had died in 1995). The model's assumptions, such as constant volatility and continuous trading without costs, fail in real markets, and practitioners adjust for that; the mathematics is a heat equation exactly.

Maximum principle and uniqueness

On a bounded region the heat equation is well posed with initial values and boundary values. Let U⊂RnU \subset \mathbb{R}^n be bounded and open, and let UT=U×(0,T]U_T = U\times(0, T] be the space-time cylinder. Its parabolic boundary ΓT\Gamma_T is the bottom and the sides, (U×{0})∪(∂U×[0,T])\big(U\times\{0\}\big)\cup\big(\partial U\times[0, T]\big): everything except the top, the part of the boundary where data are prescribed (Figure 3.2).

Theorem 3.3 The weak maximum principle

Let uu be continuous on UT‾\overline{U_T}, with utu_t and ∇2u\nabla^2u continuous in UTU_T, and ut−Δu≤0u_t - \Delta u \leq 0 there. Then

max⁡UT‾u=max⁡ΓTu.\max_{\overline{U_T}}u = \max_{\Gamma_T}u.

Proof. First suppose ut−Δu<0u_t - \Delta u < 0 strictly. If the maximum over UT‾\overline{U_T} were at a point (x0,t0)(x_0, t_0) with x0∈Ux_0 \in U and 0<t0≤T0 < t_0 \leq T, then ∇2u(x0,t0)≤0\nabla^2u(x_0, t_0) \leq 0, so Δu≤0\Delta u \leq 0 (2B.8 Calculus in Several Variables), and ut(x0,t0)≥0u_t(x_0, t_0) \geq 0 (it is 00 if t0<Tt_0 < T, and ≥0\geq 0 if t0=Tt_0 = T, since uu can only have increased to reach a maximum at the final time). Then ut−Δu≥0u_t - \Delta u \geq 0 there, a contradiction. In general apply this to v=u−εtv = u - \varepsilon t, which has vt−Δv≤−ε<0v_t - \Delta v \leq -\varepsilon < 0, and let ε→0\varepsilon \to 0.

Figure 3.2. The cylinder UTU_T and its parabolic boundary (thick): the initial time and the sides. The maximum of a solution of the heat equation is attained there, never at a later interior point or on the top, because at such a point ut≥0≥Δuu_t \geq 0 \geq \Delta u. 6A.4 Maximum Principles develops this argument in full.

Applied to the difference of two solutions and its negative, the maximum principle gives uniqueness and stability: two solutions with the same initial and boundary values agree, and if the data differ by at most ε\varepsilon the solutions differ by at most ε\varepsilon.

On all of Rn\mathbb{R}^n there is no lateral boundary, and something must replace it. Uniqueness holds among solutions that do not grow too fast: if uu solves the heat equation on Rn×(0,T]\mathbb{R}^n\times(0, T], is continuous up to t=0t = 0 with u(⋅,0)=0u(\cdot, 0) = 0, and satisfies ∣u(x,t)∣≤Aea∣x∣2|u(x, t)| \leq Ae^{a|x|^2} for some constants AA, aa, then u≡0u \equiv 0 (Evans, section 2.3.3). The growth condition cannot be dropped. Andrey Tychonoff constructed in 1935 a non-zero solution on R×(0,∞)\mathbb{R}\times(0, \infty) with zero initial values,

u(x,t)=∑k=0∞g(k)(t)(2k)!x2k,g(t)=e−1/t2,u(x, t) = \sum_{k=0}^\infty\frac{g^{(k)}(t)}{(2k)!}x^{2k}, \qquad g(t) = e^{-1/t^2},

which grows faster than any ea∣x∣2e^{a|x|^2} as ∣x∣→∞|x| \to \infty (Exercise 3.10): heat that arrives "from infinity" in no time.

Where this goes Uniqueness for Ricci flows on noncompact manifolds

The same issue arises for the Ricci flow on a complete noncompact manifold. Without some control at infinity, uniqueness can fail. The standard results assume bounded curvature: Shi constructed solutions with bounded curvature (1989), and Chen and Zhu proved in 2006 that complete solutions with bounded curvature are unique. Bounded curvature plays the role of the growth condition Aea∣x∣2Ae^{a|x|^2} here, and it is the hypothesis built into the singularity analysis of Books 11B and 12B.

Energy methods

A second route to uniqueness, which needs no maximum principle and works for systems: measure the size of a solution by an integral. For a solution on the bounded region UU with u=0u = 0 on ∂U\partial U, let e(t)=∫Uu2 dxe(t) = \int_Uu^2\,dx. Then, integrating by parts (1A.10 Divergence, Curl and the Integral Theorems),

e′(t)=2∫Uu ut dx=2∫Uu Δu dx=−2∫U∣∇u∣2 dx≤0.e'(t) = 2\int_Uu\,u_t\,dx = 2\int_Uu\,\Delta u\,dx = -2\int_U|\nabla u|^2\,dx \leq 0.

So ee never increases; if two solutions have the same data, the energy of their difference is 00 at the start and so always. More is true. By the Poincaré inequality ∫U∣∇u∣2≥λ1∫Uu2\int_U|\nabla u|^2 \geq \lambda_1\int_Uu^2 (4A.9 Sobolev Spaces), e′≤−2λ1ee' \leq -2\lambda_1e, so e(t)≤e−2λ1te(0)e(t) \leq e^{-2\lambda_1t}e(0): the solution decays exponentially at a rate set by the first Dirichlet eigenvalue of UU (4A.7 Compact Operators and Spectra). Showing that log⁡e(t)\log e(t) is a convex function of tt, by one more differentiation, proves backward uniqueness: two solutions that agree at time TT agreed at all earlier times (Evans, section 2.3.4), even though the backward problem is ill-posed (6A.1 What a PDE Is).

The heat kernel is a probability law

Proposition 3.4 Heat flow is averaging over Brownian paths

Let BtB_t be a standard Brownian motion in Rn\mathbb{R}^n, so that BtB_t is Gaussian with mean 00 and covariance tItI. Then x+B2tx + B_{2t} has density Φ(⋅−x,t)\Phi(\cdot - x, t), and the solution of the Cauchy problem is

u(x,t)=E[g(x+B2t)].u(x, t) = \mathbb{E}\big[g(x + B_{2t})\big].

Proof. A Gaussian with covariance 2tI2tI has density (4πt)−n/2e−∣y∣2/4t(4\pi t)^{-n/2}e^{-|y|^2/4t} (3A.4 Measures, Probability and Weights), which is Φ(y,t)\Phi(y, t). The formula is then Theorem 3.2, written as an expectation.

The factor 22 is a convention: probabilists normalise Brownian motion so that its generator is 12Δ\frac12\Delta, while the heat equation here has Δ\Delta. In words: the temperature at xx at time tt is the average of the initial temperature over the endpoints of random paths started at xx. The density p(⋅,t)p(\cdot, t) of a diffusing particle evolves by the heat equation, which in this role is called the Fokker–Planck equation of Brownian motion (Figure 3.3). Harmonic functions have the same interpretation, with stopping at the boundary in place of a fixed time: the solution of the Dirichlet problem at xx is the expected boundary value where a Brownian path from xx first exits (the continuous version of the gambler's ruin in 6A.2 Harmonic Functions's exercise on graphs).

Figure 3.3. Left: 20 simulated random walks approximating Brownian motion B2tB_{2t} (steps of variance 2 dt2\,dt, dt=1/1000dt = 1/1000). Right: the endpoints of 5000 such walks at t=1t = 1 against the heat kernel Φ(x,1)=(4π)−1/2e−x2/4\Phi(x, 1) = (4\pi)^{-1/2}e^{-x^2/4}, the Gaussian of variance 22 (computed, random seed fixed).
In the world Model Kelvin's age of the Earth

In 1862 William Thomson, later Lord Kelvin, used the heat equation to estimate how long ago the Earth's surface solidified. His model was a half-space of rock, initially at a uniform melting temperature VV, whose surface was suddenly held at 00. The solution is the self-similar error-function profile of 6A.1 What a PDE Is, u=Verf⁡(x/2κt)u = V\operatorname{erf}\big(x/2\sqrt{\kappa t}\big), and the temperature gradient at the surface is

G=Vπκt,sot=V2πκG2.G = \frac{V}{\sqrt{\pi\kappa t}}, \qquad\text{so}\qquad t = \frac{V^2}{\pi\kappa G^2}.

Kelvin took the then accepted increase of underground temperature, about 11 °F per 5050 feet of depth, a melting temperature of 70007000 °F, and rock diffusivities derived from measurements on Edinburgh rocks, and obtained about 9898 million years; given the uncertainties, he concluded that consolidation took place between 2020 and 400400 million years ago. The formula reproduces his central figure with a diffusivity of about 1.2×10−61.2 \times 10^{-6} m²/s, typical of rock (Exercise 3.9).

The modern age of the Earth is about 4.54.5 billion years. The mathematics was right; the model was not. Kelvin did not know about radioactive heating, discovered decades later, and, as John Perry pointed out in 1895, heat transport in a partly fluid interior by convection is much faster than conduction through solid rock. It is the classic lesson in the difference between solving an equation correctly and choosing the right equation.

Where this goes A Gaussian that runs to Perelman

Run the heat kernel backward: with τ=T−t\tau = T - t, the function (4πτ)−n/2e−∣x∣2/4τ(4\pi\tau)^{-n/2}e^{-|x|^2/4\tau} solves the backward heat equation ∂tu=−Δu\partial_tu = -\Delta u (Exercise 3.11), concentrating at the origin as t→Tt \to T. Writing it as (4πτ)−n/2e−f(4\pi\tau)^{-n/2}e^{-f} with f=∣x∣24τf = \frac{|x|^2}{4\tau}, the function ff has Hessian 12τI\frac{1}{2\tau}I. In Perelman's work this is the model: along a Ricci flow he solves the conjugate heat equation backward from a point at time TT, writes the solution as (4πτ)−n/2e−f(4\pi\tau)^{-n/2}e^{-f}, and measures how far ff is from satisfying Ric⁡+∇2f=12τg\operatorname{Ric} + \nabla^2f = \frac{1}{2\tau}g, the shrinking soliton equation; flat space with this ff is the equality case (12A.3 The 𝓦-Entropy, 12A.5 Reduced Distance and Reduced Volume). Bamler's theory of Ricci flows uses these heat kernels as probability measures on the manifold (12C.5 After Perelman), exactly as Proposition 3.4 does on Rn\mathbb{R}^n.

History

Joseph Fourier solved the heat equation on the line with what is in effect the Gaussian kernel in his 1822 Théorie analytique de la chaleur. Kelvin's estimate of the age of the Earth appeared in 1862 in "On the secular cooling of the Earth". Einstein's paper on Brownian motion was one of his papers of 1905; Marian Smoluchowski reached similar results independently in 1906; Perrin's experiments date from 1908–09. Andrey Tychonoff's non-uniqueness example appeared in 1935. The Black–Scholes and Merton papers were published in 1973.

Recall Where we stand

The heat kernel Φ=(4πt)−n/2e−∣x∣2/4t\Phi = (4\pi t)^{-n/2}e^{-|x|^2/4t} is positive, has mass 11, forms a semigroup and is an approximate identity. The Cauchy problem is solved by u=Φ(⋅,t)∗gu = \Phi(\cdot, t) * g, which is smooth for t>0t > 0 with derivative bounds Ct−k/2sup⁡∣g∣Ct^{-k/2}\sup|g|, spreads at infinite speed, and stays between inf⁡g\inf g and sup⁡g\sup g. On bounded regions the maximum is attained on the parabolic boundary, and energy decays; on Rn\mathbb{R}^n uniqueness needs a growth condition, as Tychonoff's example shows. The heat kernel is the law of x+B2tx + B_{2t}, and Einstein's ⟨x2⟩=2Dt\langle x^2\rangle = 2Dt, confirmed by Perrin, is its variance. 6A.4 Maximum Principles takes the maximum principle as far as it goes.

Exercises

Exercise 3.5 Mass and variance

(a) Show that ∫RnΦ(x,t) dx=1\int_{\mathbb{R}^n}\Phi(x, t)\,dx = 1 using ∫Re−s2ds=π\int_{\mathbb{R}}e^{-s^2}ds = \sqrt\pi (3A.5 Product Measures and Change of Variables). (b) Show that ∫x12Φ(x,t) dx=2t\int x_1^2\Phi(x, t)\,dx = 2t and ∫∣x∣2Φ(x,t) dx=2nt\int|x|^2\Phi(x, t)\,dx = 2nt.

Solution

(a) The integral factors into nn one-dimensional ones, each ∫e−s2/4tds=4πt\int e^{-s^2/4t}ds = \sqrt{4\pi t}. (b) In one dimension, ∫s2e−s2/4tds=2t4πt\int s^2e^{-s^2/4t}ds = 2t\sqrt{4\pi t}, by integrating by parts or differentiating ∫e−as2ds=π/a\int e^{-as^2}ds = \sqrt{\pi/a} in aa at a=1/4ta = 1/4t: ∫s2e−as2=π2a−3/2\int s^2e^{-as^2} = \frac{\sqrt\pi}{2}a^{-3/2}, and π2(4t)3/2/4πt=2t\frac{\sqrt\pi}{2}(4t)^{3/2}/\sqrt{4\pi t} = 2t. Summing over coordinates gives 2nt2nt.

Exercise 3.6 The kernel solves the equation

Compute ∂tΦ\partial_t\Phi and ΔΦ\Delta\Phi directly and check that they agree. (Use ∇Φ=−x2tΦ\nabla\Phi = -\frac{x}{2t}\Phi.)

Solution

∂tΦ=(−n2t+∣x∣24t2)Φ\partial_t\Phi = \big(-\frac{n}{2t} + \frac{|x|^2}{4t^2}\big)\Phi. ΔΦ=div⁡(−x2tΦ)=−n2tΦ+∣x∣24t2Φ\Delta\Phi = \operatorname{div}\big(-\frac{x}{2t}\Phi\big) = -\frac{n}{2t}\Phi + \frac{|x|^2}{4t^2}\Phi.

Exercise 3.7 Smoothing estimates

In one dimension, show that ∫∣∂xΦ(x,t)∣ dx=1πt\int|\partial_x\Phi(x, t)|\,dx = \frac{1}{\sqrt{\pi t}}, and deduce ∣ux(x,t)∣≤sup⁡∣g∣πt|u_x(x, t)| \leq \frac{\sup|g|}{\sqrt{\pi t}} for the solution of Theorem 3.2. Why must every such bound scale like t−1/2t^{-1/2}?

Solution

∫∣∂xΦ∣=2∫0∞x2tΦ dx=1t⋅14πt∫0∞xe−x2/4tdx=1t⋅2t4πt=1πt\int|\partial_x\Phi| = 2\int_0^\infty\frac{x}{2t}\Phi\,dx = \frac1t\cdot\frac{1}{\sqrt{4\pi t}}\int_0^\infty xe^{-x^2/4t}dx = \frac1t\cdot\frac{2t}{\sqrt{4\pi t}} = \frac{1}{\sqrt{\pi t}}. Then ∣ux∣=∣∫∂xΦ(x−y,t)g(y) dy∣≤sup⁡∣g∣∫∣∂xΦ∣|u_x| = |\int\partial_x\Phi(x - y, t)g(y)\,dy| \leq \sup|g|\int|\partial_x\Phi|. The bound must be invariant under the scaling u(λx,λ2t)u(\lambda x, \lambda^2t), which multiplies uxu_x by λ\lambda and tt by λ−2\lambda^{-2}, so it must be Ct−1/2sup⁡∣g∣Ct^{-1/2}\sup|g|.

Exercise 3.8 Einstein's formula with numbers

Using kB=R/NA=1.381×10−23k_B = R/N_A = 1.381 \times 10^{-23} J/K, T=293T = 293 K, η=1.0×10−3\eta = 1.0 \times 10^{-3} Pa·s and a=0.5a = 0.5 µm, compute D=kBT6πηaD = \frac{k_BT}{6\pi\eta a} and the root-mean-square displacement in one coordinate after 3030 s. If an experiment measured a mean square displacement of 3.0×10−113.0 \times 10^{-11} m² in 3030 s for these particles, what value of NAN_A would it give?

Solution

kBT=4.05×10−21k_BT = 4.05 \times 10^{-21} J; 6πηa=9.42×10−96\pi\eta a = 9.42 \times 10^{-9} kg/s; D≈4.3×10−13D \approx 4.3 \times 10^{-13} m²/s; 2D⋅30≈5.1×10−6\sqrt{2D\cdot30} \approx 5.1 \times 10^{-6} m. From ⟨x2⟩=2Dt\langle x^2\rangle = 2Dt, D=3.0×10−11/60=5.0×10−13D = 3.0 \times 10^{-11}/60 = 5.0 \times 10^{-13}, and NA=RT6πηaD=8.314×2939.42×10−9×5.0×10−13≈5.2×1023N_A = \frac{RT}{6\pi\eta aD} = \frac{8.314 \times 293}{9.42 \times 10^{-9}\times5.0 \times 10^{-13}} \approx 5.2 \times 10^{23}. Small errors in aa and η\eta change the answer proportionally, which is one reason Perrin's careful selection of uniform particles mattered.

Exercise 3.9 Kelvin's arithmetic

(a) Check that the gradient of Verf⁡(x/2κt)V\operatorname{erf}(x/2\sqrt{\kappa t}) at x=0x = 0 is V/πκtV/\sqrt{\pi\kappa t}, using erf⁡(z)=2π∫0ze−s2ds\operatorname{erf}(z) = \frac{2}{\sqrt\pi}\int_0^ze^{-s^2}ds. (b) Convert V=7000V = 7000 °F and G=1G = 1 °F per 5050 feet to SI-compatible units (a temperature difference in °F and a gradient in °F per metre suffice), and find the diffusivity κ\kappa for which t=98t = 98 million years. (c) By what factor would tt change if κ\kappa were doubled? If VV were 10,00010{,}000 °F?

Solution

(a) ∂xerf⁡(x/2κt)=2π⋅12κt\partial_x\operatorname{erf}(x/2\sqrt{\kappa t}) = \frac{2}{\sqrt\pi}\cdot\frac{1}{2\sqrt{\kappa t}} at x=0x = 0. (b) G=1/15.24≈0.0656G = 1/15.24 \approx 0.0656 °F/m; t=98×106×3.156×107≈3.09×1015t = 98 \times 10^6 \times 3.156 \times 10^7 \approx 3.09 \times 10^{15} s; κ=V2/(πtG2)≈4.9×107/(3.1416×3.09×1015×0.0043)≈1.2×10−6\kappa = V^2/(\pi tG^2) \approx 4.9 \times 10^7/(3.1416 \times 3.09 \times 10^{15}\times0.0043) \approx 1.2 \times 10^{-6} m²/s. (c) Halved; multiplied by (10/7)2≈2(10/7)^2 \approx 2, about 200200 million years, which is the figure Kelvin also gives for that melting temperature.

Exercise 3.10 Tychonoff's example, formally

Show formally (differentiating term by term) that u(x,t)=∑kg(k)(t)(2k)!x2ku(x, t) = \sum_k\frac{g^{(k)}(t)}{(2k)!}x^{2k} satisfies ut=uxxu_t = u_{xx} for any smooth gg, and that u(x,0)=0u(x, 0) = 0 if all derivatives of gg vanish at 00. (Making this rigorous needs bounds on g(k)g^{(k)} for g=e−1/t2g = e^{-1/t^2}, which is where the rapid growth in xx comes from; Fritz John's Partial Differential Equations gives the details.)

Exercise 3.11 Rehearsal: Perelman's Gaussian on flat space

Let τ=T−t\tau = T - t and u=(4πτ)−n/2e−fu = (4\pi\tau)^{-n/2}e^{-f} with f(x)=∣x∣24τf(x) = \frac{|x|^2}{4\tau}. (a) Show that uu solves the backward heat equation ∂tu=−Δu\partial_tu = -\Delta u for t<Tt < T. (b) Compute ∇f\nabla f, ∣∇f∣2|\nabla f|^2, Δf\Delta f and ∇2f\nabla^2f, and check the identities

∇2f=12τI,2Δf−∣∇f∣2+f−nτ=0.\nabla^2f = \frac{1}{2\tau}I, \qquad 2\Delta f - |\nabla f|^2 + \frac{f - n}{\tau} = 0.

The first is the shrinking soliton equation Ric⁡+∇2f=12τg\operatorname{Ric} + \nabla^2f = \frac{1}{2\tau}g on flat space, where Ric⁡=0\operatorname{Ric} = 0; the second is the flat-space case of the identity that makes Perelman's W\mathcal W-entropy constant on a shrinking soliton (12A.3 The 𝓦-Entropy).

Solution

(a) ∂t=−∂τ\partial_t = -\partial_\tau, and ∂τ((4πτ)−n/2e−∣x∣2/4τ)=Δ(⋯ )\partial_\tau\big((4\pi\tau)^{-n/2}e^{-|x|^2/4\tau}\big) = \Delta(\cdots) because it is the heat kernel in the variable τ\tau; so ∂tu=−Δu\partial_tu = -\Delta u. (b) ∇f=x2τ\nabla f = \frac{x}{2\tau}, ∣∇f∣2=∣x∣24τ2=fτ|\nabla f|^2 = \frac{|x|^2}{4\tau^2} = \frac f\tau, ∇2f=12τI\nabla^2f = \frac{1}{2\tau}I, Δf=n2τ\Delta f = \frac{n}{2\tau}. Then 2Δf−∣∇f∣2+f−nτ=nτ−fτ+fτ−nτ=02\Delta f - |\nabla f|^2 + \frac{f - n}{\tau} = \frac n\tau - \frac f\tau + \frac f\tau - \frac n\tau = 0.

© 2026 NeckPinch (www.neckpinch.com). All content in the guidebook (text, mathematics, figures and exercises) is protected by copyright. All rights reserved. No part may be copied, republished or redistributed without written permission.