Book 6A

© 2026 NeckPinch · www.neckpinch.com · All rights reserved.

Course 6Book 6A: The Heat Equation and Its RelativesChapter 10

Entropy, Information and Diffusion

Entropy along the heat flow, the log-Sobolev inequality and Li–Yau: the ancestors of Perelman’s entropy.

29 min read · Updated Oct 3, 2026

None of the Path's books covers this chapter; it is a seam. For more, read Gross, "Logarithmic Sobolev inequalities" (Amer. J. Math. 97, 1975), and Li and Yau, "On the parabolic kernel of the Schrödinger operator" (Acta Math. 156, 1986). Bakry, Gentil and Ledoux, Analysis and Geometry of Markov Diffusion Operators, is the reference for entropy methods.

In this chapter · 8 sections
  1. 10.1Milk in coffee
  2. 10.2Entropy and Fisher information
  3. 10.3The logarithmic Sobolev inequality
  4. 10.4Nash's use of entropy
  5. 10.5The Li–Yau inequality
  6. 10.6The template
  7. 10.7History
  8. 10.8Exercises

Perelman's 2002 preprint introduced an entropy for the Ricci flow, and explained that its monotonicity is "in a sense" the same as the monotonicity of entropy in statistical physics. This chapter is about the original: entropy along the heat equation. A density that diffuses becomes more spread out, and its Boltzmann–Shannon entropy H=−∫ulog⁡uH = -\int u\log u increases, at a rate equal to its Fisher information ∫∣∇u∣2u\int\frac{|\nabla u|^2}{u}. How large the entropy can be for a given amount of Fisher information is the content of the logarithmic Sobolev inequality, and with the right scale parameter that inequality is literally the statement that Perelman's W\mathcal W-entropy of flat space is non-negative. The chapter also proves the sharpest pointwise statement about positive solutions of the heat equation, the Li–Yau inequality, whose equality case is the heat kernel.

All three results share one pattern, which is the thread of the second half of this guide: a quantity that is monotone along a flow, with equality exactly for a special self-similar solution. Entropy and the Gaussian here; Huisken's density and self-shrinkers in 6A.8 Curve Shortening and the First Geometric Flows; and Perelman's W\mathcal W and reduced volume with shrinking solitons in 12A.3 The 𝓦-Entropy and 12A.5 Reduced Distance and Reduced Volume.

By the end of this chapter you will be able to:

  • compute the entropy and Fisher information of a density, and prove that entropy increases along the heat equation at the rate of the Fisher information;
  • state Gross's Gaussian log-Sobolev inequality, derive its Euclidean form with a scale parameter τ\tau, and recognise it as W≥0\mathcal W \geq 0 on flat space;
  • explain why log-Sobolev is "Sobolev in infinite dimensions";
  • state and check the Li–Yau differential Harnack inequality, and integrate it to compare values at different points and times;
  • describe the "monotone quantity plus equality case" template.

Milk in coffee

In the world Model Mixing increases entropy

Pour a little milk into still coffee. Ignoring convection, the milk's concentration uu, normalised to total mass 11, spreads by diffusion, ut=DΔuu_t = D\Delta u, and the configuration becomes steadily more mixed and never spontaneously unmixes. The quantity that measures "mixed" is the entropy

H(u)=−∫ulog⁡u dx,H(u) = -\int u\log u\,dx,

which is largest, for a given spread, when the milk is spread as evenly as possible, and −∞-\infty in the limit of a point. Along the diffusion it increases, and the rate of increase is

dHdt=D∫∣∇u∣2u dx,\frac{dH}{dt} = D\int\frac{|\nabla u|^2}{u}\,dx,

the Fisher information (Theorem 10.1). So entropy grows fastest where concentrations have sharp gradients: a thin streak of milk mixes much faster than a broad cloud, which is why stirring, which stretches the milk into thin streaks, speeds mixing so much.

This is the macroscopic face of the second law of thermodynamics, and the heat equation is a good model of it. It is not a derivation of the second law: the diffusion equation is itself an approximation to the reversible motion of molecules, and why that motion looks irreversible at large scales is a separate question in statistical mechanics. What this chapter takes from physics is the mathematics of the macroscopic description, which is exactly what Perelman borrowed.

Entropy and Fisher information

Let u>0u > 0 be a probability density on Rn\mathbb{R}^n (∫u=1\int u = 1) with enough decay for the integrals below. Its entropy and Fisher information are

H(u)=−∫ulog⁡u dx,I(u)=∫∣∇u∣2u dx=∫u ∣∇log⁡u∣2 dx=4∫∣∇u∣2 dx.H(u) = -\int u\log u\,dx, \qquad I(u) = \int\frac{|\nabla u|^2}{u}\,dx = \int u\,|\nabla\log u|^2\,dx = 4\int|\nabla\sqrt u|^2\,dx.

For the heat kernel Φ(⋅,t)\Phi(\cdot, t), a Gaussian of variance 2t2t in each coordinate (6A.3 The Heat Equation on ℝⁿ),

H(Φt)=n2log⁡(4πet),I(Φt)=n2tH(\Phi_t) = \frac n2\log(4\pi et), \qquad I(\Phi_t) = \frac{n}{2t}

(Exercise 10.5). Among all densities with a given covariance, the Gaussian has the largest entropy (3A.7 Lᵖ Spaces and Jensen’s Inequality, by Jensen's inequality). Entropy can be any real number; it increases by nlog⁡λn\log\lambda when uu is spread out by a factor λ\lambda, because HH measures a logarithm of volume.

Theorem 10.1 Entropy production

If u>0u > 0 solves ut=Δuu_t = \Delta u on Rn\mathbb{R}^n with ∫u=1\int u = 1 (and suitable decay), then

ddtH(u)=I(u)≥0.\frac{d}{dt}H(u) = I(u) \geq 0.

Proof. ddt∫ulog⁡u=∫ut(log⁡u+1)=∫Δu (log⁡u+1)=−∫∇u⋅∇log⁡u=−∫∣∇u∣2u\frac{d}{dt}\int u\log u = \int u_t(\log u + 1) = \int\Delta u\,(\log u + 1) = -\int\nabla u\cdot\nabla\log u = -\int\frac{|\nabla u|^2}{u}, integrating by parts with no boundary terms. So dHdt=I(u)\frac{dH}{dt} = I(u). (The term ∫ut=ddt∫u=0\int u_t = \frac{d}{dt}\int u = 0.)

For the heat kernel this reads ddtn2log⁡(4πet)=n2t\frac{d}{dt}\frac n2\log(4\pi et) = \frac{n}{2t}, as it must. The Fisher information itself is also monotone: it decreases, since

ddtI(u)=−2∫u ∣∇2log⁡u∣2 dx≤0\frac{d}{dt}I(u) = -2\int u\,|\nabla^2\log u|^2\,dx \leq 0

(Exercise 10.6). So HH is increasing and concave in tt. The computation behind this, a Bochner-type identity for log⁡u\log u, is the one that, on a manifold, produces a Ricci curvature term; with Ric⁡≥0\operatorname{Ric} \geq 0 the sign survives. This is the first sign of how curvature enters entropy, which is the heart of 12A.3 The 𝓦-Entropy.

Figure 10.1. A density starting as two narrow bumps at x=±1x = \pm1 (each a Gaussian of variance 0.020.02) diffusing by ut=uxxu_t = u_{xx} (computed exactly as a sum of two Gaussians, integrals by quadrature). The entropy HH increases and is concave; its slope is the Fisher information II, which decreases and approaches the heat kernel's value 12t\frac{1}{2t} (dashed) once the two bumps have merged.

The logarithmic Sobolev inequality

Entropy production says dHdt=I\frac{dH}{dt} = I. To turn this into a statement about rates of mixing, one needs an inequality comparing entropy with Fisher information. It is cleanest relative to a Gaussian.

Theorem 10.2 Gross's Gaussian log-Sobolev inequality (1975)

Let γ\gamma be the standard Gaussian measure on Rn\mathbb{R}^n, with density (2π)−n/2e−∣x∣2/2(2\pi)^{-n/2}e^{-|x|^2/2}. For every smooth gg with ∫g2 dγ=1\int g^2\,d\gamma = 1,

∫g2log⁡g2 dγ≤2∫∣∇g∣2 dγ.\int g^2\log g^2\,d\gamma \leq 2\int|\nabla g|^2\,d\gamma.

Equality holds for g=1g = 1.

Two features make this inequality special. The constant 22 does not depend on the dimension nn, unlike the constant in the Sobolev inequality (4A.10 Sobolev Embeddings and Critical Exponents), which degenerates as n→∞n \to \infty; Gross's motivation was quantum field theory, where the dimension is infinite. And it controls g2log⁡g2g^2\log g^2, a quantity only "logarithmically" better than g2g^2, which is the best one can hope for in infinite dimensions. It is equivalent to hypercontractivity of the heat flow for the Gaussian measure (Edward Nelson, 1973): the Ornstein–Uhlenbeck semigroup maps L2(γ)L^2(\gamma) into Lp(γ)L^p(\gamma) for some p>2p > 2 after a positive time, with norm 11.

Now rewrite it in the language of Perelman. Let τ>0\tau > 0, and write a probability density on Rn\mathbb{R}^n as

u=(4πτ)−n/2e−f.u = (4\pi\tau)^{-n/2}e^{-f}.

The reference density is the Gaussian of variance 2τ2\tau, which corresponds to f=∣x∣24τf = \frac{|x|^2}{4\tau}: the heat kernel at time τ\tau, or the backward Gaussian of 6A.3 The Heat Equation on ℝⁿ's rehearsal.

Theorem 10.3 The Euclidean log-Sobolev inequality with scale τ\tau

For every τ>0\tau > 0 and every smooth ff with ∫(4πτ)−n/2e−fdx=1\int(4\pi\tau)^{-n/2}e^{-f}dx = 1 (and suitable decay),

∫Rn[τ∣∇f∣2+f−n](4πτ)−n/2e−f dx≥0,\int_{\mathbb{R}^n}\Big[\tau|\nabla f|^2 + f - n\Big](4\pi\tau)^{-n/2}e^{-f}\,dx \geq 0,

with equality when f=∣x−x0∣24τf = \frac{|x - x_0|^2}{4\tau}.

Proof. Write u=(4πτ)−n/2e−fu = (4\pi\tau)^{-n/2}e^{-f} and γτ=(4πτ)−n/2e−∣x∣2/4τ\gamma_\tau = (4\pi\tau)^{-n/2}e^{-|x|^2/4\tau}, and change variables x=2τyx = \sqrt{2\tau}y to turn γτ\gamma_\tau into γ\gamma. With g2=u/γτg^2 = u/\gamma_\tau, Gross's inequality becomes (Exercise 10.7)

∫u φ dx≤τ∫u ∣∇φ∣2 dx,φ=log⁡uγτ=−f+∣x∣24τ.\int u\,\varphi\,dx \leq \tau\int u\,|\nabla\varphi|^2\,dx, \qquad \varphi = \log\frac{u}{\gamma_\tau} = -f + \frac{|x|^2}{4\tau}.

Now ∇φ=−∇f+x2τ\nabla\varphi = -\nabla f + \frac{x}{2\tau}, so

τ∫u∣∇φ∣2=τ∫u∣∇f∣2−∫u ∇f⋅x+∫u∣x∣24τ.\tau\int u|\nabla\varphi|^2 = \tau\int u|\nabla f|^2 - \int u\,\nabla f\cdot x + \int u\frac{|x|^2}{4\tau}.

Since ∇u=−u∇f\nabla u = -u\nabla f, the middle term is ∫∇u⋅x=−n∫u=−n\int\nabla u\cdot x = -n\int u = -n, integrating by parts. The left side is −∫uf+∫u∣x∣24τ-\int uf + \int u\frac{|x|^2}{4\tau}. The ∣x∣2|x|^2 terms cancel, leaving −∫uf≤τ∫u∣∇f∣2−n-\int uf \leq \tau\int u|\nabla f|^2 - n, which is the claim. Translating xx gives the equality case at any x0x_0.

The idea This is Perelman's W\mathcal W on flat space

Perelman's entropy of a metric gg on an nn-manifold, with a function ff and a scale τ\tau, is

W(g,f,τ)=∫M[τ(R+∣∇f∣2)+f−n](4πτ)−n/2e−f dV\mathcal W(g, f, \tau) = \int_M\Big[\tau\big(R + |\nabla f|^2\big) + f - n\Big](4\pi\tau)^{-n/2}e^{-f}\,dV

(12A.3 The 𝓦-Entropy). On flat Rn\mathbb{R}^n the scalar curvature RR is zero, and Theorem 10.3 says exactly that W(gRn,f,τ)≥0\mathcal W(g_{\mathbb{R}^n}, f, \tau) \geq 0, with equality for the Gaussian. Perelman's μ(g,τ)=inf⁡fW\mu(g, \tau) = \inf_f\mathcal W is therefore 00 for flat space, and μ<0\mu < 0 measures how far a manifold is, at scale τ\tau, from being Euclidean. That is how μ\mu detects collapsing, and the reason Perelman's noncollapsing theorem is a log-Sobolev inequality in disguise (12A.4 κ-Noncollapsing).

Nash's use of entropy

The log-Sobolev inequality and entropy production together control how fast a density spreads, and John Nash used this kind of argument in 1958 for a striking purpose: to prove that solutions of parabolic equations ut=∂j(aij∂iu)u_t = \partial_j(a^{ij}\partial_iu) with merely bounded measurable coefficients are Hölder continuous (6A.6 Parabolic Regularity). Among his tools was the entropy of the fundamental solution: he showed that it increases at least like n2log⁡t\frac n2\log t, just as for the heat kernel, using only the ellipticity bounds, and deduced that the fundamental solution spreads out at the parabolic rate and cannot concentrate. Entropy was a tool for regularity decades before Perelman used it for geometry, and the quantity H(u)−n2log⁡(4πet)H(u) - \frac n2\log(4\pi et), entropy measured relative to the heat kernel, is now called the Nash entropy. It reappears in the work of Hein and Naber and of Bamler on Ricci flows (12C.5 After Perelman).

In the world In use The entropy power inequality

In 1948 Claude Shannon introduced the entropy of a random signal as the measure of its information content, and stated that adding independent noise increases a signal's entropy power N(X)=12πee2H(X)/nN(X) = \frac{1}{2\pi e}e^{2H(X)/n} at least additively:

N(X+Y)≥N(X)+N(Y)N(X + Y) \geq N(X) + N(Y)

for independent random vectors XX and YY, with equality for Gaussians with proportional covariances. Aart Stam gave the first proof in 1959, using exactly the identity of this chapter: adding a small independent Gaussian is running the heat equation, and the entropy changes at the rate of the Fisher information (de Bruijn's identity). The inequality is used in information theory to bound the capacity of communication channels with noise; Shannon used it to show that, among noises of a given power, Gaussian noise is the most damaging.

The Li–Yau inequality

The Harnack inequality of 6A.2 Harmonic Functions compared a positive harmonic function at nearby points. For positive solutions of the heat equation, Peter Li and Shing-Tung Yau found in 1986 a sharp differential inequality, which integrates to a Harnack inequality across space and time.

Theorem 10.4 The Li–Yau inequality

Let u>0u > 0 solve ut=Δuu_t = \Delta u on Rn×(0,∞)\mathbb{R}^n\times(0, \infty) (or on a complete manifold with non-negative Ricci curvature). Then

∣∇u∣2u2−utu≤n2t,equivalentlyΔlog⁡u≥−n2t.\frac{|\nabla u|^2}{u^2} - \frac{u_t}{u} \leq \frac{n}{2t}, \qquad\text{equivalently}\qquad \Delta\log u \geq -\frac{n}{2t}.

Equality holds for the heat kernel Φ(x,t)\Phi(x, t).

The two forms agree because Δlog⁡u=Δuu−∣∇u∣2u2=utu−∣∇u∣2u2\Delta\log u = \frac{\Delta u}{u} - \frac{|\nabla u|^2}{u^2} = \frac{u_t}{u} - \frac{|\nabla u|^2}{u^2}. For the heat kernel, log⁡Φ=−n2log⁡(4πt)−∣x∣24t\log\Phi = -\frac n2\log(4\pi t) - \frac{|x|^2}{4t}, so Δlog⁡Φ=−n2t\Delta\log\Phi = -\frac{n}{2t} exactly. The proof applies the maximum principle to F=t(∣∇log⁡u∣2−∂tlog⁡u)F = t\big(|\nabla\log u|^2 - \partial_t\log u\big), using the Bochner identity of 6A.2 Harmonic Functions for log⁡u\log u; on a manifold the Bochner formula contributes Ric⁡(∇log⁡u,∇log⁡u)\operatorname{Ric}(\nabla\log u, \nabla\log u), which is why Ric⁡≥0\operatorname{Ric} \geq 0 is needed (9B.7 The Heat Equation on a Manifold).

Integrating it. Along any path γ(s)\gamma(s) from (x1,t1)(x_1, t_1) to (x2,t2)(x_2, t_2) with t1<t2t_1 < t_2, the Li–Yau inequality bounds the change of log⁡u\log u, and optimising over straight paths gives

u(x1,t1)≤u(x2,t2)(t2t1)n/2exp⁡(∣x2−x1∣24(t2−t1))u(x_1, t_1) \leq u(x_2, t_2)\Big(\frac{t_2}{t_1}\Big)^{n/2}\exp\Big(\frac{|x_2 - x_1|^2}{4(t_2 - t_1)}\Big)

(Exercise 10.8). This is the parabolic Harnack inequality in sharp form: a positive solution now is bounded below by its value earlier and nearby, with the Gaussian factor that the heat kernel saturates (Figure 10.2).

Figure 10.2. The Li–Yau bound checked numerically for the two-bump solution of Figure 10.1: the quantity 2t⋅max⁡x(ux2u2−utu)=2t⋅max⁡x(−(log⁡u)xx)2t\cdot\max_x\big(\frac{u_x^2}{u^2} - \frac{u_t}{u}\big) = 2t\cdot\max_x\big(-(\log u)_{xx}\big) stays below 11 (the bound n2t\frac{n}{2t} with n=1n = 1), computed exactly. It comes close to the bound early, when each bump is nearly a heat kernel, dips while the two bumps merge into a hump wider than a heat kernel of the same age, and tends to 11 again as t→∞t \to \infty.
Where this goes Harnack inequalities for the Ricci flow

Hamilton proved a matrix Harnack inequality for the Ricci flow with non-negative curvature operator (1993), modelled on Li–Yau; in its simplest trace form it says ∂tR+Rt+2∇iR Vi+2Ric⁡(V,V)≥0\partial_tR + \frac Rt + 2\nabla_iR\,V^i + 2\operatorname{Ric}(V, V) \geq 0 for every vector VV, with equality on expanding solitons. It is the key estimate for ancient solutions and for classifying singularity models (11B.2 Ancient Solutions and the Harnack Inequality). Perelman proved a Li–Yau type inequality for his conjugate heat kernel under the Ricci flow, again with equality on the Gaussian soliton, and used it in the proof of pseudolocality (12A.6 Pseudolocality).

The template

monotone quantity flow equality case where
entropy HH, Nash entropy heat equation Gaussian (heat kernel) this chapter
Li–Yau quantity Δlog⁡u+n2t\Delta\log u + \frac{n}{2t} heat equation heat kernel this chapter
Huisken's Gaussian density Θ\Theta mean curvature flow self-shrinkers 6A.8 Curve Shortening and the First Geometric Flows
Perelman's W\mathcal W and μ\mu Ricci flow + conjugate heat equation shrinking solitons 12A.3 The 𝓦-Entropy
Perelman's reduced volume Ricci flow shrinking solitons (Gaussian on flat space) 12A.5 Reduced Distance and Reduced Volume

The pattern is the one first met as "monotone and bounded sequences converge" (2A.6 Sequences), with an extra ingredient: the equality case is a self-similar solution. So a solution along which the quantity becomes constant, such as a limit of rescalings near a singularity, must be self-similar. That is how Perelman shows that singularity models are shrinking solitons, and it is why entropy is at the centre of the proof.

History

Ludwig Boltzmann's entropy dates from the 1870s; Claude Shannon's information entropy and the entropy power inequality from 1948. Ronald Fisher introduced his information in statistics in the 1920s. Aart Stam's proof of the entropy power inequality, using de Bruijn's identity, appeared in 1959. John Nash's paper on continuity of solutions of parabolic and elliptic equations, with its entropy argument, appeared in 1958. Edward Nelson proved the hypercontractivity of the Ornstein–Uhlenbeck semigroup in 1973, and Leonard Gross his logarithmic Sobolev inequality, and its equivalence with hypercontractivity, in 1975. Peter Li and Shing-Tung Yau's inequality appeared in 1986, and Hamilton's matrix Harnack inequality for the Ricci flow in 1993. Perelman's W\mathcal W-entropy appeared in his 2002 preprint.

Recall Book 6A in one paragraph

A PDE is classified by its symbol; the heat equation is parabolic, scales like (λx,λ2t)(\lambda x, \lambda^2t), is well posed forward and ill-posed backward (6A.1 What a PDE Is). Harmonic functions are averages and obey maximum principles, Harnack and Liouville (6A.2 Harmonic Functions). The heat kernel solves the heat equation on Rn\mathbb{R}^n, smooths instantly and is the law of Brownian motion (6A.3 The Heat Equation on ℝⁿ). At a first interior maximum ut≥0≥Δuu_t \geq 0 \geq \Delta u: maximum principles, comparison and invariant regions for systems follow (6A.4 Maximum Principles). Weak solutions exist by Lax–Milgram and are smooth by bootstrapping (6A.5 Weak Solutions and Elliptic Regularity); parabolic equations gain two space derivatives and one time derivative, and Bernstein's method bounds gradients by C/tC/\sqrt t (6A.6 Parabolic Regularity). Nonlinear equations exist for short time by linearising and a fixed point, and may blow up (6A.7 Nonlinear Parabolic Equations). Curve shortening and mean curvature flow show solitons, ancient solutions, neckpinches and Huisken's monotonicity (6A.8 Curve Shortening and the First Geometric Flows). Energies give Euler–Lagrange equations and gradient flows (6A.9 Calculus of Variations and Gradient Flows). Entropy increases along diffusion at the rate of the Fisher information, the log-Sobolev inequality is W≥0\mathcal W \geq 0 on flat space, and the Li–Yau inequality is sharp on the heat kernel (this chapter).

Where this goes Into Book 7A

The analysis is now ready: given a manifold, we could solve heat-type equations on it. But "manifold" has not been defined, and the Poincaré conjecture is about topology: a closed three-manifold in which every loop can be shrunk to a point is a sphere. Book 7A turns to topology, starting with only what differs from the metric spaces of Book 2B, and builds the fundamental group, covering spaces and the precise statement of the conjecture (7A.1 Topological Spaces and Quotients).

Exercises

Exercise 10.5 The Gaussian's entropy and information

For Φt=(4πt)−n/2e−∣x∣2/4t\Phi_t = (4\pi t)^{-n/2}e^{-|x|^2/4t}, compute H(Φt)=n2log⁡(4πt)+∫∣x∣24tΦt=n2log⁡(4πet)H(\Phi_t) = \frac n2\log(4\pi t) + \int\frac{|x|^2}{4t}\Phi_t = \frac n2\log(4\pi et) and I(Φt)=∫∣x∣24t2Φt=n2tI(\Phi_t) = \int\frac{|x|^2}{4t^2}\Phi_t = \frac{n}{2t}, using ∫∣x∣2Φt=2nt\int|x|^2\Phi_t = 2nt (6A.3 The Heat Equation on ℝⁿ). Check dHdt=I\frac{dH}{dt} = I.

Exercise 10.6 Fisher information decreases

For a positive solution of ut=Δuu_t = \Delta u on Rn\mathbb{R}^n, write v=log⁡uv = \log u, so vt=Δv+∣∇v∣2v_t = \Delta v + |\nabla v|^2, and I=∫u∣∇v∣2I = \int u|\nabla v|^2. Show that dIdt=−2∫u ∣∇2v∣2\frac{dI}{dt} = -2\int u\,|\nabla^2v|^2. (Differentiate, use the equations for uu and vv, and integrate by parts, using the flat Bochner identity 12Δ∣∇v∣2=∣∇2v∣2+∇v⋅∇Δv\frac12\Delta|\nabla v|^2 = |\nabla^2v|^2 + \nabla v\cdot\nabla\Delta v from 6A.2 Harmonic Functions.) Check it on the heat kernel, where ∇2v=−12tI\nabla^2v = -\frac{1}{2t}I and ddtn2t=−n2t2\frac{d}{dt}\frac{n}{2t} = -\frac{n}{2t^2}.

Solution

On the heat kernel, −2∫u∣∇2v∣2=−2⋅n4t2∫u=−n2t2-2\int u|\nabla^2v|^2 = -2\cdot\frac{n}{4t^2}\int u = -\frac{n}{2t^2}, which matches. The general computation: ddt∫u∣∇v∣2=∫Δu ∣∇v∣2+2∫u ∇v⋅∇(Δv+∣∇v∣2)\frac{d}{dt}\int u|\nabla v|^2 = \int\Delta u\,|\nabla v|^2 + 2\int u\,\nabla v\cdot\nabla(\Delta v + |\nabla v|^2); integrate the first term by parts twice to ∫u Δ∣∇v∣2\int u\,\Delta|\nabla v|^2, and use ∇u=u∇v\nabla u = u\nabla v to combine; the Bochner identity turns the result into −2∫u∣∇2v∣2-2\int u|\nabla^2v|^2.

Exercise 10.7 From Gross to the τ\tau form

(a) With uu a probability density and g2=u/γg^2 = u/\gamma (so ∫g2dγ=1\int g^2d\gamma = 1), show that ∫g2log⁡g2 dγ=∫ulog⁡uγ dx\int g^2\log g^2\,d\gamma = \int u\log\frac u\gamma\,dx and ∫∣∇g∣2dγ=14∫u ∣∇log⁡uγ∣2 dx\int|\nabla g|^2d\gamma = \frac14\int u\,|\nabla\log\frac u\gamma|^2\,dx. So Gross's inequality reads ∫ulog⁡uγ≤12∫u∣∇log⁡uγ∣2\int u\log\frac u\gamma \leq \frac12\int u|\nabla\log\frac u\gamma|^2. (b) Rescale x=2τyx = \sqrt{2\tau}y to deduce ∫ulog⁡uγτ≤τ∫u∣∇log⁡uγτ∣2\int u\log\frac{u}{\gamma_\tau} \leq \tau\int u|\nabla\log\frac{u}{\gamma_\tau}|^2 for the Gaussian γτ\gamma_\tau of variance 2τ2\tau.

Solution

(a) g=u/γg = \sqrt{u/\gamma}, so ∇g=12g∇log⁡uγ\nabla g = \frac12g\nabla\log\frac u\gamma and ∣∇g∣2γ=14uγγ∣∇log⁡uγ∣2|\nabla g|^2\gamma = \frac14\frac u\gamma\gamma|\nabla\log\frac u\gamma|^2. (b) Under x=2τyx = \sqrt{2\tau}y, densities transform with the Jacobian, ratios u/γτu/\gamma_\tau are unchanged, the left side is invariant and ∣∇y⋅∣2=2τ∣∇x⋅∣2|\nabla_y\cdot|^2 = 2\tau|\nabla_x\cdot|^2, which turns the factor 12\frac12 into τ\tau.

Exercise 10.8 Integrating Li–Yau

Let u>0u > 0 satisfy utu≥∣∇u∣2u2−n2t\frac{u_t}{u} \geq \frac{|\nabla u|^2}{u^2} - \frac{n}{2t}. Along the straight path γ(s)=x1+s(x2−x1)\gamma(s) = x_1 + s(x_2 - x_1), t(s)=t1+s(t2−t1)t(s) = t_1 + s(t_2 - t_1), 0≤s≤10 \leq s \leq 1, with ℓ=log⁡u\ell = \log u, show

ddsℓ(γ(s),t(s))≥−∣x2−x1∣24(t2−t1)−n(t2−t1)2t(s)\frac{d}{ds}\ell(\gamma(s), t(s)) \geq -\frac{|x_2 - x_1|^2}{4(t_2 - t_1)} - \frac{n(t_2 - t_1)}{2t(s)}

(use ∇ℓ⋅a+(t2−t1)∣∇ℓ∣2≥−∣a∣24(t2−t1)\nabla\ell\cdot a + (t_2 - t_1)|\nabla\ell|^2 \geq -\frac{|a|^2}{4(t_2 - t_1)} for a=x2−x1a = x_2 - x_1), and integrate to get the Harnack inequality in the text. Check that the heat kernel, from (0,t1)(0, t_1) to (x,t2)(x, t_2), makes it nearly sharp.

Exercise 10.9 Entropy power of Gaussians

For a Gaussian random vector XX in Rn\mathbb{R}^n with covariance σ2I\sigma^2I, show H(X)=n2log⁡(2πeσ2)H(X) = \frac n2\log(2\pi e\sigma^2) and N(X)=σ2N(X) = \sigma^2. Deduce that for independent such XX, YY with variances σ2\sigma^2 and ρ2\rho^2, N(X+Y)=N(X)+N(Y)N(X + Y) = N(X) + N(Y): equality in Shannon's inequality.

Solution

This is Exercise 10.5 with 2t=σ22t = \sigma^2: H=n2log⁡(4πe⋅σ22)=n2log⁡(2πeσ2)H = \frac n2\log(4\pi e\cdot\frac{\sigma^2}{2}) = \frac n2\log(2\pi e\sigma^2), so e2H/n=2πeσ2e^{2H/n} = 2\pi e\sigma^2 and N=σ2N = \sigma^2. X+YX + Y is Gaussian with covariance (σ2+ρ2)I(\sigma^2 + \rho^2)I.

Exercise 10.10 Rehearsal: W\mathcal W of the Gaussian on flat space

Let f=∣x∣24τf = \frac{|x|^2}{4\tau} and u=(4πτ)−n/2e−fu = (4\pi\tau)^{-n/2}e^{-f}. Compute directly

∫Rn[τ∣∇f∣2+f−n]u dx\int_{\mathbb{R}^n}\Big[\tau|\nabla f|^2 + f - n\Big]u\,dx

using ∫∣x∣2u=2nτ\int|x|^2u = 2n\tau, and show it is 00: equality in Theorem 10.3, and W(gRn,f,τ)=0\mathcal W(g_{\mathbb{R}^n}, f, \tau) = 0. Then show that for fλ=∣x∣24λτ+n2log⁡λf_\lambda = \frac{|x|^2}{4\lambda\tau} + \frac n2\log\lambda (a Gaussian of the wrong width, still normalised) the integral is n2(1λ−1+log⁡λ)≥0\frac n2\big(\frac1\lambda - 1 + \log\lambda\big) \geq 0, with equality only at λ=1\lambda = 1: the scale τ\tau picks out the Gaussian of variance 2τ2\tau, exactly as in Perelman's μ(g,τ)\mu(g, \tau) (12A.3 The 𝓦-Entropy).

Solution

τ∣∇f∣2=∣x∣24τ=f\tau|\nabla f|^2 = \frac{|x|^2}{4\tau} = f, so the integrand is (2f−n)u(2f - n)u, and ∫2fu=2⋅2nτ4τ=n\int2fu = \frac{2\cdot2n\tau}{4\tau} = n. For fλf_\lambda, uu is the Gaussian of variance 2λτ2\lambda\tau, with ∫∣x∣2u=2nλτ\int|x|^2u = 2n\lambda\tau. Then τ∣∇fλ∣2=∣x∣24λ2τ\tau|\nabla f_\lambda|^2 = \frac{|x|^2}{4\lambda^2\tau}, integrating to n2λ\frac{n}{2\lambda}; ∫fλu=n2+n2log⁡λ\int f_\lambda u = \frac n2 + \frac n2\log\lambda. Total: n2λ+n2+n2log⁡λ−n=n2(1λ−1+log⁡λ)\frac{n}{2\lambda} + \frac n2 + \frac n2\log\lambda - n = \frac n2\big(\frac1\lambda - 1 + \log\lambda\big), which is ≥0\geq 0 because log⁡λ≥1−1λ\log\lambda \geq 1 - \frac1\lambda, with equality only at λ=1\lambda = 1.

© 2026 NeckPinch (www.neckpinch.com). All content in the guidebook (text, mathematics, figures and exercises) is protected by copyright. All rights reserved. No part may be copied, republished or redistributed without written permission.