© 2026 NeckPinch · www.neckpinch.com · All rights reserved.
Course 6Book 6A: The Heat Equation and Its RelativesChapter 10
Entropy, Information and Diffusion
Entropy along the heat flow, the log-Sobolev inequality and Li–Yau: the ancestors of Perelman’s entropy.
None of the Path's books covers this chapter; it is a seam. For more, read Gross, "Logarithmic Sobolev inequalities" (Amer. J. Math. 97, 1975), and Li and Yau, "On the parabolic kernel of the Schrödinger operator" (Acta Math. 156, 1986). Bakry, Gentil and Ledoux, Analysis and Geometry of Markov Diffusion Operators, is the reference for entropy methods.
Perelman's 2002 preprint introduced an entropy for the Ricci flow, and explained that its monotonicity is "in a sense" the same as the monotonicity of entropy in statistical physics. This chapter is about the original: entropy along the heat equation. A density that diffuses becomes more spread out, and its Boltzmann–Shannon entropy increases, at a rate equal to its Fisher information . How large the entropy can be for a given amount of Fisher information is the content of the logarithmic Sobolev inequality, and with the right scale parameter that inequality is literally the statement that Perelman's -entropy of flat space is non-negative. The chapter also proves the sharpest pointwise statement about positive solutions of the heat equation, the Li–Yau inequality, whose equality case is the heat kernel.
All three results share one pattern, which is the thread of the second half of this guide: a quantity that is monotone along a flow, with equality exactly for a special self-similar solution. Entropy and the Gaussian here; Huisken's density and self-shrinkers in 6A.8 Curve Shortening and the First Geometric Flows; and Perelman's and reduced volume with shrinking solitons in 12A.3 The 𝓦-Entropy and 12A.5 Reduced Distance and Reduced Volume.
By the end of this chapter you will be able to:
- compute the entropy and Fisher information of a density, and prove that entropy increases along the heat equation at the rate of the Fisher information;
- state Gross's Gaussian log-Sobolev inequality, derive its Euclidean form with a scale parameter , and recognise it as on flat space;
- explain why log-Sobolev is "Sobolev in infinite dimensions";
- state and check the Li–Yau differential Harnack inequality, and integrate it to compare values at different points and times;
- describe the "monotone quantity plus equality case" template.
Milk in coffee
Pour a little milk into still coffee. Ignoring convection, the milk's concentration , normalised to total mass , spreads by diffusion, , and the configuration becomes steadily more mixed and never spontaneously unmixes. The quantity that measures "mixed" is the entropy
which is largest, for a given spread, when the milk is spread as evenly as possible, and in the limit of a point. Along the diffusion it increases, and the rate of increase is
the Fisher information (Theorem 10.1). So entropy grows fastest where concentrations have sharp gradients: a thin streak of milk mixes much faster than a broad cloud, which is why stirring, which stretches the milk into thin streaks, speeds mixing so much.
This is the macroscopic face of the second law of thermodynamics, and the heat equation is a good model of it. It is not a derivation of the second law: the diffusion equation is itself an approximation to the reversible motion of molecules, and why that motion looks irreversible at large scales is a separate question in statistical mechanics. What this chapter takes from physics is the mathematics of the macroscopic description, which is exactly what Perelman borrowed.
Entropy and Fisher information
Let be a probability density on () with enough decay for the integrals below. Its entropy and Fisher information are
For the heat kernel , a Gaussian of variance in each coordinate (6A.3 The Heat Equation on ℝⁿ),
(Exercise 10.5). Among all densities with a given covariance, the Gaussian has the largest entropy (3A.7 Lᵖ Spaces and Jensen’s Inequality, by Jensen's inequality). Entropy can be any real number; it increases by when is spread out by a factor , because measures a logarithm of volume.
If solves on with (and suitable decay), then
Proof. , integrating by parts with no boundary terms. So . (The term .)
For the heat kernel this reads , as it must. The Fisher information itself is also monotone: it decreases, since
(Exercise 10.6). So is increasing and concave in . The computation behind this, a Bochner-type identity for , is the one that, on a manifold, produces a Ricci curvature term; with the sign survives. This is the first sign of how curvature enters entropy, which is the heart of 12A.3 The 𝓦-Entropy.
The logarithmic Sobolev inequality
Entropy production says . To turn this into a statement about rates of mixing, one needs an inequality comparing entropy with Fisher information. It is cleanest relative to a Gaussian.
Let be the standard Gaussian measure on , with density . For every smooth with ,
Equality holds for .
Two features make this inequality special. The constant does not depend on the dimension , unlike the constant in the Sobolev inequality (4A.10 Sobolev Embeddings and Critical Exponents), which degenerates as ; Gross's motivation was quantum field theory, where the dimension is infinite. And it controls , a quantity only "logarithmically" better than , which is the best one can hope for in infinite dimensions. It is equivalent to hypercontractivity of the heat flow for the Gaussian measure (Edward Nelson, 1973): the Ornstein–Uhlenbeck semigroup maps into for some after a positive time, with norm .
Now rewrite it in the language of Perelman. Let , and write a probability density on as
The reference density is the Gaussian of variance , which corresponds to : the heat kernel at time , or the backward Gaussian of 6A.3 The Heat Equation on ℝⁿ's rehearsal.
For every and every smooth with (and suitable decay),
with equality when .
Proof. Write and , and change variables to turn into . With , Gross's inequality becomes (Exercise 10.7)
Now , so
Since , the middle term is , integrating by parts. The left side is . The terms cancel, leaving , which is the claim. Translating gives the equality case at any .
Perelman's entropy of a metric on an -manifold, with a function and a scale , is
(12A.3 The 𝓦-Entropy). On flat the scalar curvature is zero, and Theorem 10.3 says exactly that , with equality for the Gaussian. Perelman's is therefore for flat space, and measures how far a manifold is, at scale , from being Euclidean. That is how detects collapsing, and the reason Perelman's noncollapsing theorem is a log-Sobolev inequality in disguise (12A.4 κ-Noncollapsing).
Nash's use of entropy
The log-Sobolev inequality and entropy production together control how fast a density spreads, and John Nash used this kind of argument in 1958 for a striking purpose: to prove that solutions of parabolic equations with merely bounded measurable coefficients are Hölder continuous (6A.6 Parabolic Regularity). Among his tools was the entropy of the fundamental solution: he showed that it increases at least like , just as for the heat kernel, using only the ellipticity bounds, and deduced that the fundamental solution spreads out at the parabolic rate and cannot concentrate. Entropy was a tool for regularity decades before Perelman used it for geometry, and the quantity , entropy measured relative to the heat kernel, is now called the Nash entropy. It reappears in the work of Hein and Naber and of Bamler on Ricci flows (12C.5 After Perelman).
In 1948 Claude Shannon introduced the entropy of a random signal as the measure of its information content, and stated that adding independent noise increases a signal's entropy power at least additively:
for independent random vectors and , with equality for Gaussians with proportional covariances. Aart Stam gave the first proof in 1959, using exactly the identity of this chapter: adding a small independent Gaussian is running the heat equation, and the entropy changes at the rate of the Fisher information (de Bruijn's identity). The inequality is used in information theory to bound the capacity of communication channels with noise; Shannon used it to show that, among noises of a given power, Gaussian noise is the most damaging.
The Li–Yau inequality
The Harnack inequality of 6A.2 Harmonic Functions compared a positive harmonic function at nearby points. For positive solutions of the heat equation, Peter Li and Shing-Tung Yau found in 1986 a sharp differential inequality, which integrates to a Harnack inequality across space and time.
Let solve on (or on a complete manifold with non-negative Ricci curvature). Then
Equality holds for the heat kernel .
The two forms agree because . For the heat kernel, , so exactly. The proof applies the maximum principle to , using the Bochner identity of 6A.2 Harmonic Functions for ; on a manifold the Bochner formula contributes , which is why is needed (9B.7 The Heat Equation on a Manifold).
Integrating it. Along any path from to with , the Li–Yau inequality bounds the change of , and optimising over straight paths gives
(Exercise 10.8). This is the parabolic Harnack inequality in sharp form: a positive solution now is bounded below by its value earlier and nearby, with the Gaussian factor that the heat kernel saturates (Figure 10.2).
Hamilton proved a matrix Harnack inequality for the Ricci flow with non-negative curvature operator (1993), modelled on Li–Yau; in its simplest trace form it says for every vector , with equality on expanding solitons. It is the key estimate for ancient solutions and for classifying singularity models (11B.2 Ancient Solutions and the Harnack Inequality). Perelman proved a Li–Yau type inequality for his conjugate heat kernel under the Ricci flow, again with equality on the Gaussian soliton, and used it in the proof of pseudolocality (12A.6 Pseudolocality).
The template
| monotone quantity | flow | equality case | where |
|---|---|---|---|
| entropy , Nash entropy | heat equation | Gaussian (heat kernel) | this chapter |
| Li–Yau quantity | heat equation | heat kernel | this chapter |
| Huisken's Gaussian density | mean curvature flow | self-shrinkers | 6A.8 Curve Shortening and the First Geometric Flows |
| Perelman's and | Ricci flow + conjugate heat equation | shrinking solitons | 12A.3 The 𝓦-Entropy |
| Perelman's reduced volume | Ricci flow | shrinking solitons (Gaussian on flat space) | 12A.5 Reduced Distance and Reduced Volume |
The pattern is the one first met as "monotone and bounded sequences converge" (2A.6 Sequences), with an extra ingredient: the equality case is a self-similar solution. So a solution along which the quantity becomes constant, such as a limit of rescalings near a singularity, must be self-similar. That is how Perelman shows that singularity models are shrinking solitons, and it is why entropy is at the centre of the proof.
History
Ludwig Boltzmann's entropy dates from the 1870s; Claude Shannon's information entropy and the entropy power inequality from 1948. Ronald Fisher introduced his information in statistics in the 1920s. Aart Stam's proof of the entropy power inequality, using de Bruijn's identity, appeared in 1959. John Nash's paper on continuity of solutions of parabolic and elliptic equations, with its entropy argument, appeared in 1958. Edward Nelson proved the hypercontractivity of the Ornstein–Uhlenbeck semigroup in 1973, and Leonard Gross his logarithmic Sobolev inequality, and its equivalence with hypercontractivity, in 1975. Peter Li and Shing-Tung Yau's inequality appeared in 1986, and Hamilton's matrix Harnack inequality for the Ricci flow in 1993. Perelman's -entropy appeared in his 2002 preprint.
A PDE is classified by its symbol; the heat equation is parabolic, scales like , is well posed forward and ill-posed backward (6A.1 What a PDE Is). Harmonic functions are averages and obey maximum principles, Harnack and Liouville (6A.2 Harmonic Functions). The heat kernel solves the heat equation on , smooths instantly and is the law of Brownian motion (6A.3 The Heat Equation on ℝⁿ). At a first interior maximum : maximum principles, comparison and invariant regions for systems follow (6A.4 Maximum Principles). Weak solutions exist by Lax–Milgram and are smooth by bootstrapping (6A.5 Weak Solutions and Elliptic Regularity); parabolic equations gain two space derivatives and one time derivative, and Bernstein's method bounds gradients by (6A.6 Parabolic Regularity). Nonlinear equations exist for short time by linearising and a fixed point, and may blow up (6A.7 Nonlinear Parabolic Equations). Curve shortening and mean curvature flow show solitons, ancient solutions, neckpinches and Huisken's monotonicity (6A.8 Curve Shortening and the First Geometric Flows). Energies give Euler–Lagrange equations and gradient flows (6A.9 Calculus of Variations and Gradient Flows). Entropy increases along diffusion at the rate of the Fisher information, the log-Sobolev inequality is on flat space, and the Li–Yau inequality is sharp on the heat kernel (this chapter).
The analysis is now ready: given a manifold, we could solve heat-type equations on it. But "manifold" has not been defined, and the Poincaré conjecture is about topology: a closed three-manifold in which every loop can be shrunk to a point is a sphere. Book 7A turns to topology, starting with only what differs from the metric spaces of Book 2B, and builds the fundamental group, covering spaces and the precise statement of the conjecture (7A.1 Topological Spaces and Quotients).
Exercises
For , compute and , using (6A.3 The Heat Equation on ℝⁿ). Check .
For a positive solution of on , write , so , and . Show that . (Differentiate, use the equations for and , and integrate by parts, using the flat Bochner identity from 6A.2 Harmonic Functions.) Check it on the heat kernel, where and .
Solution
On the heat kernel, , which matches. The general computation: ; integrate the first term by parts twice to , and use to combine; the Bochner identity turns the result into .
(a) With a probability density and (so ), show that and . So Gross's inequality reads . (b) Rescale to deduce for the Gaussian of variance .
Solution
(a) , so and . (b) Under , densities transform with the Jacobian, ratios are unchanged, the left side is invariant and , which turns the factor into .
Let satisfy . Along the straight path , , , with , show
(use for ), and integrate to get the Harnack inequality in the text. Check that the heat kernel, from to , makes it nearly sharp.
For a Gaussian random vector in with covariance , show and . Deduce that for independent such , with variances and , : equality in Shannon's inequality.
Solution
This is Exercise 10.5 with : , so and . is Gaussian with covariance .
Let and . Compute directly
using , and show it is : equality in Theorem 10.3, and . Then show that for (a Gaussian of the wrong width, still normalised) the integral is , with equality only at : the scale picks out the Gaussian of variance , exactly as in Perelman's (12A.3 The 𝓦-Entropy).
Solution
, so the integrand is , and . For , is the Gaussian of variance , with . Then , integrating to ; . Total: , which is because , with equality only at .
© 2026 NeckPinch (www.neckpinch.com). All content in the guidebook (text, mathematics, figures and exercises) is protected by copyright. All rights reserved. No part may be copied, republished or redistributed without written permission.