Book 3A

© 2026 NeckPinch · www.neckpinch.com · All rights reserved.

Course 3Book 3A: Measure, Integration and LᵖChapter 4

Measures, Probability and Weights

Abstract measures, densities, and the weighted measures e^(−f) dV that Perelman uses.

23 min read · Updated Oct 2, 2026

Read with Tao, An Introduction to Measure Theory, §1.4 "Abstract measure spaces" (Boolean algebras and σ-algebras, countably additive measures, measurable functions and integration, the convergence theorems) and §2.3 "Probability spaces". The Radon–Nikodym theorem is in Tao, An Epsilon of Room I, §1.2; read its statement now and its proof when you need it.

In this chapter · 8 sections
  1. 4.1Simulating rare events
  2. 4.2Measure spaces
  3. 4.2.1Sums are integrals
  4. 4.3Pushforward measures
  5. 4.4Densities
  6. 4.4.1The Dirac mass as a limit of densities
  7. 4.5Probability as measure theory
  8. 4.6Weighted measures e−fe^{-f}e−f
  9. 4.7History
  10. 4.8Exercises

Everything in 3A.2 Lebesgue Measure and 3A.3 The Lebesgue Integral used only a few properties of Lebesgue measure: it assigns sizes to the sets of a σ-algebra, and it is countably additive. Translation invariance and boxes were needed to build it, but not to integrate against it. So the whole theory works for any countably additive measure on any set. This chapter makes that step, and it pays off at once: sums become integrals, probability becomes measure theory, and changing one measure into another becomes multiplication by a density.

The last idea is the one the guidebook needs most. Perelman never integrates against plain volume. He integrates against the weighted measure (4πτ)−n/2e−f dV(4\pi\tau)^{-n/2}e^{-f}\,dV, a probability measure written with its density in exponential form, exactly as physicists write the Boltzmann distribution e−E/kTe^{-E/kT}. The reasons are the same in both cases: energy and entropy live in the exponent, and normalising the total mass to 11 makes the measure a probability. This chapter sets up that language.

By the end of this chapter you will be able to:

  • define σ-algebras, measures and integrals on an abstract measure space, and recognise sums as integrals against counting measure;
  • compute pushforward measures and use the change-of-variables formula for them;
  • work with measures given by densities, state the Radon–Nikodym theorem, and recognise the Dirac mass as a limit of densities;
  • translate between probability and measure theory: random variables, laws, expectation;
  • reweight a measure by e−fe^{-f}, normalise it, and explain why both statistical mechanics and Perelman write densities this way.

Simulating rare events

In the world In use Importance sampling

Engineers estimating the probability that a structure fails, or banks estimating the probability of an extreme loss, often can't compute it exactly and estimate it by simulation instead: draw many random samples and count how often the bad event happens. The trouble is that the event is rare. For a quantity XX with the standard normal distribution, the probability that X>4X > 4 is p≈3.17×10−5p \approx 3.17 \times 10^{-5}. To estimate it to within 10%10\% (relative standard error) by plain sampling, you need about 3.23.2 million samples, because almost all of them land nowhere near the event.

Importance sampling draws from a different distribution instead, one that puts the samples where they matter, and corrects by a weight. Sample YY from a normal distribution centred at 44, and average w(Y) 1{Y>4}w(Y)\,1_{\{Y > 4\}}, where the weight is the ratio of the two densities,

w(y)=p(y)q(y)=e−y2/2e−(y−4)2/2=e8−4y.w(y) = \frac{p(y)}{q(y)} = \frac{e^{-y^2/2}}{e^{-(y-4)^2/2}} = e^{8 - 4y}.

The average is an unbiased estimate of pp, because ∫1{y>4}pq q dy=∫1{y>4}p dy\int 1_{\{y>4\}}\frac{p}{q}\,q\,dy = \int1_{\{y>4\}}p\,dy. Computing the variance shows that about 450450 samples now give the same 10%10\% accuracy, a saving by a factor of about 70007000. The method dates from the early days of Monte Carlo computation (Kahn and Marshall described it in 1953) and is standard in reliability engineering and financial risk. Mathematically it is a single identity: changing the measure you integrate against, and multiplying by the density of one measure with respect to the other. That density is what the Radon–Nikodym theorem below guarantees exists.

Measure spaces

Definition 4.1 σ-algebra, measure, measure space

A σ-algebra on a set XX is a collection B\mathcal{B} of subsets of XX that contains ∅\varnothing and is closed under complements and countable unions. A measure on (X,B)(X, \mathcal{B}) is a function μ:B→[0,+∞]\mu : \mathcal{B} \to [0, +\infty] with μ(∅)=0\mu(\varnothing) = 0 that is countably additive: μ(⋃nEn)=∑nμ(En)\mu\big(\bigcup_nE_n\big) = \sum_n\mu(E_n) for disjoint En∈BE_n \in \mathcal{B}. The triple (X,B,μ)(X, \mathcal{B}, \mu) is a measure space. If μ(X)=1\mu(X) = 1 it is a probability space.

The examples that matter:

  • Lebesgue measure on Rd\mathbb{R}^d with the Lebesgue measurable sets (3A.2 Lebesgue Measure), or its restriction to the Borel sets.
  • Counting measure on any set: #(E)\#(E) is the number of elements of EE (possibly ∞\infty), with every subset measurable.
  • The Dirac mass at a point x0x_0: δx0(E)=1\delta_{x_0}(E) = 1 if x0∈Ex_0 \in E, and 00 otherwise. All the mass at one point.
  • Restrictions and sums. If μ\mu is a measure and AA is measurable, μ⌞A(E)=μ(E∩A)\mu\llcorner A(E) = \mu(E \cap A) is a measure; positive combinations ∑iciμi\sum_ic_i\mu_i of measures are measures.
  • Riemannian volume. On a Riemannian manifold, the volume measure dVdV (9A.1 Riemannian Metrics and Model Spaces), built from Lebesgue measure in coordinate charts.

A function f:X→[0,+∞]f : X \to [0, +\infty] is measurable if {f>t}∈B\{f > t\} \in \mathcal{B} for all tt. The integral ∫f dμ\int f\,d\mu is defined exactly as in 3A.3 The Lebesgue Integral: by simple functions, then by supremum, then by splitting into positive and negative parts. The monotone convergence theorem, Fatou's lemma and the dominated convergence theorem hold, with the same proofs; Tao states them in this generality in §1.4.5. From now on "integrable" means ∫∣f∣ dμ<∞\int|f|\,d\mu < \infty, and L1(μ)L^1(\mu) is the space of such functions (modulo equality μ\mu-almost everywhere).

Sums are integrals

Against counting measure on N\mathbb{N}, a function is a sequence ana_n, and

∫a d#=∑nan.\int a\,d\# = \sum_na_n.

So every theorem about integrals is also a theorem about series. Monotone convergence says that sums of non-negative terms can be computed by increasing truncations. Dominated convergence gives Tannery's theorem: if an,k→aka_{n,k} \to a_k as n→∞n \to \infty for each kk, and ∣an,k∣≤Mk|a_{n,k}| \leq M_k with ∑kMk<∞\sum_kM_k < \infty, then ∑kan,k→∑kak\sum_ka_{n,k} \to \sum_ka_k. (For example, (1+xn)n=∑k(nk)xknk→∑kxkk!=ex\big(1 + \frac xn\big)^n = \sum_k\binom nk\frac{x^k}{n^k} \to \sum_k\frac{x^k}{k!} = e^x, with Mk=∣x∣kk!M_k = \frac{|x|^k}{k!}.) And 3A.5 Product Measures and Change of Variables's Tonelli theorem, applied to counting measure twice, says double series of non-negative terms can be summed in either order (2A.8 Infinite Sets).

Pushforward measures

A measurable map carries a measure from one space to another.

Definition 4.2 Pushforward

Let ϕ:X→Y\phi : X \to Y be measurable (preimages of measurable sets are measurable) and μ\mu a measure on XX. The pushforward ϕ∗μ\phi_*\mu is the measure on YY given by ϕ∗μ(A)=μ(ϕ−1(A))\phi_*\mu(A) = \mu(\phi^{-1}(A)).

Proposition 4.3 Change of variables for pushforwards

For every measurable g:Y→[0,∞]g : Y \to [0, \infty] (or g∈L1(ϕ∗μ)g \in L^1(\phi_*\mu)),

∫Yg d(ϕ∗μ)=∫Xg∘ϕ dμ.\int_Yg\,d(\phi_*\mu) = \int_Xg\circ\phi\,d\mu.

Proof. For g=1Ag = 1_A this is the definition: both sides equal μ(ϕ−1(A))\mu(\phi^{-1}(A)). By linearity it holds for simple gg, by monotone convergence for non-negative gg, and by splitting into positive and negative parts for integrable gg.

This proof pattern, indicators, then simple functions, then monotone limits, then differences, proves almost every identity about integrals. It is worth recognising on sight.

Example 4.4 Squaring a uniform random number

Let μ\mu be Lebesgue measure on [0,1][0, 1] and ϕ(x)=x2\phi(x) = x^2. For 0≤a≤10 \leq a \leq 1, ϕ∗μ([0,a])=m({x:x2≤a})=a\phi_*\mu([0, a]) = m(\{x : x^2 \leq a\}) = \sqrt a. A measure on [0,1][0, 1] whose distribution function is a\sqrt a has density ddaa=12a\frac{d}{da}\sqrt a = \frac{1}{2\sqrt a}. So if UU is uniform on [0,1][0, 1], then U2U^2 has density 12y\frac{1}{2\sqrt y} on (0,1](0, 1]: squares of uniform numbers pile up near 00.

Densities

Definition 4.5 Measure with a density

If ρ≥0\rho \geq 0 is measurable on a measure space (X,μ)(X, \mu), the formula ν(E)=∫Eρ dμ\nu(E) = \int_E\rho\,d\mu defines a measure (countable additivity is monotone convergence). We write dν=ρ dμd\nu = \rho\,d\mu and call ρ\rho the density of ν\nu with respect to μ\mu. Then ∫g dν=∫gρ dμ\int g\,d\nu = \int g\rho\,d\mu for every non-negative measurable gg.

The last identity is again proved by indicators, simple functions and monotone limits. Integrating against ρ dμ\rho\,d\mu is the same as weighting the integrand by ρ\rho. For probability densities on Rd\mathbb{R}^d, ∫ρ dx=1\int\rho\,dx = 1 and ν(E)\nu(E) is the probability that a random point lands in EE.

Which measures have densities? A necessary condition is that ν(E)=0\nu(E) = 0 whenever μ(E)=0\mu(E) = 0: a density can't put mass on a set that μ\mu doesn't see. This is called absolute continuity, ν≪μ\nu \ll \mu. The Dirac mass δ0\delta_0 is not absolutely continuous with respect to Lebesgue measure (m({0})=0m(\{0\}) = 0 but δ0({0})=1\delta_0(\{0\}) = 1), so it has no density, however tempting it is to write "δ(x)\delta(x)". The converse is the theorem.

Theorem 4.6 Radon–Nikodym theorem

Let μ\mu and ν\nu be σ-finite measures on the same σ-algebra, with ν≪μ\nu \ll \mu. Then there is a measurable ρ≥0\rho \geq 0, unique μ\mu-almost everywhere, with dν=ρ dμd\nu = \rho\,d\mu. It is written ρ=dνdμ\rho = \frac{d\nu}{d\mu}.

(σ-finite means the space is a countable union of sets of finite measure, as Rd\mathbb{R}^d is for Lebesgue measure.) The proof, in Tao's An Epsilon of Room I, §1.2, uses either the Hilbert space L2L^2 or a decomposition argument; we use only the statement. In importance sampling, the weight p/qp/q is dPdQ\frac{dP}{dQ}, and the identity that makes the method unbiased is ∫g dP=∫g dPdQ dQ\int g\,dP = \int g\,\frac{dP}{dQ}\,dQ.

The Dirac mass as a limit of densities

The Dirac mass has no density, but it is a limit of densities in a precise sense.

Proposition 4.7 Concentrating densities converge to δ\delta

Let ρ≥0\rho \geq 0 be integrable on Rd\mathbb{R}^d with ∫ρ=1\int\rho = 1, and ρε(x)=ε−dρ(x/ε)\rho_\varepsilon(x) = \varepsilon^{-d}\rho(x/\varepsilon). Then for every bounded continuous gg,

∫g ρε dx⟶g(0)=∫g dδ0(ε→0).\int g\,\rho_\varepsilon\,dx \longrightarrow g(0) = \int g\,d\delta_0 \qquad (\varepsilon \to 0).

Proof. Substituting x=εyx = \varepsilon y (the scaling rule m(λE)=λdm(E)m(\lambda E) = \lambda^dm(E) of 3A.2 Lebesgue Measure, extended to integrals in 3A.5 Product Measures and Change of Variables), ∫gρε dx=∫g(εy)ρ(y) dy\int g\rho_\varepsilon\,dx = \int g(\varepsilon y)\rho(y)\,dy. The integrand converges to g(0)ρ(y)g(0)\rho(y) and is dominated by (sup⁡∣g∣)ρ(\sup|g|)\rho, which is integrable. By dominated convergence the integral tends to g(0)∫ρ=g(0)g(0)\int\rho = g(0).

Figure 4.1. Gaussian densities with standard deviations 1,12,14,1101, \tfrac12, \tfrac14, \tfrac1{10}, each of total mass 11. Integrated against a continuous function, they give values closer and closer to the function's value at 00: they converge to the Dirac mass δ0\delta_0 (arrow) in the sense of Proposition 4.7.

The rescaled family ρε\rho_\varepsilon is the approximate identity of 2B.5 Uniform Convergence and Arzelà–Ascoli and 2B.7 Fourier Series and the First Heat Equation, now seen as a family of measures converging to a point mass. This kind of convergence, testing against continuous functions, is called weak convergence of measures. The heat kernel at time t→0t \to 0 converges to δ0\delta_0 in exactly this sense (6A.3 The Heat Equation on ℝⁿ), which is what it means for the heat equation to start from given initial data. In 12A.6 Pseudolocality, Perelman's conjugate heat kernel starts as a Dirac mass at a point and spreads backwards in time.

Probability as measure theory

Kolmogorov's 1933 axioms identify probability with measure theory of total mass 11.

Probability Measure theory
sample space Ω\Omega, events, probability P\mathbb{P} measure space with P(Ω)=1\mathbb{P}(\Omega) = 1
random variable XX measurable function X:Ω→RX : \Omega \to \mathbb{R}
law (distribution) of XX pushforward X∗PX_*\mathbb{P} on R\mathbb{R}
density of XX Radon–Nikodym derivative dX∗Pdm\frac{dX_*\mathbb{P}}{dm}
expectation E[g(X)]\mathbb{E}[g(X)] ∫g∘X dP=∫g d(X∗P)\int g\circ X\,d\mathbb{P} = \int g\,d(X_*\mathbb{P})
almost surely P\mathbb{P}-almost everywhere
independence product measure (3A.5 Product Measures and Change of Variables)

The pushforward formula, Proposition 4.3, is the statement that the expectation of g(X)g(X) can be computed either on the sample space or on the real line using the law of XX. Markov's inequality (3A.3 The Lebesgue Integral) becomes P(X≥t)≤E[X]/t\mathbb{P}(X \geq t) \leq \mathbb{E}[X]/t for X≥0X \geq 0, and applied to (X−EX)2(X - \mathbb{E}X)^2 it becomes Chebyshev's inequality, from which the weak law of large numbers follows in a line (Exercise 4.11).

Weighted measures e−fe^{-f}

A positive density can always be written as ρ=e−f\rho = e^{-f}, with f=−log⁡ρf = -\log\rho. There are good reasons to do so.

In the world Model The Boltzmann distribution

In a physical system in equilibrium with its surroundings at temperature TT, the probability of finding it in a state of energy EE is proportional to e−E/kTe^{-E/kT}, where kk is Boltzmann's constant. As a measure on the space of states, the equilibrium distribution is

dμ=1Z e−E/kT d(states),Z=∫e−E/kT d(states),d\mu = \frac{1}{Z}\,e^{-E/kT}\,d(\text{states}), \qquad Z = \int e^{-E/kT}\,d(\text{states}),

the normalising constant ZZ being the partition function. Writing the density as an exponential puts the energy in the exponent, where it adds: two independent systems have energies that add and densities that multiply. Low-energy states are exponentially favoured, and temperature sets the exchange rate. Quantities like the entropy −∫ρlog⁡ρ-\int\rho\log\rho and the free energy then become integrals of EE and log⁡Z\log Z against the measure, which is how thermodynamics is computed from statistical mechanics.

Figure 4.2. Densities ρ=e−f/Z\rho = e^{-f}/Z for three potentials ff (faint curves): a single well, a double well, and a tilted double well. Wherever ff is low, the measure is concentrated. Changing ff by a constant changes nothing after normalisation; changing it by a small function reweights the measure smoothly.
In the world In use Bayesian updating is reweighting

In Bayesian statistics, beliefs about an unknown parameter θ\theta are a probability measure, the prior π(dθ)\pi(d\theta). Observing data DD with likelihood L(D∣θ)L(D \mid \theta) replaces it by the posterior

π(dθ∣D)=L(D∣θ)∫L(D∣θ′) π(dθ′) π(dθ).\pi(d\theta \mid D) = \frac{L(D \mid \theta)}{\int L(D \mid \theta')\,\pi(d\theta')}\,\pi(d\theta).

The posterior has density L/ZL/Z with respect to the prior: Bayes' rule is a Radon–Nikodym reweighting. Writing L=e−ℓL = e^{-\ell} (with ℓ\ell the negative log-likelihood) puts the data in the exponent, and the posterior is a Boltzmann distribution with ℓ\ell as its energy, which is exactly how sampling algorithms for posteriors treat it.

Where this goes Perelman's probability measures

In 12A.2 Ricci Flow as a Gradient Flow Perelman considers a Riemannian manifold with metric gg and a function ff, and the measure e−f dVe^{-f}\,dV. His first move is to fix the measure: he lets ff evolve so that e−fdVe^{-f}dV stays constant in time, and with that normalisation Ricci flow becomes the gradient flow of the functional F(g,f)=∫(R+∣∇f∣2)e−fdV\mathcal{F}(g, f) = \int(R + |\nabla f|^2)e^{-f}dV. In 12A.3 The 𝓦-Entropy the weight becomes

dμ=(4πτ)−n/2 e−f dV,∫dμ=1,d\mu = (4\pi\tau)^{-n/2}\,e^{-f}\,dV, \qquad \int d\mu = 1,

a probability measure with a scale parameter τ\tau, and the W\mathcal{W}-entropy is an integral against it. The factor (4πτ)−n/2(4\pi\tau)^{-n/2} is the normalisation of the Gaussian (computed in 3A.5 Product Measures and Change of Variables): on flat Rn\mathbb{R}^n, with f=∣x∣2/4τf = |x|^2/4\tau, dμd\mu is exactly the Gaussian probability measure, the heat kernel. The language of this chapter (densities, normalisation, reweighting, integrals against a probability measure) is the language in which Perelman's entropy is written. Thread E runs from here through entropy and the log-Sobolev inequality (3A.7 Lᵖ Spaces and Jensen’s Inequality, 6A.10 Entropy, Information and Diffusion) to W\mathcal{W}.

History

Johann Radon proved the density theorem for measures on Rn\mathbb{R}^n in 1913, and Otto Nikodym extended it to abstract measure spaces in 1930. Maurice Fréchet had already defined integrals on abstract sets in 1915. Andrei Kolmogorov's Grundbegriffe der Wahrscheinlichkeitsrechnung (1933) founded probability on measure theory. Ludwig Boltzmann introduced the exponential distribution of energies in the 1860s and 1870s, and J. Willard Gibbs systematised it as the canonical ensemble in his Elementary Principles in Statistical Mechanics (1902). Importance sampling appears in Herman Kahn and Andy Marshall's 1953 paper on reducing sample sizes in Monte Carlo computations.

Recall Where we stand

A measure space is a σ-algebra with a countably additive measure, and integration and its convergence theorems work in that generality. Counting measure turns sums into integrals; Dirac masses put all the mass at a point; measurable maps push measures forward, and expectations can be computed on either side. A measure with a density is ρ dμ\rho\,d\mu; the Radon–Nikodym theorem gives densities for absolutely continuous measures; concentrating densities converge to Dirac masses. Probability is measure theory with total mass 11. Writing densities as e−fe^{-f} puts energy in the exponent, as in the Boltzmann distribution, Bayes' rule and Perelman's weighted measures. 3A.5 Product Measures and Change of Variables builds product measures, proves Fubini–Tonelli and the change-of-variables formula in Rn\mathbb{R}^n, and computes the Gaussian integral that normalises (4πτ)−n/2e−f(4\pi\tau)^{-n/2}e^{-f}.

Exercises

Exercise 4.8 Integrating against simple measures

Compute ∫f dμ\int f\,d\mu for (a) μ=δ2\mu = \delta_2 and f(x)=x2f(x) = x^2; (b) μ=∑n≥12−nδn\mu = \sum_{n\geq1}2^{-n}\delta_n (on R\mathbb{R}) and f(x)=xf(x) = x; (c) μ=12δ0+12m⌞[0,1]\mu = \tfrac12\delta_0 + \tfrac12m\llcorner[0, 1] and f(x)=exf(x) = e^x. Which of these measures have a density with respect to Lebesgue measure?

Solution

(a) 44. (b) ∑n2−n=2\sum n2^{-n} = 2. (c) 12+12(e−1)=e2\tfrac12 + \tfrac12(e - 1) = \tfrac e2. None of them: each puts positive mass on a Lebesgue-null set of points.

Exercise 4.9 Pushforwards

(a) Let UU be uniform on [0,1][0, 1]. Find the density of −log⁡U-\log U. (b) Let Θ\Theta be uniform on [0,2π)[0, 2\pi). Find the density of cos⁡Θ\cos\Theta on (−1,1)(-1, 1). (c) Explain why (a) is how computers generate exponentially distributed random numbers from uniform ones.

Solution

(a) P(−log⁡U>t)=P(U<e−t)=e−t\mathbb{P}(-\log U > t) = \mathbb{P}(U < e^{-t}) = e^{-t}, so the density is e−te^{-t} on [0,∞)[0, \infty), the exponential distribution. (b) 1π1−y2\frac{1}{\pi\sqrt{1 - y^2}}, the arcsine density. (c) A uniform random number pushed forward by −log⁡-\log has exactly the exponential law (inverse transform sampling).

Exercise 4.10 Tannery's theorem

Use the dominated convergence theorem for counting measure to show lim⁡n→∞∑k=0n(nk)1nk=e\lim_{n\to\infty}\sum_{k=0}^n\binom nk\frac{1}{n^k} = e, checking the domination carefully. (Show (nk)n−k≤1k!\binom nk n^{-k} \leq \frac1{k!}.)

Exercise 4.11 The weak law of large numbers

Let X1,X2,…X_1, X_2, \ldots be random variables with mean EXi=m\mathbb{E}X_i = m, variance E(Xi−m)2=σ2\mathbb{E}(X_i - m)^2 = \sigma^2, and E[(Xi−m)(Xj−m)]=0\mathbb{E}[(X_i - m)(X_j - m)] = 0 for i≠ji \neq j (for instance, independent). Let Xˉn=1n∑i≤nXi\bar X_n = \frac1n\sum_{i\leq n}X_i. Show E(Xˉn−m)2=σ2/n\mathbb{E}(\bar X_n - m)^2 = \sigma^2/n, and use Chebyshev's inequality to show P(∣Xˉn−m∣≥t)≤σ2nt2→0\mathbb{P}(|\bar X_n - m| \geq t) \leq \frac{\sigma^2}{nt^2} \to 0. Explain how this predicts that plain Monte Carlo needs about 100/p100/p samples for 10%10\% accuracy on an event of probability pp.

Solution

Expanding the square, the cross terms vanish, leaving 1n2⋅nσ2\frac{1}{n^2}\cdot n\sigma^2. For an indicator of an event of probability pp, σ2=p(1−p)\sigma^2 = p(1 - p), so the relative standard error is (1−p)/(np)\sqrt{(1-p)/(np)}; setting it to 0.10.1 gives n≈100/pn \approx 100/p.

Exercise 4.12 Absolute continuity and the chain rule

Let λ≪ν≪μ\lambda \ll \nu \ll \mu be σ-finite. Show that dλdμ=dλdνdνdμ\frac{d\lambda}{d\mu} = \frac{d\lambda}{d\nu}\frac{d\nu}{d\mu} μ\mu-almost everywhere. (Integrate an indicator against both sides.) In importance sampling, what are λ\lambda, ν\nu and μ\mu when one first changes from the true distribution to a proposal and then evaluates against Lebesgue measure?

Exercise 4.13 Rehearsal: two ways to write the same weighted energy

Let ff be smooth on Rn\mathbb{R}^n (or on a closed manifold) and u=e−f/2u = e^{-f/2}, so that u2=e−fu^2 = e^{-f}. Show that

∣∇f∣2e−f=4∣∇u∣2.|\nabla f|^2e^{-f} = 4|\nabla u|^2.

Deduce that ∫∣∇f∣2e−f dV=4∫∣∇u∣2 dV\int|\nabla f|^2e^{-f}\,dV = 4\int|\nabla u|^2\,dV. (On Rn\mathbb{R}^n, assume enough decay for the integrals to converge.) In 12A.2 Ricci Flow as a Gradient Flow this substitution turns Perelman's F(g,f)=∫(R+∣∇f∣2)e−fdV\mathcal{F}(g, f) = \int(R + |\nabla f|^2)e^{-f}dV into ∫(4∣∇u∣2+Ru2) dV\int(4|\nabla u|^2 + Ru^2)\,dV with ∫u2 dV=1\int u^2\,dV = 1, a Rayleigh quotient, so its infimum λ\lambda is the lowest eigenvalue of the operator −4Δ+R-4\Delta + R; the same substitution rewrites the W\mathcal{W}-entropy in 12A.3 The 𝓦-Entropy. The quantity ∫∣∇log⁡ρ∣2ρ\int|\nabla\log\rho|^2\rho for a probability density ρ\rho is the Fisher information, and this identity says it equals 4∫∣∇ρ∣24\int|\nabla\sqrt\rho|^2.

Solution

∇u=−12∇f e−f/2\nabla u = -\tfrac12\nabla f\,e^{-f/2}, so ∣∇u∣2=14∣∇f∣2e−f|\nabla u|^2 = \tfrac14|\nabla f|^2e^{-f}. Integrate.

© 2026 NeckPinch (www.neckpinch.com). All content in the guidebook (text, mathematics, figures and exercises) is protected by copyright. All rights reserved. No part may be copied, republished or redistributed without written permission.