© 2026 NeckPinch · www.neckpinch.com · All rights reserved.
Course 3Book 3A: Measure, Integration and LᵖChapter 4
Measures, Probability and Weights
Abstract measures, densities, and the weighted measures e^(−f) dV that Perelman uses.
Read with Tao, An Introduction to Measure Theory, §1.4 "Abstract measure spaces" (Boolean algebras and σ-algebras, countably additive measures, measurable functions and integration, the convergence theorems) and §2.3 "Probability spaces". The Radon–Nikodym theorem is in Tao, An Epsilon of Room I, §1.2; read its statement now and its proof when you need it.
Everything in 3A.2 Lebesgue Measure and 3A.3 The Lebesgue Integral used only a few properties of Lebesgue measure: it assigns sizes to the sets of a σ-algebra, and it is countably additive. Translation invariance and boxes were needed to build it, but not to integrate against it. So the whole theory works for any countably additive measure on any set. This chapter makes that step, and it pays off at once: sums become integrals, probability becomes measure theory, and changing one measure into another becomes multiplication by a density.
The last idea is the one the guidebook needs most. Perelman never integrates against plain volume. He integrates against the weighted measure , a probability measure written with its density in exponential form, exactly as physicists write the Boltzmann distribution . The reasons are the same in both cases: energy and entropy live in the exponent, and normalising the total mass to makes the measure a probability. This chapter sets up that language.
By the end of this chapter you will be able to:
- define σ-algebras, measures and integrals on an abstract measure space, and recognise sums as integrals against counting measure;
- compute pushforward measures and use the change-of-variables formula for them;
- work with measures given by densities, state the Radon–Nikodym theorem, and recognise the Dirac mass as a limit of densities;
- translate between probability and measure theory: random variables, laws, expectation;
- reweight a measure by , normalise it, and explain why both statistical mechanics and Perelman write densities this way.
Simulating rare events
Engineers estimating the probability that a structure fails, or banks estimating the probability of an extreme loss, often can't compute it exactly and estimate it by simulation instead: draw many random samples and count how often the bad event happens. The trouble is that the event is rare. For a quantity with the standard normal distribution, the probability that is . To estimate it to within (relative standard error) by plain sampling, you need about million samples, because almost all of them land nowhere near the event.
Importance sampling draws from a different distribution instead, one that puts the samples where they matter, and corrects by a weight. Sample from a normal distribution centred at , and average , where the weight is the ratio of the two densities,
The average is an unbiased estimate of , because . Computing the variance shows that about samples now give the same accuracy, a saving by a factor of about . The method dates from the early days of Monte Carlo computation (Kahn and Marshall described it in 1953) and is standard in reliability engineering and financial risk. Mathematically it is a single identity: changing the measure you integrate against, and multiplying by the density of one measure with respect to the other. That density is what the Radon–Nikodym theorem below guarantees exists.
Measure spaces
A σ-algebra on a set is a collection of subsets of that contains and is closed under complements and countable unions. A measure on is a function with that is countably additive: for disjoint . The triple is a measure space. If it is a probability space.
The examples that matter:
- Lebesgue measure on with the Lebesgue measurable sets (3A.2 Lebesgue Measure), or its restriction to the Borel sets.
- Counting measure on any set: is the number of elements of (possibly ), with every subset measurable.
- The Dirac mass at a point : if , and otherwise. All the mass at one point.
- Restrictions and sums. If is a measure and is measurable, is a measure; positive combinations of measures are measures.
- Riemannian volume. On a Riemannian manifold, the volume measure (9A.1 Riemannian Metrics and Model Spaces), built from Lebesgue measure in coordinate charts.
A function is measurable if for all . The integral is defined exactly as in 3A.3 The Lebesgue Integral: by simple functions, then by supremum, then by splitting into positive and negative parts. The monotone convergence theorem, Fatou's lemma and the dominated convergence theorem hold, with the same proofs; Tao states them in this generality in §1.4.5. From now on "integrable" means , and is the space of such functions (modulo equality -almost everywhere).
Sums are integrals
Against counting measure on , a function is a sequence , and
So every theorem about integrals is also a theorem about series. Monotone convergence says that sums of non-negative terms can be computed by increasing truncations. Dominated convergence gives Tannery's theorem: if as for each , and with , then . (For example, , with .) And 3A.5 Product Measures and Change of Variables's Tonelli theorem, applied to counting measure twice, says double series of non-negative terms can be summed in either order (2A.8 Infinite Sets).
Pushforward measures
A measurable map carries a measure from one space to another.
Let be measurable (preimages of measurable sets are measurable) and a measure on . The pushforward is the measure on given by .
For every measurable (or ),
Proof. For this is the definition: both sides equal . By linearity it holds for simple , by monotone convergence for non-negative , and by splitting into positive and negative parts for integrable .
This proof pattern, indicators, then simple functions, then monotone limits, then differences, proves almost every identity about integrals. It is worth recognising on sight.
Let be Lebesgue measure on and . For , . A measure on whose distribution function is has density . So if is uniform on , then has density on : squares of uniform numbers pile up near .
Densities
If is measurable on a measure space , the formula defines a measure (countable additivity is monotone convergence). We write and call the density of with respect to . Then for every non-negative measurable .
The last identity is again proved by indicators, simple functions and monotone limits. Integrating against is the same as weighting the integrand by . For probability densities on , and is the probability that a random point lands in .
Which measures have densities? A necessary condition is that whenever : a density can't put mass on a set that doesn't see. This is called absolute continuity, . The Dirac mass is not absolutely continuous with respect to Lebesgue measure ( but ), so it has no density, however tempting it is to write "". The converse is the theorem.
Let and be σ-finite measures on the same σ-algebra, with . Then there is a measurable , unique -almost everywhere, with . It is written .
(σ-finite means the space is a countable union of sets of finite measure, as is for Lebesgue measure.) The proof, in Tao's An Epsilon of Room I, §1.2, uses either the Hilbert space or a decomposition argument; we use only the statement. In importance sampling, the weight is , and the identity that makes the method unbiased is .
The Dirac mass as a limit of densities
The Dirac mass has no density, but it is a limit of densities in a precise sense.
Let be integrable on with , and . Then for every bounded continuous ,
Proof. Substituting (the scaling rule of 3A.2 Lebesgue Measure, extended to integrals in 3A.5 Product Measures and Change of Variables), . The integrand converges to and is dominated by , which is integrable. By dominated convergence the integral tends to .
The rescaled family is the approximate identity of 2B.5 Uniform Convergence and Arzelà–Ascoli and 2B.7 Fourier Series and the First Heat Equation, now seen as a family of measures converging to a point mass. This kind of convergence, testing against continuous functions, is called weak convergence of measures. The heat kernel at time converges to in exactly this sense (6A.3 The Heat Equation on ℝⁿ), which is what it means for the heat equation to start from given initial data. In 12A.6 Pseudolocality, Perelman's conjugate heat kernel starts as a Dirac mass at a point and spreads backwards in time.
Probability as measure theory
Kolmogorov's 1933 axioms identify probability with measure theory of total mass .
| Probability | Measure theory |
|---|---|
| sample space , events, probability | measure space with |
| random variable | measurable function |
| law (distribution) of | pushforward on |
| density of | Radon–Nikodym derivative |
| expectation | |
| almost surely | -almost everywhere |
| independence | product measure (3A.5 Product Measures and Change of Variables) |
The pushforward formula, Proposition 4.3, is the statement that the expectation of can be computed either on the sample space or on the real line using the law of . Markov's inequality (3A.3 The Lebesgue Integral) becomes for , and applied to it becomes Chebyshev's inequality, from which the weak law of large numbers follows in a line (Exercise 4.11).
Weighted measures
A positive density can always be written as , with . There are good reasons to do so.
In a physical system in equilibrium with its surroundings at temperature , the probability of finding it in a state of energy is proportional to , where is Boltzmann's constant. As a measure on the space of states, the equilibrium distribution is
the normalising constant being the partition function. Writing the density as an exponential puts the energy in the exponent, where it adds: two independent systems have energies that add and densities that multiply. Low-energy states are exponentially favoured, and temperature sets the exchange rate. Quantities like the entropy and the free energy then become integrals of and against the measure, which is how thermodynamics is computed from statistical mechanics.
In Bayesian statistics, beliefs about an unknown parameter are a probability measure, the prior . Observing data with likelihood replaces it by the posterior
The posterior has density with respect to the prior: Bayes' rule is a Radon–Nikodym reweighting. Writing (with the negative log-likelihood) puts the data in the exponent, and the posterior is a Boltzmann distribution with as its energy, which is exactly how sampling algorithms for posteriors treat it.
In 12A.2 Ricci Flow as a Gradient Flow Perelman considers a Riemannian manifold with metric and a function , and the measure . His first move is to fix the measure: he lets evolve so that stays constant in time, and with that normalisation Ricci flow becomes the gradient flow of the functional . In 12A.3 The 𝓦-Entropy the weight becomes
a probability measure with a scale parameter , and the -entropy is an integral against it. The factor is the normalisation of the Gaussian (computed in 3A.5 Product Measures and Change of Variables): on flat , with , is exactly the Gaussian probability measure, the heat kernel. The language of this chapter (densities, normalisation, reweighting, integrals against a probability measure) is the language in which Perelman's entropy is written. Thread E runs from here through entropy and the log-Sobolev inequality (3A.7 Lᵖ Spaces and Jensen’s Inequality, 6A.10 Entropy, Information and Diffusion) to .
History
Johann Radon proved the density theorem for measures on in 1913, and Otto Nikodym extended it to abstract measure spaces in 1930. Maurice Fréchet had already defined integrals on abstract sets in 1915. Andrei Kolmogorov's Grundbegriffe der Wahrscheinlichkeitsrechnung (1933) founded probability on measure theory. Ludwig Boltzmann introduced the exponential distribution of energies in the 1860s and 1870s, and J. Willard Gibbs systematised it as the canonical ensemble in his Elementary Principles in Statistical Mechanics (1902). Importance sampling appears in Herman Kahn and Andy Marshall's 1953 paper on reducing sample sizes in Monte Carlo computations.
A measure space is a σ-algebra with a countably additive measure, and integration and its convergence theorems work in that generality. Counting measure turns sums into integrals; Dirac masses put all the mass at a point; measurable maps push measures forward, and expectations can be computed on either side. A measure with a density is ; the Radon–Nikodym theorem gives densities for absolutely continuous measures; concentrating densities converge to Dirac masses. Probability is measure theory with total mass . Writing densities as puts energy in the exponent, as in the Boltzmann distribution, Bayes' rule and Perelman's weighted measures. 3A.5 Product Measures and Change of Variables builds product measures, proves Fubini–Tonelli and the change-of-variables formula in , and computes the Gaussian integral that normalises .
Exercises
Compute for (a) and ; (b) (on ) and ; (c) and . Which of these measures have a density with respect to Lebesgue measure?
Solution
(a) . (b) . (c) . None of them: each puts positive mass on a Lebesgue-null set of points.
(a) Let be uniform on . Find the density of . (b) Let be uniform on . Find the density of on . (c) Explain why (a) is how computers generate exponentially distributed random numbers from uniform ones.
Solution
(a) , so the density is on , the exponential distribution. (b) , the arcsine density. (c) A uniform random number pushed forward by has exactly the exponential law (inverse transform sampling).
Use the dominated convergence theorem for counting measure to show , checking the domination carefully. (Show .)
Let be random variables with mean , variance , and for (for instance, independent). Let . Show , and use Chebyshev's inequality to show . Explain how this predicts that plain Monte Carlo needs about samples for accuracy on an event of probability .
Solution
Expanding the square, the cross terms vanish, leaving . For an indicator of an event of probability , , so the relative standard error is ; setting it to gives .
Let be σ-finite. Show that -almost everywhere. (Integrate an indicator against both sides.) In importance sampling, what are , and when one first changes from the true distribution to a proposal and then evaluates against Lebesgue measure?
Let be smooth on (or on a closed manifold) and , so that . Show that
Deduce that . (On , assume enough decay for the integrals to converge.) In 12A.2 Ricci Flow as a Gradient Flow this substitution turns Perelman's into with , a Rayleigh quotient, so its infimum is the lowest eigenvalue of the operator ; the same substitution rewrites the -entropy in 12A.3 The 𝓦-Entropy. The quantity for a probability density is the Fisher information, and this identity says it equals .
Solution
, so . Integrate.
© 2026 NeckPinch (www.neckpinch.com). All content in the guidebook (text, mathematics, figures and exercises) is protected by copyright. All rights reserved. No part may be copied, republished or redistributed without written permission.