© 2026 NeckPinch · www.neckpinch.com · All rights reserved.
Course 6Book 6A: The Heat Equation and Its RelativesChapter 6
Parabolic Regularity
Parabolic Hölder spaces, Schauder estimates and Bernstein’s gradient method.
Read with Evans, Partial Differential Equations, section 7.1 (second-order parabolic equations: weak solutions, Galerkin approximation, regularity, maximum principles). For the Hölder-space theory, use Lieberman, Second Order Parabolic Differential Equations, chapter 4, or Krylov, Lectures on Elliptic and Parabolic Equations in Hölder Spaces, as references for statements rather than cover-to-cover reading.
6A.5 Weak Solutions and Elliptic Regularity showed that elliptic equations gain two derivatives: if and has derivatives, has . Parabolic equations gain two derivatives in space and one in time, which is the same thing measured with the parabolic scaling of 6A.1 What a PDE Is, where one time derivative counts as two space derivatives. This chapter sets up the spaces in which that statement is exact, parabolic Hölder spaces, and states the two estimates that make the gain precise, Schauder and . These are the estimates quoted in every short-time existence proof, including Hamilton's and DeTurck's for the Ricci flow.
The chapter then proves, by a method short enough to do completely, a different kind of estimate: Bernstein's gradient bound, which says that a bounded solution of the heat equation has gradient at most , by applying the maximum principle not to but to a cleverly chosen combination of and . That method is exactly how Shi's derivative estimates for the Ricci flow are proved, and the chapter's rehearsal is Shi's argument in miniature.
By the end of this chapter you will be able to:
- describe weak solutions of parabolic equations and their construction by Galerkin approximation;
- define parabolic cylinders and parabolic Hölder spaces, and explain their scaling;
- state the interior Schauder and estimates and say what each gives;
- state the De Giorgi–Nash–Moser theorem and Moser's Harnack inequality, and say why they matter for nonlinear equations;
- prove and use it to bound by Bernstein's method.
Cooling rates in a quenched bar
Steel is hardened by heating it until its crystal structure changes and then cooling it fast enough that a hard phase, martensite, forms instead of softer ones. How fast is fast enough depends on the alloy, and how deep the hardening reaches in a real part depends on how the cooling rate falls off with depth. The standard way to measure this, the Jominy end-quench test (Walter Jominy and A. L. Boegehold, 1938; now ASTM A255 and ISO 642), heats a cylindrical bar and sprays water on one end only. The cooling rate is then highest at the quenched end and falls along the bar, and the hardness measured at intervals of inch from the end records, for that steel, the hardness produced by each cooling rate.
The cooling rate is a derivative of the temperature, , and its behaviour along the bar is a statement about derivatives of solutions of the heat equation. In the simplest model, a half-line initially at temperature above the quench temperature with its end held at the quench temperature (constant properties, no latent heat), the solution is (6A.3 The Heat Equation on ℝⁿ), and the maximum cooling rate at distance occurs at time and is proportional to (Exercise 6.5). Twice as far from the quenched end, the fastest cooling is four times slower: parabolic scaling again, now for a derivative. The real test differs from the model (heat is lost from the sides, properties vary with temperature, the phase change releases heat), which is why the test is done rather than computed; but the model explains why a hardness curve measured on a bar of one size can be transferred to parts of other sizes by matching cooling rates.
Weak solutions and their regularity
Consider the problem
with a uniformly elliptic operator as in 6A.5 Weak Solutions and Elliptic Regularity. A weak solution is a function with such that, for almost every ,
and . (Functions of with values in a Banach space are Bochner integrable functions, the vector-valued version of 3A.3 The Lebesgue Integral; Evans, section 5.9.)
The construction is Galerkin's method: take the first eigenfunctions of the elliptic operator (6A.5 Weak Solutions and Elliptic Regularity), look for satisfying the equation tested against , which is a linear system of ODEs for (2B.10 Ordinary Differential Equations), and pass to the limit using the energy estimate
which holds uniformly in and comes from testing the equation with itself, exactly as the energy decay in 6A.3 The Heat Equation on ℝⁿ. Bounded sequences in Hilbert spaces have weakly convergent subsequences (4A.6 Weak Convergence and the Direct Method), and the limit is a weak solution, which is unique (Evans, section 7.1.2).
Regularity then proceeds as in the elliptic case, testing with time and space difference quotients: with smooth data the weak solution is smooth, and, as for the heat equation on , interior smoothness in space holds for every positive time, even with rough initial data (Evans, section 7.1.3). This chapter now turns to estimates in the classical spaces, which are what nonlinear problems need.
Parabolic cylinders and Hölder spaces
Parabolic scaling (6A.1 What a PDE Is) says the right space-time distance is
which is unchanged by up to the factor . Its balls are, up to constants, the parabolic cylinders
of radius in space and in time, reaching backward from , because a parabolic equation is controlled by its past (Figure 6.1).
For and a region of space-time, is the space of functions with
that is, -Hölder in space and -Hölder in time (4A.11 Hölder Spaces). The space consists of functions for which , , and are continuous and , lie in .
The time exponents are half the space exponents, as scaling demands: one time derivative weighs as much as two space derivatives. The heat operator maps to , losing exactly two units of parabolic regularity, and the Schauder estimate says it loses no more.
The Schauder and estimates
Consider the non-divergence operator
uniformly parabolic: .
If the coefficients and lie in with in that norm, then every solution satisfies
with depending only on , , and .
The proof (Lieberman, chapter 4; Krylov) freezes the coefficients at a point, compares with the constant-coefficient equation, which is the heat equation after a linear change of variables and can be estimated using the heat kernel, and controls the error by the Hölder continuity of the coefficients on small cylinders. There are versions up to the boundary and for the initial time, with compatibility conditions on the data. The form to remember is: data in , solution in , with a scaled estimate on every cylinder obtained by applying the theorem to (Exercise 6.6). Written out for , it reads : every term has the units of .
The estimates are the other standard form: under continuity of ,
They come from the theory of singular integrals of Calderón and Zygmund, and fail at and for the same reason as fails (4A.11 Hölder Spaces). estimates with large , combined with the Sobolev embedding (4A.10 Sobolev Embeddings and Critical Exponents), give Hölder continuity of , which is often the first step of a bootstrap.
Both estimates need some continuity of the leading coefficients. For nonlinear equations that is a problem: the coefficients depend on the solution, whose continuity is what one is trying to prove. The breakthrough that resolves it is the De Giorgi–Nash–Moser theorem: for divergence-form equations with merely bounded measurable, uniformly elliptic coefficients, solutions are Hölder continuous, with estimates depending only on the ellipticity constants (De Giorgi 1957 for elliptic equations, Nash 1958 for parabolic ones, Moser's simplified proofs 1960–64). Its companion is Moser's parabolic Harnack inequality (1964): a positive solution on satisfies
where is a cylinder in the earlier part of and one in the later part, separated by a time gap: positive solutions cannot vary wildly, but heat must have time to arrive (6A.2 Harmonic Functions, 6A.10 Entropy, Information and Diffusion). Nash's proof used an entropy, which is the subject of 6A.10 Entropy, Information and Diffusion.
Bernstein's method
All of the above is quoted in Ricci flow papers. The next argument is used: its structure is the structure of Shi's derivative estimates, and of many curvature estimates after them.
If on a region of , then
Proof. Derivatives commute with in flat space, so . Then for any functions, . With and summing over : . With : .
Both quantities are subsolutions: and can only decrease along the flow, in the sense of the maximum principle. The good negative term in the evolution of is what lets one control :
Let solve with , either on (bounded, with bounded derivatives on each ) or on a flat torus. Then
Proof. Let . By Proposition 6.3,
So is a subsolution, and by the maximum principle (6A.4 Maximum Principles, in its version on for bounded functions or on a torus) . Hence .
Compare the exact bound from the heat kernel (6A.3 The Heat Equation on ℝⁿ, exercises): Bernstein's constant is slightly worse, but the method uses no formula for the solution, only the equation and the maximum principle. So it works on manifolds, for variable coefficients, and for nonlinear equations, where no kernel is available. The weight handles the initial time, where no gradient bound is assumed; the term supplies the negative that cancels the derivative of the weight. Higher derivatives are handled the same way, by induction, with (Exercise 6.9).
Under the Ricci flow, the curvature satisfies a reaction–diffusion equation , and its derivative satisfies . Wan-Xiong Shi (1989) proved: if on a time interval of length at most , then
and similarly for higher derivatives, by applying the maximum principle to : exactly Bernstein's function, with the reaction terms absorbed using the curvature bound (11A.3 Short-Time Existence and Uniqueness). The consequence is that a curvature bound controls all derivatives of curvature after a short time, which is why bounded curvature is the standard hypothesis and why compactness theorems for Ricci flows need only curvature and injectivity radius bounds (11B.3 Compactness of Ricci Flows).
History
Juliusz Schauder proved his estimates for elliptic equations in 1934; the parabolic theory was developed in the following decades and assembled in the 1967 monograph of Olga Ladyzhenskaya, Vsevolod Solonnikov and Nina Ural'tseva. Alberto Calderón and Antoni Zygmund's singular integrals date from 1952. Ennio De Giorgi (1957) and John Nash (1958) independently proved Hölder continuity for equations with bounded measurable coefficients, solving Hilbert's nineteenth problem; Jürgen Moser gave new proofs and the Harnack inequalities (elliptic 1961, parabolic 1964). Sergei Bernstein introduced his method of gradient bounds at the beginning of the twentieth century in his work on nonlinear elliptic equations. Shi's derivative estimates appeared in 1989. The Jominy test dates from 1938.
Weak solutions of parabolic equations are built by Galerkin approximation with energy estimates, and are smooth in space at positive times. Parabolic cylinders and the Hölder spaces , are adapted to parabolic scaling. The Schauder estimates gain two space derivatives and one time derivative in Hölder spaces, the estimates the same in for , and De Giorgi–Nash–Moser gives Hölder continuity with only bounded measurable coefficients. Bernstein's method bounds by using and the maximum principle; it is the method of Shi's estimates. 6A.7 Nonlinear Parabolic Equations uses these tools to solve nonlinear equations.
Exercises
With , show that , that for fixed it is largest at , and that the maximum is . How does the maximum cooling rate at cm compare with that at cm?
Solution
, giving the formula. Maximise with : the derivative of is zero at . Substituting, and , so the maximum is , as stated. At cm it is of that at cm.
Let solve on , and set on . (a) Find the equation satisfied by . (b) Show and . (c) Apply Theorem 6.2 to and rewrite the result for .
Solution
(a) . (b) Chain rule; Hölder quotients pick up from . (c) .
On with zero boundary values, use the eigenfunctions . Show that the Galerkin approximation of , , has , where , and that it converges to the Fourier series solution of 2B.7 Fourier Series and the First Heat Equation as .
Prove that . Then, for a solution of the heat equation and a smooth function , show that , and deduce that is a subsolution when is convex.
Solution
and . For : and , so for convex . The case is in Proposition 6.3.
Let with . (a) Show . (b) With , show for suitable constants , , and deduce .
Solution
(a) As in Proposition 6.3, with . (b) . With and every term is . So , and .
Suppose a tensor on a closed manifold evolves by and its derivative by , where denotes bilinear expressions with bounded coefficients. Assume the consequences
If for , show that satisfies for large (depending on ), up to a term bounded by , and conclude that for . This is the argument of Shi's estimates for (11A.3 Short-Time Existence and Uniqueness).
Solution
. Since , choosing makes the coefficient of non-positive, so . By the maximum principle , using . Hence .
© 2026 NeckPinch (www.neckpinch.com). All content in the guidebook (text, mathematics, figures and exercises) is protected by copyright. All rights reserved. No part may be copied, republished or redistributed without written permission.