Book 6A

© 2026 NeckPinch · www.neckpinch.com · All rights reserved.

Course 6Book 6A: The Heat Equation and Its RelativesChapter 6

Parabolic Regularity

Parabolic Hölder spaces, Schauder estimates and Bernstein’s gradient method.

28 min read · Updated Oct 2, 2026

Read with Evans, Partial Differential Equations, section 7.1 (second-order parabolic equations: weak solutions, Galerkin approximation, regularity, maximum principles). For the Hölder-space theory, use Lieberman, Second Order Parabolic Differential Equations, chapter 4, or Krylov, Lectures on Elliptic and Parabolic Equations in Hölder Spaces, as references for statements rather than cover-to-cover reading.

In this chapter · 7 sections
  1. 6.1Cooling rates in a quenched bar
  2. 6.2Weak solutions and their regularity
  3. 6.3Parabolic cylinders and Hölder spaces
  4. 6.4The Schauder and LpL^pLp estimates
  5. 6.5Bernstein's method
  6. 6.6History
  7. 6.7Exercises

6A.5 Weak Solutions and Elliptic Regularity showed that elliptic equations gain two derivatives: if Lu=fLu = f and ff has kk derivatives, uu has k+2k + 2. Parabolic equations gain two derivatives in space and one in time, which is the same thing measured with the parabolic scaling of 6A.1 What a PDE Is, where one time derivative counts as two space derivatives. This chapter sets up the spaces in which that statement is exact, parabolic Hölder spaces, and states the two estimates that make the gain precise, Schauder and LpL^p. These are the estimates quoted in every short-time existence proof, including Hamilton's and DeTurck's for the Ricci flow.

The chapter then proves, by a method short enough to do completely, a different kind of estimate: Bernstein's gradient bound, which says that a bounded solution of the heat equation has gradient at most C/tC/\sqrt t, by applying the maximum principle not to uu but to a cleverly chosen combination of uu and ∣∇u∣2|\nabla u|^2. That method is exactly how Shi's derivative estimates for the Ricci flow are proved, and the chapter's rehearsal is Shi's argument in miniature.

By the end of this chapter you will be able to:

  • describe weak solutions of parabolic equations and their construction by Galerkin approximation;
  • define parabolic cylinders and parabolic Hölder spaces, and explain their scaling;
  • state the interior Schauder and LpL^p estimates and say what each gives;
  • state the De Giorgi–Nash–Moser theorem and Moser's Harnack inequality, and say why they matter for nonlinear equations;
  • prove (∂t−Δ)∣∇u∣2=−2∣∇2u∣2(\partial_t - \Delta)|\nabla u|^2 = -2|\nabla^2u|^2 and use it to bound ∣∇u∣|\nabla u| by Bernstein's method.

Cooling rates in a quenched bar

In the world Model The Jominy end-quench test

Steel is hardened by heating it until its crystal structure changes and then cooling it fast enough that a hard phase, martensite, forms instead of softer ones. How fast is fast enough depends on the alloy, and how deep the hardening reaches in a real part depends on how the cooling rate falls off with depth. The standard way to measure this, the Jominy end-quench test (Walter Jominy and A. L. Boegehold, 1938; now ASTM A255 and ISO 642), heats a cylindrical bar and sprays water on one end only. The cooling rate is then highest at the quenched end and falls along the bar, and the hardness measured at intervals of 1/161/16 inch from the end records, for that steel, the hardness produced by each cooling rate.

The cooling rate is a derivative of the temperature, ∣ut∣|u_t|, and its behaviour along the bar is a statement about derivatives of solutions of the heat equation. In the simplest model, a half-line x>0x > 0 initially at temperature T0T_0 above the quench temperature with its end held at the quench temperature (constant properties, no latent heat), the solution is T0erf⁡(x/2κt)T_0\operatorname{erf}(x/2\sqrt{\kappa t}) (6A.3 The Heat Equation on ℝⁿ), and the maximum cooling rate at distance xx occurs at time t=x26κt = \frac{x^2}{6\kappa} and is proportional to κT0x2\frac{\kappa T_0}{x^2} (Exercise 6.5). Twice as far from the quenched end, the fastest cooling is four times slower: parabolic scaling again, now for a derivative. The real test differs from the model (heat is lost from the sides, properties vary with temperature, the phase change releases heat), which is why the test is done rather than computed; but the model explains why a hardness curve measured on a bar of one size can be transferred to parts of other sizes by matching cooling rates.

Weak solutions and their regularity

Consider the problem

ut+Lu=f  in UT,u=0  on ∂U×[0,T],u=g  at t=0,u_t + Lu = f \ \text{ in } U_T, \qquad u = 0 \ \text{ on } \partial U\times[0, T], \qquad u = g \ \text{ at } t = 0,

with LL a uniformly elliptic operator as in 6A.5 Weak Solutions and Elliptic Regularity. A weak solution is a function u∈L2(0,T;H01(U))u \in L^2(0, T; H^1_0(U)) with ut∈L2(0,T;H−1(U))u_t \in L^2(0, T; H^{-1}(U)) such that, for almost every tt,

⟨ut,v⟩+B[u,v]=∫Ufvfor all v∈H01(U),\langle u_t, v\rangle + B[u, v] = \int_Ufv \qquad\text{for all } v \in H^1_0(U),

and u(0)=gu(0) = g. (Functions of tt with values in a Banach space are Bochner integrable functions, the vector-valued version of 3A.3 The Lebesgue Integral; Evans, section 5.9.)

The construction is Galerkin's method: take the first mm eigenfunctions w1,…,wmw_1, \dots, w_m of the elliptic operator (6A.5 Weak Solutions and Elliptic Regularity), look for um(t)=∑k≤mdk(t)wku_m(t) = \sum_{k\leq m}d_k(t)w_k satisfying the equation tested against w1,…,wmw_1, \dots, w_m, which is a linear system of ODEs for dk(t)d_k(t) (2B.10 Ordinary Differential Equations), and pass to the limit m→∞m \to \infty using the energy estimate

max⁡0≤t≤T∥um(t)∥L22+∫0T∥um∥H012 dt≤C(∥g∥L22+∫0T∥f∥L22 dt),\max_{0\leq t\leq T}\|u_m(t)\|_{L^2}^2 + \int_0^T\|u_m\|_{H^1_0}^2\,dt \leq C\Big(\|g\|_{L^2}^2 + \int_0^T\|f\|_{L^2}^2\,dt\Big),

which holds uniformly in mm and comes from testing the equation with umu_m itself, exactly as the energy decay in 6A.3 The Heat Equation on ℝⁿ. Bounded sequences in Hilbert spaces have weakly convergent subsequences (4A.6 Weak Convergence and the Direct Method), and the limit is a weak solution, which is unique (Evans, section 7.1.2).

Regularity then proceeds as in the elliptic case, testing with time and space difference quotients: with smooth data the weak solution is smooth, and, as for the heat equation on Rn\mathbb{R}^n, interior smoothness in space holds for every positive time, even with rough initial data (Evans, section 7.1.3). This chapter now turns to estimates in the classical spaces, which are what nonlinear problems need.

Parabolic cylinders and Hölder spaces

Parabolic scaling (6A.1 What a PDE Is) says the right space-time distance is

d((x,t),(y,s))=∣x−y∣+∣t−s∣1/2,d\big((x, t), (y, s)\big) = |x - y| + |t - s|^{1/2},

which is unchanged by (x,t)↦(λx,λ2t)(x, t) \mapsto (\lambda x, \lambda^2t) up to the factor λ\lambda. Its balls are, up to constants, the parabolic cylinders

Qr(x0,t0)=Br(x0)×(t0−r2,t0],Q_r(x_0, t_0) = B_r(x_0)\times(t_0 - r^2, t_0],

of radius rr in space and r2r^2 in time, reaching backward from t0t_0, because a parabolic equation is controlled by its past (Figure 6.1).

Figure 6.1. Parabolic cylinders QrQ_r and Q2rQ_{2r} ending at the same time t0t_0: doubling the radius doubles the width and quadruples the height. Interior estimates bound a solution on the inner cylinder by its size on the outer one.
Definition 6.1 Parabolic Hölder spaces

For 0<α<10 < \alpha < 1 and a region QQ of space-time, Cα,α/2(Q)C^{\alpha,\alpha/2}(Q) is the space of functions with

[f]α,α/2=sup⁡(x,t)≠(y,s)∣f(x,t)−f(y,s)∣∣x−y∣α+∣t−s∣α/2<∞,[f]_{\alpha,\alpha/2} = \sup_{(x,t)\neq(y,s)}\frac{|f(x, t) - f(y, s)|}{|x - y|^\alpha + |t - s|^{\alpha/2}} < \infty,

that is, α\alpha-Hölder in space and α2\frac\alpha2-Hölder in time (4A.11 Hölder Spaces). The space C2+α,1+α/2(Q)C^{2+\alpha,1+\alpha/2}(Q) consists of functions uu for which uu, ∇u\nabla u, ∇2u\nabla^2u and utu_t are continuous and ∇2u\nabla^2u, utu_t lie in Cα,α/2C^{\alpha,\alpha/2}.

The time exponents are half the space exponents, as scaling demands: one time derivative weighs as much as two space derivatives. The heat operator ∂t−Δ\partial_t - \Delta maps C2+α,1+α/2C^{2+\alpha,1+\alpha/2} to Cα,α/2C^{\alpha,\alpha/2}, losing exactly two units of parabolic regularity, and the Schauder estimate says it loses no more.

The Schauder and LpL^p estimates

Consider the non-divergence operator

ut−∑i,jaij(x,t)∂i∂ju−∑ibi(x,t)∂iu−c(x,t)u=f,u_t - \sum_{i,j}a^{ij}(x, t)\partial_i\partial_ju - \sum_ib^i(x, t)\partial_iu - c(x, t)u = f,

uniformly parabolic: ∑aijξiξj≥θ∣ξ∣2\sum a^{ij}\xi_i\xi_j \geq \theta|\xi|^2.

Theorem 6.2 Interior Schauder estimates

If the coefficients and ff lie in Cα,α/2(Q1)C^{\alpha,\alpha/2}(Q_1) with ∥aij∥,∥bi∥,∥c∥≤Λ\|a^{ij}\|, \|b^i\|, \|c\| \leq \Lambda in that norm, then every solution u∈C2+α,1+α/2(Q1)u \in C^{2+\alpha,1+\alpha/2}(Q_1) satisfies

∥u∥C2+α,1+α/2(Q1/2)≤C(∥f∥Cα,α/2(Q1)+sup⁡Q1∣u∣),\|u\|_{C^{2+\alpha,1+\alpha/2}(Q_{1/2})} \leq C\big(\|f\|_{C^{\alpha,\alpha/2}(Q_1)} + \sup_{Q_1}|u|\big),

with CC depending only on nn, α\alpha, θ\theta and Λ\Lambda.

The proof (Lieberman, chapter 4; Krylov) freezes the coefficients at a point, compares with the constant-coefficient equation, which is the heat equation after a linear change of variables and can be estimated using the heat kernel, and controls the error by the Hölder continuity of the coefficients on small cylinders. There are versions up to the boundary and for the initial time, with compatibility conditions on the data. The form to remember is: data in CαC^\alpha, solution in C2+αC^{2+\alpha}, with a scaled estimate on every cylinder QrQ_r obtained by applying the theorem to u(x0+rx,t0+r2t)u(x_0 + rx, t_0 + r^2t) (Exercise 6.6). Written out for ut=Δu+fu_t = \Delta u + f, it reads r2+α[∇2u]α;Qr/2≤C(sup⁡Qr∣u∣+r2sup⁡Qr∣f∣+r2+α[f]α;Qr)r^{2+\alpha}[\nabla^2u]_{\alpha;Q_{r/2}} \leq C\big(\sup_{Q_r}|u| + r^2\sup_{Q_r}|f| + r^{2+\alpha}[f]_{\alpha;Q_r}\big): every term has the units of uu.

Figure 6.2. What the estimates give. In Hölder spaces and in LpL^p for 1<p<∞1 < p < \infty, the heat operator gains exactly two space derivatives and one time derivative. In C0C^0 (and in L1L^1, L∞L^\infty) it does not: continuous data need not give C2,1C^{2,1} solutions (4A.11 Hölder Spaces).

The LpL^p estimates are the other standard form: under continuity of aija^{ij},

∥ut∥Lp(Q1/2)+∥∇2u∥Lp(Q1/2)≤C(∥f∥Lp(Q1)+∥u∥Lp(Q1)),1<p<∞.\|u_t\|_{L^p(Q_{1/2})} + \|\nabla^2u\|_{L^p(Q_{1/2})} \leq C\big(\|f\|_{L^p(Q_1)} + \|u\|_{L^p(Q_1)}\big), \qquad 1 < p < \infty.

They come from the theory of singular integrals of Calderón and Zygmund, and fail at p=1p = 1 and p=∞p = \infty for the same reason as C2C^2 fails (4A.11 Hölder Spaces). LpL^p estimates with large pp, combined with the Sobolev embedding (4A.10 Sobolev Embeddings and Critical Exponents), give Hölder continuity of ∇u\nabla u, which is often the first step of a bootstrap.

Both estimates need some continuity of the leading coefficients. For nonlinear equations that is a problem: the coefficients depend on the solution, whose continuity is what one is trying to prove. The breakthrough that resolves it is the De Giorgi–Nash–Moser theorem: for divergence-form equations ut=∂j(aij(x,t)∂iu)u_t = \partial_j(a^{ij}(x, t)\partial_iu) with merely bounded measurable, uniformly elliptic coefficients, solutions are Hölder continuous, with estimates depending only on the ellipticity constants (De Giorgi 1957 for elliptic equations, Nash 1958 for parabolic ones, Moser's simplified proofs 1960–64). Its companion is Moser's parabolic Harnack inequality (1964): a positive solution on Q2rQ_{2r} satisfies

sup⁡Q−u≤Cinf⁡Q+u,\sup_{Q^-}u \leq C\inf_{Q^+}u,

where Q−Q^- is a cylinder in the earlier part of Q2rQ_{2r} and Q+Q^+ one in the later part, separated by a time gap: positive solutions cannot vary wildly, but heat must have time to arrive (6A.2 Harmonic Functions, 6A.10 Entropy, Information and Diffusion). Nash's proof used an entropy, which is the subject of 6A.10 Entropy, Information and Diffusion.

Bernstein's method

All of the above is quoted in Ricci flow papers. The next argument is used: its structure is the structure of Shi's derivative estimates, and of many curvature estimates after them.

Proposition 6.3 The evolution of ∣∇u∣2|\nabla u|^2

If ut=Δuu_t = \Delta u on a region of Rn\mathbb{R}^n, then

(∂t−Δ)∣∇u∣2=−2∣∇2u∣2,(∂t−Δ)u2=−2∣∇u∣2.(\partial_t - \Delta)|\nabla u|^2 = -2|\nabla^2u|^2, \qquad (\partial_t - \Delta)u^2 = -2|\nabla u|^2.

Proof. Derivatives commute with ∂t−Δ\partial_t - \Delta in flat space, so (∂t−Δ)∂iu=0(\partial_t - \Delta)\partial_iu = 0. Then for any functions, (∂t−Δ)(vw)=v(∂t−Δ)w+w(∂t−Δ)v−2∇v⋅∇w(\partial_t - \Delta)(vw) = v(\partial_t - \Delta)w + w(\partial_t - \Delta)v - 2\nabla v\cdot\nabla w. With v=w=∂iuv = w = \partial_iu and summing over ii: (∂t−Δ)∣∇u∣2=−2∑i,j(∂j∂iu)2(\partial_t - \Delta)|\nabla u|^2 = -2\sum_{i,j}(\partial_j\partial_iu)^2. With v=w=uv = w = u: (∂t−Δ)u2=−2∣∇u∣2(\partial_t - \Delta)u^2 = -2|\nabla u|^2.

Both quantities are subsolutions: ∣∇u∣2|\nabla u|^2 and u2u^2 can only decrease along the flow, in the sense of the maximum principle. The good negative term −2∣∇u∣2-2|\nabla u|^2 in the evolution of u2u^2 is what lets one control ∣∇u∣2|\nabla u|^2:

Theorem 6.4 Bernstein's gradient estimate

Let uu solve ut=Δuu_t = \Delta u with ∣u∣≤M|u| \leq M, either on Rn×(0,T]\mathbb{R}^n\times(0, T] (bounded, with bounded derivatives on each [ε,T][\varepsilon, T]) or on a flat torus. Then

∣∇u(x,t)∣≤M2t.|\nabla u(x, t)| \leq \frac{M}{\sqrt{2t}}.

Proof. Let F=t∣∇u∣2+12u2F = t|\nabla u|^2 + \frac12u^2. By Proposition 6.3,

(∂t−Δ)F=∣∇u∣2−2t∣∇2u∣2−∣∇u∣2=−2t∣∇2u∣2≤0.(\partial_t - \Delta)F = |\nabla u|^2 - 2t|\nabla^2u|^2 - |\nabla u|^2 = -2t|\nabla^2u|^2 \leq 0.

So FF is a subsolution, and by the maximum principle (6A.4 Maximum Principles, in its version on Rn\mathbb{R}^n for bounded functions or on a torus) F(x,t)≤sup⁡F(⋅,0)=12sup⁡u(⋅,0)2≤12M2F(x, t) \leq \sup F(\cdot, 0) = \frac12\sup u(\cdot, 0)^2 \leq \frac12M^2. Hence t∣∇u∣2≤12M2t|\nabla u|^2 \leq \frac12M^2.

Compare the exact bound ∣∇u∣≤Mπt|\nabla u| \leq \frac{M}{\sqrt{\pi t}} from the heat kernel (6A.3 The Heat Equation on ℝⁿ, exercises): Bernstein's constant is slightly worse, but the method uses no formula for the solution, only the equation and the maximum principle. So it works on manifolds, for variable coefficients, and for nonlinear equations, where no kernel is available. The weight tt handles the initial time, where no gradient bound is assumed; the term 12u2\frac12u^2 supplies the negative ∣∇u∣2|\nabla u|^2 that cancels the derivative of the weight. Higher derivatives are handled the same way, by induction, with Fk=tk∣∇ku∣2+cktk−1∣∇k−1u∣2+…F_k = t^k|\nabla^ku|^2 + c_kt^{k-1}|\nabla^{k-1}u|^2 + \dots (Exercise 6.9).

Figure 6.3. Heat flow from the step g=sign⁡(x)g = \operatorname{sign}(x) (M=1M = 1): the exact maximal gradient max⁡x∣ux∣=1πt\max_x|u_x| = \frac{1}{\sqrt{\pi t}} (solid) and Bernstein's bound 12t\frac{1}{\sqrt{2t}} (dashed), on logarithmic axes. Both have slope −12-\frac12: each derivative decays like t−1/2t^{-1/2}, as parabolic scaling requires.
Where this goes Shi's derivative estimates

Under the Ricci flow, the curvature satisfies a reaction–diffusion equation ∂tRm⁡=ΔRm⁡+Rm⁡∗Rm⁡\partial_t\operatorname{Rm} = \Delta\operatorname{Rm} + \operatorname{Rm}*\operatorname{Rm}, and its derivative satisfies ∂t∇Rm⁡=Δ∇Rm⁡+Rm⁡∗∇Rm⁡\partial_t\nabla\operatorname{Rm} = \Delta\nabla\operatorname{Rm} + \operatorname{Rm}*\nabla\operatorname{Rm}. Wan-Xiong Shi (1989) proved: if ∣Rm⁡∣≤K|\operatorname{Rm}| \leq K on a time interval of length at most 1/K1/K, then

∣∇Rm⁡∣2≤CK2t,|\nabla\operatorname{Rm}|^2 \leq \frac{CK^2}{t},

and similarly for higher derivatives, by applying the maximum principle to F=t∣∇Rm⁡∣2+β∣Rm⁡∣2F = t|\nabla\operatorname{Rm}|^2 + \beta|\operatorname{Rm}|^2: exactly Bernstein's function, with the reaction terms absorbed using the curvature bound (11A.3 Short-Time Existence and Uniqueness). The consequence is that a curvature bound controls all derivatives of curvature after a short time, which is why bounded curvature is the standard hypothesis and why compactness theorems for Ricci flows need only curvature and injectivity radius bounds (11B.3 Compactness of Ricci Flows).

History

Juliusz Schauder proved his estimates for elliptic equations in 1934; the parabolic theory was developed in the following decades and assembled in the 1967 monograph of Olga Ladyzhenskaya, Vsevolod Solonnikov and Nina Ural'tseva. Alberto Calderón and Antoni Zygmund's singular integrals date from 1952. Ennio De Giorgi (1957) and John Nash (1958) independently proved Hölder continuity for equations with bounded measurable coefficients, solving Hilbert's nineteenth problem; Jürgen Moser gave new proofs and the Harnack inequalities (elliptic 1961, parabolic 1964). Sergei Bernstein introduced his method of gradient bounds at the beginning of the twentieth century in his work on nonlinear elliptic equations. Shi's derivative estimates appeared in 1989. The Jominy test dates from 1938.

Recall Where we stand

Weak solutions of parabolic equations are built by Galerkin approximation with energy estimates, and are smooth in space at positive times. Parabolic cylinders Qr=Br×(t0−r2,t0]Q_r = B_r\times(t_0 - r^2, t_0] and the Hölder spaces Cα,α/2C^{\alpha,\alpha/2}, C2+α,1+α/2C^{2+\alpha,1+\alpha/2} are adapted to parabolic scaling. The Schauder estimates gain two space derivatives and one time derivative in Hölder spaces, the LpL^p estimates the same in LpL^p for 1<p<∞1 < p < \infty, and De Giorgi–Nash–Moser gives Hölder continuity with only bounded measurable coefficients. Bernstein's method bounds ∣∇u∣|\nabla u| by M/2tM/\sqrt{2t} using (∂t−Δ)∣∇u∣2=−2∣∇2u∣2(\partial_t - \Delta)|\nabla u|^2 = -2|\nabla^2u|^2 and the maximum principle; it is the method of Shi's estimates. 6A.7 Nonlinear Parabolic Equations uses these tools to solve nonlinear equations.

Exercises

Exercise 6.5 Cooling rate in the end-quench model

With u=T0erf⁡(x/2κt)u = T_0\operatorname{erf}(x/2\sqrt{\kappa t}), show that ∣ut(x,t)∣=T0x2πκt−3/2e−x2/4κt|u_t(x, t)| = \frac{T_0x}{2\sqrt{\pi\kappa}}t^{-3/2}e^{-x^2/4\kappa t}, that for fixed xx it is largest at t=x26κt = \frac{x^2}{6\kappa}, and that the maximum is (6e)3/2κT02π x2\big(\frac{6}{e}\big)^{3/2}\frac{\kappa T_0}{2\sqrt\pi\,x^2}. How does the maximum cooling rate at 33 cm compare with that at 11 cm?

Solution

∂terf⁡(x/2κt)=2πe−x2/4κt⋅(−x4κt−3/2)\partial_t\operatorname{erf}(x/2\sqrt{\kappa t}) = \frac{2}{\sqrt\pi}e^{-x^2/4\kappa t}\cdot\big(-\frac{x}{4\sqrt\kappa}t^{-3/2}\big), giving the formula. Maximise t−3/2e−a/tt^{-3/2}e^{-a/t} with a=x24κa = \frac{x^2}{4\kappa}: the derivative of −32log⁡t−at-\frac32\log t - \frac at is zero at t=2a3=x26κt = \frac{2a}{3} = \frac{x^2}{6\kappa}. Substituting, t−3/2=(6κ)3/2x−3t^{-3/2} = (6\kappa)^{3/2}x^{-3} and e−a/t=e−3/2e^{-a/t} = e^{-3/2}, so the maximum is T0x2πκ(6κ)3/2x−3e−3/2\frac{T_0x}{2\sqrt{\pi\kappa}}(6\kappa)^{3/2}x^{-3}e^{-3/2}, as stated. At 33 cm it is 19\frac19 of that at 11 cm.

Exercise 6.6 Scaling the Schauder estimate

Let uu solve ut=Δu+fu_t = \Delta u + f on Qr(0,0)Q_r(0, 0), and set v(x,t)=u(rx,r2t)v(x, t) = u(rx, r^2t) on Q1Q_1. (a) Find the equation satisfied by vv. (b) Show ∇2v(x,t)=r2(∇2u)(rx,r2t)\nabla^2v(x, t) = r^2(\nabla^2u)(rx, r^2t) and [∇2v]α;Q1/2=r2+α[∇2u]α;Qr/2[\nabla^2v]_{\alpha;Q_{1/2}} = r^{2+\alpha}[\nabla^2u]_{\alpha;Q_{r/2}}. (c) Apply Theorem 6.2 to vv and rewrite the result for uu.

Solution

(a) vt=Δv+r2f(rx,r2t)v_t = \Delta v + r^2f(rx, r^2t). (b) Chain rule; Hölder quotients pick up rαr^\alpha from ∣x−y∣α|x - y|^\alpha. (c) r2sup⁡Qr/2∣∇2u∣+r2+α[∇2u]α;Qr/2≤C(sup⁡Qr∣u∣+r2sup⁡Qr∣f∣+r2+α[f]α;Qr)r^2\sup_{Q_{r/2}}|\nabla^2u| + r^{2+\alpha}[\nabla^2u]_{\alpha;Q_{r/2}} \leq C\big(\sup_{Q_r}|u| + r^2\sup_{Q_r}|f| + r^{2+\alpha}[f]_{\alpha;Q_r}\big).

Exercise 6.7 Galerkin for the heat equation on an interval

On (0,π)(0, \pi) with zero boundary values, use the eigenfunctions wk=sin⁡kxw_k = \sin kx. Show that the Galerkin approximation um=∑k≤mdk(t)sin⁡kxu_m = \sum_{k\leq m}d_k(t)\sin kx of ut=uxxu_t = u_{xx}, u(0)=gu(0) = g, has dk(t)=e−k2tg^kd_k(t) = e^{-k^2t}\hat g_k, where g^k=2π∫0πgsin⁡kx dx\hat g_k = \frac2\pi\int_0^\pi g\sin kx\,dx, and that it converges to the Fourier series solution of 2B.7 Fourier Series and the First Heat Equation as m→∞m \to \infty.

Exercise 6.8 The product rule for the heat operator

Prove that (∂t−Δ)(vw)=v(∂t−Δ)w+w(∂t−Δ)v−2∇v⋅∇w(\partial_t - \Delta)(vw) = v(\partial_t - \Delta)w + w(\partial_t - \Delta)v - 2\nabla v\cdot\nabla w. Then, for a solution uu of the heat equation and a smooth function ϕ\phi, show that (∂t−Δ)ϕ(u)=−ϕ′′(u)∣∇u∣2(\partial_t - \Delta)\phi(u) = -\phi''(u)|\nabla u|^2, and deduce that ϕ(u)\phi(u) is a subsolution when ϕ\phi is convex.

Solution

∂t(vw)=vtw+vwt\partial_t(vw) = v_tw + vw_t and Δ(vw)=vΔw+wΔv+2∇v⋅∇w\Delta(vw) = v\Delta w + w\Delta v + 2\nabla v\cdot\nabla w. For ϕ(u)\phi(u): ∂tϕ(u)=ϕ′ut\partial_t\phi(u) = \phi'u_t and Δϕ(u)=ϕ′Δu+ϕ′′∣∇u∣2\Delta\phi(u) = \phi'\Delta u + \phi''|\nabla u|^2, so (∂t−Δ)ϕ(u)=−ϕ′′(u)∣∇u∣2≤0(\partial_t - \Delta)\phi(u) = -\phi''(u)|\nabla u|^2 \leq 0 for convex ϕ\phi. The case ϕ(u)=u2\phi(u) = u^2 is in Proposition 6.3.

Exercise 6.9 The second derivative

Let ut=Δuu_t = \Delta u with ∣u∣≤M|u| \leq M. (a) Show (∂t−Δ)∣∇2u∣2=−2∣∇3u∣2(\partial_t - \Delta)|\nabla^2u|^2 = -2|\nabla^3u|^2. (b) With G=t2∣∇2u∣2+At∣∇u∣2+Bu2G = t^2|\nabla^2u|^2 + At|\nabla u|^2 + Bu^2, show (∂t−Δ)G≤0(\partial_t - \Delta)G \leq 0 for suitable constants AA, BB, and deduce ∣∇2u∣≤CMt|\nabla^2u| \leq \frac{CM}{t}.

Solution

(a) As in Proposition 6.3, with v=w=∂i∂juv = w = \partial_i\partial_ju. (b) (∂t−Δ)G=2t∣∇2u∣2−2t2∣∇3u∣2+A∣∇u∣2−2At∣∇2u∣2−2B∣∇u∣2(\partial_t - \Delta)G = 2t|\nabla^2u|^2 - 2t^2|\nabla^3u|^2 + A|\nabla u|^2 - 2At|\nabla^2u|^2 - 2B|\nabla u|^2. With A=1A = 1 and B=12B = \frac12 every term is ≤0\leq 0. So G≤sup⁡G(0)=BM2G \leq \sup G(0) = BM^2, and t2∣∇2u∣2≤12M2t^2|\nabla^2u|^2 \leq \frac12M^2.

Exercise 6.10 Rehearsal: Shi's estimate in miniature

Suppose a tensor TT on a closed manifold evolves by ∂tT=ΔT+T∗T\partial_tT = \Delta T + T*T and its derivative by ∂t∇T=Δ∇T+T∗∇T\partial_t\nabla T = \Delta\nabla T + T*\nabla T, where ∗* denotes bilinear expressions with bounded coefficients. Assume the consequences

(∂t−Δ)∣T∣2≤−2∣∇T∣2+C∣T∣3,(∂t−Δ)∣∇T∣2≤−2∣∇2T∣2+C∣T∣∣∇T∣2.(\partial_t - \Delta)|T|^2 \leq -2|\nabla T|^2 + C|T|^3, \qquad (\partial_t - \Delta)|\nabla T|^2 \leq -2|\nabla^2T|^2 + C|T||\nabla T|^2.

If ∣T∣≤K|T| \leq K for 0≤t≤1/K0 \leq t \leq 1/K, show that F=t∣∇T∣2+β∣T∣2F = t|\nabla T|^2 + \beta|T|^2 satisfies (∂t−Δ)F≤0(\partial_t - \Delta)F \leq 0 for β\beta large (depending on CC), up to a term bounded by CβK3C\beta K^3, and conclude that ∣∇T∣2≤C′K2t|\nabla T|^2 \leq \frac{C'K^2}{t} for 0<t≤1/K0 < t \leq 1/K. This is the argument of Shi's estimates for T=Rm⁡T = \operatorname{Rm} (11A.3 Short-Time Existence and Uniqueness).

Solution

(∂t−Δ)F≤∣∇T∣2+CtK∣∇T∣2−2β∣∇T∣2+CβK3(\partial_t - \Delta)F \leq |\nabla T|^2 + Ct K|\nabla T|^2 - 2\beta|\nabla T|^2 + C\beta K^3. Since tK≤1tK \leq 1, choosing β≥1+C2\beta \geq \frac{1 + C}{2} makes the coefficient of ∣∇T∣2|\nabla T|^2 non-positive, so (∂t−Δ)F≤CβK3(\partial_t - \Delta)F \leq C\beta K^3. By the maximum principle F≤sup⁡F(0)+CβK3t≤βK2+CβK2=C′′K2F \leq \sup F(0) + C\beta K^3t \leq \beta K^2 + C\beta K^2 = C''K^2, using t≤1/Kt \leq 1/K. Hence t∣∇T∣2≤C′′K2t|\nabla T|^2 \leq C''K^2.

© 2026 NeckPinch (www.neckpinch.com). All content in the guidebook (text, mathematics, figures and exercises) is protected by copyright. All rights reserved. No part may be copied, republished or redistributed without written permission.