Book 2B

© 2026 NeckPinch · www.neckpinch.com · All rights reserved.

Course 2Book 2B: Spaces, Functions and ChangeChapter 7

Fourier Series and the First Heat Equation

Fourier’s solution of heat flow on a ring: smoothing, energy decay and the first heat kernel.

37 min read · Updated Oct 2, 2026

Read with Tao, Analysis II, chapter "Fourier series": periodic functions, inner products on periodic functions, trigonometric polynomials, periodic convolutions (where the Fejér kernel appears), and the Fourier and Plancherel theorems. The heat equation is not in Tao; this chapter adds it.

In this chapter · 7 sections
  1. 7.1Seasons underground
  2. 7.2Periodic functions and their coefficients
  3. 7.3Does the Fourier series converge?
  4. 7.4Fourier's solution of the heat equation on a ring
  5. 7.4.1What the formula shows
  6. 7.5The Gibbs phenomenon
  7. 7.6History
  8. 7.7Exercises

In 1807 Joseph Fourier submitted a memoir to the Paris Academy claiming that the temperature in a solid could be found by writing the initial temperature as a sum of sines and cosines, and letting each one decay at its own rate. The claim that an arbitrary function could be written that way was greeted with scepticism, and making sense of it occupied analysts for a century: it is one of the main reasons the concepts of function, convergence and integral in Books 2A and 2B were made precise at all. Fourier's book, Théorie analytique de la chaleur, appeared in 1822.

This chapter does Fourier's analysis in the setting of periodic functions, and then solves Fourier's problem: heat flow on a ring. Here the heat equation, the equation behind thread H of this guidebook and, eventually, behind Ricci flow, is solved completely and explicitly. Every property of heat flow that later chapters prove with harder tools can be seen here in the formula: high frequencies die fastest, so the solution becomes smooth instantly; an energy decreases; the maximum can only fall; and the solution is an average of the initial data against a kernel. When Hamilton described Ricci flow as a heat equation for the metric (11A.1 The Equation and Its First Solutions), these are the properties he meant.

By the end of this chapter you will be able to:

  • compute Fourier coefficients, and use orthonormality of e2πikxe^{2\pi ikx} to find best approximations;
  • explain why Fourier series of continuous functions need not converge, and prove that their Fejér means do;
  • prove Plancherel's identity and use it to sum series such as ∑1/k2\sum 1/k^2;
  • solve the heat equation on a ring by Fourier series, and prove instant smoothing, energy decay, the maximum principle and exponential convergence to the mean;
  • explain the Gibbs phenomenon, and why positive kernels avoid it.

Seasons underground

In the world Model Why a cellar is cool in summer

Over a year, the temperature at the ground's surface rises and falls roughly like a sine wave. Below the surface, heat moves by conduction, so the temperature u(z,t)u(z, t) at depth zz obeys the heat equation

∂u∂t=κ ∂2u∂z2,\frac{\partial u}{\partial t} = \kappa\,\frac{\partial^2 u}{\partial z^2},

where κ\kappa is the soil's thermal diffusivity. If the surface temperature is T0+Acos⁡(ωt)T_0 + A\cos(\omega t), with ω=2π/(1 year)\omega = 2\pi/(1 \text{ year}), a solution is

u(z,t)=T0+A e−z/dcos⁡ ⁣(ωt−zd),d=2κωu(z, t) = T_0 + A\,e^{-z/d}\cos\!\Big(\omega t - \frac zd\Big), \qquad d = \sqrt{\frac{2\kappa}{\omega}}

(Exercise 7.11). The temperature wave travels downwards and dies out as it goes. At depth zz its amplitude is reduced by the factor e−z/de^{-z/d} and it is delayed by the fraction z2πd\frac{z}{2\pi d} of a year. The damping depth dd sets the scale.

For a soil diffusivity of κ=5×10−7\kappa = 5 \times 10^{-7} m²/s, a typical value for moist soils, d≈2.2d \approx 2.2 m. At a depth of πd≈7\pi d \approx 7 m, the wave is delayed by half a year, so the seasons there are reversed: the warmest time is midwinter. But by then the amplitude is only e−π≈4%e^{-\pi} \approx 4\% of the surface swing, so the ground at that depth is close to the annual mean temperature all year round. A cellar a few metres down lags the seasons and swings much less; ground-source heat pumps exploit the same near-constant temperature; and the depth to which permafrost thaws each summer is governed by the same law. For the daily temperature cycle, ω\omega is 365365 times larger and dd is 365≈19\sqrt{365} \approx 19 times smaller, about 1212 cm: a day's heat barely gets below the topsoil.

The decisive feature is that each frequency decays at its own rate, faster for higher frequencies: the damping depth is proportional to 1/ω1/\sqrt\omega. On a ring of metal, where heat flows around instead of down, the same principle gives the complete solution of the heat equation. To use it we need to break an arbitrary periodic function into frequencies.

Periodic functions and their coefficients

Work with functions of period 11: f(x+1)=f(x)f(x + 1) = f(x). Equivalently, functions on the circle R/Z\mathbb{R}/\mathbb{Z}, the real line with xx and x+1x + 1 identified (2A.2 Sets, Functions and Equivalence). Let C(R/Z)C(\mathbb{R}/\mathbb{Z}) be the continuous periodic functions with complex values, a complete metric space with the sup metric (2B.5 Uniform Convergence and Arzelà–Ascoli).

Definition 7.1 Inner product and L2L^2 norm

For f,g∈C(R/Z)f, g \in C(\mathbb{R}/\mathbb{Z}),

⟨f,g⟩=∫01f(x) g(x)‾ dx,∥f∥2=⟨f,f⟩1/2=(∫01∣f∣2)1/2.\langle f, g\rangle = \int_0^1 f(x)\,\overline{g(x)}\,dx, \qquad \|f\|_2 = \langle f, f\rangle^{1/2} = \Big(\int_0^1|f|^2\Big)^{1/2}.

This behaves like the dot product on Cn\mathbb{C}^n: it is linear in the first slot, conjugate-symmetric, and ⟨f,f⟩>0\langle f, f\rangle > 0 unless f=0f = 0 (by the positivity argument of 2B.1 Metric Spaces). The Cauchy–Schwarz inequality ∣⟨f,g⟩∣≤∥f∥2∥g∥2|\langle f, g\rangle| \leq \|f\|_2\|g\|_2 and the triangle inequality for ∥⋅∥2\|\cdot\|_2 follow exactly as in 2B.1 Metric Spaces. So d2(f,g)=∥f−g∥2d_2(f, g) = \|f - g\|_2 is a metric, the root-mean-square distance.

The building blocks are the characters ek(x)=e2πikxe_k(x) = e^{2\pi ikx}, k∈Zk \in \mathbb{Z}, which go ∣k∣|k| times around the unit circle as xx runs over [0,1][0, 1].

Proposition 7.2 Orthonormality

⟨ek,el⟩=1\langle e_k, e_l\rangle = 1 if k=lk = l and 00 if k≠lk \neq l.

Proof. ⟨ek,el⟩=∫01e2πi(k−l)x dx\langle e_k, e_l\rangle = \int_0^1 e^{2\pi i(k - l)x}\,dx. For k=lk = l the integrand is 11. For m=k−l≠0m = k - l \neq 0 it has antiderivative e2πimx2πim\frac{e^{2\pi imx}}{2\pi im}, which takes the same value at 00 and 11.

A trigonometric polynomial is a finite sum ∑∣k∣≤Nckek\sum_{|k| \leq N} c_ke_k. Its coefficients can be recovered by taking inner products: ck=⟨P,ek⟩c_k = \langle P, e_k\rangle. For a general ff this suggests a definition.

Definition 7.3 Fourier coefficients and partial sums

The Fourier coefficients of f∈C(R/Z)f \in C(\mathbb{R}/\mathbb{Z}) are f^(k)=⟨f,ek⟩=∫01f(x)e−2πikx dx\hat f(k) = \langle f, e_k\rangle = \int_0^1 f(x)e^{-2\pi ikx}\,dx. Its Fourier series is ∑k∈Zf^(k)e2πikx\sum_{k \in \mathbb{Z}}\hat f(k)e^{2\pi ikx}, and the partial sums are SNf=∑∣k∣≤Nf^(k)ekS_Nf = \sum_{|k| \leq N}\hat f(k)e_k.

For real ff, f^(−k)=f^(k)‾\hat f(-k) = \overline{\hat f(k)}, and the terms pair up into real cosines and sines: f^(k)ek+f^(−k)e−k=2Re⁡(f^(k)e2πikx)\hat f(k)e_k + \hat f(-k)e_{-k} = 2\operatorname{Re}\big(\hat f(k)e^{2\pi ikx}\big).

The partial sum SNfS_Nf is the best approximation to ff by trigonometric polynomials of degree at most NN, in the root-mean-square sense. This is Pythagoras: f−SNff - S_Nf is orthogonal to every eke_k with ∣k∣≤N|k| \leq N, so for any trigonometric polynomial PP of degree ≤N\leq N,

∥f−P∥22=∥f−SNf∥22+∥SNf−P∥22≥∥f−SNf∥22.\|f - P\|_2^2 = \|f - S_Nf\|_2^2 + \|S_Nf - P\|_2^2 \geq \|f - S_Nf\|_2^2.

Taking P=0P = 0 gives Bessel's inequality: ∑∣k∣≤N∣f^(k)∣2=∥SNf∥22≤∥f∥22\sum_{|k|\leq N}|\hat f(k)|^2 = \|S_Nf\|_2^2 \leq \|f\|_2^2 for every NN. In particular f^(k)→0\hat f(k) \to 0 as ∣k∣→∞|k| \to \infty.

Does the Fourier series converge?

The natural hope is that SNf→fS_Nf \to f. Write the partial sum as an average:

SNf(x)=∑∣k∣≤N∫01f(y)e2πik(x−y) dy=∫01f(y) DN(x−y) dy,DN(x)=∑∣k∣≤Ne2πikx=sin⁡((2N+1)πx)sin⁡(πx).S_Nf(x) = \sum_{|k|\leq N}\int_0^1 f(y)e^{2\pi ik(x - y)}\,dy = \int_0^1 f(y)\,D_N(x - y)\,dy, \qquad D_N(x) = \sum_{|k|\leq N}e^{2\pi ikx} = \frac{\sin\big((2N+1)\pi x\big)}{\sin(\pi x)}.

Expressions of this kind are periodic convolutions, (f∗g)(x)=∫01f(y)g(x−y) dy(f * g)(x) = \int_0^1 f(y)g(x - y)\,dy, and SNf=f∗DNS_Nf = f * D_N. If DND_N were an approximate identity in the sense of 2B.5 Uniform Convergence and Arzelà–Ascoli, non-negative with integral 11 and concentrating at 00, we would have SNf→fS_Nf \to f uniformly. It has integral 11 and its peak at 00 grows. But it is not non-negative: it oscillates, with side lobes that decay slowly (Figure 7.1). Its total absolute mass ∫01∣DN∣\int_0^1|D_N| grows like log⁡N\log N, and that is enough to wreck convergence: there are continuous functions whose Fourier series diverge at a point (du Bois-Reymond, 1873).

The cure is to average the partial sums.

Definition 7.4 Fejér means and the Fejér kernel

The Fejér mean is σNf=1N(S0f+S1f+⋯+SN−1f)=f∗FN\sigma_Nf = \frac{1}{N}\big(S_0f + S_1f + \cdots + S_{N-1}f\big) = f * F_N, where

FN(x)=1N∑n=0N−1Dn(x)=∑∣k∣<N(1−∣k∣N)e2πikx=1N(sin⁡(Nπx)sin⁡(πx))2.F_N(x) = \frac1N\sum_{n=0}^{N-1}D_n(x) = \sum_{|k| < N}\Big(1 - \frac{|k|}{N}\Big)e^{2\pi ikx} = \frac{1}{N}\left(\frac{\sin(N\pi x)}{\sin(\pi x)}\right)^2.

The last formula (Exercise 7.10) is the important one: FN≥0F_N \geq 0. Averaging has cancelled the oscillation.

Figure 7.1. The Dirichlet kernel D8D_8 (oscillating, with negative lobes) and the Fejér kernel F8F_8 (never negative). Both have integral 11. Only FNF_N is an approximate identity, which is why Fejér means converge where partial sums can fail.
Theorem 7.5 Fejér's theorem

For every f∈C(R/Z)f \in C(\mathbb{R}/\mathbb{Z}), σNf→f\sigma_Nf \to f uniformly as N→∞N \to \infty.

Proof. FNF_N is an approximate identity. Mass 11: integrate the middle expression term by term; only k=0k = 0 survives. Non-negative: the last expression. Concentration: for δ≤∣x∣≤12\delta \leq |x| \leq \tfrac12, sin⁡2(πx)≥sin⁡2(πδ)\sin^2(\pi x) \geq \sin^2(\pi\delta), so FN(x)≤1Nsin⁡2(πδ)→0F_N(x) \leq \frac{1}{N\sin^2(\pi\delta)} \to 0 uniformly there. Now repeat the proof of the Weierstrass theorem (2B.5 Uniform Convergence and Arzelà–Ascoli): with M=sup⁡∣f∣M = \sup|f| and δ\delta chosen from the uniform continuity of ff so that ∣f(x−y)−f(x)∣≤ε/2|f(x - y) - f(x)| \leq \varepsilon/2 when ∣y∣<δ|y| < \delta,

∣σNf(x)−f(x)∣≤∫−1/21/2∣f(x−y)−f(x)∣ FN(y) dy≤ε2+2M⋅1Nsin⁡2(πδ),|\sigma_Nf(x) - f(x)| \leq \int_{-1/2}^{1/2}|f(x - y) - f(x)|\,F_N(y)\,dy \leq \frac\varepsilon2 + 2M\cdot\frac{1}{N\sin^2(\pi\delta)},

which is less than ε\varepsilon for large NN, for every xx.

Three consequences follow, and they are what make Fourier series usable.

Corollary 7.6 Density, uniqueness and Plancherel

For f,g∈C(R/Z)f, g \in C(\mathbb{R}/\mathbb{Z}):

  1. Trigonometric polynomials are dense in C(R/Z)C(\mathbb{R}/\mathbb{Z}) in the sup metric.
  2. If f^(k)=0\hat f(k) = 0 for every kk, then f=0f = 0. So ff is determined by its Fourier coefficients.
  3. (Plancherel) ∥f−SNf∥2→0\|f - S_Nf\|_2 \to 0, and ∥f∥22=∑k∈Z∣f^(k)∣2\displaystyle \|f\|_2^2 = \sum_{k\in\mathbb{Z}}|\hat f(k)|^2, and more generally ⟨f,g⟩=∑kf^(k)g^(k)‾\langle f, g\rangle = \sum_k \hat f(k)\overline{\hat g(k)}.
  4. If ∑k∣f^(k)∣<∞\sum_k|\hat f(k)| < \infty, then SNf→fS_Nf \to f uniformly.

Proof. (1) Each σNf\sigma_Nf is a trigonometric polynomial. (2) Then every σNf=0\sigma_Nf = 0, and σNf→f\sigma_Nf \to f. (3) Given ε\varepsilon, pick a trigonometric polynomial PP of some degree MM with sup⁡∣f−P∣≤ε\sup|f - P| \leq \varepsilon, so ∥f−P∥2≤ε\|f - P\|_2 \leq \varepsilon. For N≥MN \geq M, SNfS_Nf is the best approximation of degree NN, so ∥f−SNf∥2≤∥f−P∥2≤ε\|f - S_Nf\|_2 \leq \|f - P\|_2 \leq \varepsilon. Then ∥SNf∥22=∑∣k∣≤N∣f^(k)∣2→∥f∥22\|S_Nf\|_2^2 = \sum_{|k|\leq N}|\hat f(k)|^2 \to \|f\|_2^2. The formula for ⟨f,g⟩\langle f, g\rangle follows by polarisation. (4) By the M-test the series converges uniformly to some continuous hh; its coefficients are those of ff (integrate term by term, 2B.5 Uniform Convergence and Arzelà–Ascoli), so h=fh = f by (2).

How fast f^(k)\hat f(k) decays measures how smooth ff is. Integrating by parts (2A.11 The Riemann Integral, with no boundary terms because ff is periodic), f′^(k)=2πik f^(k)\widehat{f'}(k) = 2\pi ik\,\hat f(k). So if ff is C2C^2, then ∣f^(k)∣=∣f′′^(k)∣4π2k2≤sup⁡∣f′′∣4π2k2|\hat f(k)| = \frac{|\widehat{f''}(k)|}{4\pi^2k^2} \leq \frac{\sup|f''|}{4\pi^2k^2}, which is summable, and the Fourier series converges uniformly. Smoothness of ff is decay of f^\hat f. That sentence, made quantitative, is the theory of Sobolev spaces (4A.9 Sobolev Spaces).

Fourier's solution of the heat equation on a ring

Now Fourier's problem. A thin ring of metal of circumference 11 has temperature u(x,t)u(x, t) at position x∈R/Zx \in \mathbb{R}/\mathbb{Z} and time tt. Choosing units so the diffusivity is 11,

∂tu=∂x2u,u(x,0)=f(x).\partial_t u = \partial_x^2 u, \qquad u(x, 0) = f(x).

Separate the frequencies. Try u(x,t)=∑kak(t)e2πikxu(x, t) = \sum_k a_k(t)e^{2\pi ikx}. Since ∂x2ek=−4π2k2ek\partial_x^2e_k = -4\pi^2k^2e_k, the equation decouples into one ordinary differential equation per frequency:

ak′(t)=−4π2k2 ak(t),ak(0)=f^(k)⟹ak(t)=f^(k) e−4π2k2t,a_k'(t) = -4\pi^2k^2\,a_k(t), \quad a_k(0) = \hat f(k) \qquad\Longrightarrow\qquad a_k(t) = \hat f(k)\,e^{-4\pi^2k^2t},

by the uniqueness in 2B.6 Power Series, Exponentials and Bump Functions. So the candidate solution is

u(x,t)=∑k∈Zf^(k) e−4π2k2t e2πikx.(∗)u(x, t) = \sum_{k\in\mathbb{Z}}\hat f(k)\,e^{-4\pi^2k^2t}\,e^{2\pi ikx}. \tag{$*$}

In words: each frequency decays exponentially, at a rate proportional to the square of the frequency. The k=0k = 0 term, the mean temperature, never decays. Frequency 11 decays like e−39.5te^{-39.5t}; frequency 1010 like e−3948te^{-3948t}.

Theorem 7.7 Heat equation on a ring

Let f∈C(R/Z)f \in C(\mathbb{R}/\mathbb{Z}) and define uu by (∗)(*) for t>0t > 0, and u(⋅,0)=fu(\cdot, 0) = f. Then:

  1. For t>0t > 0, the series and all its term-by-term derivatives in xx and tt converge uniformly on {t≥t0}\{t \geq t_0\} for each t0>0t_0 > 0, so uu is infinitely differentiable for t>0t > 0 and satisfies ∂tu=∂x2u\partial_tu = \partial_x^2u.
  2. u(⋅,t)→fu(\cdot, t) \to f uniformly as t→0t \to 0.
  3. uu is the only solution of the problem that is continuous on R/Z×[0,∞)\mathbb{R}/\mathbb{Z} \times [0, \infty) and smooth for t>0t > 0.

Proof. (1) ∣f^(k)∣≤∥f∥2≤sup⁡∣f∣|\hat f(k)| \leq \|f\|_2 \leq \sup|f| by Bessel. A term differentiated pp times in xx and qq times in tt is bounded by sup⁡∣f∣ (2π∣k∣)p+2qe−4π2k2t0\sup|f|\,(2\pi|k|)^{p+2q}e^{-4\pi^2k^2t_0} on t≥t0t \geq t_0, and these bounds are summable over kk, because the Gaussian factor beats every power. By the M-test and 2B.5 Uniform Convergence and Arzelà–Ascoli, all derivatives may be taken term by term, and each term solves the equation.

(2) Write u(⋅,t)=f∗Htu(\cdot, t) = f * H_t with the periodic heat kernel

Ht(x)=∑k∈Ze−4π2k2te2πikx.H_t(x) = \sum_{k\in\mathbb{Z}}e^{-4\pi^2k^2t}e^{2\pi ikx}.

HtH_t has integral 11 (only k=0k = 0 survives integration). It is positive: this follows from the identity Ht(x)=∑n∈Z14πte−(x−n)2/4tH_t(x) = \sum_{n\in\mathbb{Z}}\frac{1}{\sqrt{4\pi t}}e^{-(x - n)^2/4t}, which says that the ring's heat kernel is the Gaussian heat kernel of the line wrapped around the circle. (The identity is an instance of the Poisson summation formula, proved in 4A.5 The Fourier Transform; take it on trust here, or see Exercise 7.13 for a proof of positivity that avoids it.) From the Gaussian form, HtH_t concentrates at 00 as t→0t \to 0. So HtH_t is an approximate identity, and the proof of Fejér's theorem gives f∗Ht→ff * H_t \to f uniformly.

(3) If vv is another solution, w=u−vw = u - v solves the heat equation with w(⋅,0)=0w(\cdot, 0) = 0. Its energy E(t)=∫01∣w∣2E(t) = \int_0^1|w|^2 satisfies E′≤0E' \leq 0 (by the computation in the next subsection) and E(t)→0E(t) \to 0 as t→0t \to 0, so E≡0E \equiv 0.

Figure 7.2. Heat on a ring, computed from (∗)(*), starting from a square wave (+1+1 on half the ring, −1-1 on the other half). At t=0.001t = 0.001 the corners are already smooth; by t=0.01t = 0.01 the higher frequencies are gone and the profile is nearly a single sine wave; by t=0.05t = 0.05 its amplitude has fallen to about 0.180.18. The mean, 00, never changes.

What the formula shows

Each property below is visible in (∗)(*), and each is a theme that later books prove for much harder equations, where no formula is available.

1. Instant smoothing. However rough ff is (here, merely continuous; in Course 3, merely square-integrable), u(⋅,t)u(\cdot, t) is infinitely differentiable for every t>0t > 0, because e−4π2k2te^{-4\pi^2k^2t} crushes the high frequencies. The heat equation cannot be run backwards in general: running it backwards would multiply f^(k)\hat f(k) by e+4π2k2te^{+4\pi^2k^2t}, and only very special ff survive that (Exercise 7.12). Smoothing is the reason Ricci flow is useful at all: it improves a metric, and the improvement is quantified by Shi's derivative estimates (11A.3 Short-Time Existence and Uniqueness), the analogue of the bound in part 1 of the proof.

2. Energy decreases. By Plancherel, ∫01∣u(x,t)∣2dx=∑k∣f^(k)∣2e−8π2k2t\int_0^1|u(x, t)|^2dx = \sum_k|\hat f(k)|^2e^{-8\pi^2k^2t}, which visibly decreases in tt. Directly, as in the rehearsal of 2A.11 The Riemann Integral, integrating by parts with no boundary terms,

ddt∫01u2 dx=2∫01u ∂x2u dx=−2∫01(∂xu)2 dx≤0.\frac{d}{dt}\int_0^1 u^2\,dx = 2\int_0^1 u\,\partial_x^2u\,dx = -2\int_0^1(\partial_xu)^2\,dx \leq 0.

Likewise the Dirichlet energy ∫(∂xu)2=∑4π2k2∣f^(k)∣2e−8π2k2t\int(\partial_xu)^2 = \sum 4\pi^2k^2|\hat f(k)|^2e^{-8\pi^2k^2t} decreases. In fact the heat equation is the direction of steepest descent of the Dirichlet energy, the gradient flow of 12∫∣∂xu∣2\tfrac12\int|\partial_xu|^2 (Exercise 7.15). This is thread V's first appearance: Ricci flow is, after Perelman's modification, the gradient flow of his F\mathcal{F}-functional (12A.2 Ricci Flow as a Gradient Flow).

3. Convergence to equilibrium, at a rate set by the first frequency. Let fˉ=f^(0)\bar f = \hat f(0) be the mean. Then

∥u(⋅,t)−fˉ∥22=∑k≠0∣f^(k)∣2e−8π2k2t≤e−8π2t ∥f−fˉ∥22.\|u(\cdot, t) - \bar f\|_2^2 = \sum_{k \neq 0}|\hat f(k)|^2e^{-8\pi^2k^2t} \leq e^{-8\pi^2t}\,\|f - \bar f\|_2^2.

The temperature evens out exponentially fast, at the rate 4π24\pi^2 of the lowest non-zero frequency. This spectral gap is equivalent to an inequality of Poincaré type (Exercise 7.14), and it reappears as the first eigenvalue of the Laplacian on a manifold (9B.7 The Heat Equation on a Manifold).

4. The maximum principle. Since u(⋅,t)=f∗Htu(\cdot, t) = f * H_t with Ht≥0H_t \geq 0 and ∫Ht=1\int H_t = 1, each value u(x,t)u(x, t) is an average of values of ff. So min⁡f≤u(x,t)≤max⁡f\min f \leq u(x, t) \leq \max f: the heat equation creates no new maxima or minima, and in fact max⁡xu(x,t)\max_x u(x, t) is non-increasing in tt (apply the same argument starting from any time ss). For the heat equation on a ring this follows from the positivity of a kernel; in 6A.4 Maximum Principles it is proved for general heat equations by the second-derivative test (2A.10 Derivatives, 2B.8 Calculus in Several Variables), and in 11A.4 Maximum Principles under Ricci Flow Hamilton extends it to tensors.

Where this goes The first heat kernel

The kernel HtH_t is the first appearance of the heat kernel, the fundamental solution of the heat equation. On the line it is the Gaussian 14πte−x2/4t\frac{1}{\sqrt{4\pi t}}e^{-x^2/4t} (6A.3 The Heat Equation on ℝⁿ); on a Riemannian manifold it is defined abstractly and estimated geometrically (9B.7 The Heat Equation on a Manifold). As a function of tt (at x=0x = 0), Ht(0)=∑ke−4π2k2tH_t(0) = \sum_k e^{-4\pi^2k^2t} is a theta function, an object of number theory and complex analysis (Book 5A). And in Perelman's work the conjugate heat equation, run backwards in time, carries a heat-kernel-like density (4πτ)−n/2e−f(4\pi\tau)^{-n/2}e^{-f} whose behaviour encodes the geometry (12A.2 Ricci Flow as a Gradient Flow, 12A.6 Pseudolocality). The factor (4πτ)−n/2(4\pi\tau)^{-n/2} there is the same normalisation as the 14πt\frac{1}{\sqrt{4\pi t}} here.

The Gibbs phenomenon

Fourier series of functions with jumps converge badly near the jumps. For the square wave (+1+1 on (0,12)(0, \tfrac12), −1-1 on (12,1)(\tfrac12, 1)), the partial sums overshoot the jump by a fixed amount that does not shrink as NN grows; the overshoot just moves closer to the jump (Figure 7.3). In the limit, the partial sums rise to 2π∫0πsin⁡tt dt≈1.179\frac2\pi\int_0^\pi\frac{\sin t}{t}\,dt \approx 1.179, overshooting the value 11 by about 9%9\% of the jump of size 22.

The overshoot is caused by the negative lobes of the Dirichlet kernel. The Fejér means, which average against a positive kernel, never overshoot: by the argument of property 4, they stay between −1-1 and 11. Positivity of the averaging kernel is what makes both the maximum principle and the absence of overshoot work.

Figure 7.3. The Gibbs phenomenon near a jump of the square wave. Partial sums of degree 99, 2929 and 9999 all overshoot to about 1.181.18; higher degree only narrows the overshoot. The Fejér mean (dashed) stays below 11.
In the world In use JPEG throws away high frequencies

A JPEG image is compressed block by block. Each 8×88 \times 8 block of pixels is transformed by the discrete cosine transform, a finite cousin of the Fourier series that writes the block as a combination of 6464 cosine patterns of increasing frequency. The coefficients are then rounded, coarsely for the high frequencies and finely for the low ones, and many high-frequency coefficients become zero and cost almost nothing to store. This works because, as the decay of f^\hat f for smooth ff suggests, most of the content of a typical photograph is in the low frequencies. Where an image has a sharp edge, the discarded high frequencies were needed, and the result is the ringing near edges in a heavily compressed JPEG: a two-dimensional relative of the Gibbs phenomenon.

In the world Model Why a violin doesn't sound like a flute

A sustained musical note is (nearly) periodic, so it is a Fourier series: a fundamental frequency, which sets the pitch, plus harmonics at integer multiples of it. Two instruments playing the same pitch have the same fundamental and differ in the sizes of the coefficients ∣f^(k)∣|\hat f(k)| for k≥2k \geq 2. That difference is what we hear as timbre. A flute's tone is close to a pure sine wave, with weak harmonics; a bowed violin string, which is dragged and released in a stick–slip motion, produces a waveform close to a sawtooth, whose harmonics decay only like 1/k1/k (Exercise 7.8), and it sounds correspondingly brighter. Smoothness of the waveform is decay of the coefficients, heard.

History

Fourier's 1807 memoir was judged by Lagrange, Laplace, Monge and Lacroix; its prize-winning revision (1811) and the 1822 book established the method, and Chapter IV of the book treats the movement of heat in a ring. Peter Gustav Lejeune Dirichlet gave the first rigorous convergence theorem in 1829, for piecewise monotone functions. Paul du Bois-Reymond constructed a continuous function whose Fourier series diverges at a point in 1873. In 1900, aged twenty, Lipót Fejér proved that the Cesàro means of the Fourier series of every continuous function converge uniformly. The overshoot at jumps was described by Henry Wilbraham in 1848 and again by J. Willard Gibbs in 1899, whose name it carries. The questions Fourier raised, what a function is and in what sense a series converges, led to Riemann's integral (1854), to Cantor's set theory (which began with sets of uniqueness for trigonometric series), and to Lebesgue's integral (1902), the subject of Course 3.

Recall Where we stand

Periodic functions have Fourier coefficients f^(k)=⟨f,ek⟩\hat f(k) = \langle f, e_k\rangle with respect to the orthonormal characters e2πikxe^{2\pi ikx}. Partial sums are best approximations, but may fail to converge for continuous ff, because the Dirichlet kernel is not positive. Fejér means average against a positive kernel and converge uniformly, which gives density of trigonometric polynomials, uniqueness and Plancherel's identity. Smoothness of ff is decay of f^\hat f. The heat equation on a ring is solved by letting each coefficient decay like e−4π2k2te^{-4\pi^2k^2t}: the solution is smooth at once, its energy decreases, it converges exponentially to the mean, and it obeys the maximum principle because it is an average against the positive heat kernel. 2B.8 Calculus in Several Variables turns to functions of several variables, where the Laplacian ∂x2\partial_x^2 becomes the trace of the Hessian.

Exercises

Exercise 7.8 Two waveforms

(a) For the square wave f=1f = 1 on (0,12)(0, \tfrac12) and −1-1 on (12,1)(\tfrac12, 1), show that f^(k)=2πik\hat f(k) = \frac{2}{\pi ik} for odd kk and 00 for even k≠0k \neq 0, so ff "==" 4π∑k odd>0sin⁡(2πkx)k\frac4\pi\sum_{k \text{ odd} > 0}\frac{\sin(2\pi kx)}{k}. (b) For the sawtooth g(x)=x−12g(x) = x - \tfrac12 on [0,1)[0, 1), show g^(k)=i2πk\hat g(k) = \frac{i}{2\pi k} for k≠0k \neq 0. (These functions are not continuous, but the integrals defining f^(k)\hat f(k) still make sense.)

Exercise 7.9 Plancherel sums a series

Apply Plancherel's identity to the sawtooth of Exercise 7.8 (it holds for piecewise continuous functions too, by approximating them in ∥⋅∥2\|\cdot\|_2 by continuous ones), and deduce ∑k=1∞1k2=π26\sum_{k=1}^\infty\frac1{k^2} = \frac{\pi^2}{6}.

Solution

∥g∥22=∫01(x−12)2dx=112\|g\|_2^2 = \int_0^1(x - \tfrac12)^2dx = \frac1{12}, and ∑k≠0∣g^(k)∣2=2∑k≥114π2k2\sum_{k\neq0}|\hat g(k)|^2 = 2\sum_{k\geq1}\frac{1}{4\pi^2k^2}. Equate: ∑1k2=4π22⋅112=π26\sum\frac1{k^2} = \frac{4\pi^2}{2}\cdot\frac{1}{12} = \frac{\pi^2}6.

Exercise 7.10 The Fejér kernel is positive

With z=e2πixz = e^{2\pi ix}, show that Dn(x)=zn+1/2−z−n−1/2z1/2−z−1/2D_n(x) = \frac{z^{n+1/2} - z^{-n-1/2}}{z^{1/2} - z^{-1/2}}, and that ∑n=0N−1Dn=(zN/2−z−N/2)2(z1/2−z−1/2)2\sum_{n=0}^{N-1}D_n = \frac{(z^{N/2} - z^{-N/2})^2}{(z^{1/2} - z^{-1/2})^2}. Deduce FN(x)=1N(sin⁡Nπxsin⁡πx)2F_N(x) = \frac1N\big(\frac{\sin N\pi x}{\sin\pi x}\big)^2.

Exercise 7.11 Seasons underground

(a) Verify that u(z,t)=e−z/dcos⁡(ωt−z/d)u(z, t) = e^{-z/d}\cos(\omega t - z/d) solves ut=κuzzu_t = \kappa u_{zz} when d=2κ/ωd = \sqrt{2\kappa/\omega}. (Write u=Re⁡eiωt−(1+i)z/du = \operatorname{Re}e^{i\omega t - (1+i)z/d}.) (b) With κ=5×10−7\kappa = 5 \times 10^{-7} m²/s, compute dd for the annual and the daily cycle. (c) At what depth is the annual swing reduced to 1%1\% of its surface value, and how far behind the surface is it there?

Solution

(a) With v=eiωt−(1+i)z/dv = e^{i\omega t - (1+i)z/d}: vt=iωvv_t = i\omega v and vzz=(1+i)2d2v=2id2vv_{zz} = \frac{(1+i)^2}{d^2}v = \frac{2i}{d^2}v, so vt=κvzzv_t = \kappa v_{zz} exactly when ω=2κ/d2\omega = 2\kappa/d^2. Take real parts. (b) Annual: ω=2π/(3.156×107 s)\omega = 2\pi/(3.156\times10^7\text{ s}), d≈2.24d \approx 2.24 m. Daily: d≈0.117d \approx 0.117 m. (c) e−z/d=0.01e^{-z/d} = 0.01 at z=dln⁡100≈10.3z = d\ln 100 \approx 10.3 m, with a lag of ln⁡1002π≈0.73\frac{\ln 100}{2\pi} \approx 0.73 of a year, about 8.88.8 months.

Exercise 7.12 The heat equation can't be reversed

Suppose uu solves the heat equation on the ring for t∈[−1,0]t \in [-1, 0], continuously, with u(⋅,0)=fu(\cdot, 0) = f. Show that ∣f^(k)∣≤Ce−4π2k2|\hat f(k)| \leq C e^{-4\pi^2k^2} for some CC. Deduce that if ff is merely continuous (for example, f^(k)=1/k2\hat f(k) = 1/k^2), no such backward solution exists.

Solution

The coefficients of u(⋅,t)u(\cdot, t) satisfy ak′=−4π2k2aka_k' = -4\pi^2k^2a_k, so f^(k)=ak(0)=ak(−1)e−4π2k2\hat f(k) = a_k(0) = a_k(-1)e^{-4\pi^2k^2}, and ∣ak(−1)∣≤sup⁡∣u(⋅,−1)∣|a_k(-1)| \leq \sup|u(\cdot, -1)|. A function with f^(k)=1/k2\hat f(k) = 1/k^2 (k≠0k \neq 0) is continuous by the M-test, but 1/k21/k^2 is not O(e−4π2k2)O(e^{-4\pi^2k^2}).

Exercise 7.13 The heat kernel is positive, without Poisson summation

Let f≥0f \geq 0 be continuous and periodic, and u=f∗Htu = f * H_t. Show that u≥0u \geq 0 for all t>0t > 0, by considering m(t)=min⁡xu(x,t)m(t) = \min_x u(x, t): at a point where the minimum is attained, ∂x2u≥0\partial_x^2 u \geq 0 (2A.10 Derivatives), so uu can't be decreasing there; make this rigorous with uε=u+εtu_\varepsilon = u + \varepsilon t (or u+ε(1+t)u + \varepsilon(1 + t)) as in Hamilton's trick (2A.10 Derivatives). Deduce Ht≥0H_t \geq 0 by taking ff to be a narrow bump.

Hint

Suppose uε=u+ε(1+t)u_\varepsilon = u + \varepsilon(1 + t) first becomes 00 at a time t1>0t_1 > 0 and point x1x_1. There ∂tuε≤0\partial_tu_\varepsilon \leq 0 and ∂x2uε≥0\partial_x^2u_\varepsilon \geq 0, but ∂tuε−∂x2uε=ε>0\partial_tu_\varepsilon - \partial_x^2u_\varepsilon = \varepsilon > 0. For the last step, (f∗Ht)(x)=∫f(y)Ht(x−y)dy(f * H_t)(x) = \int f(y)H_t(x - y)dy, and if Ht(x0)<0H_t(x_0) < 0 a bump concentrated near 00 makes this negative at x=x0x = x_0.

Exercise 7.14 Rehearsal: the spectral gap as a Poincaré inequality

(a) For C1C^1 periodic uu with mean uˉ\bar u, prove Wirtinger's inequality ∫01∣u−uˉ∣2≤14π2∫01∣u′∣2\int_0^1|u - \bar u|^2 \leq \frac{1}{4\pi^2}\int_0^1|u'|^2, using Plancherel and u′^(k)=2πiku^(k)\widehat{u'}(k) = 2\pi ik\hat u(k). When is it an equality? (b) Use it to prove the decay estimate of property 3 without Fourier series: if E(t)=∫∣u−uˉ∣2E(t) = \int|u - \bar u|^2 for a solution of the heat equation, show E′=−2∫∣ux∣2≤−8π2EE' = -2\int|u_x|^2 \leq -8\pi^2E, and conclude E(t)≤e−8π2tE(0)E(t) \leq e^{-8\pi^2t}E(0) (compare 2A.10 Derivatives's comparison argument). This route, "an energy identity plus a functional inequality gives exponential decay", is the one that survives on manifolds and for nonlinear equations, where Fourier series don't exist. Perelman's entropy arguments use a log-Sobolev inequality in exactly this role (6A.10 Entropy, Information and Diffusion, 12A.3 The 𝓦-Entropy).

Solution

(a) ∫∣u−uˉ∣2=∑k≠0∣u^(k)∣2≤∑k≠04π2k24π2∣u^(k)∣2=14π2∫∣u′∣2\int|u - \bar u|^2 = \sum_{k\neq0}|\hat u(k)|^2 \leq \sum_{k\neq0}\frac{4\pi^2k^2}{4\pi^2}|\hat u(k)|^2 = \frac{1}{4\pi^2}\int|u'|^2. Equality iff u^(k)=0\hat u(k) = 0 for ∣k∣≥2|k| \geq 2: u=uˉ+acos⁡2πx+bsin⁡2πxu = \bar u + a\cos 2\pi x + b\sin 2\pi x. (b) The mean is constant (2A.11 The Riemann Integral), so E′=2∫(u−uˉ)ut=2∫(u−uˉ)uxx=−2∫ux2≤−8π2EE' = 2\int(u - \bar u)u_t = 2\int(u - \bar u)u_{xx} = -2\int u_x^2 \leq -8\pi^2E. Then (e8π2tE)′≤0(e^{8\pi^2t}E)' \leq 0.

Exercise 7.15 The heat equation as a gradient flow

For smooth periodic uu and vv, let D(u)=12∫01ux2D(u) = \tfrac12\int_0^1 u_x^2. Show that ddsD(u+sv)∣s=0=−∫01uxx v\frac{d}{ds}D(u + sv)\big|_{s=0} = -\int_0^1 u_{xx}\,v. So, measuring changes by the inner product ⟨⋅,⋅⟩\langle\cdot,\cdot\rangle, the direction in which DD decreases fastest is v=uxxv = u_{xx} (normalised), and the heat equation ut=uxxu_t = u_{xx} moves uu in that direction. This is the first instance of the variational thread V; it is developed in 6A.9 Calculus of Variations and Gradient Flows and reaches Ricci flow in 12A.2 Ricci Flow as a Gradient Flow.

© 2026 NeckPinch (www.neckpinch.com). All content in the guidebook (text, mathematics, figures and exercises) is protected by copyright. All rights reserved. No part may be copied, republished or redistributed without written permission.