The Privilege of the Inverse Square: Two Proofs of the Shell Theorem

In the last post I borrowed a point from Feynman: the mathematical form of a law carries far more physics than the sentence it seems to be saying. This post is a worked example. Nominally we are proving the shell theorem. What we are really after is a more pointed question — why is gravity $1/r^2$, and not $1/r^{1.9}$ or $1/r^3$?


0. Stating the problem properly

Take a thin spherical shell of uniform density: radius $R$, total mass $M$, surface density $\sigma = M/4\pi R^2$. Newton proved two things about it:

  • Theorem A (outside). The shell attracts an external point mass exactly as if all of $M$ were concentrated at the centre.
  • Theorem B (inside). The shell exerts zero force on a point mass anywhere inside it — not just at the centre, but at every interior point.

Theorem B is the counterintuitive one. Stand inside the shell, off to the left. The left-hand wall is much closer to you. Why doesn’t it win?

The stock answer is that the near patch is closer but smaller. That is only half an answer. The real question is: the near patch shrinks and the far patch grows — why should those two effects cancel exactly?

To see what is going on, let’s do something Newton didn’t do but which is worth doing: rewrite the force law as

$$ F = \frac{G\,m_1 m_2}{l^{\,n}}, $$

let the exponent $n$ float freely, and redo everything. Under what conditions do Theorems A and B survive?

Spoiler: only for $n=2$. Not “approximately only” — exactly and uniquely.


1. The geometric proof: seeing the answer without calculus

【Figure 1: a double cone from an interior point $P$, cutting patches $\mathrm{d}A_1$ and $\mathrm{d}A_2$ at distances $l_1$ and $l_2$.】

Let $P$ be a point inside the shell. From $P$, draw a very narrow double cone — two opposed cones with the same opening angle. It cuts two small patches out of the shell, of mass $\mathrm{d}m_1$ and $\mathrm{d}m_2$, at distances $l_1$ and $l_2$. The whole shell can be tiled by such double cones, so if every pair cancels, the total force is zero.

1.1 The step that matters: both patches are tilted equally

Here is the point most popular accounts skip, and it happens to be the only place where “sphere” enters the argument at all.

Let the cone subtend solid angle $\mathrm{d}\Omega$. If a patch were perpendicular to the line of sight, its area would be $l^2\,\mathrm{d}\Omega$. But a patch on a sphere is generally tilted; if its normal makes angle $\psi$ with the line of sight, the true area is

$$ \mathrm{d}A = \frac{l^2\,\mathrm{d}\Omega}{\cos\psi}. $$

Now look at the chord $AB$ through $P$. Since $OA = OB = R$, the triangle $OAB$ is isosceles, so $\angle OAB = \angle OBA$. The surface normals at $A$ and $B$ are along $OA$ and $OB$, so the chord meets the surface at the same angle at both ends:

$$ \psi_1 = \psi_2 . $$

That is the whole secret. Because it is a sphere, the obliquity correction is identical at both ends, and therefore

$$ \frac{\mathrm{d}m_1}{\mathrm{d}m_2} = \frac{\mathrm{d}A_1}{\mathrm{d}A_2} = \frac{l_1^2}{l_2^2}. $$

The masses are in the ratio of the squared distances — and note where that “2” came from. It came from area growing as the square of distance in three-dimensional space. It is a fact of geometry and has nothing to do with gravity.

1.2 Now let the exponent in

The two patches pull on $P$ with forces in the ratio

$$ \frac{F_1}{F_2}=\frac{\mathrm{d}m_1/l_1^{\,n}}{\mathrm{d}m_2/l_2^{\,n}} =\frac{l_1^2}{l_2^2}\cdot\frac{l_2^{\,n}}{l_1^{\,n}} =\left(\frac{l_1}{l_2}\right)^{2-n}. $$

Everything follows from that single exponent $2-n$:

Exponent$\left(l_1/l_2\right)^{2-n}$Physical consequence
$n<2$far patch winsnet force points toward the far wall, i.e. back through the centre
$n=2$identically 1; pairs cancelnet force is zero everywhere
$n>2$near patch winsnet force points toward the nearest wall, away from the centre

So “no field inside a shell” and “inverse square” are not merely related — they imply each other. The 2 in the force law is exactly the 2 in “area grows as the square”, cancelled off with perfect precision.

A pretty corollary comes for free. In a $n<2$ universe the centre of the shell is a stable equilibrium; in a $n>2$ universe it is unstable; only at $n=2$ does the entire interior become a region of neutral equilibrium. Our universe sits precisely on the dividing line.


2. The integral proof: turning that intuition into a formula

The geometric argument is elegant, but it only handles the interior, and the phrase “pairs cancel” needs a little care to be made rigorous. So let’s just do the integral — it settles both theorems at once, and for arbitrary $n$.

【Figure 2: shell centred at $O$, field point $P$ at distance $r$; a ring bounded by $\theta$ and $\theta+\mathrm{d}\theta$; a point on the ring is a distance $l$ from $P$, and $\alpha$ is the angle between that line and $PO$.】

2.1 Setting up

Slice the shell into thin rings by polar angle $\theta$, measured from the axis $OP$. A ring has radius $R\sin\theta$ and width $R\,\mathrm{d}\theta$, so

$$ \mathrm{d}M=2\pi R\sin\theta\cdot R\,\mathrm{d}\theta\cdot\sigma=2\pi R^{2}\sigma\sin\theta\,\mathrm{d}\theta . $$

Every point of a given ring is the same distance $l$ from $P$, and by symmetry the components perpendicular to $OP$ cancel around the ring. Only the axial component survives:

$$ \mathrm{d}g=\frac{G\,\mathrm{d}M}{l^{\,n}}\cos\alpha . $$

(Convention: $g>0$ means the field points toward the centre $O$.)

2.2 Choosing the substitution: use $l$, not $\cos\theta$

Most textbooks substitute $x=\cos\theta$ here. It works, but the integrand is ugly and you end up consulting a table of integrals. Integrating over $l$ instead is much cleaner, and the limits come out carrying physical meaning.

By the law of cosines,

$$ l^{2}=R^{2}+r^{2}-2Rr\cos\theta,\qquad \cos\alpha=\frac{r^{2}+l^{2}-R^{2}}{2rl}. $$

Differentiating the first: $2l\,\mathrm{d}l=2Rr\sin\theta\,\mathrm{d}\theta$, that is,

$$ \sin\theta\,\mathrm{d}\theta=\frac{l\,\mathrm{d}l}{Rr}. $$

Substituting, most factors of $l$ cancel:

$$ \boxed{\; g=\frac{\pi R\sigma G}{r^{2}}\int_{l_-}^{l_+} \Big[\,l^{\,2-n}+(r^{2}-R^{2})\,l^{-n}\Big]\mathrm{d}l \;} $$

and the limits are simply the nearest and farthest distances from the field point to the shell:

$$ l_-=|r-R|,\qquad l_+=r+R . $$

Both terms are now plain powers — no integral table needed — and the exponent $n$ sits out in the open where we can experiment with it.

2.3 $n=2$: two pieces of magic

Set $n=2$. The integrand becomes $1+(r^{2}-R^{2})l^{-2}$, with antiderivative

$$ \Phi(l)=l+\frac{R^{2}-r^{2}}{l}. $$

Inside ($r<R$). Here $l_-=R-r$ and $l_+=R+r$, so

$$ l_-l_+=(R-r)(R+r)=R^{2}-r^{2}, $$

which means the antiderivative is exactly

$$ \Phi(l)=l+\frac{l_-l_+}{l}. $$

This function is invariant under $l\mapsto l_-l_+/l$ — and that map sends $l_+$ to $l_-$. So

$$ \Phi(l_+)=l_++l_-=\Phi(l_-)\quad\Longrightarrow\quad g=0 . $$

This is my favourite step in the whole post. Theorem B is not a coincidence that emerges from grinding algebra; the antiderivative possesses an exact symmetry that swaps the two endpoints. And only $n=2$ produces it. This is the analytic counterpart of “pairs cancel” from Section 1.

Outside ($r>R$). Now $l_-=r-R$, $l_+=r+R$, $l_-l_+=r^{2}-R^{2}$, and $\Phi(l)=l-l_-l_+/l$, giving

$$ \Phi(l_+)-\Phi(l_-)=(l_+-l_-)-(l_–l_+)=2(l_+-l_-)=4R, $$

$$ g=\frac{\pi R\sigma G}{r^{2}}\cdot 4R=\frac{4\pi R^{2}\sigma G}{r^{2}}=\frac{GM}{r^{2}} . $$

Both theorems, in one calculation:

$$ g=\begin{cases} GM/r^{2}, & r>R,\\[4pt] 0, & r<R. \end{cases} $$

2.4 Trying other exponents

For general $n$ (excluding $n=1$ and $n=3$, where logarithms appear), the antiderivative is

$$ \Phi(l)=\frac{l^{\,3-n}}{3-n}+(r^{2}-R^{2})\frac{l^{\,1-n}}{1-n}, $$

which no longer has the endpoint-swapping symmetry, and $\Phi(l_+)-\Phi(l_-)$ does not vanish.

Take $n=4$ as a concrete case. Writing $u=r/R<1$, the interior field works out to

$$ g=-\,\frac{2GM}{3R^{4}}\cdot\frac{u}{\left(1-u^{2}\right)^{2}} . $$

The minus sign means the force points away from the centre, toward the nearest wall — exactly what Section 1 predicted for $n>2$. It also diverges as $u\to1$: the closer you drift to the wall, the harder it pulls. A spherical shell would be a dangerous place to be inside of in such a universe.

The exterior is worth a look too. Still with $n=4$,

$$ g=\frac{GM}{2r^{2}}\left[\frac{1}{r^{2}-R^{2}}+\frac{3r^{2}+R^{2}}{3\left(r^{2}-R^{2}\right)^{2}}\right], $$

which approaches $GM/r^{4}$ only when $r\gg R$. In other words: “treat the sphere as a point mass at its centre” is a long-range approximation for $n\neq2$, and an exact identity only for $n=2$.

That deserves more weight than it usually gets. Newton wanted to test his law against the motion of the Moon, and that test presupposes he could legitimately replace both the Earth and the Moon by points. If the exponent were anything but 2, that replacement would carry an error, and the whole test would degrade from a precision check into an order-of-magnitude estimate. Theorem A is not a corollary of the law of gravitation. It is the precondition for the law being testable.


3. Gauss’s law: the same result in three lines

Nothing in Section 2 was hard, but it was undeniably laborious. What follows may make you wonder whether those pages were wasted.

3.1 Gauss’s law, for readers who haven’t met it

Start with a picture of field lines. Every mass $m$ is a sink for gravitational field lines (gravity attracts, so the lines run into the mass), and the number of lines it swallows is proportional to $m$. Define the flux through a small patch of area $\Delta S$ as

$$ \Delta\Phi = g\,\Delta S\cos\theta , $$

where $\theta$ is the angle between the field and the patch’s normal. Flux is just the number of field lines crossing that patch.

Now the obvious objection: why is this picture consistent at all? Why don’t field lines evaporate on the way out?

Take a point mass $m$ and draw a sphere of radius $r$ around it. On that sphere the field is $g=Gm/r^{2}$ everywhere, the area is $4\pi r^{2}$, and the total flux is

$$ \Phi = \frac{Gm}{r^{2}}\cdot 4\pi r^{2}=4\pi Gm . $$

The $r$ has dropped out. That is the entire justification for the conservation of field lines: the field thins out as $1/r^2$, the sphere grows as $r^2$, and the two cancel precisely.

Replace the law by $1/r^{n}$ and the same calculation gives $\Phi\propto r^{2-n}$. The flux would depend on the radius, meaning field lines multiply or die in empty space. In such a universe “field line” would not be a useful concept, and Gauss’s law would not exist.

For an arbitrary closed surface and an arbitrary mass distribution one can show (we’ll take it on faith here):

$$ \oint \boldsymbol{g}\cdot\mathrm{d}\boldsymbol{S}=-4\pi G\,M_{\text{enc}}, $$

with the outward normal taken as positive; the minus sign is only there because gravity points into masses. The electrostatic version is $\oint\boldsymbol{E}\cdot\mathrm{d}\boldsymbol{S}=Q_{\text{enc}}/\varepsilon_0$ — identical in form, because Coulomb’s law is also inverse-square.

3.2 The three lines

【Figure 3: two concentric Gaussian spheres, one inside the shell ($r<R$) and one outside ($r>R$).】

Step 1 (don’t skip this one). The field of a uniform shell must be radial, with a magnitude depending only on $r$. The reason is symmetry: rotate the shell about its centre by any angle and nothing about the system changes, so the field cannot change either. The only vector field invariant under all such rotations is a radial one whose magnitude depends only on $r$.

Step 2. Take a concentric sphere of radius $r$ as the Gaussian surface. By Step 1, $g$ has the same value everywhere on it and is perpendicular to it, so the flux is just $g\cdot4\pi r^{2}$. Hence

$$ g\cdot 4\pi r^{2}=4\pi G\,M_{\text{enc}}\quad\Longrightarrow\quad g=\frac{G\,M_{\text{enc}}}{r^{2}} . $$

Step 3. Inside, $M_{\text{enc}}=0$, so $g=0$. Outside, $M_{\text{enc}}=M$, so $g=GM/r^{2}$. Done.

3.3 But this isn’t a free lunch

Gauss’s law shortens the proof to the point of anticlimax, and that fact deserves some thought. Is it the more fundamental method?

I’d say it isn’t a free lunch so much as a change of bookkeeping. We saw in 3.1 that the only reason Gauss’s law holds is the cancellation between $1/r^2$ and $4\pi r^2$. The inverse-square law has been packed into the tool in advance; small wonder that using the tool afterwards is easy.

This is not circular. The chain of implication is perfectly clean:

$$ \text{inverse square}\;\Longrightarrow\;\text{conservation of field lines (Gauss)}\;\Longrightarrow\;\text{shell theorem}, $$

and every arrow is a genuine derivation. But it is a reminder that a “simpler proof” often just relocates the complexity into the definitions and the machinery. Which is precisely the theme of the previous post: a good mathematical form — here, “flux” and “field lines” — is powerful because it compresses physical content into its own structure, so that later you can draw on it in whole blocks.

The integral is laborious but lets you watch the exponent 2 do its work, step by step. Gauss’s law is effortless because the 2 has already been hidden inside the tool. Both roads are worth walking once.


4. Running it backwards: the most precise null experiment in physics

So far we have been proving “inverse square $\Rightarrow$ no field inside”. The table in Section 1 says the converse holds too, and that turns out to be enormously useful.

Physicists love a null experiment. Rather than measuring how big a quantity is, measure that it is zero — because “we could not detect any deviation” typically buys far more precision than “we measured the value to be such-and-such”.

That is exactly what happened historically. Priestley noticed that there is no electric field inside a charged metal container and inferred from it that the electric force must be inverse-square. Cavendish then turned the observation into a quantitative experiment: if the force goes as $1/r^{2+q}$, the residual field inside the shell is proportional to $q$, so failing to detect a residual field puts a bound on $q$.

The same strategy has been refined ever since, gaining more than a dozen orders of magnitude. The current bound on the deviation is somewhere around $10^{-16}$ — which doubles as one of the strongest pieces of evidence that the photon is massless, since a massive photon would turn the Coulomb potential into a Yukawa potential $e^{-\mu r}/r$, break the inverse-square law, and leave a residual field inside the shell.

【Note to self: check the specific figures against primary sources before publishing — Cavendish $q\lesssim 0.02$, Maxwell $q\lesssim 1/21600$, Williams–Faller–Hill (1971) $q\sim(2.7\pm3.1)\times10^{-16}$.】

So the shell theorem is not just a pretty exercise. It is the main reason we know the exponent is 2 and not 2.000001.


5. What else is hiding in that “2”

While we’re here, let’s finish the thought. The exponent 2 carries a great deal more than the shell theorem:

  • It is telling you space is three-dimensional. Conservation of field lines requires $g\times(\text{area of a sphere of radius } r)$ to be constant. In $d$ dimensions the area of a sphere goes as $r^{d-1}$, so the exponent must be $d-1$. When we write $1/r^2$ we are really writing $3-1=2$.
  • It is why orbits close. Bertrand’s theorem: among all central forces, only the inverse square and the linear restoring force (Hooke’s law) make every bound orbit a closed curve. Nudge the exponent and planetary orbits precess slowly; the ellipse never quite finishes. Kepler’s first law is another child of that 2.
  • It forbids electrostatic levitation (Earnshaw’s theorem). You cannot hold a charged particle in stable suspension using electrostatic fields alone. The reason is the flip side of the shell theorem: the inverse-square law makes the potential extremum-free in any source-free region, so there is no potential well to sit in. “Zero everywhere inside a shell” is the extreme case of “no extremum”.

Feynman was right. The mathematical form of a law says vastly more than the sentence it appears to state. On its face $F=Gm_1m_2/r^2$ is only a claim about two point masses attracting each other. But written into that unassuming superscript are the dimensionality of space, the closure of orbits, the hollowness of shells, and the masslessness of the photon.


Next up: why planetary orbits are ellipses — another consequence of the same 2.