TOPIC 19 · PHASE C

Power Laws and Heavy Tails

When the average is a trap

2026-08-05 · Self-organisation and criticality

You can compute a country's average height, and you can compute the average energy released by an earthquake. The first number means something. The second does not merely mislead — it does not exist. Not for want of data. It does not exist.

You are doing something entirely ordinary: taking an average. How many copies a book sells. How many households a blackout affects. How much revenue a new user brings in. Pull the history, divide, get a number, plan around it.

For heights, weights and commute times this is fine. For those three it walks you into a ditch — and into the particular kind of ditch where trying harder makes it worse: collect another year of data and the number will not settle down, it may double. Collect another year and it may go back.

The problem is not sample size. The problem is that "take an average" quietly assumes something nobody ever told you — that a typical value exists. For a large class of systems it does not, and that class happens to contain the things you most want to predict: crashes, breakout hits, accidents, epidemics, fortunes.

The issues in this phase each ask a different question about criticality. Topic 16 asked why microscopic detail stops mattering at the critical point. Topic 17 asked how connectivity suddenly spans a system. Topic 18 asked who tuned the system to criticality. This one is not about a mechanism but about the output — what the statistical signature these systems emit actually looks like, and which of your habitual moves it cancels.

01One Hundred Draws, the Same Total

Look at the thing before naming it.

Below are two sets of one hundred numbers whose totals are identical. Think of them as two ways of splitting the same sum of money, the same pool of traffic, the same audience. Both are sorted largest-first and drawn as bars against the same vertical scale.

100 draws each · identical totals · same vertical scale Normal (height-like) Power law (sales-like) mean largest = 1.6 × the mean 53 fall below the mean largest = 12 × the mean 71 fall below the mean · top 10 take 44%
Same count, same total, two ways of splitting it. The bar hitting the ceiling on the right is not an outlier — it is what this distribution looks like when working normally.

On the left, "the average is 6.7" is a sentence with content. Nearly every bar sits close to it and the tallest is only 1.6 times as high. Tell someone the average and the picture in their head is roughly right.

On the right the same sentence carries almost nothing. Seventy-one of the hundred numbers fall below the mean. The median is 3.1, less than half the mean, while the largest is twelve times the mean and the top ten together take 44% of everything. "Average 6.7" describes no actual bar: most are far below it, a few are far above, and the region around it is the emptiest part of the distribution.

The shape on the right is a power law. Its definition can be stated plainly: the number of things exceeding size x is proportional to x raised to some negative power. Written out, P(exceeding x) ∝ x−β, and β is the tail index. Smaller β means a heavier tail and more monsters. → ref · Identifying and Misidentifying Power Laws

That formula comes with an immediately usable reading: put both axes on a scale where each step multiplies by ten — log-log axes — and a power law is a straight line whose slope is −β. A bell-shaped normal distribution on the same paper curves down and hits the floor: the number of events it predicts beyond five times the mean is essentially zero.

The reason for the two shapes lies in how they are generated. A normal distribution comes from many small factors added together — height is hundreds of genes and nutritional conditions summed, none of them decisive alone. Power laws typically come from multiplication or contagion: a book that already sells well is seen by more people and so sells better still; a grid node that is already connected carries more load and so fails more readily. Addition grinds differences flat. Multiplication amplifies them.

One overused claim deserves handling here. The "80/20 rule" — twenty percent of the items hold eighty percent of the total — is not the definition of a power law. It is what one particular tail index happens to produce: β ≈ 1.16 gives exactly 80/20. Change β and the split changes completely (smaller β might give 90/10, larger β only 60/20). Treating 80/20 as a universal law is treating a parameter as a principle.

🎯 DECISION

Given any set of "size" data — sales, blast radius of an incident, customer value, loss per accident — the first move is not to average it. It is to sort it and compute what fraction of the total the largest 1% holds. If that fraction exceeds 10%, drop the mean from your reporting and report the median together with the top-1% share instead. This step needs no statistics and takes five minutes.

🌀 Economics and institutions · why statistical agencies report the median National statistics offices publish median household income rather than the mean precisely because income has the shape on the right. Follow the mechanism and a less obvious conclusion appears: any policy target framed "per capita" over a heavy-tailed quantity describes not the situation of most people but the size of the tail. Floor space per capita, bandwidth per capita, books per capita all suffer this — the indicator rises, and all that may have happened is that the head stretched a little further while the middle did not move at all.

02The Average That Never Settles

Saying the average "describes nobody" is only a complaint about description. This section makes a harder claim: you may not be able to measure that average at all.

You probably half-believe a rule: more data means more stability. Flip a coin a hundred times and the fraction of heads jumps around; flip it a million times and it sits on 50%. That rule is the law of large numbers, and it underwrites nearly every piece of advice that begins "collect more data".

It has preconditions, and in the heavy-tailed world — the broad class of distributions whose extreme events decay more slowly than exponentially, of which the power law is the cleanest member — those preconditions can fail. The figure below was actually run: draw a hundred thousand random numbers from each of two distributions, recompute the running average after every single draw, and watch where it goes. The vertical axis is the running average divided by the theoretical mean, so 1.0 means converged.

Running mean ÷ theoretical mean — 100,000 draws 1.0 2.0 3.0 Power law — not converged Normal — converged fast Draw no. 21,668: one number lifts the mean to 2.9 × 1 10 100 1k 10k 100k Draws so far (each step ×10) → The orange series does have a finite mean — 100,000 draws have not found it
The blue line lies down within a few hundred steps. The orange line is lifted almost threefold by one draw at twenty thousand, then spends the remaining eighty thousand steps crawling back, ending at 1.3.

Note the detail that matters about the orange line: it does not wobble, it steps. For twenty thousand draws it sits near 0.7, and had you stopped there you would have believed you had measured it. Then one number arrives and lifts the whole average to 2.9 times the truth in a single step. That number is not an error, not dirty data, not something to strip out — it is what this distribution is supposed to do. And more such jumps are coming, just ever more rarely.

So "collect more data" fails here completely: your average does not creep toward the truth, it sits low for long stretches and is occasionally over-corrected by a single jump. No individual snapshot is trustworthy, and you have no way of knowing where between two jumps you currently stand.

Behind this is a clean mathematical fact worth memorising, because it turns "can this quantity be averaged" into something you can look up. For a power law with tail index β:

The tail index β decides which familiar statistics are still alive β = 1 β = 2 no mean no variance mean exists no variance mean exists variance exists tail index β → earthquake energy ≈ 0.67 word frequency / city size ≈ 1 wealth (the 80/20 point) ≈ 1.16 Normally distributed height has no place on this axis — its tail decays faster than any power law
Where β falls decides which of your statistical habits still apply. This is not a question of precision but of existence.

The mean exists only if β exceeds 1; the variance — that is, the very concept of "typical fluctuation" — only if β exceeds 2. "Does not exist" does not mean "is very large". It means the integral diverges and there is no limiting value: the sample average you compute is a random number wandering with sample size, converging to nothing.

(A convention trap lurks here: many papers quote the exponent α of the probability density instead, related by α = β + 1, so the same fact is written there as "the mean exists if α > 2, the variance if α > 3". Before using any quoted number, check which convention it uses, or you will be off by exactly one.)

This is not a mathematician's toy. Earthquakes live in the leftmost band. The Gutenberg-Richter law says each unit of magnitude cuts the count roughly tenfold; converted into released energy, the tail index is about 2/3 — less than 1. Which means the question "how much energy does an average earthquake release" has no answer within the model. The only reason a number can be computed in practice is that the Earth is finite and fault lengths are bounded, which chops the tail off. The average you compute is set by where the truncation falls, not by the physics.

🎯 DECISION

Subject any average you plan to decide on to a running-mean test: add the samples one at a time in order, recompute the average after each, and plot the curve. If it lies flat, the mean is usable. If it is still stepping, your average is currently determined by the few largest samples — report the median and quantiles instead. Ten lines of code, and it should come before any formal statistical test.

🌀 Biology and medicine · one R₀, two completely different epidemics Epidemiology summarises spread with the basic reproduction number R₀, the average number of people one case infects. But measurements of SARS-CoV-2 transmission showed extreme unevenness between individuals: Endo and colleagues estimated the dispersion parameter k at about 0.1 in 2020, corresponding to roughly ten percent of cases producing eighty percent of transmission. The same R₀ = 2 can mean "everyone infects two" or "ninety percent infect nobody and a few infect dozens". The averages are identical; the correct response is opposite. The first calls for broadly reducing contact, the second for targeting crowded settings and backward contact tracing. Two worlds with the same mean and different distributions need different interventions.

03What Scale-Free Actually Means

"Scale-free" is often used as an adjective, as though it were a literary way of saying "big and very uneven". It is in fact a specific, checkable statement.

Start with its opposite. Suppose adult male height is bell-shaped with a mean of 1.75 m and a standard deviation of 7 cm. Then 24% clear 1.80 m, 1.6% clear 1.90 m, 0.018% clear 2.00 m and 0.0000287% clear 2.10 m. Each extra 10 cm cuts the population to 1/15, then 1/90, then 1/619 of what it was. The divisor itself grows fast: harder as you go up, and the difficulty accelerates.

A power law does nothing of the kind. The probability of doubling is independent of where you currently are. Whether a book has sold ten thousand copies or a million, its chance of doubling again is the same number: 2−β. At β = 1.16 that is 45%, always 45%.

How fast the odds of "one more step up" collapse Normal · each extra 10 cm of height Power law · each doubling of size > 1.80 m 24% > 1.90 m 1.6% > 2.00 m 0.018% > 2.10 m 0.00003% ÷ 15 ÷ 90 ÷ 619 the divisor is accelerating > 10k copies 100% > 20k copies 45% > 40k copies 20% > 80k copies 9% ÷ 2.2 ÷ 2.2 ÷ 2.2 constant this is "scale-free"
The three ÷2.2 on the right are identical, and that is no coincidence: it is the literal content of "scale-free" — change the ruler and the pattern is unchanged.

That is what scale-free means: no position on the distribution is special. There is no normal range, therefore nothing outside the normal range, therefore no typical scale to anchor on. Switch the horizontal units from copies to tens of thousands of copies and the picture is identical.

This property has a direct and thoroughly counterintuitive consequence. Under a power law, already large is evidence of larger still. Precisely: if you know a quantity has already exceeded x, the expectation of its final value is x times the fixed factor β/(β−1) — the expected value is proportional to the current value. At β = 1.16 that factor is 7.25: a book already at 100,000 copies has an expected final figure of 725,000; one at a million expects 7.25 million. It never tops out.

A normal distribution does the opposite. Given that a man clears 1.90 m, his expected height is 1.925 m; given that he clears 2.00 m, 2.017 m. The conditional expectation hugs the threshold — raise the threshold and the expectation barely moves. This is where regression to the mean lives: the distribution has a centre, and an observation far from it is usually followed by one closer to it.

So "it has run up so far, it is due for a pullback" is not general wisdom. It is a claim with a precondition, and the precondition is that the quantity has a centre. For scale-free quantities it fails, and it fails in the opposite direction.

🎯 DECISION

Classify the quantity before choosing your intuition. Thin-tailed quantities — height, reaction time, daily commute — obey regression to the mean, so expect extremes to be followed by a fallback. Heavy-tailed quantities — sales, reach, loss per incident, revenue per customer — do the reverse: size reached so far is the best available predictor of further growth. Concretely: in your post-mortem template, replace the question "is this an outlier we should exclude" with "which class is this quantity in". The first question guarantees that you throw away your most informative data whenever the tail is heavy.

🌀 Engineering history · which process to migrate In 1997 Harchol-Balter and Downey measured the distribution of UNIX process lifetimes and found it heavy-tailed. That overturned the prevailing scheduling intuition — "this process has been running a long time, it must be nearly done". Under a heavy tail the reverse holds: the longer a process has lived, the longer its expected remaining life. Their load-balancing rule follows directly: migrate the processes that have already lived longest, because only those are likely to survive long enough to repay the cost of migrating them. The same conditional-expectation property, turned into one executable scheduling line.

04Where This Breaks Down

Power laws are the most abused tool in complexity science, for a simple reason: they look like a straight line, and straight lines make people relax. Four limits before use.

First, "it looks straight" is not evidence. Log-log axes are an extremely forgiving way to draw: they compress the vertical range across several orders of magnitude, and the eye's sensitivity to curvature collapses along with it. A log-normal distribution — what you get when many small factors multiply — plotted on the same paper is nearly inseparable from a power law across two or three decades. In 2009 Clauset, Shalizi and Newman re-examined a batch of widely cited power-law data sets under one uniform statistical procedure; several failed outright, and among those that passed, many could not be distinguished from a log-normal. → ref · Identifying and Misidentifying Power Laws

Second, real tails are always truncated. No book sells to more people than exist; no earthquake exceeds the crust. So a strictly non-existent mean never shows up in real data as infinity. It shows up as the curve in section 02: computable, non-converging, and with a value effectively set by where the truncation falls. That does not weaken the conclusion, it makes it operational — if you must report a mean, report alongside it how long your observation window was and what the largest observed value is.

Third, and most important: the same straight line can be emitted by entirely different mechanisms.

Many to one: several mechanisms emit the same line preferential attachment proportional growth + a floor clusters at a critical point one straight line on log-log axes infer backwards line observed → "some amplifying mechanism is running": valid line observed → "this particular mechanism is running": invalid
Rich-get-richer, random proportional growth, critical connectivity — three unrelated processes, one power law. The line cannot serve as evidence for any specific mechanism.

This directly bounds how far the previous issue's conclusion travels. Seeing a power law entitles you to say "some mechanism in this system amplifies across scales". It does not entitle you to conclude "therefore this system is in a self-organised critical state", because proportional growth, preferential attachment and extremal aggregation all emit the same line and none of them involve criticality. Inferring mechanism from shape is inverting a many-to-one map, and that inverse is not unique.

Fourth, heavy-tailed is not the same as power law. This issue uses "heavy tail" for the broad class in which extreme events decay more slowly than exponentially; the power law is only its cleanest member. Log-normal and stretched-exponential distributions are also in the class, yet all their moments exist. So the section-02 results about non-existent means apply only to the power-law branch and cannot be carried wholesale onto "heavy tails". Saying a quantity is heavy-tailed and saying it follows a power law are claims of very different strength.

🎯 DECISION

Never treat "fits a power law" as evidence for a mechanism. If you want to claim that some specific mechanism is at work — rich-get-richer, self-organised criticality, contagion — you must produce another testable consequence of that mechanism: an ordering in time it predicts, an intervention effect it predicts, the shape of some other quantity it predicts. With the distribution shape as your only evidence, the defensible statement is "there is amplification across scales", and you stop there.

🌀 Literature and the arts · authorship attribution in stylometry Deciding who actually wrote a text from word-frequency statistics works precisely because shared structure carries no information: the word frequencies of every natural-language text follow roughly the same power law, so "it obeys Zipf's law" is useless for attribution and only the residual departures help — the rate of function words, preferences among rare collocations. This is isomorphic to the many-to-one map above: a shape shared by all generating processes is the greatest common divisor of those processes, and therefore exactly the thing that cannot distinguish them. Every practice that treats conformity to a universal regularity as identifying evidence is offering a common divisor as a fingerprint.

🎒 Scenarios · BigCat

  1. Engineering and system designA latency dashboard reading "mean response 180 ms" is very nearly a report about the slowest tenth of a percent of requests, and neither its rises nor its falls can be interpreted. Read three numbers instead: median, P99 (the latency 99% of requests beat), and what fraction of total waiting time the slowest 0.1% accounts for. More valuable still is putting section 03 to work: since a request that has already waited a long time has a longer expected remaining wait, "time out and reissue" is an effective tactic under heavy-tailed latency rather than wasted work — the reissued request is a fresh draw from the distribution, with a shorter expectation than continuing to wait. Concretely: give idempotent reads a timeout-and-retry, with the threshold near P99 rather than at some multiple of the mean.
  2. Health and energyGiving yourself a pass because "I averaged seven hours this week" is the private version of the section-01 error: it smears two nights of four hours and five nights of eight into one number, and the body does not keep its books that way. Recovery is governed by the worst few nights. Track three counts instead: nights under five hours this month, how many times two such nights fell back to back, and the longest unbroken run of deficit. The move to stop: grading yourself on a weekly average.
  3. ParentingCounting "we argued three times this week" carries almost no information about conflict, because conflict severity is heavy-tailed — one genuine rupture outweighs dozens of squabbles combined. What can change is what gets recorded: not the count, but how bad the worst one was, and how long it took to get back to talking normally. A lengthening recovery time is a far more serious signal than a rising count; fewer incidents with slower recovery usually means conflict is being suppressed rather than resolved.

🌀 Crossings

Going Deeper

If the mean is unreliable, why not switch to the median everywhere?

Because many questions are about totals, and totals are necessarily set by the tail. An insurer owes the sum of all claims, not the typical claim; a server must absorb the total wait across all requests. On such questions the median is safe nonsense — stable, but answering a different question. The fix is not one substitute number but a set: the median for the typical case, quantiles for the shape of the tail, and the largest observation to record how long you have been watching.

Given a quantity, how do you tell whether "already big means bigger" or regression to the mean applies?

Look for an intrinsic ceiling mechanism. Height is bounded by skeleton and metabolism, exam scores by the maximum mark; such quantities have a centre and regress toward it. Sales, reach and wealth have no such mechanism and do have positive feedback — selling well means being seen means selling better — so the tail stays open. A crude but serviceable test: could this quantity go up tenfold without hitting any physical or institutional wall? If yes, do not assume regression.

Beyond "stop using the mean", is there anything positive to do with a heavy tail?

There is more to do than in the thin-tailed world, not less. If outcomes are decided by a few draws, then increasing the number of attempts, keeping the attempts genuinely different from one another, and bounding the downside of each becomes more effective than raising the per-attempt success rate. That is why venture funding, drug screening and creative work all take the shape of portfolios — not gambling, but design fitted to the shape of the distribution.

Does the tail index change over time?

It does, and that often matters more than its level. Markets have different tail shapes in calm and in crisis; a social network's spread distribution changes when the recommender is rewritten. The trouble is that estimating a tail index requires many tail samples, and tail samples are inherently scarce — so "the tail index has shifted" is usually measurable only long after the shift. That is a genuine weakness of this issue's toolkit; do not expect real-time warning from it.

Further Reading