TOPIC 37 · PHASE G

Risk in a Heavy-Tailed World

Ruin is not an expected-value question

2026-08-23 · Fragility, Resilience & Early Warning

Here is a gamble whose expected return is +5% every single round, and in which more than eight players out of ten end up losing most of what they had. No bad luck is required. The word "average" is quietly pointing at two different things, and their answers run in opposite directions.

When you judge something you will do over and over — a strategy, a line of business, a way of pulling a deadline forward — you almost always start with "what does it come to on average". The move is so natural that nobody checks the conditions under which it works.

There are two. First, no single number in the pile can be large enough to set the average by itself. Second, you have to be able to repeat the thing many times and still be there after each one. Break either condition and the average does not give you a slightly wrong answer. It gives you an answer that points the other way.

Two other faces of heavy tails already have their own issues here: Topic 18 on how such a distribution grows by itself (slow accumulation plus threshold release), Topic 19 on what the distribution looks like and why its variance may not exist, Topic 36 on monitoring a system near a threshold. This one asks a single question: given that the world is like this, how should risk actually be computed and bets actually be sized — and why the standard machinery does not err at random, but underestimates systematically.

01Tails Are Not Outliers — They Are the Ledger

Measure 120 people, add up their heights, and ask what share the tallest person contributes. About 1.1%. Take 1,200 people and the share drops towards a fraction of a percent. On the question "how much height is there in this room", the tallest person has almost no say.

Now swap heights for wealth, for the loss from a single accident, for the area burned by a single wildfire. Out of 120 draws, the largest can be 13% of the total, and the top three a quarter of it. Double the sample and that share does not politely shrink — it waits for a bigger draw to show up and then hauls the total over to itself again.

That is what a heavy tail (also called a fat tail) actually means. Not "a big number turns up now and then", but: the total is decided by a handful of observations, and that does not ease off as the sample grows.

Same 120 draws, same total — the whole difference is one barThin-tailed — human heightHeavy-tailed — wealth, losseslargest = 1.1% of totalnothing here can move the totallargest = 13% of totaltop three = 24%Both panels sum to the same total. Shared vertical scale: bar height = that draw’s share of the total.
Same 120 draws, same total. On the left the tallest bar moves nothing; on the right the tallest bar is the whole ledger.

Which gives you a one-line test: divide the largest item by the sum, then double the sample size and compute it again. If the ratio falls, you are in a thin-tailed world. If it refuses to move, you are in a heavy-tailed one. No distribution needs to be fitted, and nothing needs to look like a straight line on a log-log plot — which, incidentally, is the single most common way this gets misdiagnosed → ref · identifying power laws.

The difference then changes what you do with data all the way down. In a thin-tailed world an outlier is usually noise: a mistyped entry, a sensor glitch. Trim it and what remains is cleaner. In a heavy-tailed world the thing you just trimmed is the thing you were trying to explain — the "typical behaviour" computed after removing the largest few describes a system that has never existed.

Earthquakes are the cleanest calibration available. One step up in magnitude is roughly 31.6× the energy released (101.5). So "how big is an average earthquake" is answered almost entirely by the largest few in the record: change the observation window and you get a different answer. That is not imprecision. The question has no stable answer here.

🎯 THE DECISION

In any table you actually make decisions from, put a column for "largest item as a share of the total" next to the mean, and compute it over two window lengths. If the share falls as the window grows, the mean is usable. If it does not, treat every average, standard deviation and year-on-year figure in that table as decoration — do not set budgets, leverage or headcount from them.

🌀 Art & letters · "Nobody knows anything" William Goldman's 1983 line is usually read as industry cynicism. Under this section's mechanism it is a statistical claim: box-office revenue is concentrated in a handful of films, and the largest single item can settle a studio's year. That yields a conclusion which is not obvious — raising the studio's average hit rate barely moves total revenue, because total revenue is not set by the median film. Only two things move it: place more bets, and do not sell the upside early when one lands (sequels, rights, participation).

02Why Risk Models Systematically Underestimate

The claim is not that models are sometimes wrong. It is that four independent biases all push the same way: they make danger look smaller.

First, "never seen" is not "very rare". Medical statistics has a crude and excellent rule called the rule of three: if an adverse event has not been observed once in n patients, the 95% upper bound on its rate is about 3/n. A phase-three trial with 3,000 participants and not a single serious adverse event guarantees only that the rate is probably under one in a thousand — and one in a thousand, across a million users, is a thousand cases. Zero observations carry far less information than intuition assumes.

Second, means and volatilities simply do not converge under heavy tails. The chart below does the same thing twice: draw one more sample, recompute the average, and see how far it still sits from the truth.

Sample mean ÷ true mean, as draws accumulate1.01010010002number of draws (log scale)sample mean ÷ true meannormal (thin tail)power law α = 1.3 (heavy tail)one record-breaking drawlifts the whole mean by 45%The blue line settles by ~50 draws. The pink one is still being thrown around by single draws at 2,000 — and the next jump has no schedule.
Sample mean divided by true mean. The blue line (thin tail) converges within fifty draws; the pink one (heavy tail, α = 1.3) is still being hauled around by single draws at two thousand.

The thin-tailed line settles after fifty-odd draws. The heavy-tailed line is still being dragged 40% off course by a single new draw at two thousand — and the next drag has no schedule. Which means the standard deviation computed from such data is an unstable number: change the stretch of history and it changes value. Every risk measure built on top of it inherits that.

Third, the most widely used measure is blind to the tail by construction. VaR (value at risk) says: "on 99% of days, the loss will not exceed X". Notice what it says about the remaining 1% — nothing at all. Whether the loss out there is 1.1×X or 50×X, VaR reports the same number. And in a heavy-tailed world, the difference inside that 1% is exactly the difference between a painful day and being out of the game.

Fourth, the largest value you have seen is a biased lower bound. "Worst drawdown on record", "peak traffic since launch", "hundred-year flood" — all of them mean "so far". In a thin-tailed world records get harder to break. In a heavy-tailed one they keep breaking, on no schedule. Designing to the historical maximum assumes you have already seen the worst.

And one mechanism amplifies all four: correlations rise in the tail. Things that normally move independently — different assets, different suppliers, different data centres — move together in the extreme. Diversification fails on precisely the day you needed it.

Stack those and you get the famous line from August 2007, when Goldman Sachs CFO David Viniar explained a week of enormous losses in the firm's funds to the Financial Times: "We were seeing things that were 25 standard deviation moves, several days in a row." Under a normal distribution, a 25-sigma event should not occur once in the age of the universe. So what the sentence really reports is not that markets went mad: when a model announces an absurdly small probability, the thing being falsified is the model, not the world.

🎯 THE DECISION

Replace "standard deviation / 95% confidence interval / VaR" in your risk reporting with two quantities: the average loss given that the threshold was breached (expected shortfall, also called CVaR) and how big a single event puts me out. The first forces you inside the tail; the second forces you to write the absorbing barrier as an actual number. If that is too much, at least annotate every probability estimate with whether it was counted or extrapolated.

🌀 Engineering history · Feynman's Appendix F After Challenger, Feynman found that NASA management put the probability of catastrophic failure at one in 100,000 while working engineers put it around one in 100. He did the arithmetic: one in 100,000 means you could launch a shuttle every day for 300 years and lose one. The actual record is 135 flights and two losses, about 1 in 67. The conclusion this supports is harder than "management lied": when a probability is small enough that no attainable frequency record could ever test it, the number reports the confidence of a model rather than the risk of a system. "One in 100,000" should be read as "we have no data".

03Expected Value Counts Ten Thousand People — You Have One Life

Play a gamble. You start with 100. Flip a fair coin: heads, your money grows by 50%; tails, it shrinks by 40%. The expected return per round is (+50% − 40%) ÷ 2 = +5%. It looks like free money.

Run it 4,000 times, sixty rounds each:

+50% / −40% on a fair coin — expected +5% every round100×10×0.1×0.01×0.001×015304560roundsmultiple of the starting stake (log)average of 4,000 paths(the ensemble average)the median path (time average)×38×0.0483% of the 4,000 paths end below the starting stake; 55% end below one tenth of it.
Four thousand paths through the same gamble. The copper line (ensemble average) climbs, the teal line (median path) falls — and both numbers are correct.

The average across 4,000 paths reaches 38× the stake. The median path holds 4% of it. 83% of players finish below where they started and 55% finish below a tenth. And the reason that handsome average keeps rising is that the top 1% of paths — forty of them — hold 93% of all the wealth in the ensemble.

Both numbers are right; they average over different things. "+5% per round" averages sideways: ten thousand people standing at the same instant, their money summed and divided by headcount — the ensemble average. What you actually live through averages downwards: one person, round after round — the time average.

Going downwards, money multiplies, it does not add. So the relevant quantity is the geometric mean: √(1.5 × 0.6) = 0.9487, which is −5.1% per round. In the same gamble, "expected +5%" and "−5% every round" are both true statements.

When are the two averages equal? Physics has a name for that condition: ergodicity. Boltzmann's assumption was that the long-run behaviour of one gas molecule matches the behaviour of many molecules at one instant. For a gas that broadly holds, because molecules do not drop out. For money it fails, because you can hit zero — and zero is an absorbing barrier: nothing follows it, so every later round and every bit of its positive expectation vanish with you. Expected value is counting rounds you are no longer present for.

The repair is not abstinence; it is sizing. Stake only a fraction f of what you have and the same gamble changes character immediately:

All in (f = 100%): −5.1% per round | half (f = 50%): exactly break-even | f = 25%: +0.62%, the optimum | f = 10%: +0.40%

Position size decides whether this gamble is slow suicide or a compounding machine, and the expected return never changed at any point. That optimum has a name: the Kelly criterion (1956), which maximises the logarithmic growth rate — that is, the time average. Almost nobody bets full Kelly in practice: the drawdowns are brutal, and it assumes your probability estimates are correct, an assumption usually more dangerous than the gamble.

If you want to know what your own path looks like, the honest method is to run the rules ten thousand times and look at the whole distribution rather than at one number produced by a formula → ref · Monte Carlo & ensemble simulation.

🎯 THE DECISION

Ruin cannot be priced with expected value. Three moves: ① rewrite any repeated decision in multiplicative form — multiply the outcomes together and take the nth root; if that number is below 1, do not do it, however pretty the expectation; ② give any single exposure an absolute cap (not a cap set as a fraction of expected gain), which pulls the multiplicative process back towards an additive one; ③ for irreversible things — zero, health, legal standing, reputation — forbid expected-value arithmetic entirely: they do not enter the average, they end it.

🌀 Military history · Rome and Carthage were not symmetric In the Second Punic War Hannibal annihilated a Roman army at Cannae; Rome refused to negotiate and raised new legions the following year. Carthage lost once, at Zama, and the war was over. Comparing the two by win rate or tactical skill is computing an ensemble average — flattening every battle together. What decided the outcome was where the absorbing barrier sat: Rome could absorb a catastrophic defeat and continue existing, Carthage could not. The conclusion generalises to any contest: an advantage is sometimes not "playing better" but pushing your own absorbing barrier further away — depth, allies, a manpower pool, cash.

04Where This Breaks Down

Heavy-tail language is very easy to wear as a universal coat. Five limits follow, two of which are enough to overturn most of what you have just read.

First, not everything is heavy-tailed. Quantities built by adding many small bounded contributions — heights, body temperature, measurement error, a day's commute — are thin-tailed, the central limit theorem applies as advertised, and the mean and standard deviation are good rulers. Heavy tails come from specific mechanisms: multiplicative growth, preferential attachment (rich get richer), contagion coupling, winner-take-all. The test is whether the mechanism is present, not whether the plot looks straight. And most things that look straight fail a proper test → ref · identifying power laws. Applying heavy-tail thinking where it does not belong has a real price: you hoard defences against disasters that will never arrive, and spend the budget on fear.

Second, the ergodicity argument has a hard precondition: you cannot pool and you cannot reset. For an entity that can pool a large number of near-independent risks — an insurer, an index fund, an institution with continuing outside capital — the ensemble average really is the right object, because it genuinely is all ten thousand paths at once. So "never use expected value" is wrong. The right question is not whether expected value is legitimate but "in this particular case, am I the path or the ensemble?"

Third, one formula, two rival justifications — do not treat them as one thing. In 1738, working on the St Petersburg paradox at the Petersburg Academy, Daniel Bernoulli proposed that people weigh not money but its utility, which grows roughly with the logarithm of wealth: a psychological explanation. Ole Peters has recently defended the same logarithm from the time average: not a matter of preference at all, but a property of repeated multiplicative processes. Economists have not accepted the substitution — Doctor, Wakker and Wang replied directly in Nature Physics in 2020, arguing that it repackages an old problem. You need not take a side, but you should know where the two justifications diverge: exactly at the question of whether risk can be pooled — the psychological account lets preferences change, the dynamic account says preference has nothing to do with it.

Fourth, heavy tails are not fatalism. What you give up is point prediction, not intervention. The distribution itself is structural and can be edited: per-bet caps, stop-losses, modular firebreaks, insurance, splitting one large bet into many small ones — each of these truncates the tail. Saying "it is unpredictable anyway" and doing nothing moves a conclusion about prediction onto a question about structure, where it does not belong.

Fifth, and hardest to hold to: do not use "black swan" as a retroactive excuse. An honest heavy-tail claim has to be sayable in advance: events of what size, arriving about how often, absorbed by what structure. If you cannot say those three things beforehand, saying "heavy tail" afterwards is only a more dignified name for the failure.

Before using any of this, answer three questionsIs this quantity addedup, or multiplied / contagious?Additive, bounded → thin tail: the mean is a good rulerMultiplicative, preferential attachment, contagion → heavy tailIs thisparticular loss reversible?Reversible, repeatable → expected value is fineZero, out, irreversible → expected value cannot price ruinAre you onepath, or an ensemble?Can pool and reset (insurers, index funds) → the ensemble averageOnly this one path (a person, a firm) → time average / log growthOnly after these three does the choice of model matter.
Answer these three before using anything from this issue. When all three answers sit on the top branch, none of it applies.
🎯 THE DECISION

Before using anything from this issue, answer the three questions above: added or multiplied · is this particular loss reversible · am I the path or the ensemble. When all three answers land on the top branch (additive, reversible, ensemble), use expected value with a clear conscience — none of this applies there, and forcing it on will have you building sea walls where there is no sea.

🌀 Medicine · "average life extension" and number needed to treat Clinical practice has a measure called NNT (number needed to treat): how many people must be treated to prevent one event. It has to exist precisely because of this section's dividing line — the benefit of a preventive measure at population scale is an ensemble property, while each person faces one path, and the benefit is concentrated in the few who would otherwise have had the event. So "extends life by an average of x days" carries almost no information for an individual: public health is natively an ensemble decision and personal choice is natively a path decision, and one expected-value number answering both will mislead both.

🎒 In Practice · BigCat

  1. Investing & position sizingYou see a handsome annualised return or backtest Sharpe and your first instinct is "could this take some leverage" — that is sizing one path from an ensemble average. Compute it differently: multiply the strategy's period returns together and take the root to get the geometric mean; then ask "after three of the worst months in this record arrive back to back, am I still here". The move to stop: setting a leverage ceiling from the worst historical drawdown — that number is a biased lower bound, it only ever means "so far".
  2. Health & energyThe recurring situation: two nights burned for a deadline, the cost afterwards feels acceptable, so it happens again next time. The expected cost of any one instance really is small, but this is a multiplicative process with irreversible items hidden in it (injury, chronic conditions, immunity). The quantity to watch is not "how tired was I" but recovery speed — after a week of the same intensity, how many days back to baseline, and is that number growing or shrinking (Topic 36's critical-slowing signal, turned on yourself). The inference to stop making: treating "nothing went wrong last time" as evidence of safety — by the rule of three, ten clean rounds only bound the risk at about 30%.
  3. Leading a teamThe recurring situation: a decision comes up and the whole room debates whether the expected payoff is big enough, while nobody asks whether it can be undone. Sort decisions into reversible / irreversible first and set separate bars: lower the bar for reversible ones, move faster, allow failures (this class should be made more often, not less); route irreversible ones through a slower path that must state "if this is wrong, how long to get back to today, and at what cost". One sentence you can use as-is: "if this decision is wrong, can we be back where we are today by next quarter?" — if nobody can answer, it is not ready to be made.

🌀 Crossings

Going Deeper

If sample means are unreliable under heavy tails, how do insurers price anything?

Not by the accuracy of a point estimate. By three other things: layering (the catastrophe layer is carved out and sold on to reinsurers), limits (every policy writes a cap, truncating the tail by contract), and capital (reserves cover the tail of the distribution, not its mean). They did not get heavy tails right; they rebuilt themselves into a structure that survives being wrong. It is a mature example of structural intervention beating point prediction.

When is "take more shots and keep the upside" the wrong advice?

When each attempt carries a risk of going to zero and the attempts cannot be pooled. More attempts raise the probability of hitting at least once, and also the probability of being knocked out before you hit; which grows faster depends on whether the single-attempt loss is capped. That is why barbell-shaped betting always pairs "take many shots" with "cap the loss on each" — pull the two apart and it turns over.

Does the time-average argument prove long-termism?

No. It proves that in a multiplicative process the quantity to maximise is the logarithmic growth rate, which requires you to still be present. It is entirely silent on what is worth pursuing — someone can rationally choose a single all-in path if the thing they want is not compounding. Reading a dynamical result straight off as life advice is the easiest overreach in this issue.

Where is an organisation's absorbing barrier, and how would you measure it?

Try writing it as a falsifiable sentence: after what event can this organisation not return to where it is today? How many months of broken cash flow, how many key people leaving at once, a data incident touching how many users. Being unable to state a threshold usually means nobody owns the question — not that the barrier is absent.

Why do people overreact to plane crashes and terrorism yet stay numb to cumulative risk?

One explanation: people are sensitive to single events that can be pictured, and have no intuition for processes that multiply. That runs exactly counter to this issue's mechanism — the things that deserve the alarm are usually the second kind (compounding debt, slow health erosion, systems whose coupling keeps tightening), and they offer no picture at all.

Further Reading