TOPIC 18 · PHASE C

Self-Organized Criticality

Big events need no big cause

2026-07-20 · Self-Organization & Criticality

After every crash and every outage we go looking for the culprit. But there is a whole class of system whose largest disasters are set off by the most ordinary thing imaginable — and in those systems, the act of hunting for a culprit carries no information at all.

A magnitude-8 earthquake and a magnitude-3 earthquake begin the same way: somewhere a piece of rock stops holding and slips. No great earthquake is caused by a correspondingly great slip. The difference isn't in the beginning. It's in how much else lets go afterwards.

That sounds like the familiar line about small causes having large effects. The genuinely counterintuitive part comes next: systems like this climb by themselves to the state where a touch can bring everything down, and then stay there. Nobody tunes a parameter. Nothing unusual arrives from outside. Stability is the thing that would need explaining; being one grain from collapse is the resting state.

Five issues on this site deal with criticality, from different angles: Topic 10 is bifurcation and tipping (the system jumps to another stable state), Topic 16 is phase transitions and universality (why details stop mattering at the critical point), Topic 17 is percolation (connectivity suddenly spanning), Topic 36 is critical slowing down (how to monitor it). This issue asks one question only: who tuned the system to the critical point? The answer is that nobody did — it walked there on its own.

01The story of one grain

In 1987 three physicists — Per Bak, Chao Tang and Kurt Wiesenfeld — wrote down a set of rules simple enough to be faintly ridiculous → ref · sandpile model:

Picture a sheet of graph paper with a few grains of sand in each square. Pick a square at random and add one grain. If that square now holds more than three grains it topples: four grains leave, one going to each of its four neighbours. A neighbour that now holds more than three topples in turn. Grains that spill off the edge are gone.

That's the entire model. No physics, no friction coefficient, no gravity. Yet run it and this happens: the grain you just added may do nothing at all, or it may set off a chain of topplings that sweeps across half the sheet. And you cannot tell the two apart in advance, because the grain you added was identical both times.

Same rule, same grain — outcomes three orders of magnitude apart 1. add a grain this cell is now full 2. one toppling one grain to each neighbour 3. cascade 19 cells drawn in the next grain may move just 1 cell Shade = grains in the cell; orange = cells that toppled during this avalanche
One avalanche in the sandpile model. The action that triggered it is identical to the action that does nothing.

The most important sentence of this issue is already buried in there: in systems like this there is no correspondence between the triggering event and the size of the event. So when you ask what caused this great avalanche, the answer is always an ordinary grain of sand — true, and completely unilluminating.

🌀 Engineering · dynamite at ski resorts Standard avalanche control is not about holding the snow back. It is about setting off strings of small avalanches with explosives while the snowpack is still thin. The logic is that rule read backwards: since the trigger is neither controllable nor informative, stop managing triggers and manage accumulation — hold the system further from its critical point. Prevention stops being "predict which grain" and becomes "don't let the pile reach that height."

02The system walks to the cliff edge by itself

In the phase-transition picture, the critical point is something you tune to: cool the magnet to precisely the right temperature and it sits at criticality. Miss by a little and it doesn't. This is why physics long treated criticality as rare and delicately prepared.

What's strange about the sandpile is that nobody is tuning anything. You add grains one at a time. Past a certain point the system settles onto a particular slope — too shallow and grains pile up and steepen it, too steep and a large avalanche unloads it. Both sides push inward.

Criticality here isn't a target to aim at; it's an attractor. The system's own dynamics carry it there and then it stays. That is the whole content of the word "self-organized": nobody organized it into criticality, it went by itself.

Along the slope axis, both sides are pushed to the same value shallow steep slope of the pile → critical slope θc add grains → pile up → steepen avalanche → unload → flatten No one sets the system to θc — it is the one place neither side pushes away from
Criticality isn't a point being aimed at, it's an attractor. That is what "self-organized" means in self-organized criticality — SOC for short, as it is used below.

This changes the character of the whole subject. If criticality had to be finely prepared, critical systems in the real world would be vanishingly rare. If systems climb there on their own, then anything built out of slow accumulation plus threshold release drifts automatically into a state where something big could happen at any moment — crustal stress, forest fuel load, leverage in a financial system, the technical debt nobody in the org wants to raise.

🎯 DECISION LINE

Stop asking what triggered it; ask how far it currently sits from criticality. The trigger is neither predictable nor informative, but accumulation is measurable — snowpack depth, fuel load, leverage ratio, the queue of unreviewed changes. Move your monitoring budget off catching triggers and onto measuring build-up.

🌀 Institutions · the reversal in forest fire policy The US Forest Service spent decades suppressing every fire it could, and the result was undergrowth accumulating untouched for generations until a fire became Yellowstone 1988. Policy eventually reversed toward controlled burns — because suppressing small events was driving the system toward criticality. The inference isn't obvious: intuition says fewer events means safer, when in a self-organized critical system the small events you prevent are stored up as the magnitude of a future large one.

03There is no such thing as a typical size

Run the sandpile ten thousand times, record how many cells each avalanche touched, and plot the distribution. You get a power law: enormous numbers of tiny avalanches, very few huge ones, no gap in between — and no typical value.

Compare human height. Height has a typical value, around 1.7 metres, with almost everyone clustered nearby. You have never met a three-metre person and never will. That's a bell curve; it has a scale.

Avalanche size has nothing of the kind. The longer you measure, the larger the largest avalanche you've seen, without limit. Asking how big a typical avalanche is, is like asking how big a typical earthquake is — the question has no mathematical answer.

avalanche size (log) → frequency (log) power law straight = each 10× in size, a fixed drop in frequency normal all of this: the bell curve says "cannot happen" One chart. For the events on the right, a bell curve predicts a rate of about zero — and they happen every year
On log-log axes a power law is a straight line: no break point, and no "typical size".

That straight line is everywhere in the real world. The Gutenberg-Richter law puts it most bluntly for earthquakes: raise the magnitude by one and the number of such quakes falls to roughly a tenth. The relation holds across several orders of magnitude and a century has not overturned it. Blackout sizes, forest fire areas, and market drawdowns all produce comparable lines.

So "big events need no big cause" acquires an exact meaning: large and small events are two ends of one process, not two kinds of thing. To go looking for a special cause commensurate with the largest one is to assume it belongs to a separate category — and on the distribution there is no such dividing line to be found.

🌀 History & war · Richardson's statistics of deadly quarrels The meteorologist Lewis Fry Richardson spent decades counting the dead in conflicts of every size, and found that from street brawls to world wars the magnitudes follow a power law with no break. That yields a conclusion historians tend to dislike: if great wars and small skirmishes sit on the same straight line, then "what caused the First World War" — a demand for a unique cause worthy of its scale — may have been the wrong question from the start. The Archduke was the grain of sand.
🎯 DECISION LINE

Stop naming a culprit for every crash, outage, or wave of resignations. Do two things instead: plot the size distribution of comparable past events and check whether it's a straight line; if it is, shift resources from stopping the next trigger toward reducing coupling and accumulation. In a power-law world, blame carries essentially zero information.

04Where this stops being true

Now the other side. Self-organized criticality is among the most abused ideas in complexity science, and you should know its limits before you use it.

First, real sand doesn't really obey it. This is a little awkward — around 1990 people actually ran the experiment, building real piles and measuring avalanches, and found that beyond a certain size what appeared were periodic large collapses rather than a power law. The sandpile model is a fine model; it just isn't a model of sandpiles.

Later work in Chicago and Oslo gave a more honest answer. In the rice-pile experiment published in Nature in 1996, elongated grains produced a power law and near-spherical grains did not. What mattered was whether grains could interlock and pass stress along. That's a specific precondition, and not everything satisfies it.

Second, "looks like a power law" usually isn't one. On log axes the eye readily reads a gentle curve as a straight line, and log-normal or exponentially truncated distributions look much the same when plotted. In 2009 Clauset, Shalizi and Newman re-examined a set of widely cited power-law datasets with a uniform statistical test, and a substantial share failed. Looking straight is not evidence.

Third, and most important: none of this is an excuse. "Big events need no big cause" says the magnitude needs no matching cause. It does not say there was no cause, and it does not say nobody is responsible. Every toppling in the sandpile has an exact causal chain you can trace cell by cell. SOC claims one thing only: the length of those chains is distributed without a scale, so the identity of the chain's starting point does not explain the size of the outcome. Reading it as "so no one can be blamed" swaps a claim about statistics for a claim about responsibility.

🎯 DECISION LINE

Test before you conclude (Clauset's method has ready implementations), and state explicitly whether your system satisfies the three preconditions: slow accumulation, threshold release, local coupling. Missing any one of them, don't tell SOC stories about it — least of all to get someone off the hook.

🎒 Scenarios · BigCat

  1. TEAMS & ORGEvery incident review lands on "so-and-so pushed a change without review" — that is naming the grain of sand for an avalanche. Try the other view: pull this quarter's change frequency, review queue length, and rollback rate and look at the trend lines. If all three are climbing, it would have gone the same way whoever made that change. Concrete move: delete the "responsible party" field from the review template and add a line for "how this quarter's accumulation metrics moved versus last".
  2. INVESTINGReviewing a drawdown as "it was that piece of news" is the same grain-hunting. What is measurable is the build-up: rising correlation (diversification quietly failing), leverage levels, thinning liquidity depth — these are the snowpack. And the real decision isn't predicting the trigger; it's whether position size survives the tail. With no typical drawdown, you cannot set leverage from an average historical one.
  3. WRITING & THIS SITEDon't judge topic direction by average response per issue. Reach is almost certainly heavy-tailed — total readership is decided by a handful of issues, and the mean carries next to no information. So the goal isn't for every issue to clear a bar; it's to keep the tail possible: hold the range of topics open, don't imitate last issue's winning formula. A concrete test: if the last ten issues look increasingly alike, you have already cut off your own tail.

🌀 Crossings

Going deeper

If suppressing small events builds up big ones, when is suppression the right call?

It turns on whether the accumulation genuinely transfers. Forest fuel accumulates, so putting out small fires has a price; early cases of an infectious disease being contained do not get "stored" into a larger future epidemic, because the susceptible population isn't monotonically increasing the way fuel is. The test: whatever the event you prevented would have released — did it stay in the system, or did it actually go away?

Why might tighter financial regulation make crises larger?

One explanation is exactly SOC's: regulation removes small local failures, institutions therefore crowd into the permitted configuration, correlation rises and coupling tightens — the accumulation continues, only its release is postponed and synchronised. But this is easily abused into "so don't regulate", which it does not imply: an equally available reading is that the wrong thing was regulated (triggers rather than coupling).

Does "scale-free" mean disasters can't be defended against?

No. What it rules out is predicting the size of the next one, not lowering the overall level of risk. Changing coupling strength or inserting modular firebreaks can shift the entire power-law line downward or introduce an exponential cutoff — which is exactly what grid islanding and financial firewalls do. You give up point prediction, not intervention.

Can an organisation deliberately hold itself near criticality to stay adaptive?

Edge-of-chaos advocates say yes, and that it's the best place for innovation. Be careful: as it stands that is more metaphor than testable claim — you would first have to say what this organisation's "slope" is and how you'd measure it. Criticality without a stated unit is just an attractive word.

Further reading