Steady is not the same as sturdy
2026-08-25 · Fragility, resilience and early warning
One company posts eight straight quarters within 2% of plan. Another lurches up and down. You will almost certainly say the first one can take a hit — and for a whole class of systems that judgement is backwards.
Steadiness can come from two completely different things. Either the system really is tough and nothing you do knocks it apart, or it has squeezed out every scrap of spare capacity, locked every step to the next one, and given each person exactly one job — which does smooth out the day-to-day wobble, and in the same motion drops the largest shock it can survive to almost nothing. On a dashboard the two look identical.
In 1973 the Canadian ecologist C. S. Holling pointed out that we have been using one word for two things. How fast it snaps back and how far you can push it before it doesn't are separate quantities that can move in opposite directions. Managers, engineers and annual physicals habitually measure only the first.
Several issues here circle fragility, each asking something different. Topic 23 asked about structural asymmetry — why the same network shrugs off random failure and dies to a targeted one. Topic 25 asked how failure propagates. Topic 36 asked whether you can see the collapse coming. Topic 37 asked what the losses are distributed like. This one asks two things only: what actually buys a system the ability to take a hit, and why it goes brittle precisely during its best-looking stretch.
Separate the two quantities first. Each has a name, and running them together is the source of about half the confusion in this field.
Engineering resilience: how quickly the system returns to its normal state after a disturbance. The unit is speed. It assumes there is exactly one normal state and that a shock merely pushes you off it for a while.
Ecological resilience: how large a disturbance it can absorb before it is pushed into a different state and can no longer come back. The unit is magnitude. It assumes there is more than one stable state, and that a shock might carry you across.
A bowl and a marble makes it obvious. The marble rests at the bottom; the shape of the bowl decides everything → ref · phase space & attractors. A deep, narrow bowl: nudge the marble and it whips straight back — very fast recovery — but push a little harder and it goes over the rim and is gone. A shallow, wide bowl: the marble takes ages to settle — slow recovery — but getting it out of that bowl takes far more work.
The bowl has a technical name: the basin of attraction, meaning every starting point from which the system returns to that state. Recovery speed depends on how steep the walls are; the shock you can absorb depends on how wide the bowl is. Nothing whatsoever requires those two to improve together.
This also patches last issue's early-warning signal. The critical slowing down of Topic 36 — recovery getting sluggish after a knock — measures exactly how steep the walls are, which is to say engineering resilience. As a warning it works: flattening walls mean a shallowing bowl. The converse does not hold. Fast recovery tells you the walls are steep and says nothing at all about how wide the bowl is. A deep narrow bowl gives you a beautiful check-up every single time.
Stop using "time to return to normal" as your resilience metric. Ask two numbers instead: how hard a push does it take before it doesn't come back, and when did you last actually push that hard. The second question usually has no answer — in which case the stability you are holding is not data, it is simply an untested system.
Holling later worked with Lance Gunderson to push the observation one step further: a system doesn't just sit in some bowl. The bowl itself deforms on a cycle. They called it the adaptive cycle.
Two variables draw it. Potential: how much usable stuff the system has accumulated — biomass, capital, skill, working code, earned trust. Connectedness: how tightly the parts are locked to one another — how many steps assume the other steps won't change. A third quantity hides off the axes: resilience.
Four phases. The letters are an old habit borrowed from ecology:
Exploitation (written r): open ground. Plenty of room, cheap to claim a spot, everyone grows their own thing without much reference to anyone else. Low potential, low connectedness, high resilience — collapse costs little because there is little to lose.
Conservation (K): the ground is taken and the competition is now about efficiency. So things specialise, spare capacity gets trimmed, interfaces get standardised, inventory drops toward zero. Potential climbs to its maximum and connectedness climbs with it — and resilience bottoms out right here.
Release (Ω): a not-especially-large disturbance takes the locked structure apart in one go. It is fast, usually one to two orders of magnitude shorter than the phases before it.
Reorganization (α): the resources are lying loose, constraints are at their weakest, uncertainty is at its highest. This is when new things can most easily get in — and when outsiders can most easily take the ground.
(r and K come from ecology's r-selection and K-selection, the strategy for claiming open ground and the strategy for competing in a full one. Ω and α are Holling's own: the last and first letters of the Greek alphabet, meaning end and beginning.)
Here is the sentence that matters most this issue: efficiency and resilience are not two goals you can optimise separately; they are two ends of the same dial. The moves that raise efficiency — cut the standby capacity, let everyone do only what they are best at, freeze the interfaces, drive inventory to zero — are individually unimpeachable, and together they add up to raising connectedness. So K is not the result of bad management. It is the result of getting every step right.
Usefully, none of these three quantities is vague. Connectedness is countable: how many dependencies point at each unit, how many single points the system contains. Potential is countable: the stock accumulated. And there is a more sensitive one still: what fraction of your resources is uncommitted — the people, budget and time not yet spoken for, usually called slack. K has a recognisable signature: output still rising while slack falls and dependency counts climb.
Give "which phase are we in" two numbers you can actually pull: fraction of uncommitted resources and count of single-point dependencies. First falling, second rising, output still hitting records — that is K. It is the window for arranging a release on your own terms, not the window for adding leverage. Put both numbers next to the output metrics; on their own nobody looks at them.
Release looks like pure loss, but it does one irreplaceable job in the cycle: it turns the capital that conservation had bound up back into something usable. A forest fire returns decades of nutrients locked in wood to the soil in a single afternoon — not a metaphor, a mass flow you can weigh.
What actually decides the outcome is reorganization. Constraints are weakest and first-comers take the ground, so everything turns on one thing: memory — where the seeds for rebuilding come from. The seed bank in the soil, the trees left standing, the mycorrhizal network that didn't burn, decide whether this patch grows back into forest or into scrub.
Holling later gave the whole arrangement a bigger name: panarchy. Its claim is that a system is not one cycle but a stack of them, fast and slow nested together. Needles turn over in weeks, canopy gaps in years, crowns in decades, forest and soil in centuries. Every layer runs its own adaptive cycle, at tempos that differ by orders of magnitude.
Between the layers there are exactly two channels, and Holling named them too:
Revolt: a small fast layer entering release ignites the large slow one above it. The crucial part — this channel is only open while the layer above is sitting in K. A match in a wet forest is nothing at all; with the layer above fully loaded and fully connected, the same match is a crown fire.
Remember: the large slow layer feeds seeds, soil, norms and old hands downward when the layer below enters reorganization. It decides whether what grows back is the thing that was there before or something else.
Put the two channels side by side and you get a conclusion that is not visible from either alone: how bad a collapse gets depends on which phase the layer above is in; what grows back depends on whether this layer can still reach the memory above it. Neither answer lives at the scale where the incident happened. Which means a post-mortem confined to the layer that broke cannot answer either question — and confining it to that layer is exactly what post-mortems do by default.
Schedule releases that are cheap, regular and unannounced: rotate people, kill instances at random, actually pull one dependency each quarter and run without it. Note what this is for — it is not "validating the runbook", since an announced drill can only validate the runbook. It is for clearing the hidden coupling that quietly grew during K, the coupling nobody ever wrote down, which only appears when something is pulled. And do one more thing separately: write down where this system's memory lives, in whom, and check that it does not die together with the part you are drilling.
Four things genuinely buy it. Each blocks a different failure and each carries a specific bill — and confusing which one blocks what is how you spend the money and still go down.
Redundancy: several units can do the same job. It blocks single-point failure. Bill: paying for capacity that produces nothing.
Modularity: cut the system into blocks that are tightly coupled inside and loosely coupled between. It blocks propagation — one block going down doesn't drag its neighbours with it. Bill: anything crossing a boundary gets slower and dearer (Topic 3's near-decomposability is the other face of this).
Response diversity: the same job carried by units that react differently. It blocks common-cause failure — precisely the class redundancy cannot touch. It is the first of the four to get cut, because under normal conditions it looks like waste: both options work, so why keep the awkward one.
Loose coupling: leave time, buffer and interruptibility between units. It buys speed. Charles Perrow's normal accident theory puts it bluntly: what tight coupling really takes away is not fault tolerance but the time you needed to notice you were wrong. Bill: inventory, delay, apparent inefficiency.
The crux is the difference between redundancy and response diversity. A redundancy count degrades in the face of a common cause. You have three backups — from one supplier, in one building, reading one config file, resting on one person's judgement. Against a common-cause shock the number three is a one. Multiplying three identical components as independent events in an availability calculation is a systematic overestimate: independence was assumed in, and it is almost never measured.
When auditing redundancy, don't count copies — count common causes. Five questions: do these copies share a supplier? a building or a power feed? a config? one person's judgement? one assumption? Any yes and the effective multiplicity gets discounted, and the discounted number goes in the doc. Turning that list into a template is far cheaper than adding another copy.
Now the other side. "Resilience" currently appears in ecology papers, annual reports and government white papers at the same time, which by itself should make you cautious.
First, unqualified, the word is empty. The title of Carpenter and colleagues' 2001 paper is simply a question: resilience of what, to what? Leave those two blanks unfilled and any system can be called resilient or not, so the claim excludes nothing — and a description that excludes nothing carries no information. Most "improve resilience" in policy documents stops right there.
Second, the adaptive cycle is a heuristic, not a model. Holling and Gunderson use that word themselves. The boundaries between the four phases are extremely hard to draw on real data, and after the fact any stretch of history can be fitted into four phases — the textbook symptom of unfalsifiability. Worse, plenty of systems never complete the cycle at all: stuck in K is a rigidity trap, stuck in α and unable to grow is a poverty trap. Treating the cycle as seasons that must come round keeps people waiting for a spring that isn't scheduled.
Third, resilience is value-neutral. Dictatorships, drug-resistant bacteria, and a process everyone hates but nobody can change are all extremely resilient. "Increase resilience" as a goal contains no good or bad — that comes entirely from what you put in those two blanks.
Fourth, the trade-off is denominated in real money, not in "we need balance". Redundancy, loose coupling and response diversity all have prices: capacity carried, inventory held, time spent. The honest form is a premium — to survive this class of shock, I am willing to pay this much per year. If you can't name that figure, designing for uncertainty hasn't started.
Before anyone says "improve resilience", finish the sentence: keep ____ functioning / against shocks of type ____ / tolerating losses of ____ / at a cost of ____ per year. If all four blanks can't be filled, the project either hasn't been thought through or is really doing something else. Those four blanks double as the acceptance criteria — without them, nobody can tell afterwards whether the money went to the right place.
When the system really does have exactly one stable state and the shocks are known to be bounded, engineering resilience is resilience and recovery speed is the right metric — plenty of mechanical and control systems are like this. So the test isn't "stability is bad", it is a prior question: on what grounds are you sure there is only one bowl? In ecosystems, organisations and minds that question usually has no good answer. On a servo motor it usually does.
Because every individual step is locally rational. Removing one piece of spare capacity or adding one dependency yields a benefit that is immediate, visible and attributable, and can go in this quarter's report; the cost is diffuse, delayed, and lands in somebody else's tenure. It is a process where each step is right and the sum is wrong, and it cannot be stopped by reminding everyone to value resilience. Stopping it requires a role explicitly accountable for slack — whose output always looks like zero, which is why it is the first thing optimised away.
Check two things: whether the memory survived, and whether the constraints genuinely loosened. A great many "reorganizations" are the same people rebuilding the same thing under the same rules — that is not α, it is a tremor inside K, and in a few months everything is as before. A crude but serviceable test: after this collapse, did any person or practice get in that could not have got in before? If not, nothing loosened.
Mostly you can't — this is panarchy at its most honest and least usable. What you can do is find proxies for the layer above: sector-wide leverage, supply-chain concentration, how tight the regulator is, how much slack your peers are carrying. All indirect, all lagging. But knowing that part of your safety is not yours to manage does change position sizes and the scale of what you commit to, and that is most of what the framework has to offer.