TOPIC 36 · PHASE G

Critical Slowing Down and Early Warning Signals

It gets slow before it breaks

2026-08-22 · Fragility, Resilience & Early Warning

A system about to fail looks exactly the way it always looks. The quantity that is actually changing is how long it takes to come back after a nudge — and almost nobody records that one.

Think of a door with a spring closer. When it was new you pushed it open, let go, and it snapped shut. A few years on, the same push, the same release, and it drifts — swinging, hesitating, easing slowly closed. Note that the door's position has not changed: left alone it sits shut, exactly as it did when new. Only the speed of return has changed. And one day it stops closing at all.

This is the fifth time this site has taken up criticality, but the first four were about how a system flips. This one asks something else: can you see the flip coming? The counterintuitive part is that if there is a signal, it is usually not in the number you are watching. Water clarity, share price, mood, error rate — these tend to be suspiciously calm right up to the transition. The signal hides in the way those numbers wobble, and wobble is the first thing a monthly average destroys.

The five criticality topics look at the same cliff from different angles: Topic 10 is bifurcation and tipping (the system jumps to another steady state), Topic 16 is phase transitions and universality (why details stop mattering at the critical point), Topic 17 is percolation (connectivity going through in one step), Topic 18 is self-organized criticality (the system climbs to the critical point by itself). This one asks only two things: is there a measurable precursor, and when is there none at all.

01Draw Stability First: a Bowl and a Ball

Saying a system is "stable" means something quite specific: nudge it and it comes back by itself. Draw that and you get a ball sitting in a bowl. The nudge pushes the ball up the wall; coming back is the ball rolling to the bottom.

The shape of the bowl sets how fast it returns. Steeper walls mean the same push lifts the ball less and it rolls back quickly; a flatter bowl means the same push sends it further and it takes ages to settle. Physics calls that speed the recovery rate; its inverse is the recovery time — how long it takes to get home. → ref · Phase space & attractors

Now connect Topic 10. When external conditions creep on — phosphorus accumulating in a lake, a metal part flexed over and over, someone short of sleep for months — the system inches toward a bifurcation point: the point past which it jumps to another state and does not come back. Mathematically, approaching that point is the flattening of the bowl. At the bifurcation the bottom is completely flat and the recovery rate has fallen to zero.

The phenomenon has a name: critical slowing down. Its prediction is admirably concrete: the closer to the tipping point, the slower the recovery from a disturbance.

The flatter the bowl, the slower the ball returns — while its position barely moves recovery rate: high nudge it, it snaps back lower nudge it, it lingers → 0 nudge it, it barely returns external conditions creep on (nutrients, heat, leverage, fatigue) → the bowl flattens deviation from baseline time after the nudge → fast return slower return barely returns one nudge, identical every time
Three bowls, three depths. The ball has barely moved — if all you record is where the ball is, the three panels look much the same. The whole difference lives in the return curve.

This is not just blackboard work. In 2012 a group of Dutch researchers drove a population of cyanobacteria — bacteria that live by photosynthesis — toward a real tipping point in the lab: above a critical light level, photo-inhibition collapses the whole population, and it does not come back. Before the collapse they repeatedly applied identical small disturbances and timed the return. The closer to that critical light level, the slower the recovery. It remains one of the cleanest demonstrations that slowing down really happens in a living system.

🎯 DECISION LINE

For every system you actually care about, define a bounce-back measure: after a small disturbance of comparable size, how long does it take to return to baseline? Then track the trend of that time rather than the state value itself. The test is not "how long this time" but "is the same disturbance taking longer this month than last".

🌀 Medicine · how fast your pulse drops after a run A 1999 study in the New England Journal of Medicine followed 2,428 adults: those whose heart rate fell by 12 beats per minute or less in the first minute after exercise had a six-year mortality of 19%, against 5% for those who recovered normally. Note what was measured — not the heart rate itself, which can be perfectly normal at rest and at peak. It was the same mechanism as above: apply a standard disturbance, time the return. A check-up form is dense with state values; the hardest predictor on it is a slope.

02You Don't Have to Push It — the Noise Already Does

Section 1 leaves a practical problem: you cannot nudge a lake every week, nor take a company and shove it to see how fast it steadies. Fortunately you don't have to. Wind, rain, temperature, random buy and sell orders, the small mess of daily life — something is pushing all the time. The system is never actually at rest; it is constantly being knocked off and sliding back.

So the slowdown writes itself into that wobble, in two quantities you can compute directly.

First, consecutive moments start to resemble each other. If the bounce-back is fast, the deviation created now is gone by the next reading, so consecutive readings have almost nothing to do with each other. If the bounce-back is slow, most of this moment's deviation survives into the next — and the readings start to look alike. There is a standard way to measure that: lag-1 autocorrelation (pair every reading with the one after it and see how tidily those pairs line up). Slowing down drives it toward 1.

Second, the swings get wider. Deviations that decay slowly stack on top of each other, so the same size of push produces a bigger excursion. How wide the swings are is the variance — the average squared distance from the mean; bigger means more spread out.

These are two faces of one thing, and they fit in a single line of near-trivial arithmetic. Write the system's behaviour near baseline as "next moment = α × this moment + a bit of random push". That α is the surviving fraction: α near 0 means the system forgets instantly (fast recovery), α near 1 means it remembers almost everything (slow recovery). Statisticians call this line a first-order autoregression, written AR(1); its continuous-time twin is the Ornstein-Uhlenbeck process. α is a direct translation of the recovery rate, and the width of the swings scales as 1 ⁄ √(1−α²) — take α from 0.2 to 0.95 and the same push produces swings three times as wide. → ref · AR(1) & the OU process

Nobody nudges it — ambient noise does, and the slowdown is written into the noise observed state variable external conditions creep on (time →) rolling variance time → lag-1 autocorrelation time → Both indicators rise — both are noisy and both lag: a longer window is smoother and later
The top line is just a series of readings shoved around by random pushes — the pushes never change size. The only thing changing is how fast it forgets. The two indicators below are computed from that line alone.

The approach has been tested for real, on a whole lake. Over three years, researchers gradually stocked largemouth bass into Peter Lake in Wisconsin, deliberately pushing its food web toward a flip, while monitoring a neighbouring untouched lake as a control. The result: variance and autocorrelation in the manipulated lake began rising more than a year before the food web actually shifted.

There is a striking human counterpart. A 2014 study using high-frequency mood diaries — several prompts a day — reported that autocorrelation and variance in emotions rise before both the onset and the lifting of depression. That result drew a published objection at the time: critics argued the data showed a difference between groups rather than individuals demonstrably approaching a tipping point. It is best treated as a direction worth pursuing, not a settled finding.

🎯 DECISION LINE

If you are going to monitor something, keep the raw high-frequency data; don't store only aggregates. Variance and autocorrelation are both quantities about adjacent readings, and taking a monthly mean destroys them by definition — ten years of monthly reports contain none of this signal at all. There is a computable floor for sampling: sample several times faster than the system's recovery time. If recovery takes three days, monthly is useless.

🌀 Economics & institutions · reporting frequency decides which risks are visible Regulatory filings, board packs, quarterly reviews: nearly all of them carry means and period-end snapshots. But the two indicators above are high-frequency quantities — aggregate them to a month or a quarter and they cease to exist in the data. Which yields an uncomfortable conclusion: the risks an organisation can see are set by its reporting frequency, not by its risk appetite. To see a new class of risk, change the sampling first and the attitude second; the other order does nothing.

03Rather Than Wait for the Wobble, Tap It

Reading the recovery rate out of ambient noise is free, and also the dirtiest option: you need a long stretch of series and you have to hope nothing else changed during it. There is a faster, cleaner way — disturb it yourself and start the clock.

You have seen this done. A cardiac stress test is exactly a deliberate perturbation: the doctor does not wait for the day you happen to go running, they put you on a treadmill, apply a controlled and comparable push, and watch how fast you come back. Engineering calls the same idea a tap test — strike the structure, see how quickly the vibration dies away. Software calls it chaos engineering or a game day — kill a machine on purpose and time how long the system takes to return to normal.

The load-bearing word is comparable. The size, location and timing of each perturbation must be as alike as you can make them, or what you measure is the difference between the pushes rather than the difference in the system.

Tap it three times: same push, longer and longer return baseline (normal state) halfway back probe ① halfway back in 1× probe ② 2.4× probe ③ 6.4× It still returns to the same baseline every time — the state stays “normal”, only the speed changes
Three identical taps; the return takes 1, then 2.4, then 6.4. And every time it does get all the way back to the baseline — if all you ask is "did it return to normal", the answer is yes three times out of three.

One class of system gets a bonus for free: systems spread out in space. Dryland vegetation, coral reefs, forests — a single satellite image holds thousands of readings. So instead of waiting for time to pass you can compute variance and neighbour-similarity across space: patches getting coarser and neighbouring cells becoming more alike is the same phenomenon as rising variance and autocorrelation in time, cut a different way. One snapshot can occasionally stand in for years of waiting.

🎯 DECISION LINE

Change the output of a drill from "pass / fail" to a number: recovery time. And require the perturbation to be comparable (same injection, same load band), or the number cannot be compared across runs. The value of a drill is not proving you can survive it — it is giving you a curve with a readable trend. Surviving but taking twice as long as last quarter is the single most important thing the drill produced.

🌀 Engineering & the history of technology · tap a bridge and you often measure the weather Structural health monitoring uses the decay of vibration after a tap to judge damage. In practice a nuisance turned up: stiffness drift caused by temperature is frequently larger than the change caused by real damage — you think you are measuring a crack and you are measuring how cold last night was. The lesson transfers directly: any active-probe method must record its environmental covariates, or it measures the environment instead of the system. Applied to drills: if you don't log traffic, staffing and whether a release just went out, your recovery-time curve is not comparable and its trend is fiction.

04The Arithmetic of an Alarm

Suppose you did all of the above right: high-frequency data, comparable probes, indicators climbing. Turning that into an alarm that actually rings runs into a very plain piece of arithmetic.

First, three easy traps.

Trap one: the wrong variable. Slowing down shows up only in quantities coupled to the direction that is about to lose stability. A system has dozens of metrics and most of them are unrelated; watching those, you will wait forever. Choosing the variable requires a guess about the mechanism — so this method does not free you from understanding the system, only from having an exact model of it.

Trap two: window length and detrending are free parameters. Computing variance means choosing how long the sliding window is, and usually removing a slow trend first — and how long, and how removed, are up to you. That is enough freedom that trying a few settings will often produce whichever conclusion you were hoping for. The only honest procedure is to fix and write down the parameters before you look at the result.

Trap three: the test should be a trend, not a threshold. "Alarm if autocorrelation exceeds 0.8" is meaningless — baseline values differ wildly between systems. What is meaningful is whether the indicator is persistently rising over a stretch (usually quantified with a rank correlation test).

And then the arithmetic. Transitions are rare events, and rare events dilute even a good detector:

A rather good early-warning detector, doing arithmetic on a rare event 200 monitoring windows · white dot = this window raised an alarm nothing happened false alarm a real transition alarm raised sensitivity 75% → 3 of the 4 real ones false alarms 10% → 20 of 196 calm windows so 23 alarms in total real alarms = 3 / 23 ≈ 13% No amount of sensitivity changes that ratio What you can change is the cost an alarm triggers
This does not say early warning is useless. It says that when the alarm rings, what you hold is a message that the probability went from 2% to 13% — not a verdict.

So the only actions an early warning signal can drive are the cheap, reversible ones: raise the monitoring frequency, thicken a buffer, take leverage down a notch, postpone an irreversible decision by a few weeks. It cannot drive "shut the plant down" — not because the signal is bad, but because an 87% false-alarm rate multiplied by the cost of a shutdown is unpayable, and after the second one somebody quietly switches the detector off.

🎯 DECISION LINE

Before building a detector, do the multiplication: expected base rate × sensitivity (how many real ones you catch) against false-positive rate × the cost of each false alarm. If the product doesn't work, the fix is not a more sensitive detector — it is the other road: raise reversibility, so that a transition, when it comes, costs little enough that you didn't need to know in advance. Warning and reversibility are substitutes; on a tight budget the second is usually the sounder buy.

🌀 History & the military · the signals and the noise at Pearl Harbor Roberta Wohlstetter's 1962 study of Pearl Harbor established a conclusion that has never really been overturned: the warning signals were all there beforehand, drowned in an equal quantity of signals pointing elsewhere; they are obvious afterwards only because we know the answer. That is the arithmetic above, in history — the bottleneck was never sensitivity, it was the base rate and the cost of responding. The unobvious consequence: adding sensors to an intelligence system raises false alarms in proportion. What can actually be changed is making it cheap to act on a warning.

05Where This Breaks Down

Now the other side, and this section matters more than the four before it. Early warning signals — EWS for short — are among the most oversold ideas in complexity science of the past fifteen years. There is something real in them, and there are a great many ways for them to fail.

First, only one kind of transition gives you a window. The entire argument rests on a premise: that the system is being carried toward a bifurcation gradually by slowly creeping external conditions. At least two kinds of transition violate it. One is pushed over by noise: the bowl never got shallower, one shove was simply large enough to kick the ball over the ridge. The other is conditions changing too fast: the bowl really is flattening, but no slower than the transition itself, so the slowdown never accumulates into a measurable signal.

All three “collapse suddenly” — only the first one gives you a window ① slow fold approach fluctuations keep growing the other state warning window exists variance and autocorrelation rise early ② noise kicks it over nothing odd before the jump the other state no warning window the bowl never flattened; the kick was huge ③ too fast to warn slowdown had no time the other state window too short to use slowing down never built into a signal
All three end identically. Only the first leaves a readable trace beforehand.

This is not a theoretical worry. The Dansgaard-Oeschger events of the last ice age — abrupt warmings of a dozen degrees within decades — were checked for slowing down and it is not there; the best reading is that they were noise-triggered, hence unforecastable in principle. Ecology has its own version: a 2010 analysis showed that a whole family of perfectly standard ecological models produces no leading indicators before a regime shift.

Second, a signal is not proof that a tipping point is ahead. Rising variance and autocorrelation establish that recovery has slowed — that resilience has fallen. But falling resilience need not mean there is a cliff in front of you; the system may simply be degrading smoothly toward a worse but continuous state. A 2025 review in climate science put it bluntly: reading rising EWS as "a tipping point lies ahead" packs into the signal something the signal does not contain.

Third, the statistical power is far worse than people assume. Reliably detecting a rising trend inside noise takes a long, high-frequency series over a period in which nothing else moved — real data routinely satisfies none of those. A 2012 study devoted to this question concluded that at realistic data volumes the reliability of these tests is too low to carry a decision.

Fourth, retrospective success is not success. This affliction is universal in forecasting and EWS did not escape it. Seizure prediction has been studied for over thirty years, retrospective "successes" have been abundant, and prospective blinded validation stayed out of reach for decades — only recently have long-term implanted recordings produced limited genuine forecasting. The well-known 2023 paper warning of a collapse of the Atlantic meridional overturning circulation (AMOC) supplied a fresh example: in 2025 its authors issued a correction acknowledging coding errors in the estimation procedure. None of this voids the research programme, but it is a reminder: between a model that "would have warned of" a past collapse and a usable alarm lies an entire regime of prospective validation.

Fifth, a monitored system reacts. This is the Topic 26 mechanism showing up here: once variance and autocorrelation become targets, the cheapest way to hit them is not to become more resilient but to flatten the indicator — lengthen the reporting interval, smooth what gets submitted, absorb small fluctuations internally. Same system, better numbers.

Sixth, self-organized critical systems have no "distance to the tipping point" at all. The systems of Topic 18 — sandpiles, earthquakes, faults — are not approaching a critical point; they sit on one, throwing off events of every size continuously. Asking how far such a system is from its transition is the wrong question, because there is no bifurcation being approached. → ref · The sandpile model

🌀 Earth science · Haicheng and Tangshan The February 1975 Haicheng earthquake in Liaoning was preceded by a dense foreshock sequence; the local authorities evacuated on that basis, and it is still cited as one of the few successful earthquake predictions on record. The Tangshan earthquake eighteen months later had no recognisable foreshocks and no warning. Under this section's taxonomy the two are not in conflict: Haicheng was read off a specific precursor that happened to occur, not off critical slowing down — the crust belongs to the family that sits permanently near criticality, with no approaching bifurcation to monitor. The consequence generalises to every claimed forecasting success: a hit that cannot say which precursor it read is not a transferable method, only a case where the foreshocks happened to be there.
🎯 DECISION LINE

When reporting any early warning indicator, report three things with it: the failure conditions (what makes you think this system is the slow-bifurcation kind rather than the noise-triggered kind), the base rate (how often this class of transition has occurred in your history), and when the window and detrending parameters were fixed (before the result, or after). Missing any of the three, the signal should not enter a decision. And the safer stance throughout: treat EWS as an input to vigilance, and reversibility as the actual defence.

🎒 In Practice · BigCat

  1. engineering & system designAfter every release, error rate or latency twitches and then settles. What people watch is whether it crossed a threshold; if it settled, it passed. Watch how long the settling took instead: for releases of comparable size, has the return-to-baseline time gone from 4 minutes to 12 over a quarter? It is far more sensitive than incident counts, because it is already climbing while nothing has broken yet. One concrete change: post-release monitoring records not just the peak but the timestamp of the return to baseline, and both go on the same trend chart.
  2. parentingAfter a meltdown the usual post-mortem asks what set it off. Measure something else: how long from the meltdown to being able to talk normally again. The trigger is different every time and carries almost no information; the trend in recovery time is readable. If comparable small setbacks are taking steadily longer to come back from over a few weeks, something underneath is draining the reserve — sleep, school pressure, a particular relationship — and that is what to address, not today's incident. One concrete change: stop counting meltdowns, start timing recoveries.
  3. investing & position sizingDrawdown reviews are almost always about depth — how far it fell and why. Move the attention to time: after shocks of comparable size, has the number of days for the portfolio to regain its previous high been lengthening over the past two years? Lengthening usually means one of two concrete things: correlations among your holdings are rising (diversification failing), or what you hold is getting less liquid. Both are directly checkable. One concrete change: put "days to recovery" alongside "maximum drawdown" as a standing metric, and only move leverage when the former lengthens materially — not on every decline.

🌀 Crossings

Going Deeper

If slowing down only appears on a slow approach to a bifurcation, how would you know in advance that this is what you face?

Strictly, you can't know for certain. Two usable approximations: has the system ever shown hysteresis (the same parameter value supporting two different states, which implies bistability and therefore a bifurcation)? And is the driver changing far more slowly than the system's own recovery time? When both hold, the premise of EWS roughly stands; when one fails, downweight it. Note that both tests rely on history — and a kind of transition that has never occurred in your history is invisible to them.

Once a warning indicator is published, the monitored party optimises it. Should warnings be published at all?

This is Goodhart's law from Topic 26, and there is no clean solution. One partial answer is two layers: indicators used to raise the alarm must be public (or nobody responds), while indicators used to audit the detector itself should be held back and sampled. Another is to prefer indicators where flattening them requires actually improving — recovery time, for instance, is hard to shorten without genuinely adding slack. Unfortunately there are not many like that.

Could you run it in reverse — deliberately hold a system near criticality to buy sharper responsiveness?

Some people argue exactly that ("the edge of chaos"), on the grounds that response to small inputs is maximal near the critical point. But the cost is symmetric: large response also means poor rejection of disturbance, and slowing down means recovery takes longer — the sensitivity is bought with recoverability. Doing it deliberately requires continuously measuring how far from criticality you are, and Section 5 explains why that is often impossible.

Why is "it came back to baseline" so reassuring and so nearly information-free?

Because it is a binary judgement while the mechanism emits a continuous signal. Before the transition, the system returns to baseline every single time — until the last time, when it doesn't. A binary test outputs "normal" across the entire window in which warning was possible, and its first alarm coincides with the event. Every monitor that binarises a continuous quantity has this problem, not just this one.

Further Reading