TOPIC 11 · PHASE B

Two Limits of Predictability

One wall is measurement; the other is that no shortcut exists

2026-07-29 · Dynamics & Unpredictability

You probably assume forecasts miss because the instruments aren't good enough and the data isn't plentiful enough. Half the time that's right. The other half is something else entirely — some things cannot be computed, not because you know too little, but because there is no way to find out short of letting them happen at full length.

Tomorrow's weather comes out pretty well. Next weekend's is shakier. Two weeks out it is barely better than guessing. That boundary is not a boundary of technique — make the supercomputer a hundred times faster and it barely moves. It is a wall.

The counterintuitive part is that there is more than one such wall, and the two look nothing alike. The first is you can't measure well enough: the system is exquisitely sensitive to its initial state, your measurement error is amplified exponentially, and the forecast horizon is capped by how fast that error grows. Money prises this one open a crack. The second wall is a different animal. Suppose you knew the initial state exactly, and the rules were exact with no fuzziness anywhere. There may still be no way of learning the outcome faster than running the thing one step at a time. That wall does not open — it is not about how much you know, it is about whether a shortcut exists at all.

Three issues here deal with unpredictability, and they divide the work. Topic 8 is chaos itself (deterministic does not mean predictable); Topic 9 is how one parameter walks a system all the way into chaos. This issue asks one thing: given that you can't predict it, which kind of "can't" is this, and what is still available in each case. The two limits have to be read against each other, because their remedies are nearly opposite — misdiagnose the wall and every dollar goes to the wrong place.

01Error has a doubling time

In 1961 the meteorologist Edward Lorenz wanted to continue a forecast run he had already computed, so he typed an intermediate state back in off the printout → ref · Lorenz system. The printout carried three decimal places; the machine had been holding six. The number he typed differed from the original by less than one part in a thousand — and the weather that came out was an entirely different weather.

What matters is not the slogan "a tiny difference makes a big difference", but how much, and how fast. In systems like this, two trajectories starting almost on top of each other pull apart exponentially: the gap doubles every fixed interval. That interval is the error doubling time; the growth rate behind it is the Lyapunov exponent, and the time it takes for an error to grow from negligible to unignorable is the Lyapunov time.

Exponential growth has a brutal consequence that only shows up when you run it backwards: to buy one more doubling time of forecast horizon, you have to halve your initial error. To buy ten, you have to cut it to a thousandth (2 to the tenth is about 1000). Investment in precision is multiplicative; the horizon it returns is additive. That is why weather forecasting advances at roughly a day per decade rather than doubling whenever compute doubles.

In modern operational forecasting the doubling time for synoptic-scale error is about a day and a half. In 2019 a group (Zhang and colleagues) used 9 km and 3 km models to estimate the limit directly: today's skilful lead time is around 10 days, and cutting current initial-condition error by a full order of magnitude buys at most 5 more days — after which you hit a line at roughly two weeks. That line is the ceiling assuming a perfect model, not the current state of the art. Weather running out at two weeks is not meteorologists slacking.

Error grows exponentially — the vertical axis is logarithmic 0 7 14 21 forecast lead time (days) → gap between two trajectories (log) saturation: no better than guessing intrinsic limit ≈ 2 weeks 10× finer won't cross it ≈10 days buys only 4–5 days today's observation error initial error cut to 1/10 Precision is multiplicative, horizon is additive — and at some point there is a line you cannot cross
The vertical axis is logarithmic, so exponential growth draws as a straight line. Tenfold better measurement shifts the whole line down a notch and right by four or five days.
🎯 DECISION

Set your replanning cycle by the error doubling time, not by the calendar. How: pull your last 20 predictions (delivery dates, sales, capacity, recovery times), group them by how far ahead each was made, and find the lead time at which the miss goes out of control — that is the doubling time for this thing, and your review interval should be half of it. The action to stop: turning the estimation meeting into two hours in pursuit of getting it right the first time.

🌀 Engineering & the history of technology · midcourse correction A spacecraft is never launched on a trajectory computed once and then left alone; correction manoeuvres are scheduled into the flight in advance. The reason is exactly this section's mechanism: orbit-determination error grows as the flight proceeds, so the later you correct, the larger the deviation to cancel and the more propellant it burns. The real design variable is therefore not how precisely you determined the initial orbit but how densely the correction windows are spaced. The inference transfers to any plan: when error grows exponentially, shortening the review interval is far cheaper than improving the initial estimate — and human instinct runs the other way, always reaching for "let's think it through properly this time".

02Don't predict the trajectory, predict the distribution

Having hit the first wall, meteorologists did not give up. They changed the deliverable — and that move is worth stealing by anyone who forecasts anything.

Since the 1990s the major centres no longer run the model once. They perturb the initial state a few dozen times — each perturbation within what observational error allows, so every perturbed version is equally "entitled" to be the true current state — and run the forecasts in parallel. The European Centre for Medium-Range Weather Forecasts (ECMWF) runs 51: one control forecast plus 50 perturbed members. The technique is called ensemble forecasting.

What it delivers is not a number but a crowd of numbers. And how widely that crowd spreads is itself the most important product: if all 51 say rain on Saturday, you can say rain with real confidence; if 30 say rain and 21 say no, the correct output is "about a 60% chance", not a forced pick. The "70% chance of rain tomorrow" on your phone comes from precisely this.

The same 12 members — the spread itself is the product ① members bunched ② members spread out rain threshold rain threshold 12/12 above the line → just say "it will rain" small circle = initial states within observation error 7/12 above the line → "about a 60% chance" picking one member throws the information away
Both panels start equally uncertain and end up completely different in how much you may say. A tight ensemble earns you a statement; a wide one earns you a probability.

Note that this is not uncertainty repackaged as a disclaimer. Probability forecasts are testable: take every day in the year on which you said 70%, and it should have rained on close to seven-tenths of them. This is calibration, and it is a scorable quantity. Point forecasts have no such property — when a single "rain tomorrow" turns out wrong, you cannot separate a bad model from bad luck, so nobody has to answer for anything.

🎯 DECISION

Change the deliverable from one number to an interval plus a probability, and then score the interval. Concretely: attach an 80% interval to every number you report; each quarter, count what fraction of outcomes landed inside. The target is 80% — far more than that means you widened the interval to protect yourself, far less means you are fooling yourself. What gets held to account shifts from "did you call it" to "how well calibrated are your intervals". This step needs no new model, only a new table.

🌀 Economics & institutions · the central bank's fan chart From the mid-1990s the Bank of England's inflation reports began replacing the single projected number with a fan chart: a widening shaded region marked out in probability bands. It looks like hedging, and it is the opposite — it turns calibration into something you can be held to. If the 90% band only contained six-tenths of actual outcomes, that can be computed and put to the committee. The inference is mildly counterintuitive: publishing an interval is easier to falsify than publishing a point, and therefore more accountable, not more slippery. An institution that reports one target number has placed itself where it cannot be checked.

03Same system, some questions do have answers

The second thing matters more and gets dropped most often: predictability is not a property of the system, it is a property of the question.

The temperature in Beijing on 17 March next year: unanswerable. The mean temperature in Beijing next March, and roughly how many days will top 25 °C: answerable, and rather well. Same atmosphere, same equations, same pile of measurement error. The only difference is that the first asks about one point on the trajectory and the second asks where that trajectory tends to sit, and how often.

That is exactly the difference between weather and climate. Climate is not "weather much later"; climate is the statistics of weather. In the language of Topic 7: the trajectory wanders on an attractor, and which point it wanders to is unanswerable, but the shape of the attractor is stable — stable enough that if you recount over a different stretch of time, you get very nearly the same distribution. → ref · phase space & attractors

One simulation — the left question has no answer, the right one is solid enough to act on trajectory: where next? distribution: how often, where thousands of steps, always inside these two wings — but left wing or right wing next? no answer teal bars = an early stretch (300k steps) gold line = a much later stretch (300k steps) recount over a different stretch and the shape nearly coincides — that is climate versus weather
Chaos destroys long-range prediction of the trajectory. It does not destroy prediction of the statistics — and the statistics are what most real decisions actually ask for.

So "this system is chaotic, it can't be predicted" is an incomplete sentence, and often a lazy one. How high to build the levee, how much redundancy to hold, how large a position to take — every one of those asks for a statistic. Changing the question is far cheaper than changing the tooling, and it is often the only move that works.

🎯 DECISION

When someone says "that can't be computed", ask one question back: do you want a trajectory (the state at a particular moment), an end state (where it finally settles), or a statistic (the long-run distribution, rate, or proportion)? Their computability is completely different. Rewrite every question that can be rewritten as a statistic — turn "which week next quarter will the outage happen" into "roughly how many outages next quarter, and how long is the longest" — the second is answerable, and it is the number your staffing and redundancy actually need.

🌀 Biomedicine · you can't predict one heartbeat, you can predict the rhythm The interval between successive beats on an ECG varies, and the variation is not noise — a healthy heart is irregular beat to beat by design. Predicting which millisecond the next beat lands on is hopeless; but turn a few hundred intervals into a distribution (clinically, heart rate variability) and its width is stable enough to serve as a prognostic marker: abnormally reduced variability is associated with elevated cardiovascular risk, a finding replicated many times over. Same heart, unanswerable beat by beat and answerable in the distribution — and the clinical information lives in the distribution layer. The beat-by-beat layer carries an illusion.

04The second wall: no shortcut

Now demolish the first wall entirely and see what is left standing behind it.

Suppose your system is discrete: cells, integers, exact rules. No decimals, no measurement error, no rounding. You know the initial state to the last bit. What could possibly stop you now?

Try the plainest system there is: a one-dimensional cellular automaton. A row of cells, each black or white; at every step all cells update at once, and a cell's new colour depends only on its own old colour and those of its two neighbours. Three cells give eight possible neighbourhoods, each assigned an outcome, so there are 256 possible rules, numbered 0 to 255. → ref · Game of Life & cellular automata

Start from a single black cell in the middle. Rule 90 grows a Sierpiński triangle, as regular as woven cloth. How regular? Whether the cell in row n, column k is black has a closed form — it equals Pascal's triangle (the binomial coefficients) modulo 2, so you plug in n and k and compute it directly without running a single step. This is reducible: a shortcut exists.

Rule 30 is a few bits away and grows something like scattered sand: some striping on the left, and in the middle a broad region where nobody has found any pattern at all. Wolfram used it as a pseudorandom generator for decades; in 2019 he offered $30,000 for three fairly basic questions about it, one of them being whether the centre column ever becomes periodic. The prize is still unclaimed. To know the colour of the middle cell at step one million, the only known method is to run a million steps. The property is called computational irreducibility.

Rule 90 — reducible Rule 30 — irreducible cell (n,k) = Pascal's triangle mod 2 no pattern found in the centre column same start: one black cell row one million? plug into the formula row one million? run a million steps The rules differ by a few bits and neither side has any measurement error — this wall is not about precision
Both rules fit in eight lines, both are exact, both are noiseless. The only difference is whether a path shorter than the process exists.

Be precise about how this differs from the first wall, because it is the sentence that matters most here: irreducibility is not because you measured badly (there is no measurement), not because the rule is complicated (eight lines), and not because noise was amplified (there is no noise). It says that the process is its own shortest description. To jump to the answer you would need a path shorter than the process, and for the overwhelming majority of such rules that path does not exist.

There is a harder layer still: Rule 110 has been proved Turing-complete (Matthew Cook, 2004) — it can simulate any computer. General questions about its long-run behaviour therefore sit at the level of the halting problem: not hard to compute, but undecidable, with no algorithm that answers for all inputs.

🎯 DECISION

Once you have established that the process in front of you is irreducible, stop spending on "computing it more accurately" — that money buys nothing. Spend on three things instead: make one step cheaper (shorten the cycle of one real trial), make the result of a step reversible, and make the observation between steps denser. In a world without shortcuts the only available speed-up is trial throughput, not predictive precision. This is the real argument for small fast steps — not a cultural preference, a consequence of the wall.

🌀 Mathematics · the nth digit of π does have a shortcut "You can only run it at full length" is never a claim you may make in advance. In 1995 Bailey, Borwein and Plouffe found a formula that computes the nth hexadecimal digit of π directly, without computing the preceding n−1 digits — something widely believed impossible until then. So "no shortcut" describes a state in which no counterexample has yet been found, and counterexamples do occasionally turn up. The practical inference: whoever says a thing "just has to be run" carries the burden of proof, and the payoff from hunting for a shortcut is a jump, not an increment.

05Which wall are you hitting

The two walls have nearly opposite remedies, so the diagnosis is everything. Three questions do it.

First: would an order of magnitude more measurement precision lengthen the horizon? Yes → precision wall, and buying observations, data and sensors pays (only additively, but it pays). No → stop buying; that budget belongs somewhere else.

Second: would an order of magnitude more compute let you see further? Under the precision wall: yes, because compute is what makes bigger ensembles and finer grids affordable. Under the computation wall: compute only makes the same process run faster — a million steps is still a million steps, just in less time. It shrinks a constant, not the wall.

Third: do I want a trajectory, an end state, or a statistic? This one dissolves a great many laments about unpredictability, and it works against both walls at once: the shape of a chaotic system's attractor is stable, and the density of black cells in an irreducible automaton is usually stable and estimable too. Reword the question and the wall may no longer be on your route.

① Precision wall ② Computation wall finer measurement: helps (additively) more compute: helps (bigger ensembles) more waiting: no help remedy: ensembles + intervals, replan on the doubling time finer measurement: useless (no error) more compute: shrinks a constant running it out: the only way remedy: make one step cheaper, take only reversible steps shared exit: ask for a statistic, not a trajectory
Diagnosis dictates the prescription. The same sentence — "we need better forecasts" — means buy data on the left and don't buy data on the right.

One combined case needs care: real systems often hit both walls at once. Economies, ecosystems and organisations are all unmeasurable and irreducible together. The order is then fixed — ask the third question first (can this become a statistic), then the first (is precision worth buying), and only then concede the second wall. Done the other way round, you spend the money first and discover afterwards that it went to the wrong wall.

🎯 DECISION

In any meeting about "we need better forecasts", run the three questions and write the conclusion down: which wall we are hitting, and therefore what we will stop funding. Without that last clause, the forecasting budget grows forever along the "buy more data" path — and that path only pays in front of the first wall.

🌀 Eastern thought · the questions to be set aside A passage in the Aṅguttara Nikāya sorts questions into four ways of answering: those to be answered directly, those requiring analysis before an answer, those met with a counter-question, and those to be set aside. The last is not evasion; it is a ruling that the form of the question is itself defective. This section does the same work, with one thing added that the taxonomy lacks: a criterion. "Unknowable" gets split into three distinct things — can't measure, no shortcut, wrong object — each pointing at a different next step. A taxonomy without a criterion can only be applied by authority; which is why "that one's just unknowable" is always said in a meeting by the most senior person in the room.

06Where this breaks down

Now the other side. Every claim above has a place where it fails, and you should know them before using any of it.

First, the Lyapunov exponent is an average, and real predictability changes day to day. "Doubling every day and a half" is a long-run average along the trajectory. Some atmospheric situations are remarkably sturdy — error grows slowly and a ten-day forecast holds; others fall apart at a touch and collapse in three days. Forecasters call this flow-dependent predictability, and the ensemble spread is precisely how they measure it on the day. So do not treat "two weeks" as a hard line that holds every day; it is a ceiling in order of magnitude, not a schedule.

Second, model error and initial-condition error are different things, and this issue's framework only covers the latter. Ensembles perturb the initial state. If the model itself has the physics wrong, all fifty members err in the same direction while the spread still looks narrow — and you get a confident, wrong probability. Climate projection faces mainly this kind of error. It is not a chaos problem but a structural one: adding members does nothing, you have to change the model. A narrow interval is not reliability, and this is the most common misreading of ensembles.

Third, computational irreducibility is not currently a theorem. Wolfram hangs it on a larger conjecture he calls the principle of computational equivalence: that almost all non-trivial rules are equivalent in computational power. There are many supporting cases and no proof. Rule 110's Turing-completeness is a rigorous result; "therefore any specific question is unpredictable" is not — a Turing-complete system can still contain plenty of specific questions that are decidable or even have closed forms. Using "irreducible" as a general-purpose shield for "so it can't be computed" is the easiest mistake to make here, and one of the most criticised uses of complexity science.

Fourth, and most important: neither wall endorses "so don't decide". The limits of predictability constrain statements about a specific future state. They do not constrain changing the distribution. You can still thin the bad tail, speed up recovery, and size the bet so it survives the whole distribution. Giving up point prediction is not giving up intervention — it is what frees the budget for it.

🎒 In Practice · BigCat

  1. Engineering & systemsThe sentence that recurs most during an incident is "how long until it's back?" You give a time, then slip it, and every slip costs more trust than the last. Change the deliverable: give two numbers and a moment — "optimistic 20 minutes, pessimistic 2 hours, next update in 20 minutes" — and set the update interval to the time it takes your estimate error to double (the further apart your last two estimates were, the shorter it should be). The sentence to retire is "let me look again, maybe X minutes". Afterwards, track exactly one thing: what fraction of your stated ranges contained the actual recovery time.
  2. Practice & mindThe recurring situation is setting a timetable for your own practice — by such a date there should be such progress — and declaring the method useless when the date arrives. That asks the wrong object. A single day's state is a trajectory (whether you sit well today depends on sleep, weather, one email; unanswerable); the distribution is what's measurable. Change the record from "how was today" to "over the past 30 days, on what fraction did I sit the full length", and judge only on that ratio. Stop the daily self-assessment: it measures noise, and it makes you switch methods on noise — and switching resets the statistic to zero.
  3. ParentingWhen a child shows a change you don't understand, the first reflex is usually to gather more information: ask the teacher, ask other parents, check the phone. That reflex assumes a precision wall. The test can be applied before you move: write down "if the answer is A, my next step is ___; if B, ___" — and if the two next steps are the same, don't gather. That energy belongs to the side that runs it out together, not to the side that measures. Only one kind of information is worth buying: the kind that changes your next step.

🌀 Crossings

Going Deeper

If statistics are predictable and trajectories are not, why does "the long run" so often feel more unnerving?

Because predictable statistics rest on a precondition: the rules generating the data did not change over that stretch, and the attractor holds its shape. Climate change is exactly the case where the attractor itself is moving, and market structure moves too. The move offered here — estimate the future distribution from historical statistics — is what fails first. The test is plain: did the rules change over this period? If they did, the distribution in your hand belongs to the system that used to exist.

Can an irreducible system contain reducible sub-questions?

Yes, and these are the most useful findings of all. The whole trajectory may be out of reach while some conserved quantity, some mean density, or some yes/no question about eventual extinction has a closed form. The difference between Rule 90 and Rule 30 is the reminder: reducibility is a property of the question. So the work worth doing is hunting for the sub-questions that happen to be reducible, rather than complaining that the whole is unpredictable.

Can an ensemble mislead you?

It can, and quite specifically: if the perturbations are chosen badly — only along unimportant degrees of freedom — the members stay bunched and hand you false confidence. Spread is meaningful only once the perturbation scheme itself has been validated. When you see a very narrow interval, the first question is not "what's the conclusion" but "how was this interval generated".

Is computational irreducibility the same thing as NP-hardness?

No, and merging them produces wrong conclusions. NP-hardness is about how solution cost grows with problem size, and usually comes with the property that an answer is easy to verify. Irreducibility says that even at fixed size there is no path faster than stepping through, and the answer is typically not easy to verify — you cannot check it without running it. Both can hold at once, but they are different walls with different remedies.

If predictive ambition has to be scaled back, what is left of "forecasting" as a personal skill?

Three trainable things: judging which wall a problem sits behind, rewriting the question into an answerable form, and calibrating your own intervals. All three can be scored and all three improve with practice. Calling specific outcomes, in many domains, cannot be trained — its ceiling is set by the system, not by you.

Further Reading