TOPIC 24 · PHASE D

Spreading and Contagion

Ideas are not viruses — don't copy the virus playbook

2026-08-10 · Networks

"How many people does one patient infect on average?" — not one term in that number belongs to the virus alone. It is the product of the pathogen, your behaviour, and the network you happen to sit in. Which is why the same disease can be two different diseases in two cities.

A rumour goes round the office, an app is suddenly on everyone's phone, a flu season, a clip that won't stop resurfacing. We call all of it "spreading", and we borrow the word viral for all of it. That word comes from epidemiology — and the borrowing leaves the most important premise behind.

The premise left behind is this: a virus only needs to touch you once. Anything that asks you to change what you do — switch tools, drop a habit, be the first to say the plan is wrong — usually needs several different people to do it in front of you first. That one-word difference flips the whole playbook: the structure that makes viruses fly (long-range shortcuts, wide weak ties) is exactly the structure that stalls behaviour.

The six network topics divide the ground: T20 is structure itself (degree, paths, clustering), T21 is shortcuts and small worlds, T22 is how hubs grow, T23 is what that structure resists and what breaks it, T25 is cascading failure through load redistribution. This one asks a single question: once something is running across the network, what decides whether it stops?

01How Many Does One Person Infect

In 1927 two Scottish researchers, Kermack and McKendrick, took a shortcut that turned out to be the whole field: track nobody in particular, and sort the population into three piles instead. → ref · the SIR model

Pile one is the susceptible: never had it, will catch it on contact. Pile two is the infectious: currently passing it on. Pile three is the removed: recovered or dead, either way out of the game. Their initials give the model its name — SIR. There are two rules: a susceptible who meets an infectious becomes infectious with some probability, and an infectious eventually becomes removed.

Out of those two rules falls one number, the basic reproduction number, written R₀: drop a single infectious person into a population where everyone is still susceptible, and count how many they infect on average. Above 1 it grows, below 1 it dies out on its own. The line is almost suspiciously clean.

What deserves a pause is what R₀ is made of. It is a product of three things: how many people you contact per unit time × the chance a contact transmits × how long you stay infectious. Only the middle term belongs mostly to the pathogen; the other two are behaviour and environment. So "this disease has an R₀ of 3" is an incomplete sentence — it needs "where, when, and among whom".

Three piles, two arrows, one number S Susceptible I Infectious R Removed β contact, catch it γ no longer spreading R₀ = contact rate people met per day × per-contact chance the pathogen's own term × infectious period how long you can pass it on The other two are behaviour and environment — one pathogen, two cities, two values of R₀
All of SIR: three piles, two transitions. R₀ is a product of three factors, two of which are not the virus.

One number gets used even more than R₀: the effective reproduction number, R_eff — how many one infectious person is passing it to right now. It equals R₀ times the fraction of people still susceptible. That sentence looks unremarkable, and it is the seed of everything below: spreading eats its own fuel. Every person infected is a person removed from the susceptible pile, and the same virus with the same behaviour starts running slower.

🎯 DECISION LINE

"Reduce transmission" is not one action, it is three separately priced ones. Split R₀ into contact rate, per-contact probability and infectious period, and ask of each: what does cutting this one by a third cost? Masks and encryption cut the middle term; capacity limits and isolation cut the first; fast detection and fast takedown cut the third. Get all three price lists on the table before choosing — "we must control the spread" has no units.

🌀 Engineering history · SQL Slammer In January 2003 a 376-byte worm scanned the internet by picking IP addresses at random, doubled every 8.5 seconds, and owned roughly 90% of vulnerable hosts in under ten minutes. It could be reconstructed so precisely afterwards precisely because the textbook assumption — contact a target chosen at random — was literally true of it: it knew nobody and had no social circle. Which gives an unflattering corollary: human epidemics are hard to forecast not mainly because viruses are complicated, but because people do not mix at random.

02Why It Stops By Itself

If R_eff = R₀ × (fraction still susceptible), then once that fraction drops below 1/R₀, R_eff falls under 1 and the outbreak contracts. That gate has a name — the herd immunity threshold, equal to 1 − 1/R₀. At R₀ = 3 it is 67%.

This is exactly what Kermack and McKendrick set out to explain: why does an epidemic recede while a large crowd of never-infected people is still standing there? The fashionable answer at the time was that the pathogen weakened. They showed you need no such assumption — burn part of the fuel and the fire shrinks by itself.

But there is a turn here that nearly everyone drops. The epidemic does not end when the threshold is crossed. Crossing it only means new cases start falling. A large number of people are still mid-infection, and they do not stop breathing because a statistical line was crossed. They keep transmitting, susceptibles keep falling, and the system sails well past the threshold before it truly halts. That extra stretch is overshoot.

The numbers are worse than the intuition: at R₀ = 3 the threshold is 67%, but letting it burn out unmitigated infects about 94% of the population. All 27 extra percentage points are overshoot. That is the arithmetic problem with "let it reach herd immunity naturally" — it treats the threshold as the finish line when it is only the turning point.

How much fuel is left when the fire goes out (R₀ = 3) share never infected herd immunity threshold — 33% left susceptible burnt out: only 6% left susceptible (94% infected) susceptible S/N R_eff is exactly 1 here new cases peak — this is not the end overshoot 27% extra burnt time →
The threshold is where new cases peak, not where it ends. Those 27 points inside the red band burn after the point where it "should" have stopped.

Overshoot also settles a very practical argument: when to let go is a real question. Lift an intervention exactly when R_eff has just crossed 1 and the susceptible share is still hugging the threshold — the moment it lifts, R_eff bounces back above 1 and the whole thing restarts. Any "stop at the target" plan needs headroom for overshoot, which means the exit condition should never be a date; it should be a stock level held below a line for a stated duration.

🌀 Demography · Keyfitz and population momentum In 1971 Nathan Keyfitz asked a question most people thought they knew: if fertility dropped to exact replacement level tomorrow, would population growth stop? It would not — he calculated that high-fertility countries would keep growing to roughly 1.6 times their then-current size before levelling off, because past high fertility had already loaded a large cohort into childbearing age and that stock still has to run through. Same mechanism as overshoot: pushing the driver past its threshold turns the flow over, it does not turn the stock around. Hence a test that applies to every "stop when we hit the target" policy: if you are acting on a flow but care about a stock, the moment you hit target is never the moment to stop.

03Where the Average Stops Working

Both sections above smuggled in an assumption called homogeneous mixing: everyone is equally likely to meet everyone, like a well-stirred soup. It makes the maths beautiful, and it puts the conclusions some distance from reality.

People are not soup. Some meet two hundred people a day; some meet five a week. T22 covered why contact degree distributions are so often heavy-tailed → ref · spotting and misreading power laws. Once degrees are uneven, the quantity that decides whether something takes off changes: it is no longer the mean degree ⟨k⟩ but ⟨k²⟩ / ⟨k⟩ — the mean of the square. Squaring is brutally sensitive to the tail: someone with 200 contacts weighs four hundred times as much as someone with ten.

In 2001 Pastor-Satorras and Vespignani carried this to its conclusion, and it startled the field: if the degree distribution is a power law with exponent at or below 3, then ⟨k²⟩ diverges as the network grows and the epidemic threshold goes to zero. There is no safe band of "too weakly transmissible to take off". Any transmissibility at all can establish itself.

Same mean number of contacts, two completely different networks homogeneous everyone has about the same number of contacts heavy-tailed a few people have very many; the mean hides it threshold of the homogeneous net heavy-tailed threshold collapses toward 0 transmissibility λ → that whole left band — "too weak to spread" — does not exist on the right-hand network
The threshold is not set by the mean number of contacts but by the mean of the square. The few people in the tail nearly define the fate of the whole network.

There is a layer beyond "some people have more contacts": the distribution of secondary infections itself is heavy-tailed. The term is overdispersion. SARS-CoV-2's dispersion parameter was estimated around 0.1, meaning roughly 10% of cases produced about 80% of transmission. So the person described by "three infections on average" may not exist at all — most infect nobody, a few infect dozens.

That changes the shape of intervention outright. Under homogeneous mixing, an intervention must be universal to work, because every person contributes equally. Under overdispersion, transmission concentrates in a few settings — enclosed, poorly ventilated, long duration, loud — so setting-targeted measures are absurdly cheap: a tiny footprint buys most of the mass in the tail.

🎯 DECISION LINE

Whenever a spreading metric is quoted "per person", ask about the shape of its distribution first. If secondary transmission is heavy-tailed, the mean carries almost no information and the thing to measure is the tail: what fraction of events produced what fraction of the consequences. The action changes to match — stop shaving everyone uniformly and instead enumerate the high-transmission settings and handle them one by one. Blunt test: sort past events by consequence; if the top 10% account for more than half the total, uniform intervention is buying you the wrong thing.

🌀 History and medicine · Typhoid Mary Mary Mallon was a New York cook who never fell ill herself, yet was linked to at least 51 typhoid cases and three deaths; she was forcibly quarantined on North Brother Island twice, some 26 years in total, until her death in 1938. The case is usually taught as public-health ethics, but this section's mechanism adds a cold criterion: only when transmission is strongly overdispersed can an intervention aimed at one individual have a system-level effect. Under homogeneous mixing the benefit of isolating any single person is necessarily negligible, and therefore unjustifiable. So half the answer to "should we target this person" lives in the shape of the distribution, not only in ethical intuition — and the same accounting states the price out loud: 26 years.

04Is Once Enough

Here is the most useful distinction in this topic. A virus needs one contact: you meet it, you catch it with some probability, and how many carriers you met last week is irrelevant. That is simple contagion.

A great deal of social behaviour is not like that. Quitting to start something; abandoning a tool you've used for ten years; being the first in the room to say the plan is wrong — the bar for these is several different people around me already did it. One person doing it is noise; three unrelated people doing it is signal. That is complex contagion, named by Centola and Macy in 2007 and prefigured by Granovetter's 1978 threshold model: everyone carries a private number — "I'll do it once n others have" — and those numbers vary across a population.

The point is that this distinction inverts the value of network structure. T21 showed that a handful of long-range shortcuts collapses a network's diameter → ref · the small-world model, which is wonderful for simple contagion: one shortcut suffices to carry a spark across. But complex contagion is not short of distance, it is short of multiple independent signals landing on the same person. A shortcut delivers exactly one and then it is spent; a dense circle of mutual acquaintances delivers the same thing from three directions, which is precisely what a threshold requires. The structure that makes viruses fly is the structure that stalls behaviour.

Centola turned this into an experiment in 2010. Over 1,500 participants were randomly assigned to one of two artificial social networks: a clustered lattice (your neighbours are also each other's neighbours, so ties are highly redundant) or a random network (long ties, short paths). The behaviour to spread was registering for a health forum. The result ran against small-world intuition: the clustered network spread it faster and further — 54% adoption versus 38%, at roughly four times the speed.

Same network, same single shortcut, two outcomes ① simple contagion — one contact is enough the only shortcut one signal is enough to light the far side infected all of them fall ② complex contagion — needs 2 different neighbours first the same shortcut delivers 1 signal, threshold is 2 — it stops here adopted nobody moves
Shortcuts save distance; complex contagion is short of repetition. The structure that makes viruses fly is the one that stalls behaviour.
🎯 DECISION LINE

Before funding any "rollout", settle one yes-or-no question: is one exposure enough? Yes (see it and you'll share it, install it and it works) → buy reach: shortcuts, cross-cutting ties, headcount covered. No (it demands a changed practice or a taken risk) → buy density: saturate one small circle, then move to the next, and never spread thin. The two spend in opposite directions; splitting the difference lands you short on both. The check is concrete: ask people who already adopted, "how many different people did you hear it from?" Mostly one — you have a simple contagion. Mostly two or three — you have a complex one, and every dollar priced per head reached is being wasted.

🌀 Economics and institutions · why cold starts go city by city Ride-hailing, payments, collaboration tools: the adoption bar is "how many people I know already use it" — textbook complex contagion. National advertising (reach, one exposure) barely moves it, while saturating one city or one campus at a time does, because saturation is exactly what manufactures repeated exposure from several acquaintances of the same person. Which yields a hard budgeting corollary: under complex contagion, splitting a budget across ten cities is not one tenth as effective as concentrating it on one — it can be zero, because each city stalls below threshold, and spending below threshold accumulates nothing.

05Where This Breaks Down

Contagion models are the most heavily borrowed toolkit in complexity science, and therefore the most frequently broken in transit. Here is where they come apart.

One: R₀ is not a constant you can look up. Measles has long been quoted at 12–18, nearly as common knowledge. In 2017 Guerra and colleagues ran a systematic review, gathering 58 estimates from 18 studies, and concluded that the actual spread of estimates is far wider than that familiar range, and strongly dependent on population structure, era and estimation method. No surprise — two of R₀'s three factors were behaviour and environment all along. Quoting an R₀ without saying which network and which behaviours produced it is quoting a number with no units.

Two: the compartmental assumptions are load-bearing, not decorative. The model requires homogeneous mixing (dismantled in section 3), discrete irreversible transitions (catch it once, immune for life), and a population whose composition doesn't shift over the period. Real pathogens have latent periods (add an E compartment), reinfection, mutation, seasonality. The model gives you a skeleton, not a body.

Three: applied to ideas, what is missing is a definition of infection. When exactly has someone "accepted" an idea? On hearing it? On agreeing? On changing behaviour? Without a decidable moment of transition, R₀ cannot be measured and the whole apparatus degrades into metaphor. Memetics ran aground precisely here — it treated ideas as self-replicating units, yet decades on, "replication" still has no observable counterpart; the Journal of Memetics ceased publication in 2005, and what it lacked was never data.

Four: complex contagion has its own boundary. Centola's elegant experiment ran on networks the researchers built: assignment was randomised, the behaviour was a single one (registering for a health forum), and participants were strangers. It demonstrated forcefully that social reinforcement is real; it did not demonstrate that all social spreading is complex contagion. Recent work also identifies a trade-off between reach and reinforcement: too clustered and reinforcement is ample but reach is slow; too random and reach is fast but reinforcement is thin. The optimum sits between them, which means "build denser small circles" is a prescription with a dose, not a direction.

🎯 DECISION LINE

Before using any spreading model, write down three things: ① which observable event counts as your infection (a download? a payment? two straight weeks of use?); ② who the edges of your network actually connect; ③ whether one exposure is enough. If any of the three won't fill in, don't report an R₀ — report something you can actually measure, such as the distribution of "how many distinct sources had this new user heard of us from". A crude number you can fill in beats an elegant parameter you cannot.

🎒 Scenarios · BigCat

  1. writing and this site itselfThe recurring situation: an issue goes out, you check views and shares the next day, and the high-numbered issue counts as "got it right". But what this site is trying to transmit — a changed way of framing a problem — is, if anything, a complex contagion: adoption needs the same person to run into the same mechanism in several different settings. View counts measure the simple-contagion side: a single exposure. What to change: make the headline number return depth instead of per-issue reach — over the past three months, how many people read three or more issues; and whether anyone can name a decision where they used a specific criterion from a specific issue. What to stop: rewriting headlines to lift the share rate. Share rate optimises the number of shortcuts, and what you're short of is repetition.
  2. parentingThe recurring situation: you want a habit to take (reading, exercise, finishing before playing), so you say it more often and explain it better. That is the simple-contagion playbook — more exposures from the same source. By section 4, complex contagion wants several different, mutually independent sources: your tenth repetition still counts as one person in the threshold tally. What to change: don't add repetitions of your own, add to the count of people they actually see doing the thing — and ideally people who don't know each other, or it's one source on replay. Checkable criterion: besides you, how many people does the child genuinely see doing this? The answer is usually zero or one, and the threshold is usually two or three.
  3. investing and position sizingThe recurring situation: something is growing fast (users, stores, penetration) and you extrapolate the last few months of growth rate forward. Section 1 explains why that is systematically optimistic: before the susceptible pool is meaningfully drawn down, R_eff hasn't started falling and the curve is pure exponential — and a pure exponential stretch carries no information about the final size. Its shape is completely insensitive to the ceiling. What to change: don't look at the growth rate; write down the denominator first (how many people, households or firms are genuinely reachable), work out what share has already been consumed, then watch a signal that leads growth — whether new additions divided by installed base has begun to decline, which is the first visible sign of R_eff turning. What to stop: compounding the last three months forward by twenty-four.

🌀 Crossings

Going Deeper

What if something spreads by both simple and complex routes at once?

Most real things are mixed: hearing about it is simple contagion, adopting it is complex. The interesting part is that the two layers have opposite optimal structures — the awareness layer wants shortcuts, the adoption layer wants density. That suggests a division of labour: use reach to build a stock of "have heard of it", then use density to convert it locally. Worth pressing on: can saturating the awareness layer actually raise the adoption threshold — everyone has heard of it, nobody has gone first, and waiting itself becomes a signal?

Why does flattening the curve reduce the total, rather than just spreading it out?

Intuition says an intervention only delays. Overshoot gives the other answer: intervention lowers R_eff, and final size is a function of R_eff — push it from 3 to 1.5 and the final attack rate falls from about 94% to about 58%. Flattening genuinely changes the total, because it changes the momentum with which the system charges past the threshold. The catch is that it must be sustained until susceptibles are actually drawn down, or lifting it produces a rebound.

Is overdispersion good news or bad news?

Both. Bad: the mean stops predicting, small outbreaks have enormous variance, and early data barely extrapolates. Good: since most transmission concentrates in a few settings, cheap local measures can remove most of the risk — an opportunity that simply does not exist under homogeneous mixing. Which is why "is this system overdispersed?" belongs at the top of intervention design rather than in a statistical footnote.

Is there a "removed" pile for ideas?

This is the hardest part of porting SIR to ideas. Someone who has had the disease is no longer susceptible, but someone who heard an argument and declined it can still accept it later — arguably more easily, since thresholds accumulate. That suggests the dynamics of ideas resemble a memoried SIS (repeat infection possible) more than SIR, and in a process with memory the notion of "threshold" needs redefining. Whoever gives that pile an observable definition moves this whole question forward.

Further Reading