The hard question was never "why did they start?" — it's "why keep going once it stopped feeling good?"
Say "addiction" and most people picture something that feels too good to stop. Ask anyone who has been stuck for years and you tend to get the opposite answer: it stopped being good a long time ago, and they still do it. Someone who has smoked for twenty years rarely describes the cigarette as enjoyable — they describe not smoking as unbearable. This isn't a story about willpower. It's a story about circuits: the brain runs wanting and liking on two different systems, and addiction inflates one of them past the point of distortion while the other quietly withers.
First, a myth has to go: dopamine is not the "pleasure molecule." The misunderstanding runs deep — plenty of popular science still repeats it. Dopamine is one of the small signalling molecules the brain uses to talk to itself (these are called neurotransmitters; think of them as different broadcast channels). What it actually governs, in plain terms, is going after things — getting up, reaching, fixating, lunging. Whether the thing feels good is a separate question.
The cleanest evidence comes from a set of classic experiments. Strip a rat's brain of nearly all its dopamine and it stops going after sugar water — it will sit next to food and need to be fed by hand. But drip that same sugar water into its mouth and the "this tastes good" facial reactions (tongue protrusions, lip licking) are indistinguishable from a normal rat's. Liking, fully intact. Wanting, gone.
It runs the other way too. People with Parkinson's often take drugs that directly imitate dopamine, and a minority of them abruptly develop pathological gambling, compulsive shopping or binge eating. Turn wanting up pharmacologically and behaviour follows. And what they frequently report is this: the gambling wasn't enjoyable.
So who handles the pleasure? A much smaller and far more fragile system: a few patches inside the nucleus accumbens and ventral pallidum known as hedonic hotspots, running on the brain's own opioid and cannabinoid molecules. They're small and hard to amplify — which is exactly why pleasure has a ceiling and craving doesn't. (The full circuit, reward circuit; who does what among the chemicals, neurotransmitter systems)
Machine learning has a family of methods called reinforcement learning: a program tries things in an environment, gets a score when it does well (the reward — the number that lands right now), and learns to estimate how much it can collect from here on out (the value). The crucial part: what actually drives behaviour is value, not reward. Liking is like reward; wanting is like value. Addiction looks like an agent whose value estimate has been inflated far past reality while the actual reward has gone to zero: the ledger still says "this is worth a fortune," so the policy keeps executing, even though nothing arrives.
Normally dopamine is a very economical teacher. In the classic experiments, a monkey gets juice: unexpected juice makes dopamine neurons fire a burst. But if a light comes on before every delivery, then once the monkey has learned the association the burst moves entirely onto the light — the juice itself no longer produces one. And if the light comes on but no juice arrives, dopamine dips below baseline. That dip is what disappointment looks like physiologically.
What the signal actually carries is one sentence: "how much better than I expected was that?" Better than expected, a burst; exactly as expected, silence; worse, a dip. Which means natural pleasure comes with its own brake built in: once something becomes predictable, the signal goes to zero. Nobody gets excited about the third bite of dinner. (Where this teaching signal lands and how it rewrites actions, basal ganglia)
What makes addictive substances so effective is that they skip the prediction machinery and raise dopamine chemically. Cocaine blocks the pump that clears dopamine away, so it lingers where it is. Amphetamine is cruder still — it runs the pump backwards and forces dopamine out. Nicotine directly excites the neurons that make it. Opioids release the brake: they silence the inhibitory neurons that normally hold those cells down. Alcohol works several of these angles at once.
The consequence is nasty. A natural reward can be predicted away and its signal shuts off; the pharmacological hit cannot. Every single use still leaves a "better than expected" signal behind, so the learning never saturates. Everything bound to it — that street, that glass, the gesture of unlocking a phone — keeps getting painted with "this matters," brighter each time. And the process sensitizes: run into the cue long after quitting and the pull can be stronger than it ever was, even though the pleasure is long gone.
Reinforcement learning has a basic requirement: the error has to be able to reach zero, or the value estimates never settle. Now suppose someone welds a constant positive number onto one action's reward channel — do it and you score, regardless of what the environment actually delivers. The error never zeroes, the value estimate inflates without bound, and the policy collapses onto that one action. In AI this is reward channel tampering, a species of reward hacking: the program found a shortcut that cranks the score while bypassing the real goal. A drug does almost exactly this to a brain — it doesn't make the world better, it counterfeits the signal that says the world got better.
Addiction isn't built in a day, and it follows a specific migration route. Early on, the behaviour is run by the ventral part of the striatum (the deep structure that decides whether to act), which computes "is this worth it" — you're still choosing. After enough repetitions, control shifts dorsally, into the region that stores habits: this situation, therefore this action, no longer costed against outcome. Like driving home on autopilot when you meant to go somewhere else. The behaviour has come unhooked from its purpose.
Meanwhile the other end loosens. Imaging studies keep finding that a class of dopamine receptor (D2) in the striatum becomes less available in addiction — and the more it drops, the lower the metabolic activity in the prefrontal cortex, the region in charge of "hold on, think about what happens next" (prefrontal cortex). More accelerator, less brake.
Laid out whole, the process is a loop in three parts. Stage one is bingeing and intoxication — this is the part that's about the high. Stage two is withdrawal and negative affect: to counter a pleasure system that keeps getting yanked upward, the brain actively recruits an opposing system — stress signalling in the extended amygdala, running on the same molecules as the stress response (HPA axis). Remove the drug and only the opposing system is left running: anxiety, irritability, nothing feels like anything. Stage three is craving and anticipation: the prefrontal cortex stops blocking and starts supplying reasons — "just this once," "today is different."
Which finally explains the cruellest paradox: the motive switches from "to feel good" to "to stop feeling bad." The pleasure shrinks, the pain becomes the baseline, and with every lap the floor — how you feel when you're doing nothing at all — settles one notch lower.
Reinforcement learning splits decision-making into two gears. Model-based: simulate "what happens if I do this," then choose — slow, expensive, but updatable. Model-free: just look up a cached table, "this state, that action" — fast, cheap, and blind to the fact that the cache has expired. It's long been observed that with enough repetition systems slide from the first toward the second: computing costs more than looking up. Addiction is that slide running on a biological brain — from "is it worth it" to "this situation, therefore this action" — with a long-invalid high score still sitting in the cache and nobody updating it.
The single most useful thing to know about relapse: extinction is not erasure. When you've been abstinent a long time and the cues stop producing action, the brain has not wiped the original "there is a reward here" memory. It has laid a new "there isn't anymore" inhibitory memory on top of it. Both memories coexist — and the new one is picky about place: it's bound to the context you were in while quitting. Go back to the old street, hit real stress, or take a single taste, and the old layer surfaces again (synaptic plasticity; who tags the context, hippocampus).
Stranger still is the time course of craving. Intuition says it declines steadily from day one of abstinence. What actually gets measured is a rise before the fall: over the first weeks to months, cue-induced craving climbs to a peak, and only then slowly recedes. This is the incubation of craving, seen repeatedly in animals and in people. The practical implication is blunt — the most dangerous moment is usually not the first week, but the stretch where you've decided you're fine.
Environment carries far more weight than most people assume. Heroin use was widespread among US troops in Vietnam, and roughly one in five returning soldiers considered themselves addicted — by the clinical experience of the time, that cohort was doomed to relapse. What actually happened: only about five percent relapsed during their first year home. Same brains; an entirely different set of cues, companions and daily structure — and an entirely different trajectory.
So the useful moves point in a clear direction: change the environment before you test the will. Removing cues from your life costs far less than resisting them where they are, and rebuilding structure and relationships answers the real question — what fills the space afterwards. Use medication where it exists: agonist therapy with methadone or buprenorphine cuts opioid mortality to roughly half; alcohol has naltrexone and others; nicotine has varenicline and replacement therapy. Among behavioural programmes, contingency management — an immediate, tangible reward for every clean test — remains the most effective option for stimulants. None of this is a morality exam. If you're stuck, get professional help.
"Extinction isn't erasure" has an almost point-for-point counterpart in large language models. Post-training a model to refuse a category of question does not delete the underlying capability from the weights — it presses a "don't answer that way" layer on top. Change the framing and the suppressed thing resurfaces: that's jailbreaking. Hence a whole research problem called machine unlearning: making a model genuinely forget something is remarkably hard, because what it learned was never a record you could delete — it's a tendency smeared across countless connection weights. On the biological side the answer is identical: a memory doesn't live in a slot, can't be deleted, and can only be covered over by new learning.
The split between wanting and liking wasn't discovered in the twentieth century — earlier observers just had no electrodes, only themselves to watch: