CHAPTER DEEP READ · ACCELERATE · CH 3
Accelerate: The Science of Lean Software and DevOps · Ch 3 · Forsgren, Humble & Kim · 2018
"Culture" is the hardest word in any company to pin down and the one everybody uses. Someone says "we need to change the culture," the meeting ends, and nobody knows what to do differently tomorrow. Chapter 3 of Accelerate sets out to turn that adjective into a gauge you can read — and then tells you the reading can be changed, by a route most people would not guess.
If you want to know whether a family gets along, asking "is your family close?" gets you nowhere; everyone says yes. Ask instead: when a kid fails a test, does the test come home, or does it go in the bin? That one answer tells you more than a wall of framed family mottoes.
This chapter makes exactly that swap: stop asking what the culture is like, and watch how bad news travels.
There were only two roads, and both were dead ends. One was mysticism: culture is "vibe," unmeasurable, so "culture work" collapses into posters, offsites and leadership speeches — impossible to prove useful, impossible to prove useless. The other was borrowing an employee-satisfaction survey, which measures whether you are happy, not whether news gets out when something breaks.
Worse, culture can't be seen directly. What you can see is the surface — the slogans, the seating plan, the stories a company tells about itself — and the surface is the part that lies: the companies most eager to assign blame are often the ones with "we embrace failure" on the wall.
The chapter borrows a classification drawn from studies of plane crashes and medical accidents. The difference between organizations that come to grief and organizations that don't isn't the thickness of the rulebook — it's whether a piece of bad news can get from the person who spotted it to the person who needs it. On that basis there are three kinds of organization.
One: the messenger gets punished. So next time nobody speaks up — the news is buried, or polished a little on its way up. Two: the messenger gets passed around. The news gets stuck on a department wall — "not our problem" — and the thing gets half-fixed while the same bug lives on elsewhere. Three: the messenger is prized. The organization actively goes looking for bad news, and when something breaks it asks "what in the system let this happen?" before it asks "who did this?"
The counterintuitive part of the third kind: it looks like it has more problems — its incident-report count is usually higher. Not because more goes wrong, but because everyone else's incidents are being hidden.
This is where the chapter earns its keep. The standard approach to changing culture is the all-hands, the values talk, the new slogan — trying to change how people think first. This chapter's answer runs the other way: change how people work, and the thinking follows. To make a room full of people tidier, moving the bin within arm's reach beats a three-hour lecture on hygiene. In software: make releases small and automatic, and failures surface early and cheaply — a change of forty lines is too small an incident to be worth blaming anyone for, so speaking up stops being frightening.
Exactly one honest cost: this gauge measures how people feel. It runs on an anonymous survey, and the moment it is signed, or used to rank teams against each other, people answer with what management wants to hear and the gauge stops working.
Don't ask "how is our culture" — watch whether a piece of bad news can reach the person who needs to see it: is the messenger punished, passed around, or prized? And the order of change is backwards from what you'd expect: you don't win hearts and then change practices, you change practices and the hearts follow.
Want the seven survey items, the statistics that prove the gauge measures something real, and what culture actually predicts? → Switch to the deep read
Chapter 3 takes the vaguest word in DevOps and nails it down as a measurable, reproducible quantity. The pivot is this: don't ask about culture, ask about information flow — the observable face of organizational culture is whether bad news reaches the person who needs it. On that basis the chapter adopts Ron Westrum's three-way typology of organizational culture, turns it into a gauge with seven survey items, and delivers the book's single most actionable claim: culture isn't changed by changing minds, it's changed by changing practices — which makes culture both a cause of performance and a consequence of engineering work.
Third chapter of Part I, "What We Found," directly after Chapter 2. Chapter 2 established the dependent variable (the four measures); this chapter establishes the key variable on the book's other main thread — organizational culture — and along the way demonstrates the methodology the whole book runs on: how to rigorously measure something you cannot observe directly. Chapters 4 (technical practices), 9 (sustainable work) and 11 (leadership) all reuse it. In real life it maps to that meeting: "we're doing a DevOps transformation, so step one is culture work" — a sentence this chapter says is in the wrong order.
DevOps has put culture first since the word was coined, yet culture that can't be described can't be measured, and what can't be measured can neither be proven useful nor shown to be improving. So "culture work" degrades into posters, offsites and CEO speeches — the problem isn't insincerity, it's that none of it is falsifiable. Worse: if culture can't be measured, Chapter 1's causal chain from culture to delivery performance to organizational performance snaps in the middle, and everything downstream is left hanging.
Take a concrete anchor. Two e-commerce companies, same cloud, same CI/CD pipeline, same container orchestration:
5% canary sees checkout success rate drop two percentage points, calls a halt in the channel and rolls back 7 minutes later. The postmortem's output is "add an automated checkout-success comparison to the canary checklist."5% to everyone.Identical tooling, two orders of magnitude difference in time to restore. The difference isn't in the pipeline; it's in whatever sits between those 7 minutes and those 4 hours.
What happens if you don't solve it? You spend the whole budget on tools and the numbers don't move. The subtler consequence: the organization keeps running in a state where telling the truth is risky, and your dashboards still look great — because the bad news is filtered out before it ever reaches a dashboard. You don't have fewer problems; you have less information about them.
There is a classic layering in the culture literature, which this chapter follows: culture exists on three levels at once, and observability is inversely related to truth.
Basic assumptions are culture's real body — defaults so taken for granted that nobody articulates them, such as "obviously you sandbag the estimate when leadership asks." The people inside them don't know they're there, so they're close to unmeasurable. Values are what members can say out loud about how things ought to be. They can be surveyed, and this is the level the chapter works at; the cost is that people report the "ought," not the "is." Artifacts — slogans, the office layout, process documents, the war stories a company tells — are the easiest to observe and the hardest to read: a "we embrace failure" banner may come from the most open company in the industry or the most punitive one. What's on the wall is often what's missing in the room.
The chapter's key move is a change of question. Ron Westrum studied aviation and healthcare — domains where getting it wrong kills people. His observation was that the difference between organizations that fail and organizations that don't lies not in the thickness of the rulebook but in how information flows, and that culture is what governs that flow. Google Cloud, restating him on their own blog, put it this way: information flow "is both influential and indicative of how parts or all of an organization will behave when trouble arises."
Hence the three types:
Westrum also gave three criteria for good information, more easily missed than the typology itself: it answers the question the receiver actually needs answered; it is timely; it is presented so the receiver can use it. Note the third — "it was sent" is not "it arrived." A company can run 300 Slack channels and a thousand alerts a day and still have blocked information flow, because nobody can fish the one that matters out of the stream. This is where tool count is most misleading.
Why does the model fit software so well? Because software delivery is exactly the kind of activity it was built for: high-frequency, cross-functional, and partial failure is the normal state. One release needs dev, test, ops, security and product aligned at once, and something will always break.
Culture can't be observed directly, so it is treated as a latent construct: measure not the thing, but a set of observable things that should all show up together if the thing is real. A set, not a single item — one question is too noisy, and wording, mood or the last incident can all drag it. The chapter used seven 7-point Likert items (in substance):
Seven items aren't enough on their own; the gauge has to be shown to work. The chapter runs three checks, in plain terms: are these seven really measuring one thing (convergent validity), does it separate from other constructs (discriminant validity — is this just job satisfaction wearing a hat?), and is the reading stable across samples (reliability). All three pass.
Here is the contribution that's easy to read past. Westrum's typology was originally a theoretical classification: insight, generalized, without large-sample validation. This chapter finds that the seven items genuinely cluster — which means the classification is not a handy metaphor but corresponds to a real, measurable quantity. Promoting a framework from "sounds right" to "can be measured" is worth as much as any single finding in the book.
The predicting side: Westrum culture predicts three things — software delivery performance (Chapter 2's four measures), organizational performance (profitability, productivity, market share), and job satisfaction. The third gets overlooked, but it's what makes the whole thing self-sustaining: where the culture is good, people don't leave, and improvement gets continuity.
The side that gets changed is where the chapter's practical value lies. The conventional route to culture change is to change how people think first. The book's research model points the other way: continuous delivery and lean management practices push culture toward generative. They do it because they change information flow directly rather than arguing with anyone:
Put differently: you can't turn the culture dial directly, but culture is the sum of a pile of practices, and practices can be changed directly. Lean manufacturing arrived at the same idea long ago — reviewing the Toyota–GM joint venture at NUMMI, John Shook wrote that the way to change culture is not to change how people think first but to start by changing what they do (MIT Sloan Management Review, 2010).
One word on the loop's dark side: it turns backwards just as readily. One blame-shaped incident review makes the next report half an hour later; half an hour later makes the next incident bigger; a bigger incident draws heavier blame. Decay and improvement run the same loop in opposite directions.
Table 1 · Westrum's three types × six characteristics — the chapter's skeleton, and the most usable self-check on the page
| Characteristic | Pathological (power) | Bureaucratic (rules) | Generative (performance) |
|---|---|---|---|
| Cooperation | Low | Modest | High |
| Messengers | are shot | are neglected | are trained |
| Responsibilities | Shirked | Narrow | Risks are shared |
| Bridging between teams | Discouraged | Tolerated | Encouraged |
| Failure leads to | Scapegoating | Justice, by the book | Inquiry |
| Novelty | Crushed | Treated as a problem | Implemented |
The row most often misread is "justice, by the book." It sounds like fairness; it is the signature bureaucratic symptom — what gets pursued is "who failed to follow the process," never "why did the process allow this." The bureaucratic failing was never having rules; it's rules that serve a department's turf instead of the mission.
Table 2 · Symptom you observe → likely type → the lever to pull (and its honest cost)
| What you observe | Likely | Lever | Cost / caution |
|---|---|---|---|
| Few incidents reported, plenty of production trouble | Pathological | Put blameless postmortems into the process; require the output to be a system change, never a personnel action | Reported incidents will go up at first — tell leadership in advance that this is the good sign, or it reads as "things got worse" |
Every release needs 3 departments to sign off; windows are two weeks apart | Bureaucratic | Replace the CAB with peer review plus automated gates; return decision rights to whoever holds the information | Regulated industries can't do this bluntly — swap in an evidence trail auditors accept, don't remove the evidence |
| Dev says "shipped, not mine"; ops says "I can't rescue this code" | Bureaucratic | Shared monitoring and shared on-call; put the four delivery measures on one dashboard both sides read | Handing dev the pager without the right to change and ship the code just deepens the standoff |
| Survey scores are high, but nobody dissents in meetings | Suspect | First check whether the survey was signed, and whether the direct manager sent it | An identified survey measures what people think they should answer — here data is more dangerous than no data |
| Teams score two types apart | Normal | Diagnose per team; find what the high-scoring team does and spread it | "Company culture" is a misleading phrase — the measure is team-level |
Table 3 · Four ways to measure culture, and when each is right
| Method | Good for | Cost |
|---|---|---|
| Single item (e.g. "would you recommend this team?") | A quick trend line | Noisy, easily dragged by one event; not a construct |
| Multi-item latent survey (this chapter) | Tracking yourself over time; finding your weakest item | Needs anonymity and enough responses; measures perception, not behaviour |
| Interviews and observation | Explaining why it is this way | Expensive, slow, not reproducible; people tell you what you want to hear |
| Behavioural proxies (names in postmortems, incident report volume, cross-team review rate) | Checking that what the survey says matches what people do | Hard to standardize; instantly gameable once used for evaluation |
The pragmatic combination is the second plus the fourth: the survey gives direction, the behavioural proxies check it. If "failures improve the system" scores high while the last three postmortems produced only "raise awareness, enforce the process," the survey is lying to you.
Three places this lands directly. First, treat the seven items as a quarterly checkup: compare only against your own last quarter, never rank teams against each other (same reasoning as Chapter 2), and take next quarter's improvement target from the lowest-scoring item. Second, read the direction of incident-report volume correctly: it usually rises as culture improves, because what used to be buried has surfaced — the single finding most likely to be misread upward as "why are there suddenly more incidents?" Third, interviews and architecture reviews: asked "how do you do DevOps culture," posters and offsites are below the bar; "we measure information flow with seven items, we compare only against ourselves, and this quarter's item is replacing the CAB with peer review plus automated gates" is the answer of someone who read the chapter.
1 · The chapter breaks the deadlock of "culture can't be described, so it can't be measured" by changing the question: not what the culture is like, but whether bad news reaches the person who needs to see it.
2 · Culture sits on three levels — artifacts (visible but most misleading; the wall shows what's missing), values (surveyable, this chapter's level, but "ought" ≠ "is"), basic assumptions (the real body, nearly unmeasurable).
3 · Westrum's three types: pathological (power-oriented — messengers punished, information hidden or distorted), bureaucratic (rule-oriented — stuck at the department wall, information ignored), generative (performance-oriented — information sought out and routed to whoever can use it).
4 · Three criteria for good information: it answers the question the receiver needs answered, it's timely, it's presented so it can be used — "sent" is not "arrived," and 300 Slack channels prove nothing about flow.
5 · The measurement is a latent construct with seven items (information sought / messengers safe / responsibilities shared / cross-functional work rewarded / inquiry before culprit / new ideas welcome / failure improves the system), validated for convergent validity, discriminant validity and reliability.
6 · The contribution most often read past: Westrum's typology began as theory, and this chapter shows the seven items cluster statistically — making it a measurable construct rather than a useful metaphor.
7 · Culture predicts three things: software delivery performance, organizational performance (profit / productivity / market share), and job satisfaction.
8 · The most actionable claim: the order of culture change is backwards — not minds first and practices second, but practices first (small batches, automated tests, peer review over CAB, shared monitoring, blameless postmortems), with minds following, because those practices change information flow directly.
9 · The causal loop runs backwards too: one blame-shaped review makes the next report half an hour later, which makes the next incident bigger, which draws heavier blame. Hence three guardrails you can't skip: compare only to yourself, stay anonymous, measure at team level — sign it or rank with it and the gauge fails on the spot, at which point data is worse than no data.