"Add a cache — splitting into microservices is too heavy." You handed over a specific solution. He'll come back next time; you've become the bottleneck. And because he never understood why, the next similar call stalls too.
Intent: "What I want is P99 latency under 200ms by end of quarter, without adding more than one person's worth of ops burden."
Boundaries: "Two constraints — no new database, and the change has to be canary-able and reversible."
Hand over the how: "Inside those lines, you judge. Come back to me Marquet-style: 'I intend to add a cache, because…' — state your reasoning, and as long as it's within the lines, I nod."
"How could you miss something this basic? Run the checklist before you ship next time." He remembers the shame, learns no judgment, and next time just hides more.
"Facts first, no judgment: this ship missed monitoring. Let's reverse-engineer it together — what did the checklist in your head look like before you shipped?"
(After listening) "So monitoring wasn't in your default checklist. That's not a memory problem — it's a checklist that needs upgrading. How would you change it so next time a gap like this is caught by a mechanism, not by someone remembering?"
"You lead writing that improvement, and present it to the whole team next retro — turn your lesson into a team asset."
"Friday team outing — escape room plus dinner!" Short-term fun, but the root causes (an unreasonable schedule, invisible purpose) go untouched. They come back to the same meat grinder.
Order of sacrifice (Leaders Eat Last): "This week I cut my own two meetings to sit in and clear bugs with you; I'll write the status report so you can stay focused." You give up your own thing first.
Rebuild task cohesion: "Let's align — why is this release worth this push? When it's done I commit two comp days for the whole team. Whatever's riskiest, we look at together — nobody carries it alone."
(After the release) then celebrate. Celebration is the glue that follows winning together — not a substitute for it.
"Hold on — let's fully nail the root cause before we decide whether to roll back." Information is rising, but so are the losses. By the time you're 100% sure, the users are gone.
Stop the bleeding with a good-enough move: "Info's at 50% — that's enough. Roll back to the last stable version to stop the bleeding — it's reversible, cost is bounded. Investigate root cause in parallel."
Iterate fast: "We sync the situation every 15 minutes and adjust the call as new info lands. Don't aim to get it right in one shot — aim to spin faster than the incident."
Orient on experience: "Last time this kind of latency spike was a maxed-out connection pool — check that hypothesis first, switch instantly if wrong." (recognition-primed decision)