"Sure, I'll write it." — you've just put your name on a number you can't audit.
"I can't verify that number." — pure resistance with no substitute; it reads as "not supportive of AI," and the slide still gets written, just by someone else.
"I'll write two lines. Funnel ledger: 6 AI use cases, 2 in production, 40 daily users. P&L ledger: those 2 cut average code-review wait from 11 hours to 6; over the same period cloud and licence costs rose about $8k/quarter.
I can claim cycle-time reduction; I can't claim labour savings — nobody left and nobody was reassigned. If this slide needs the 30%, let me define how it's calculated, so next year we can report the same number the same way."
"Understood, I'll write 30%. I'll also add one line of definition in the appendix: derived from cycle-time reduction in 2 production use cases; excludes added cost and headcount change." — leadership wants a narrative; you don't have to fight it. All you need is one line you can point to later.
Action: list every AI use case you own in two columns: funnel status on the left, the matching financial line on the right. Leave the right column blank where it is blank — don't fill it with a story.
Reflection: if you could report exactly one number to the CFO, which would it be? Why isn't it the one you're reporting now?
"I'm not sure those 12 headcount were really saved." You're not challenging a data point, you're challenging someone's promotion packet. The outcome is fixed: the number doesn't change, you get filed under "has reservations about AI," and you stop receiving information from that line.
"I'm for doubling it. To deliver this one cleanly I'd like to reuse last year's definition: how were those 12 headcount calculated? Hours priced at whose rate, and was the window the full year or the peak month? I'll write this year's target on the same basis so the two years' numbers don't contradict each other." Once the method is on the table, the shortfall surfaces on its own — and not because you pointed at it. That's the only low-cost path.
"On last year's basis, this year's target = 8 FTE-equivalent hours removed; settled at the end of Q3, and if the actual is only 3, I'll report it in that same email and cut the Q4 budget." — writing down in advance what failure will look like is the only way to make the settlement actually happen.
Action: find the most recent public "AI saved us X" in your organization and ask its author for the method — method only, no assessment.
Reflection: if your team underdelivers this year, in what forum, to whom, and in what words would you volunteer it? If you can't name the forum, you're rolling too.
"That number isn't credible." — you have no evidence and an obvious motive (it looks like defending your own roadmap).
"I want to do it, but I need to align on method first, or my 4.2x won't be the same object as theirs. I'll ask Marco six things: does the denominator include eval and rework hours; is the baseline 'nothing' or the existing lint rules; is the 4.2x over all 12 use cases or the 3 that survived; hours at whose rate; peak-month annualized or measured across the year; is the data a survey or cycle time. Then I'll give you an estimate on the same basis — possibly still very good, only this time auditable."
"I'd like to copy your method for that 4.2x rather than invent my own and end up with clashing numbers. How did you set the denominator and the window?" — "I want to copy your method" is the least adversarial way to request a methodology: it puts you in the learner's seat, and to make his own number hold up he'll volunteer the details.
Action: take one AI benefit number you currently cite and mark where each of the six dials sits. See which ones are generous.
Reflection: if every AI proposal had to use strict settings, which of your projects dies first? Should it?
"Expected to significantly improve test coverage and developer efficiency." No definition, no threshold, no exit — the only possible ending is narrative: at quarter end it's judged on mood, and usually judged a success, because admitting failure has no upside.
Metric ① merged AI-generated tests ≥ 25% of new tests (denominator includes rejected ones).
Metric ② those tests have caught ≥ 3 defects that would have reached production (per revert / hotfix records, not self-report).
Review date: last week of Q4, already fixed.
Both met → request expansion to 10 people. One met → hold at 3, no expansion. Neither → stop in Q1, and I write the conclusion in the same document.
Full cost side: 3 person-quarters + roughly $6k inference and tooling + roughly 15% extra review hours.
"I didn't write an ROI multiple, because any multiple I write today is a definition I chose myself. What I wrote is thresholds and exit conditions — so at the end of Q4 there's nothing to negotiate, you just read the table." In an environment where everyone reports multiples, criteria that can convict you are extremely scarce credibility.
Action: add a six-line criteria block (metric / definition / threshold / date / action if met / action if missed) to your largest AI investment, and send it to the sponsor for a one-line confirmation.
Reflection: do you have a project that is already past the point where it should have been stopped? Is it still alive because of evidence, or because nobody wants to declare the settlement?