Day 71 · 2026.07.30

Two Sets of Books on the AI Budget: Funding This Round on Savings That Never Arrived

Topic: AI Investment & ROI Discipline·4 principles
"For a successful technology, reality must take precedence over public relations, for nature cannot be fooled." — Richard P. Feynman, 1986
How this differs from earlier issues: Day 50 covered metric traps in general; Day 54 covered managing an AI-augmented team. This one is only about money: when the case for this round of AI investment rests on savings the last round promised and never booked. Four moves: keep the two ledgers apart; recognize the expectation rollover; the six dials behind any ROI; set the bar before you fund it.
PRINCIPLE 01

Keep the funnel ledger and the P&L ledger apart Shipped Is Not the Same as Saved

paired indicatorsshipped ≠ savedwho signs off
An AI project has two kinds of "success": it shipped and people use it (the funnel ledger), and a financial line item actually went down (the P&L ledger). Most reporting presents the first as the second, and not out of malice — the funnel ledger moves fast and is visible, while the P&L ledger takes quarters and requires finance to agree. Once the two are merged onto one slide, they can never be separated again.
"Because indicators direct one's activities, you should guard against overreacting. This you can do by pairing indicators, so that together both effect and counter-effect are measured." — Andy Grove, High Output Management, ch. 2 "Managing the Breakfast Factory"
Two ledgers: one project, two independent tracks
Funnel ledger · how far the use case got (fast, visible, easy to stack up)
Proposedsomeone wants it
POCdemo works
In prodreal users
A year onstill used
P&L ledger · how far the money got (slow, needs someone else to agree)
saving measured
signed off by someone other than you
baked into next year's baseline
cloud bill / contract / headcount actually down
Cell 3 of the funnel ledger routinely gets reported as cell 4 of the P&L ledger — the single most common misalignment in AI reporting.
Situation: before the quarterly review, your boss says: "Put 'AI made the team 30% more efficient' on the slide."
✗ Two common reactions, both bad

"Sure, I'll write it." — you've just put your name on a number you can't audit.
"I can't verify that number." — pure resistance with no substitute; it reads as "not supportive of AI," and the slide still gets written, just by someone else.

✓ Better: write two lines and split the ledgers on the spot

"I'll write two lines. Funnel ledger: 6 AI use cases, 2 in production, 40 daily users. P&L ledger: those 2 cut average code-review wait from 11 hours to 6; over the same period cloud and licence costs rose about $8k/quarter.
I can claim cycle-time reduction; I can't claim labour savings — nobody left and nobody was reassigned. If this slide needs the 30%, let me define how it's calculated, so next year we can report the same number the same way."

✓ If your boss says "just write 30%, don't overcomplicate it"

"Understood, I'll write 30%. I'll also add one line of definition in the appendix: derived from cycle-time reduction in 2 production use cases; excludes added cost and headcount change." — leadership wants a narrative; you don't have to fight it. All you need is one line you can point to later.

  • Was the AI benefit I last reported a funnel-ledger or a P&L-ledger number?
  • For the "saving" I claimed, which financial line (headcount / cloud bill / contract) actually fell in the same period?
  • Does my benefit number have a paired indicator measuring the counter-effect (added cost, extra review hours, defect rate)?
  • Treating POC count as progress: a POC is the cheapest form of success, which is exactly why it is easy to mass-produce. What's scarce is "shipped and still alive a year later."
  • Jumping from "time saved" to "cost saved": hours freed that don't convert into fewer people or more output become slack. Slack has value, but it cannot be booked in the P&L ledger — booking it promises the organization cash that does not exist.

Action: list every AI use case you own in two columns: funnel status on the left, the matching financial line on the right. Leave the right column blank where it is blank — don't fill it with a story.

Reflection: if you could report exactly one number to the CFO, which would it be? Why isn't it the one you're reporting now?

PRINCIPLE 02

Funding round N+1 with savings round N never delivered The Expectation Rollover

Bain 2026rolloverask for the method, not the truth
Bain's survey of 951 companies with over $100M in revenue found that roughly 44% of large enterprises justify their next round of AI investment with savings from earlier automation projects — savings that systematically underdelivered; and within the underdelivering subgroup, about nine in ten are still increasing spend. This isn't collective stupidity, it's a structure: the shortfall from the last round is never settled, it gets folded into the assumptions of the next one. It doesn't disappear; it rolls forward.
"For a successful technology, reality must take precedence over public relations, for nature cannot be fooled." — Richard P. Feynman, Report of the Presidential Commission on the Space Shuttle Challenger Accident, Appendix F (1986)
Situation: in the budget meeting your VP says, "Last year's CI automation saved 12 headcount, so we're doubling the AI platform budget." You know those 12 never vanished from any org chart.
✗ Don't say

"I'm not sure those 12 headcount were really saved." You're not challenging a data point, you're challenging someone's promotion packet. The outcome is fixed: the number doesn't change, you get filed under "has reservations about AI," and you stop receiving information from that line.

✓ Better: ask for the method, not the truth

"I'm for doubling it. To deliver this one cleanly I'd like to reuse last year's definition: how were those 12 headcount calculated? Hours priced at whose rate, and was the window the full year or the peak month? I'll write this year's target on the same basis so the two years' numbers don't contradict each other." Once the method is on the table, the shortfall surfaces on its own — and not because you pointed at it. That's the only low-cost path.

✓ Step two (put it in your own proposal)

"On last year's basis, this year's target = 8 FTE-equivalent hours removed; settled at the end of Q3, and if the actual is only 3, I'll report it in that same email and cut the Q4 budget." — writing down in advance what failure will look like is the only way to make the settlement actually happen.

  • Can I find the original deck for the savings the last round promised? Was it ever settled?
  • Is the money I'm asking for justified by new evidence, or by an unsettled promise from the last round?
  • Does my own proposal state what counts as failure?
  • Questioning a number's honesty in public: it converts a technical problem into a personal one, and the cost far outweighs the gain. Same subject, two postures: asking about method is collaboration; asking about truth is audit.
  • The inverse error: refusing to invest because the method looks generous. Underdelivery doesn't mean no value — the use cases genuinely in use are genuinely in use. The difference is one thing only: don't let them serve as collateral for the next round.

Action: find the most recent public "AI saved us X" in your organization and ask its author for the method — method only, no assessment.

Reflection: if your team underdelivers this year, in what forum, to whom, and in what words would you volunteer it? If you can't name the forum, you're rolling too.

PRINCIPLE 03

The six dials behind any ROI number Ask for the Dial Settings Before the Value

Goodhartinterrogation listself-reported vs measured
For the same project, legitimately turning six dials moves the ROI from 0.4x to 5x without a single false statement. So when you hear an ROI number, ask for the dial settings before you ask for the value — the value is a conclusion; the dials are the content.
"Any observed statistical regularity will tend to collapse once pressure is placed upon it for control purposes." — Charles Goodhart, "Problems of Monetary Management: The U.K. Experience" (1975), later generalized as Goodhart's Law
Six dials: two legitimate ways to compute the same project
Dial
Generous (flattering)
Strict (auditable)
Denominator
tokens and licences only
plus rework, eval sets, migration and ops hours
Baseline
compared with doing nothing
compared with a serious non-AI alternative
Inclusion
only the surviving use cases
denominator includes every attempt, including the dead
Labour rate
hours priced at average or senior rates
priced at the actual owner's rate, counting only what's recoverable
Window
best month × 12
measured across the full year, learning curve and regression included
Source
asking users "how much time do you feel you saved?"
measured cycle time / throughput / defect rate
All six turned generous can legitimately report an unprofitable project at 5x. Interrogation order: source → denominator → inclusion.
Situation: a peer team announces 4.2x ROI on AI code review. Your skip-level asks you: "Why aren't you doing this?"
✗ Don't say

"That number isn't credible." — you have no evidence and an obvious motive (it looks like defending your own roadmap).

✓ Better: I want to do it, but let's align the method first

"I want to do it, but I need to align on method first, or my 4.2x won't be the same object as theirs. I'll ask Marco six things: does the denominator include eval and rework hours; is the baseline 'nothing' or the existing lint rules; is the 4.2x over all 12 use cases or the 3 that survived; hours at whose rate; peak-month annualized or measured across the year; is the data a survey or cycle time. Then I'll give you an estimate on the same basis — possibly still very good, only this time auditable."

✓ To Marco (peer to peer — must not read as an audit)

"I'd like to copy your method for that 4.2x rather than invent my own and end up with clashing numbers. How did you set the denominator and the window?" — "I want to copy your method" is the least adversarial way to request a methodology: it puts you in the learner's seat, and to make his own number hold up he'll volunteer the details.

  • Does the denominator include rework and eval-set hours?
  • Is the baseline "doing nothing," or a serious non-AI alternative?
  • Is the success-rate denominator every attempt, or only the live use cases?
  • Are the hours saved self-reported or measured? Is the window peak-annualized or full-year?
  • Interrogating others but not yourself: your own proposal is probably on generous settings too, because generous settings get approved. Audit yourself first and the questions carry standing.
  • A note on gender: once ROI depends on self-report (dial 6), it becomes a self-promotion contest. Exley and Kessler found that at nearly identical objective performance, women describe their own performance to others systematically less favourably than men do (Christine L. Exley & Judd B. Kessler, "The Gender Gap in Self-Promotion," Quarterly Journal of Economics 137(3), 2022). And self-report is the default in AI benefit accounting. Two things you can do: move the headline benefits to measurement; and when you must collect self-reports, supply a fixed template and unit ("which specific actions did you not have to do this week?") instead of free text.

Action: take one AI benefit number you currently cite and mark where each of the six dials sits. See which ones are generous.

Reflection: if every AI proposal had to use strict settings, which of your projects dies first? Should it?

PRINCIPLE 04

Decide which definition counts as success before you fund it Set the Bar Before You Fund It

Klein premortemexit conditionssix-line bar
Arguing after the fact about whether an AI project succeeded never resolves — because the criteria get picked backwards out of the result. The only fix is to freeze the criteria at the moment you approve the money: which metric, on what definition, at what threshold, reviewed on what date, and what happens if it misses. An investment with no exit condition never fails; it just extends forever.
"Unlike a typical critiquing session, in which project team members are asked what might go wrong, the premortem operates on the assumption that the 'patient' has died." — Gary Klein, "Performing a Project Premortem," Harvard Business Review, September 2007
Situation: you're requesting 3 people for a quarter to build AI-assisted test generation.
✗ How it usually gets written

"Expected to significantly improve test coverage and developer efficiency." No definition, no threshold, no exit — the only possible ending is narrative: at quarter end it's judged on mood, and usually judged a success, because admitting failure has no upside.

✓ Better: the criteria section of the proposal (six lines, frozen)

Metric ① merged AI-generated tests ≥ 25% of new tests (denominator includes rejected ones).
Metric ② those tests have caught ≥ 3 defects that would have reached production (per revert / hotfix records, not self-report).
Review date: last week of Q4, already fixed.
Both met → request expansion to 10 people. One met → hold at 3, no expansion. Neither → stop in Q1, and I write the conclusion in the same document.
Full cost side: 3 person-quarters + roughly $6k inference and tooling + roughly 15% extra review hours.

✓ One line when you present it

"I didn't write an ROI multiple, because any multiple I write today is a definition I chose myself. What I wrote is thresholds and exit conditions — so at the end of Q4 there's nothing to negotiate, you just read the table." In an environment where everyone reports multiples, criteria that can convict you are extremely scarce credibility.

  • Does the proposal contain a specific number for "this much counts as failure"?
  • Are the criteria measured or self-reported? Is the miss action written into the proposal, and who executes it?
  • Is the cost side complete (rework, eval sets, extra review hours, inference)?
  • Is the review date fixed, or is it "when we get around to it"?
  • Criteria written too soft: "continue if we see positive signal" is the same as writing nothing. Only criteria that can convict are criteria.
  • Success criteria with no exit condition: the project survives as "give it one more quarter" — the most expensive form of waste in AI investment, because it never spends a failure while permanently occupying your best people.
  • A trade-off worth admitting: strict criteria will block some exploration that genuinely deserves to happen. The honest move is to label it "learning investment, no ROI committed, capped at X person-months" rather than inventing an ROI to get it through review. Giving exploration the right name protects it better than giving it a flattering number.

Action: add a six-line criteria block (metric / definition / threshold / date / action if met / action if missed) to your largest AI investment, and send it to the sponsor for a one-line confirmation.

Reflection: do you have a project that is already past the point where it should have been stopped? Is it still alive because of evidence, or because nobody wants to declare the settlement?

Deeper questions

If I insist on strict settings while everyone else uses generous ones, don't I just lose?
Short-term, yes, and that shouldn't be glossed over. The workable middle path is two lines side by side: report externally on the organization's prevailing definition (otherwise you're minting a currency nobody accepts), and add a second line on strict settings, noting which dial accounts for the gap. Two years later, when generous numbers start attracting accountability, you'll be the only one holding credible historical data. The dividing line: strict settings belong on large, long-horizon investments; running full accounting on a three-week pilot isn't worth it.
Won't "learning investment" become a shield against any accountability?
It can, so give it three hard constraints: a spend cap, a time cap, and one non-ROI verifiable output (an eval set that changes the next decision, a technical hypothesis actually falsified). Its criterion isn't whether it saved money, it's this: what do we now know that we didn't three months ago, and which specific decision did it change? No answer means it wasn't learning, it was delay.
Does any of this hold at a startup?
Partly, with different weights. Startup AI spend is often not about saving money but about not dying, and strict ROI gates would kill correct bets; but the two ledgers must still stay apart — one of the most common ways startups die is recording "the demo was stunning" as "the product is useful," then hiring and raising against the latter. What scale really changes is the settlement cycle: quarterly at a big company, weekly at a small one — and at a small one the person settling is usually the person who proposed it.
Leadership knows the numbers are on generous settings. Why report them anyway?
Because in capital markets and internal competition, "having an AI strategy" carries value independent of ROI: it moves valuation, budget allocation, and who gets seen as the future. This isn't pure deception; it's two real objectives stacked onto one number. Understanding that saves you a great deal of futile correction — you can't change the external narrative, but you can defend the internal ledger. The line is: you may go along with the narrative externally; you must not use the narrative as an input internally. The moment it drives staffing and schedules, your team works real overtime for a benefit that doesn't exist.