MIT researchers found that about 95% of corporate generative-AI pilots deliver no measurable P&L impact. Vendors read that as an adoption problem. We read it as a measurement problem: most pilots fail because nobody defined what success would cost or pay before the spend started.
AI deserves the same discipline you would apply to a new baler, a new truck, or a new hire. Here is what that discipline looks like before a dollar leaves the building.
The five numbers to establish first
1. The baseline. Whatever the tool is supposed to improve, measure it now, for at least a month, honestly. Hours per invoice processed. Quotes turned around per day. First-call resolution rate. Collection days. If you cannot state the baseline, you have already guaranteed the pilot cannot prove anything, because there will be nothing to compare against.
2. The fully loaded cost. License fees are the visible fraction. Add implementation time, the internal hours of the people configuring and correcting it, training, and the ongoing review of outputs. A $500-a-month tool that consumes ten hours a week of a $90K employee is a $2,800-a-month tool.
3. The dollar value of the improvement. Convert the operational metric to money before the pilot, not after. If quote turnaround drops from three days to one, what does that do to win rate, and what is a point of win rate worth? If the answer is "we're not sure," that is not a reason to skip the estimate. It is the estimate to make. A pilot whose best case is $15K a year should not consume $40K of attention.
4. The owner of the number. One person accountable for the metric moving, with the authority to change the process around the tool. Tools do not own outcomes. In every failed pilot we have examined, the tool had a champion but the number had no owner.
5. The kill criteria. Decide in advance what result, by what date, ends the experiment. Ninety days is usually right. Without a kill date, weak pilots survive on sunk cost and optimism, and the real price is the attention they keep consuming.
Where the returns actually are
For the operations-heavy businesses we serve, the payoffs cluster in unglamorous places: document handling (invoices, tickets, PODs, contracts), quote and proposal drafting, collections correspondence, customer-service triage, and summarizing the operational data you already generate. The pattern is consistent. AI pays fastest where the work is high-volume, text-shaped, and currently done by someone whose time has a better use.
Notice what is not on that list: anything customer-facing without review, anything safety-critical, and anything where an error compounds silently. Those need governance before they need software, which is its own discipline.
The 90-day shape of a pilot that can actually succeed
Month one is baseline measured, cost loaded, value estimated, owner named, kill criteria written. Month two is the tool in production on real volume, with the owner adjusting the process weekly. In month three you read the number against the baseline and decide: scale it, fix it, or kill it, on the criteria you wrote when you were sober about it.
That is it. No transformation roadmap, no steering committee. One process, one number, one owner, one decision.
Finance's job isn't just to report the score. It's to help design the operating system that creates it. That applies to AI spend exactly as it applies to pricing or throughput: the businesses that get returns are the ones that measure like owners instead of buying like fans.
Want to know if your business is set up to get a return before you spend? The free AI-Readiness Scorecard takes two minutes and tells you where the foundations are missing.