Three years into enterprise generative AI, the "we are still learning" grace period is over. Finance is asking what the training budget bought. The honest answer, based on the best available evidence, is that almost nobody can tell them, and the numbers people do collect are less reliable than they look.
Start with the uncomfortable baseline
McKinsey's State of AI 2025 surveyed 1,993 respondents across 105 countries. Only 39% attribute any EBIT impact at all to AI, and among those who do, most put it below 5% of enterprise earnings.
Deloitte's State of AI in the Enterprise 2026, covering 3,235 senior leaders in 24 countries, finds the gap in sharper relief: 74% hope AI will increase revenue, 20% say it has. Meanwhile 66% report productivity or efficiency gains. Saved hours are not automatically money, and the space between those two numbers is where most AI business cases quietly fail.
The most honest number I found is from ISACA's 2026 AI Pulse Poll of 3,400 professionals. 22% say ROI met or exceeded expectations. But 23% say it is too early to tell and another 22% simply do not know. Nearly half the room is saying "we cannot answer that." That, not the 22%, is the headline.
The mechanical reason is unglamorous. As recently as mid-2024, McKinsey found fewer than one in five organisations tracking well-defined KPIs for their generative AI work. You cannot report a return you never instrumented.
Why the numbers you do collect are suspect
This is the finding that should change how you run evaluations. In a randomised trial by METR, experienced open-source developers were 19% slower when using AI tools. Afterwards, the same developers estimated AI had made them 20% faster. They had predicted a 24% speedup going in.
A 39-point gap between believed and measured performance. If your AI training evaluation is a post-course survey asking people whether it made them more productive, that is the size of error you may be reporting to your board.
Be fair to the study: 16 developers, working in codebases they knew intimately, which is the setting least favourable to AI. And in February 2026 METR published a redesign note saying their follow-up was compromised by selection effects, with confidence intervals crossing zero. They published evidence against their own headline finding, which is more than most researchers manage. The lesson survives regardless: self-reported productivity is not evidence.
Where the credible gains actually are
The strongest evidence in this entire field is peer-reviewed and consistent on one point. In Generative AI at Work, published in the Quarterly Journal of Economics, roughly 5,200 customer support agents saw a 14% average productivity gain, rising to 34% for novice and low-skilled workers.
The same shape appears in software. Three randomised trials across 4,867 developers at Microsoft, Accenture and a Fortune 100 manufacturer found junior developers gaining 21% to 40% while seniors gained 7% to 16%. Short-tenure staff gained 27% to 39%, long-tenure 8% to 13%.
Two different industries, different research teams, same pattern: the return concentrates in your least experienced people. Most AI enablement budgets are spent in exactly the opposite direction, on senior staff and leadership awareness sessions. Worth noting that two of the three software firms are AI vendors or major AI partners, and that all of this measures output volume rather than quality.
The gap nobody is filling
Here is what I could not find, after a genuinely thorough search: any credible published data on the ROI of AI training as distinct from AI tooling. Every number available measures what happened when a tool was deployed. None isolates what the enablement contributed.
That is a real hole in the evidence base, and anyone claiming otherwise is quoting a vendor. It also means the honest position for an L&D leader right now is not "here is our proven return." It is "here is the measurement design that will let us answer this in two quarters."
What a defensible measurement design looks like
Pick one workflow, not a portfolio. Measure the output before training, using something the business already counts rather than something you invented. Split by experience level, because the evidence says that is where the variance lives. Do not use a self-report survey as your primary measure; use it only to explain what the objective numbers show. And set the review date before you start, so the answer is not negotiated afterwards.
That is slower than a satisfaction score. It is also the only version that survives a finance review.
On sourcing. Every figure links to its source. Two famous statistics are deliberately absent. The claim that 95% of generative AI pilots fail is everywhere right now, but the underlying MIT NANDA document was not retrievable, so every version circulating is second-hand. And Gartner's prediction that 30% of projects would be abandoned by end of 2025 was an analyst forecast with no stated survey basis, about a period that has now passed without published verification. Neither belongs in a board paper.