The headline numbers on AI return contradict each other because they average two very different populations: a small group capturing real, measurable value, and a large group spending to look modern. Pool them and you get a muddle. Separate them and a clear pattern appears, and it has almost nothing to do with which model anyone picked.
The numbers only contradict each other if you pool them
Look at the two most-cited 2025 studies side by side. An MIT analysis of enterprise generative AI found that despite $30 billion to $40 billion of investment, about 95 percent of organizations were getting no measurable return, while roughly 5 percent were extracting real value (an interview-based figure its authors call only directionally accurate). BCG, surveying 1,250 executives, landed in nearly the same place from a different angle: only 5 percent of companies were capturing AI value at scale, and a full 60 percent reported no material value at all. These are not the same metric, so do not read the matching 5 percent as one statistic; read it as two independent instruments pointing at the same shape. And the shape is not “AI does not work.” McKinsey’s late-2025 survey found 88 percent of organizations now use AI in at least one function, up from 78 percent a year earlier. Usage is nearly universal. It is the value that is concentrated: in that same survey, only around 39 percent could attribute any enterprise-level impact on profit to AI, and most of those put it below five percent. The contradiction dissolves the moment you separate who is using AI from who is getting paid for it.
They picked a process, not a capability
The winners did not “adopt AI.” They picked one expensive, repetitive process and rebuilt it. The MIT study is precise about where the losses pile up: general assistants like ChatGPT and Copilot are everywhere (more than 80 percent of firms have tried them) and they do help individuals, but they rarely move the P&L, because they do not adapt to a specific workflow. Custom, process-specific systems are where the money is, and they are a funnel: 60 percent of organizations evaluated one, only 20 percent reached a pilot, and just 5 percent reached production. BCG names the same failure from the executive seat: too many companies experiment too widely, spreading effort across scores of use cases instead of focusing end to end on a few important functions. Breadth feels like progress. Depth is what pays.
They redesigned the workflow, not just bolted on a model
This is the habit the data singles out most sharply. McKinsey, dissecting what its high performers do differently, found that fundamentally redesigning workflows has the single biggest effect on whether a company sees any earnings impact from generative AI, and that its high performers (about 6 percent of respondents) are roughly three times more likely to have done it. Bolting a model onto an unchanged process buys you a faster version of the old bottleneck. Rebuilding the process around the model, with a human at the edge, is what moves the number.
Return on AI is mostly return on process discipline. The model is necessary; it is almost never the thing that was missing.
They measured the before, so they could prove the after
You cannot show return on a baseline you never recorded. The disciplined teams wrote down the cost of the old process, the hours, the error rate, the cycle time, before they changed anything; the MIT researchers defined success itself as deployment beyond the pilot with measurable KPIs, checked six months later. The undisciplined teams are now reconstructing a baseline from memory to justify a renewal, which fools no one in finance. If you cannot state the before, you cannot claim the after, and “it feels faster” is not a number a CFO will fund twice.
They treated the model as a swappable input
The counterintuitive habit: the teams with real return spend most of their effort on the workflow, the data, and the evaluation, and treat the model as a commodity they can swap. They also tend to buy rather than build the AI layer itself; the MIT data shows partnerships with specialized vendors reach deployment about twice as often as in-house builds. The common thread is the same: they judge tools on business outcomes, not on benchmark scores, and they refuse to make the model the protagonist. The teams chasing a bigger model as the answer are optimizing the one variable the evidence says matters least.
What this looks like at population scale
Zoom out and the concentration is visible even in government data. The US Census Bureau finds only about 18 percent of firms use AI in any business function, and 57 percent of those that do confine it to three functions or fewer; adoption skews hard to large firms (37 percent among those with 250 or more employees, versus under 20 percent of the smallest). The Federal Reserve makes the gap vivid: weight by firm and adoption sits around 18 percent, weight by employment and it runs far higher, because the big players move first. So if you are hunting for return in your own numbers, run the test backwards. Name one process, measure it honestly, rebuild it with a human gate, and keep the model swappable. If you cannot name the process and the baseline, you do not have an AI strategy. You have an AI budget, and the difference shows up two quarters later, in exactly the place everyone is now looking.