The 95 percent number, and what it actually says
MIT's NANDA researchers found that 95 percent of enterprise generative AI pilots produce no measurable return. Read closely, that is not a verdict on the technology.
In July 2025, MIT's Project NANDA published a report called The GenAI Divide: State of AI in Business 2025. One number from it travelled everywhere: roughly 95 percent of enterprise generative AI pilots produce no measurable return on the profit and loss statement. Only about 5 percent make it into production with results anyone can point at.
The number came from a multi method study: a systematic review of more than 300 publicly disclosed AI initiatives, plus structured interviews and surveys, with a research period from January to June 2025. Against that sits an estimated 30 to 40 billion dollars of enterprise spending on generative AI. So the honest summary is not that the tools do not work. It is that the money went in faster than the organizations could absorb it.
I keep quoting this number in rooms in Belgrade, and it always lands the same way. Someone relaxes, because their own stalled initiative is suddenly normal rather than embarrassing. That relief is useful for about a minute. Then the question becomes what the other 5 percent did differently.
The report's own answer is not about model choice. The successful cases tended to be narrow, embedded in one workflow, and built with the people who do that work rather than presented to them. The failing cases tended to be broad, tool led, and owned by whoever bought the licences. That matches what I see in practice with a precision that is almost boring.
Two cautions about the statistic itself, because I would rather you use it well than repeat it loudly. First, it measures pilots, and a pilot with no P&L effect is not automatically a failed experiment: some of them are cheap ways to learn that a task is not worth automating yet. Second, it is preliminary research on publicly disclosed initiatives, so companies that quietly succeed and companies that quietly fail are both underrepresented. The direction is credible. The decimal point is not the point.
What the number is genuinely good for is budget conversations. If 95 percent of pilots return nothing, then the default assumption for your next pilot should be that it returns nothing, and the interesting question is how cheaply you can find that out. That reframes the request from a platform purchase to a two week experiment with a named owner and a number to hit.
Here is the version I now use. Pick one repeated task. Write down what it costs today in minutes or in items per day. Rebuild it with the tooling inside the loop and the people who own it in the room. Re measure after four weeks. If it did not move, stop and say so out loud, because the 5 percent is not made of companies with better instincts, it is made of companies that were willing to kill things.
The gap the report calls a divide is not really between companies with AI and companies without it. It is between companies that can change a process and companies that can only buy software. That was true before generative AI existed. This is just the most expensive demonstration of it we have had so far.
If you have a pilot that stalled, I would genuinely like to hear where it stopped. The failures are where the useful detail lives, and almost nobody publishes them.
Reply
If any of this is wrong, or right in a way you can add to, I would rather hear it. Write to antanaskoviczarko@gmail.com or find me on LinkedIn.