Hype Index: '95% of AI Pilots Fail'
Aug 6, 2026 · 5 min read · by Jordan Kwan
TL;DR: The "95% of AI pilots fail" stat is real as a citation and weak as a measurement: it comes from a preliminary MIT working paper whose 95% rests on 52 convenience-sampled interviews and a definition of failure that means "no measurable ROI within six months," not "the project died." We score it 70% noise / 30% signal: enterprise AI pilots genuinely struggle, but the number doing the rounds does not mean what the people quoting it think it means.
You have heard this one in a boardroom by now. Ninety-five percent of AI pilots fail, MIT said so. It moved markets in August 2025, it has anchored a thousand vendor decks since, and in 2026 it is still the single most-quoted number in the AI bubble debate. Time to rate it.
What is the claim?
Stated the way it circulates:
An MIT study found that 95% of enterprise AI pilots fail.
The primary source is The GenAI Divide: State of AI in Business 2025, a preliminary working paper from MIT Project NANDA, July 2025. The sentence people are quoting reads: "95% of organizations are getting zero return," following a line about $30-40 billion in enterprise GenAI investment.
What did the report actually measure?
Three things, per its own methodology section: a review of 300+ publicly disclosed AI initiatives, structured interviews with people at 52 organizations, and surveys of 153 senior leaders, gathered around industry conferences. That is a convenience sample, and the report says so in its own way, conceding its figures are "directionally accurate based on individual interviews rather than official company reporting."
The definition is the load-bearing detail. "Success" meant deployment beyond pilot with measurable ROI impact roughly six months after the pilot. So "fail" does not mean the project was canceled or the tech did not work. It means nobody could show a P&L number within six months, a bar that most eighteen-month enterprise transformations would also miss.
Even the amplification was sloppy: Fortune's launch article described the methodology as 150 interviews and a 350-person survey, numbers that do not match the report's own 52 and 153, and secondhand posts still repeat Fortune's version today. When Wharton's Ethan Mollick got the report, his reaction was the whole audit in one sentence: "I am not sure how generalizable the findings are based on the methodology (52 interviews, convenience sampled, failed apparently means no sustained P&L impact within six months but no coding explanation)." Paul Roetzer of the Marketing AI Institute was blunter: "Please don't put any weight into this study. This is not a viable, statistically valid thing."
What is actually true?
The direction. Strip the 95% and the pile of independent evidence that enterprise AI pilots struggle is real:
- RAND's 2024 report opens with "by some estimates, more than 80 percent of AI projects fail," twice the failure rate of non-AI IT projects. Note the honesty of "by some estimates": RAND is citing prior figures, and its own contribution was interviewing 65 data scientists about why projects die.
- S&P Global's 2025 survey data found 42% of businesses scrapped most of their AI initiatives, up from 17% a year earlier, and the average organization abandoned 46% of proof-of-concepts before production.
- Gartner predicts over 40% of agentic AI projects will be canceled by end of 2027, a claim this series rates separately. Its analyst Anushree Verma's diagnosis: "Most agentic AI projects right now are early stage experiments or proof of concepts that are mostly driven by hype and are often misapplied."
Pilots stall, ROI is slow, and hype writes checks that deployments cannot cash in six months. That part of the story survives every audit, and it is why the agent demo economy looks the way it does.
What is not true?
That MIT measured 95% of AI pilots failing. The number is one preliminary paper's convenience-sampled estimate of a six-month ROI bar, and notice that 95, 80, 42, and 40 are not even measurements of the same thing: zero return in six months, an estimate of project failure, companies scrapping most initiatives, and a prediction about future cancellations. In 2026 coverage they get mashed into one interchangeable doom stat, which is how a zombie statistic feeds.
Here is my favorite part, and my original contribution to the pile. I pulled the top 20 accessible search results for the stat on August 5, 2026: 16 of 20 name the MIT report, but only 10 of 20 link to a primary copy. Four cite it entirely secondhand, usually via Fortune. And the canonical MIT URL for the report now redirects to a generic lab overview page: the most-quoted AI statistic on earth no longer has an official copy at its own source. A stat this load-bearing should not require an archive hunt to read.
What should you do instead?
Interrogate any AI failure stat with three questions: what was the denominator, what did "fail" mean, and over what window. Under that lens, "95% fail" becomes "most pilots could not show P&L impact in six months," which is a very different input to your planning. Budget for slow ROI, pick one pain point, and measure something specific, which is where the credible productivity research lands too. Even the report's lead author, Aditya Challapally, framed the successes that way: the winners "pick one pain point, execute well, and partner smartly."
Verdict
Enterprise AI pilots really do stall at rates that should embarrass the industry; the specific number everyone quotes is a preliminary estimate with a definition nobody reads, misreported by its own amplifiers, orphaned by its own publisher. Like the PhD-level reasoning claim, it went viral precisely because nobody checked.
Verdict: 70% noise / 30% signal. Quote the struggle, not the stat.
Written by Jordan Kwan, founder of Reachium.
I build Reachium, the LinkedIn outreach platform behind the tactics you just read. Same brain, live product.
See what Reachium does ↗