Research
Research

I Traced the 95%, the 80%, and the 40% to Their Sources. One of Them Doesn't Have One.

Aug 15, 2026 · 7 min read · by Jordan Kwan

TL;DR: The three stats measure three different things. MIT's 95% is the share of surveyed organizations reporting no measurable P&L impact from generative AI roughly six months after a pilot, off 52 interviews. Gartner's 40% is a prediction about agentic AI cancellations by end-2027, published with no method. The "80% of AI projects fail" figure has no primary source at all: it traces through RAND's 2024 report, which hedges it as "by some estimates," to a 2022 Fortune column citing unnamed "recent surveys." I took the first 10 search results quoting each stat on August 15, 2026, and 9 of 30 restated their own stat accurately.

Every "AI is failing" argument runs on one of three numbers, swapped mid-paragraph as though they were readings of the same instrument. Two of them are not even measurements.

What do the three stats actually measure?

I fetched each primary document. What is actually in them:

Stat Primary source What it is
95% MIT Project NANDA, The GenAI Divide Preliminary measurement: organizations getting zero return on generative AI
80% None found Folklore
40% Gartner, June 25, 2025 Prediction: agentic projects canceled by end-2027

The MIT report's own words are "95% of organizations are getting zero return," off 52 structured interviews and 153 surveys collected at four conferences. Its methodology defines success as "deployment beyond pilot phase with measurable KPIs," with "ROI impact measured 6 months post-pilot." Its canonical URL at nanda.media.mit.edu/ai_report_2025.pdf still redirects to a generic Media Lab overview page, which I re-confirmed today. This site rated the 95% separately, and nothing here changes that verdict.

Gartner's is one line on a public page: "Over 40% of agentic AI projects will be canceled by the end of 2027, due to escalating costs, unclear business value or inadequate risk controls." No survey sits behind it. The January 2025 poll of 3,412 webinar attendees appearing two paragraphs later measured investment posture, not cancellations. Full rating here.

Where does the 80% actually come from?

Nowhere. This is the finding.

The figure is usually credited to Gartner, and no Gartner publication says it. The closest is a February 2018 prediction that "through 2022, 85 percent of AI projects will deliver erroneous outcomes", which is about output quality, not project death.

What the citers point to is RAND. The Root Causes of Failure for Artificial Intelligence Projects, by James Ryseff, Brandon De Bruhl and Sydne Newberry, published August 13, 2024, contains the sentence: "By some estimates, more than 80 percent of AI projects fail." RAND was honest. It wrote "by some estimates," never claimed to have measured the rate, and spent its 65 interviews on why projects fail, not how many.

So I followed RAND's footnote. It cites Jeremy Kahn, "Want Your Company's A.I. Project to Succeed? Don't Hand It to the Data Scientists, Says This CEO," Fortune, July 26, 2022, a profile of an AI vendor's CEO. The relevant line: "That's borne out in a slew of recent surveys, where business leaders have put the failure rate of A.I. projects at between 83% and 92%."

No survey is named. The range is not 80%. And what those unnamed surveys captured was business leaders' perception of a failure rate, not project outcomes. The chain ends there: a hedged range, in a vendor profile, sourced to nothing, rounded down by RAND and hardened into a measurement by everyone downstream.

How did I run the count?

On August 15, 2026 I ran three DuckDuckGo queries: "95% of AI pilots fail", "80% of AI projects fail", "40% of agentic AI projects" canceled 2027. I took the first 10 organic results for each, excluding the Gartner release from the third. All 30 resolved, four only through a browser after returning 401 or 403 to a plain fetch.

I scored each article's fullest statement of the stat against three tests, accurate only if it passes all three:

  1. Source. Does it name the actual origin, and represent it as what it is?
  2. Population. Right denominator: generative AI or all AI, organizations or pilots?
  3. Outcome. Right verb: "zero measurable return" is not "fail," and "canceled" is not "failed."

What did the count show?

Stat Body accurate Headline accurate Main distortion
95% (MIT) 2 of 10 0 of 10 Drops "generative" (8 of 10)
80% (unsourced) 1 of 10 0 of 10 Credits RAND as measurer (6 of 10)
40% (Gartner) 6 of 10 8 of 10 "canceled" to "fail" (3 of 10)
Total 9 of 30 8 of 30

Specifics, because unnamed subjects are uncitable. On the 95%, Fortune's own follow-up still says the study "interviewed 150 executives, surveyed 350 employees," against the report's actual 52 and 153.

On the 80%, the misattribution is not the one I expected. Zero of the ten credited Gartner. Six credited RAND as though RAND had produced the number. Tom's Hardware says "according to research by the RAND Corporation, over 80% of these AI projects will fail" and swaps RAND's comparator, non-AI IT projects, for "non-AI technology-related startups." One claims five separate studies converge on 80%. The University of Queensland Business School sources it to an academic's book. And CIO's "Why 80% of AI projects fail" contains the number exactly once, in its own headline; the body never mentions it. The one article that passed said outright that RAND's study is qualitative and the figure is "best read as 'the large majority fail,' not a precise percentage."

On the 40%, Reuters and both Forbes contributors reproduced it cleanly. The four failures: three canceled-to-failed verb swaps, and one article that invented a provenance by describing the 3,412-attendee webinar poll as "the survey" behind the 40%.

Why does the weakest stat get quoted the most accurately?

Because accuracy tracks how easy the source is to open, not how good it is.

Gartner's 40% is the weakest of the three as evidence: a forecast with no published math about a category that barely existed when it was written. It still scored three times better than MIT's, because it is one unambiguous sentence on a free public page that a writer can copy without making a decision. MIT's 95% measures a narrow thing, and its paper has to be hunted through mirrors, so writers quote the headline they read somewhere else. The 80% has no page to open at all, which is why it mutates freely: there is nothing to check it against.

Same pattern as the AI SDR churn figure that traces to an article which does not contain it. A stat detaches from its referent, the hedge is stripped, and the number outlives its sentence.

So which one is true?

None of them supports "most AI projects fail," which is the sentence all three are used to prove.

The 40% cannot be true or false yet: a prediction about 2027 with no disclosed method, best read as a base rate rather than a verdict. The 80% is not a finding and should not be quoted. The 95% is the only one anchored to a real measurement, and its honest restatement is too narrow to carry a doom headline: 95% of organizations in a convenience sample of 52 interviews reported no measurable P&L impact from generative AI within about six months of piloting it.

What does this not prove?

Ten results per query on one engine on one day is a snapshot, and a different week reshuffles the articles. My three tests are strict by design: an article saying "95% of AI pilots fail" while linking the MIT report is compressing rather than lying, and a looser scorer would pass several. That raises 9 of 30. It does not reverse the ranking, which is the part that matters.

None of this shows enterprise AI is going well. RAND's interviews, the MIT report's direction and Gartner's forecast all point the same way, and the case that a lot of AI spending produces nothing measurable is strong. It is just not made by any of these three numbers, and the one with no source is quoted hardest.

Written by Jordan Kwan, founder of Reachium.

I build Reachium, the LinkedIn outreach platform behind the tactics you just read. Same brain, live product.

See what Reachium does ↗