Do 95% of AI Projects Fail? What the MIT Study Actually Measured

Do 95% of AI Projects Fail? What the MIT Study Actually Measured

The claim that 95% of AI projects fail overstates what the MIT NANDA study established. Its preliminary 2025 report found that roughly 95% of organizations in its evidence base had not achieved measurable P&L impact from generative AI initiatives within the study’s criteria and time window. It did not show that 95% of all AI projects are technically broken or permanently unsuccessful.

AI

The claim that 95% of AI projects fail overstates what the MIT NANDA study established. Its preliminary 2025 report found that roughly 95% of organizations in its evidence base had not achieved measurable P&L impact from generative AI initiatives within the study’s criteria and time window. It did not show that 95% of all AI projects are technically broken or permanently unsuccessful.

Key findings

• The source was a preliminary July 2025 report, not a peer-reviewed census of all enterprise AI projects.

• Its evidence base included 300+ publicly disclosed initiatives, structured interviews with representatives from 52 organizations, and survey responses from 153 senior leaders at four industry conferences.

• The study defined success as deployment beyond pilot with measurable KPIs. It assessed ROI six months after the pilot, which can undercount longer enterprise implementations.

• The result concerned measurable P&L impact, not model accuracy, user satisfaction, time savings, or technical completion.

• Other evidence points in the same direction without proving the exact 95% rate: McKinsey found significant enterprise value concentrated among a small minority, and ONS found extensive use remained uncommon among adopters.

• The useful lesson is that tool access and pilot activity do not automatically create captured financial value.

Methodology and definitions

The primary source is The GenAI Divide: State of AI in Business 2025, produced in collaboration with MIT Project NANDA. The document describes itself as “Preliminary Findings from AI Implementation Research” and covers research conducted from January through June 2025.

Its multi-method design included:

• A systematic review of more than 300 publicly disclosed AI initiatives.

• Structured interviews with representatives from 52 organizations.

• Survey responses from 153 senior leaders collected at four major industry conferences.

The report’s methodology section defined success as deployment beyond pilot with measurable KPIs. ROI impact was measured six months after pilot and adjusted for department size. The report says bootstrap confidence intervals were used where applicable.

It also lists important limitations: potential selection bias, incomplete representation of enterprise segments and regions, inconsistent success metrics across organizations and industries, reliance on interview responses for some build-versus-buy estimates, confounding from concurrent improvements, and a six-month observation period that may be too short for complex deployments.

For this fact check, failure should mean a project did not meet a defined success criterion by a defined time. It should not be used as a synonym for no enterprise P&L, no production deployment, no user benefit, or a technical error. Those are different outcomes.

What the 95% headline means

The NANDA report’s central claim was that only a small minority of enterprise generative AI initiatives crossed what it called the “GenAI Divide” and produced measurable financial impact. The widely repeated formulation was that roughly 95% had zero measurable P&L impact.

That is a strong warning about value capture. It is not a universal failure probability.

A project can deliver a working chatbot, save employees time, improve draft quality, or reach a limited production deployment without producing measurable enterprise earnings within six months. Under a strict P&L criterion, that project may belong on the no-impact side even though it created another type of benefit.

The reverse is also possible. A project can report a positive KPI without covering full implementation, review, training, integration, and maintenance costs. A headline success rate is only as useful as its metric.

The most accurate editorial wording is: “A preliminary MIT NANDA report found that roughly 95% of organizations in its studied evidence base had not achieved measurable P&L impact from enterprise GenAI initiatives under its criteria.” Avoid “MIT proved 95% of all AI projects fail.”

Why the methodology deserves caution

First, the report analyzes public initiatives as part of its evidence. Public announcements contain uneven detail and may overrepresent large, visible programs. Quiet internal successes and abandoned experiments may both be missed.

Second, organizations and leaders willing to discuss implementation can differ systematically from those that decline. The report itself acknowledges that the sample could tilt toward more experimental or more cautious adopters.

Third, success metrics vary. A support automation may be evaluated on resolution time and customer satisfaction, while a research assistant may be evaluated on analyst capacity. Converting both to P&L in six months requires assumptions.

Fourth, six months is short for initiatives that need data access, procurement, security review, integration, training, process redesign, and adoption. The window is useful for detecting rapid value, but it can classify slow returns as no return.

Finally, the document is a preliminary report. Its findings are valuable enough to investigate, but the exact percentage should not be treated with the precision of a representative national statistic.

What other 2025-2026 evidence shows

McKinsey’s 2025 State of AI report offers a larger but still self-reported survey. Of 1,993 participants, 88% said their organizations regularly used AI in at least one function. Only 39% reported any enterprise-wide EBIT impact, and 109 respondents met the report’s high-performer definition of significant value plus more than 5% EBIT contribution.

McKinsey’s phase data also show the adoption gap. Among AI-using organizations, 32% were experimenting, 31% piloting, 30% scaling, and 7% fully scaled. Those results do not validate NANDA’s exact 95%, but they support the claim that broad use is much more common than deep financial impact.

The UK ONS July 2026 analysis found that 35% of businesses with at least 10 employees used at least one AI technology. Among adopters, only 10% reported extensive use. The average number of AI technologies per adopting business had risen only from about 1.4 to 1.6 since 2023.

US Census AI diffusion research found that 18% of firms used AI during the November 2025 to January 2026 reference period. Among adopters, 57% used AI in three or fewer functions and 65% in three or fewer worker tasks. Breadth of integration was positively correlated with commercial performance, but the study does not establish that breadth caused the performance.

Together, these sources support a careful conclusion: experimentation has diffused rapidly, while extensive integration and enterprise financial impact remain much less common.

Why projects miss P&L

The evidence points to implementation patterns rather than one universal model problem:

• The project begins with a tool instead of a measurable workflow.

• No baseline exists for time, quality, cost, conversion, or errors.

• AI output is added without changing downstream work.

• Review and exception costs absorb gross time savings.

• Data access and system integration arrive late.

• Ownership is split between technology and business teams.

• Employees can ignore the workflow without consequence or feedback.

• The pilot has no scale, stop, or redesign criteria.

• Benefits create capacity that the company never redeploys.

These are hypotheses supported by recurring survey patterns, not a causal decomposition of the NANDA percentage.

Practical SMB implications

An SMB should not use the 95% headline as a reason to avoid AI or rush toward a supposedly failure-proof vendor. It should use the headline to demand a falsifiable pilot.

Define one workflow, one owner, one baseline, and one decision date. Set a minimum quality threshold, maximum review burden, target adoption rate, and conservative value range. Include a stop condition.

Use a realistic AI implementation timeline that includes data access, process mapping, real-case testing, user adoption, and measurement. A 30-day pilot can validate a bounded workflow, but it may not establish enterprise P&L.

Report outcomes precisely: “reduced median handling time by X,” “met quality threshold on Y of Z cases,” or “did not recover operating cost.” Do not substitute “successful AI transformation” for a workflow metric.

Limitations

This article relies on the report version available for review and on public methodological notes. The original NANDA document’s hosting has changed, so a preserved report copy may be needed for access.

McKinsey and ONS use different denominators and cannot validate NANDA’s exact rate. No cited source provides a representative global failure rate for all AI project types.

What would change our view

A peer-reviewed follow-up with a preregistered success definition, representative sampling, project-level outcomes, full cost data, and observations beyond six months could strengthen or revise the 95% estimate. Evidence that no-impact projects later reached durable production value would lower the inferred failure rate; audited long-term zero returns would strengthen it.

FAQs

Did MIT prove that 95% of AI projects fail?

No. A preliminary MIT NANDA report found roughly 95% of organizations in its evidence base had not achieved measurable P&L impact under its criteria. It was not a representative census of every AI project.

What counted as success in the MIT NANDA report?

The methodology described success as deployment beyond pilot with measurable KPIs. ROI was assessed six months after the pilot.

Does no measurable P&L mean an AI tool had no benefit?

No. It may have saved time or improved a task without creating material enterprise earnings. It may also have had costs that offset the benefit. The outcomes must be reported separately.

Why do AI pilots fail to scale?

Common patterns include weak workflow definition, missing baselines, poor integration, unmeasured review costs, unclear ownership, limited adoption, and no criteria for scaling or stopping.

How can an SMB avoid a no-impact pilot?

Choose one measurable workflow, capture the manual baseline, test real cases, include all costs, define human review, set adoption and quality thresholds, and make an explicit scale, redesign, or stop decision.

Get a 20-Minute AI Workflow Audit

AI Operator can turn one AI idea into a bounded pilot with a baseline, controls, success threshold, and stop rule before implementation spend expands.

Start the 20-minute AI workflow audit

Newsletter

You read this far, might as well sign up.

AI Operator

Newsletter

You read this far, might as well sign up.

AI Operator

Newsletter

You read this far, might as well sign up.

AI Operator