The number gets quoted a lot and usually misread. "95% of AI pilots fail" gets taken to mean the technology underdelivers. The study found something else.
MIT's researchers looked at 300 public enterprise deployments alongside 52 executive interviews and surveys of 153 leaders. Their finding was narrower and more useful than the headline: 95% produced no measurable P&L impact. The models mostly did their jobs. Nobody could demonstrate that the work showed up in the accounts.
Those are very different failures, and only one of them is fixable by buying a better model.
The measurement problem comes first
One detail reframes the whole statistic. Most pilots launch without predefined success criteria.
A pilot with no success criteria can't succeed. That's a definitional claim, not a pessimistic one. If nobody agreed in advance which number should move and by how much, there's no test left to pass at the end. The technology performs exactly as specified, the pilot still lands in the 95%, and measurable impact was never something anyone set up to measure.
A pilot with no success criteria doesn't fail. It just ends.
Sit with that before blaming capability. It means some unknown share of that 95% is deployments that did work and couldn't prove it.
Where the work actually is
The second finding explains why the gap between demo and production runs so much wider than teams expect. Roughly 80% of the effort of moving from pilot to production goes to data engineering, governance, workflow integration and measurement infrastructure.
Which leaves the model somewhere in the remaining 20%. The part that gets evaluated, benchmarked, argued about on social media and picked in a bake-off is the small slice.
| Work | Share of effort | Share of discourse |
|---|---|---|
| Data engineering, governance, integration, measurement | ~80% | low |
| Model selection and prompt work | ~20% | most of it |
Effort share from the MIT analysis. Discourse share is a characterisation rather than a measured figure. Scan any AI newsletter this week and judge for yourself.
It also explains a related survey result. 65% of CEOs report misalignment with their own CFO on AI's long-term value. They aren't arguing about whether the technology is impressive. One of them is looking at a demo and the other is looking for a line in the accounts that moved.
Only two sectors show real transformation
Of nine major sectors examined, only two showed material business transformation from generative AI: technology and media.
Those happen to be the two sectors whose core product is the thing the technology produces: text, code, images, media. Where the output is the product, value lands directly. Where AI has to be threaded through a physical process, a regulated workflow or a legacy system, the 80% of unglamorous integration work reasserts itself.
That's a reasonable prior for anyone forecasting their own return. If you don't run a text-or-code business, two-of-nine sits closer to your base rate than the technology-sector case studies you're being sold.
What separates the 5%
- A number defined before starting. A specific metric, a baseline, and a threshold agreed with whoever owns the P&L. "Improve productivity" doesn't qualify.
- Budget weighted toward integration over model access. A plan that's mostly API spend is funding the 20%.
- Friction accepted rather than avoided. The MIT commentary is blunt about this. Pilots fail because organisations dodge the hard integration work, not because the models fall short.
- Measurement built in advance. Instrument at the start. Retrofitting attribution to a finished pilot is how a working deployment ends up counted as a failure.
- A better model won't do it. That lives in the 20%, and by the study's own framing it's rarely the constraint.
The test that costs nothing
$2.52 trillion in AI spending against a $600 billion ROI gap is an abstraction at that scale. Shrunk to a single organisation it gets sharper.
If you can't say right now which number should move, by how much, and who checks, you're already in the 95%. You can know that on day one, before spending a single token. It's the cheapest test available.
Method & sources
This is analysis of published research, not original measurement. Unlike the work on this site drawn from first-party logs and invoices, every figure here belongs to someone else and inherits their methodology and limits.
The 95% figure, the 300-deployment sample, the 52 executive interviews, the 153 leader surveys, the 80% integration-effort share and the two-of-nine sector finding are all from the MIT study of enterprise GenAI deployments. The $2.52 trillion spend and $600 billion ROI gap are from industry aggregation reporting on that study. The 65% CEO–CFO misalignment figure comes from associated reporting.
Two caveats. "No measurable P&L impact" is a stricter bar than "failed", so some deployments in that 95% likely worked and could not be attributed. And self-reported enterprise surveys carry a known bias: organisations are more willing to describe a pilot as inconclusive than as a failure.
No commercial relationship with any vendor named. No affiliate arrangement on anything in this piece.