The number gets quoted a lot and usually misread. "95% of AI pilots fail" is taken to mean the technology underdelivers. That is not what the study found.
MIT's researchers looked at 300 public enterprise deployments alongside 52 executive interviews and surveys of 153 leaders. The finding was narrower and more useful: 95% produced no measurable P&L impact. Not that the models failed at their tasks — that nobody could demonstrate the work showed up in the accounts.
Those are very different failures, and only one of them is fixable by buying a better model.
The measurement problem comes first
The detail that reframes the whole statistic: most pilots launch without predefined success criteria.
A pilot with no success criteria cannot succeed. Not "is unlikely to" — cannot, definitionally. If nobody agreed in advance what number should move and by how much, then at the end there is no test to pass. The technology performs exactly as specified and the pilot still lands in the 95%, because measurable impact was never something anyone set up to measure.
A pilot with no success criteria doesn't fail. It just ends.
This is worth sitting with before blaming capability. It means some unknown share of that 95% contains deployments that did work and could not prove it.
Where the work actually is
The second finding explains why the gap between demo and production is so much wider than teams expect. Roughly 80% of the effort to move from pilot to production is data engineering, governance, workflow integration and measurement infrastructure.
Which means the model — the part that gets evaluated, benchmarked, argued about on social media and chosen in a bake-off — is somewhere in the remaining 20%.
| Work | Share of effort | Share of discourse |
|---|---|---|
| Data engineering, governance, integration, measurement | ~80% | low |
| Model selection and prompt work | ~20% | most of it |
Effort share from the MIT analysis. Discourse share is my characterisation, not a measured figure — but scan any AI newsletter this week and judge for yourself.
This also explains a related survey result: 65% of CEOs report misalignment with their own CFO on AI's long-term value. That is not a disagreement about whether the technology is impressive. It is what happens when one person is looking at a demo and the other is looking for a line in the accounts that moved.
Only two sectors show real transformation
Of nine major sectors examined, only two — technology and media — showed material business transformation from generative AI.
The uncomfortable reading is that these are the two sectors whose core product is the thing the technology produces: text, code, images, media. Where the output is the product, value lands directly. Where AI has to be threaded through a physical process, a regulated workflow or a legacy system, the 80% of unglamorous integration work reasserts itself.
That is a reasonable prior for anyone forecasting their own return. If your business is not a text-or-code business, the two-of-nine figure is closer to your base rate than the technology-sector case studies you are being sold.
What separates the 5%
- A number defined before starting. Not "improve productivity" — a specific metric, a baseline, and a threshold agreed with whoever owns the P&L.
- Budget weighted toward integration, not model access. If the plan is mostly API spend, it is funding the 20%.
- Friction accepted rather than avoided. The MIT commentary is direct: pilots fail because organisations avoid the hard integration work, not because the models fall short.
- Measurement built in advance. Instrumented at the start, because retrofitting attribution to a finished pilot is how a working deployment ends up counted as a failure.
- Not: a better model. It is in the 20%, and by the study's own framing it is rarely the constraint.
The number worth holding onto
$2.52 trillion in AI spending against a $600 billion ROI gap is an abstraction at that scale. The version that transfers to a single organisation is smaller and sharper:
If you cannot say now which number should move, and by how much, and who checks — you are already in the 95%. That is knowable on day one, before a single token is spent, and it is the cheapest test available.
Method & sources
This is analysis of published research, not original measurement. Unlike the research on this site drawn from my own logs and invoices, every figure here is someone else's, and inherits their methodology and limits.
The 95% figure, the 300-deployment sample, the 52 executive interviews, the 153 leader surveys, the 80% integration-effort share and the two-of-nine sector finding are all from the MIT study of enterprise GenAI deployments. The $2.52 trillion spend and $600 billion ROI gap are from industry aggregation reporting on that study. The 65% CEO–CFO misalignment figure comes from associated reporting.
Two caveats worth stating. "No measurable P&L impact" is a stricter bar than "failed" — some deployments in that 95% likely worked and could not be attributed. And self-reported enterprise surveys carry a known bias: organisations are more willing to describe a pilot as inconclusive than as a failure.
No commercial relationship with any vendor named. No affiliate arrangement on anything in this piece.