Analysis·AI ROI

95% of enterprise AI pilots return nothing. The failure isn't the model.

MIT looked at 300 enterprise AI deployments and found 95% delivered no measurable P&L impact. The striking part is what the surviving 5% did differently — and it has almost nothing to do with which model they picked.

RA
Raunak A.
4 August 2026 · 9 min read · No vendor relationships
A sealed matte black box on a warm off-white backdrop, one edge catching light.
Illustration generated with AI. Every figure, chart and table in this piece is measured, not generated.

The verdict

MIT's study — 52 executive interviews, 153 leader surveys, 300 public deployments — found 95% of enterprise GenAI pilots delivered no measurable P&L impact. Set against roughly $2.52 trillion of AI spending and a $600 billion ROI gap. The common thread in the failures is not model choice or capability: it is that roughly 80% of the work in getting to production is data engineering, governance, workflow integration and measurement — and most pilots launch with no success criteria defined at all.

Pilots with no P&L impact
95%
Deployments analysed
300
ROI gap
$600bn
Sectors transformed
2of 9

The number gets quoted a lot and usually misread. "95% of AI pilots fail" is taken to mean the technology underdelivers. That is not what the study found.

MIT's researchers looked at 300 public enterprise deployments alongside 52 executive interviews and surveys of 153 leaders. The finding was narrower and more useful: 95% produced no measurable P&L impact. Not that the models failed at their tasks — that nobody could demonstrate the work showed up in the accounts.

Those are very different failures, and only one of them is fixable by buying a better model.

The measurement problem comes first

The detail that reframes the whole statistic: most pilots launch without predefined success criteria.

A pilot with no success criteria cannot succeed. Not "is unlikely to" — cannot, definitionally. If nobody agreed in advance what number should move and by how much, then at the end there is no test to pass. The technology performs exactly as specified and the pilot still lands in the 95%, because measurable impact was never something anyone set up to measure.

A pilot with no success criteria doesn't fail. It just ends.

This is worth sitting with before blaming capability. It means some unknown share of that 95% contains deployments that did work and could not prove it.

Where the work actually is

The second finding explains why the gap between demo and production is so much wider than teams expect. Roughly 80% of the effort to move from pilot to production is data engineering, governance, workflow integration and measurement infrastructure.

Which means the model — the part that gets evaluated, benchmarked, argued about on social media and chosen in a bake-off — is somewhere in the remaining 20%.

Where pilot-to-production effort goes vs. where attention goes
Work Share of effort Share of discourse
Data engineering, governance, integration, measurement~80%low
Model selection and prompt work~20%most of it

Effort share from the MIT analysis. Discourse share is my characterisation, not a measured figure — but scan any AI newsletter this week and judge for yourself.

This also explains a related survey result: 65% of CEOs report misalignment with their own CFO on AI's long-term value. That is not a disagreement about whether the technology is impressive. It is what happens when one person is looking at a demo and the other is looking for a line in the accounts that moved.

Only two sectors show real transformation

Of nine major sectors examined, only two — technology and media — showed material business transformation from generative AI.

The uncomfortable reading is that these are the two sectors whose core product is the thing the technology produces: text, code, images, media. Where the output is the product, value lands directly. Where AI has to be threaded through a physical process, a regulated workflow or a legacy system, the 80% of unglamorous integration work reasserts itself.

That is a reasonable prior for anyone forecasting their own return. If your business is not a text-or-code business, the two-of-nine figure is closer to your base rate than the technology-sector case studies you are being sold.

What separates the 5%

  • A number defined before starting. Not "improve productivity" — a specific metric, a baseline, and a threshold agreed with whoever owns the P&L.
  • Budget weighted toward integration, not model access. If the plan is mostly API spend, it is funding the 20%.
  • Friction accepted rather than avoided. The MIT commentary is direct: pilots fail because organisations avoid the hard integration work, not because the models fall short.
  • Measurement built in advance. Instrumented at the start, because retrofitting attribution to a finished pilot is how a working deployment ends up counted as a failure.
  • Not: a better model. It is in the 20%, and by the study's own framing it is rarely the constraint.

The number worth holding onto

$2.52 trillion in AI spending against a $600 billion ROI gap is an abstraction at that scale. The version that transfers to a single organisation is smaller and sharper:

If you cannot say now which number should move, and by how much, and who checks — you are already in the 95%. That is knowable on day one, before a single token is spent, and it is the cheapest test available.

Method & sources

This is analysis of published research, not original measurement. Unlike the research on this site drawn from my own logs and invoices, every figure here is someone else's, and inherits their methodology and limits.

The 95% figure, the 300-deployment sample, the 52 executive interviews, the 153 leader surveys, the 80% integration-effort share and the two-of-nine sector finding are all from the MIT study of enterprise GenAI deployments. The $2.52 trillion spend and $600 billion ROI gap are from industry aggregation reporting on that study. The 65% CEO–CFO misalignment figure comes from associated reporting.

Two caveats worth stating. "No measurable P&L impact" is a stricter bar than "failed" — some deployments in that 95% likely worked and could not be attributed. And self-reported enterprise surveys carry a known bias: organisations are more willing to describe a pilot as inconclusive than as a failure.

No commercial relationship with any vendor named. No affiliate arrangement on anything in this piece.

What AI actually costs, weekly.

Releases, research and analysis — with the money question asked every time. Five mornings a week, about ten minutes each.

Free · No vendor sponsorship of editorial · Unsubscribe anytime

AI ROI Enterprise Analysis