Practical guides
Why enterprise AI implementations fail
The causes repeat. Almost none of them are about model quality, and most are settled before the first line of code.
Enterprise AI has a high failure rate and a well-documented one. What is less discussed is how predictable the failures are. Across programmes, the same seven causes recur, and they are visible early enough to act on.
This is not a list of technical risks. Model quality is rarely the binding constraint by 2026. The constraints are organisational, and they are the kind that a steering committee can fix in a fortnight if someone is willing to name them.
What the research establishes
MIT's The GenAI Divide: State of AI in Business 2025 found that around 95% of generative AI pilots deliver no measurable impact on the profit and loss statement, with roughly 5% reaching production with real value. McKinsey's State of AI survey finds most organisations still experimenting, with about a third beginning to scale. The UK government's AI adoption research found that among firms already using AI, only about half felt ready to scale it.
Read together: adoption is common, value is rare, and the failure is concentrated at the point where a working system has to change how the business operates.
Seven failure modes
The failure that presents as success
There is a worse outcome than a pilot that stops: a pilot that is declared a win and changes nothing.
People report saving time. Everyone agrees the tool is useful. It gets counted. But unless the saved time turns into throughput, decision quality, customer outcomes, or margin, it never reaches the profit and loss statement. The UK government research found exactly this pattern, with businesses reporting improved workforce productivity alongside no change in revenue.
Efficiency at the task level and value at the business level are different claims. Most failed programmes proved the first and were funded on the second.
We traced the specific places the money leaks in the real cost of AI is not the budget, it is the value gap.
What the survivors do differently
The programmes that reach production are unglamorous and they look alike.
One workflow
A single named workflow with a current cost or cycle time, rather than a portfolio of pilots across departments.
One number
A threshold agreed before launch that would justify scaling, expressed in the currency the business already reports.
One name
The person already accountable for that number owns the outcome, and has authority to change the process.
One exit
A stated condition under which the work stops, agreed while nobody is defending anything.
MIT's research adds two details worth holding. Mid-market companies often reach full implementation faster, in around ninety days, because they can redesign a process without a committee. And builds involving an external partner succeeded roughly twice as often as internal-only builds, largely because an outside team is forced to be explicit about the workflow and the outcome before starting. Our version of that discipline is written up in how we work.
How to tell in the first month
You do not need to wait for a programme to fail to know it will. Four questions, asked in week three, are diagnostic.
Can someone state the number this is supposed to move, without hedging? Can someone name the person accountable for it? Does an evaluation set of real cases exist yet? And has anyone written down what would make you stop? Four yeses is not a guarantee. Fewer than three is close to one.
The failures in this list are not caused by immature technology. They are caused by starting to build before deciding what the build is for, which is a much older mistake. The strategy version of this argument is where it begins.
Most of these are fixable in a fortnight.
If a programme is stalling and you cannot say which of the seven it is, that is usually the finding. Discovery exists to produce it quickly.