AI TRANSFORMATION RISK · THE EVIDENCE AND THE FIX

Why AI transformation fails —
and how to make the risk carryable.

Every study agrees on the failure rate and every commentator can diagnose it. What almost nobody ranking on this question can do is prevent it contractually. This page sets out the evidence, the four structural causes, the four delivery choices that remove them, and what a contract looks like when the vendor rather than the buyer carries the risk.

The fix Read the full paper

What is the AI pilot failure rate?

The AI-specific figures are now familiar, and they agree with each other.

95%
of GenAI pilots: no measurable P&L impact
MIT Project NANDA, 300+ deployments
~6%
of firms attribute >5% of EBIT to AI
McKinsey State of AI, ~1,993 organisations
40%+
of agentic-AI projects cancelled by end-2027
Gartner forecast, June 2025

The figure that matters more is older. McKinsey and the University of Oxford studied some 5,400 large IT projects and found that, on average, they ran 45% over budget and delivered 56% less value than predicted, with one project in six overrunning by more than 200%. That study predates the current AI cycle by more than a decade. Its subject is large custom software generally, and its findings are almost indistinguishable from the AI rows above.

45%
average budget overrun
McKinsey–Oxford, ~5,400 projects, 2012
56%
less value than predicted
same study
1 in 6
overrun by more than 200%
same study

The conclusion is uncomfortable but useful: the failure is not new and it is not about models. AI did not create the pattern. It made the pattern measurable.

Why do AI transformation projects fail?

Three things happen in nearly every failed engagement, and all three are structural rather than accidental.

The buyer was asked what to build. So every gap in the specification became the buyer’s fault, and the system arrived technically correct and operationally wrong.

It ran over, or never ended. Every discovery became a change request, and the commercial model rewarded duration. A consultant paid by the day, a development shop paid by the headcount-month, an AI vendor paid by the seat — each is paid whether or not the number moves.

The buyer accepted less. A compromise on the original problem, signed off because the money was already spent.

Downstream of those three sit the failure modes most post-mortems name: the wrong use case, chosen from the technology backwards; no governance layer tracking the metric after go-live; and a strategy that drifts between quarterly reviews while sprints close and dashboards turn green. They are real. They are also symptoms of the three above.

What the buyer wanted to know was simple: will it fix this, will it pay back, will it be done in time. The market has only ever answered: probably, we hope, eventually.

Why do outcome-based contracts fail as well?

Outcome-based pricing proposes to move the risk. The proposal is correct. But moving a risk to a party that cannot control it does not make the risk smaller; it makes the party hedge.

That is why so many “outcome-based” contracts contain projected returns produced before any baseline exists, milestone payments tied to documents rather than results, warranty periods measured in weeks, and change-request clauses that quietly restore effort billing through the side door. The vendor is not being dishonest. They are being rational: if their delivery model gives them no way to make the outcome predictable, they cannot afford to be paid only on it. The pricing is not the problem to solve first.

How do you de-risk an AI transformation?

Four delivery choices. Each has a wrong answer that is the industry default, and each is a precondition for a vendor being able to carry the risk at all.

1 · Baseline before build

In week one, before anything is specified: cycle time per case end to end, error rate against the governing rules, throughput at constant headcount. Observable by both parties, fixed before the build, operational rather than financial. Done is then judged against the baseline, not against a slide.

2 · Burden on the system, not the user

A model, retrieval or a graph behind a chat window leaves the work on the user: prompt, judge, verify, re-key, absorb the blame. Only above that divide — where the process runs itself — can cycle time, error rate and throughput move without adding people.

3 · A deterministic layer with the final say

Rules, entitlements, calculations and the audit trail in code that runs first. The language model drafts, extracts and explains and is never the last word on a figure. This is what makes an error-rate commitment possible.

4 · One specification, no handoffs

Strategy to analyst to architect to developer is four translations, and meaning degrades at each. One specification that runs from diagnosis to running software, approved once, is why the outcome is right — not a speed claim.

Those four are what make a hard gate possible: thirty days from approved specification to an operation running end to end. Not a demo, not a pilot. A case flows through the system and comes out the other side, or it does not, and both parties can see which. That gate is where the delivery model hands over to the commercial model — and only at that point does outcome pricing stop being a hedge and start being a commitment. What the commercial model looks like when it follows from the delivery model →

What should you do when an AI pilot has already failed?

Choose the costliest problem, not the strategic one. One operational constraint leadership already knows is expensive, with a countable unit, that a small group can approve and that can be baselined in a week.

Demand the baseline before the proposal. Return any proposal that contains a projected return before the operation has been measured. Ask for the week-one measurement plan: what will be counted, how, and by whom.

Ask the four delivery questions. Where does the burden sit after go-live — with our people, or with the system? What has the final say over a decision that touches money, entitlement or compliance — a rule, or a model? How many handoffs sit between the problem we describe and the code that runs, and what artefact survives them? What is the go-live gate, when is it, and what happens to payment if it is not met?

Read the commercial terms for the side door. A rate card, billable change requests, milestone payments tied to documents rather than a live operation, warranty in weeks, any distinction between defects and changes. Each is effort billing re-entering. The full checklist for an AI RFP →

Hold everything else back. Whatever the vendor’s wider platform offers, it is worth nothing until the first system is live. Judge them on one live operation, and choose what comes next only after the number has moved.

Questions

What is the AI pilot failure rate?

MIT's Project NANDA found that 95% of generative-AI pilots produced no measurable P&L impact across more than 300 deployments. McKinsey's State of AI survey of roughly 1,993 organisations found about 6% attribute more than 5% of EBIT to AI. Gartner forecasts that over 40% of agentic-AI projects will be cancelled by end-2027, citing cost, unclear value and inadequate controls. Older and more telling: McKinsey with the University of Oxford studied some 5,400 large IT projects and found they ran 45% over budget and delivered 56% less value than predicted, with one in six overrunning by more than 200%.

Why do AI transformation projects fail?

Three structural causes recur in nearly every failed engagement, and none is about model quality. The buyer was asked what to build, so every gap in the specification became the buyer's fault. It ran over or never ended, because every discovery became a change request and the commercial model rewarded duration. And the buyer accepted less — a compromise on the original problem signed off because the money was already spent. Add the three failure modes downstream of that: the wrong use case chosen from the technology backwards, no governance model tracking the metric after go-live, and adoption treated as a workstream that starts after the software ships.

Why do outcome-based contracts fail as well?

Because most vendors offering them change only the invoice. Moving a risk to a party that cannot control it does not make the risk smaller; it makes the party hedge — projected returns produced before any baseline exists, milestone payments tied to documents rather than a live operation, warranty periods measured in weeks, and change-request clauses that restore effort billing through the side door. The pricing is not the problem to solve first; the delivery model underneath it is.

How do you de-risk an AI transformation?

Four delivery choices, each with a wrong answer that is the industry default. Baseline before build: measure cycle time per case, error rate against the governing rules, and throughput at constant headcount in week one, before anything is specified. Put the burden on the system, not the user — above the level-3/4 divide, where a process runs itself rather than a chat window answering questions. Give a deterministic layer the final say over any decision touching money, entitlement or compliance, so the language model can draft and explain but never decide. And run one specification from diagnosis to running software without handoffs. Only once those four exist can a vendor carry the outcome risk instead of hedging it.

What should you do when an AI pilot has already failed?

Do not run another pilot. Choose the costliest operational problem leadership already knows is expensive — one with a countable unit that a small group can approve and that can be baselined in a week. Demand the week-one measurement plan before any proposal containing a projected return. Ask the four delivery questions: where the burden sits after go-live, what has the final say over a regulated decision, how many handoffs separate the problem from the code, and what the go-live gate is and what happens to payment if it is missed. Then read the commercial terms for the side door — a rate card, billable change requests, milestone payments tied to documents — and judge the vendor on one live operation before committing to anything wider.

What does 'the vendor carries the risk' actually mean in a contract?

The build is funded by the vendor. Payment is triggered by go-live and nothing before it: under a usage model the meter starts the day the operation runs and nothing is owed if it never does; under a licence model the final tranche falls due only at go-live and if it is not met the buyer keeps what was built and never pays it. Scope changes, change requests, timesheets and rework are not invoiced. And the vendor keeps changing the system until it fits, for as long as the operation stays the same, without distinguishing a bug from a change. Any of those terms missing means the risk has moved back to the buyer, whatever the contract is called.

Bring us the pilot that failed.

We will baseline the operation it was meant to fix, in week one, and tell you honestly whether it is one a vendor can carry the risk on.

Book a Discovery Download the paper