AI Procurement · For boards, CIOs and heads of procurement

Most AI transformations fail in the RFP, before anyone writes code.

Roughly 95% of AI pilots produce no measurable P&L impact. The usual explanation is the technology. It is not. By the time an RFP is issued, the solution has already been chosen — usually by the party least equipped to choose it — and every supplier who responds is competing to deliver the wrong thing well. This page is how to write an AI RFP that does not do that, and what to settle before you issue one.

The Evidence

It is not your team. It is the shape of the deal.

Three independent research programmes arrive at the same place. Read together, they describe a procurement failure rather than a technology failure.

95%of AI pilots produced no measurable P&L impact. MIT NANDA, across more than 300 deployments. Not 95% technically broken — 95% that ran, demoed, and changed nothing that reached the accounts.
~6%of firms attribute more than 5% of EBIT to AI. McKinsey, approximately 1,993 organisations. The strongest single correlate of EBIT impact across around 25 attributes tested was workflow redesign — yet only about 21% of adopters had redesigned any workflow at all.
40%+of agentic AI projects forecast to be cancelled by end-2027. Gartner. Cancellation, not underperformance: the programme is stopped before it reaches production.
30 yrsThis predates AI. McKinsey–Oxford, roughly 5,400 large IT projects: about 45% over budget and delivering some 56% less value than promised. Large custom systems have behaved this way for three decades. AI only made the gap easy to measure.

If pilots failed because models were weak, better models would have fixed it by now. They have not, because the binding constraint sits earlier — in how the work was specified and how the supplier was paid.

The Mechanism

An RFP is a specification. That is exactly the problem.

A well-run procurement takes a written specification and finds the best price for it. That process is sound when you know what you need — a building, a fleet, a payroll system. It breaks for AI transformation for three structural reasons.

1. The buyer is asked to specify a technology they have not yet operated

You know your business better than any supplier ever will. What a system could do for that business is a different question, and it depends on operating experience you do not yet have — that is the whole reason you are buying. Asking you to name the platform, the model and the integration scope inverts the expertise. It should have come from the supplier; instead the supplier is handed your guess and asked to price it.

2. Specifying the means transfers design risk to you and leaves delivery risk with you

The moment the RFP names a solution, responsibility for that choice is yours. If the system is built exactly as specified and the operation does not improve, no supplier has breached anything. You carry the design risk because you specified, and the delivery risk because the contract pays for effort. This is why programmes end with what the deck calls an accepted compromise: a settlement on the original problem, signed off because the money was already spent.

3. Effort-based pricing makes overrun structural, not accidental

Consultants bill days. Development shops bill headcount-months. AI vendors bill seats and credits. Every one of those models earns more when the scope grows, so every one of them must price change requests — and every discovery becomes one. The overrun is not a failure of the supplier's professionalism. It is what the contract rewards.

Fixing this after the RFP is issued is close to impossible: the specification is now the contract. It has to be fixed before.

Before You Issue

Four things to settle before the RFP exists.

01The problem, not the solution. Name the operation that costs you most — the bottleneck, in one sentence, in business terms. If the sentence contains a product category, it is still a solution.
02The baseline, measured. Cycle time per case end to end, error rate counted against the governing rules, throughput at constant headcount. Measure it before anyone builds anything. Quoting an improvement figure before a baseline exists is the behaviour that produced the 95%.
03Who carries the risk if the number does not move. Decide this before you see pricing, because it determines which suppliers can honestly bid. It is the single most revealing question in an AI procurement.
04What must stay deterministic. The ledger, entitlements, calculations, compliance decisions and the audit trail. Write down what a language model is never permitted to alter. In regulated and public-sector environments this is the difference between a system that can be signed off and one that cannot.

Talking to capable suppliers while you do this is not a probity problem. Speak to several, document what you learn, and put the resulting problem statement into the RFP rather than any one vendor's answer. What you must not do is let the first supplier you speak to write the specification you then tender.

The Checklist

What to put in the AI RFP — and what to leave out.

ASKThe operation, the metric, the baseline and the date. State the business outcome that defines success and let suppliers propose the architecture.
ASKWhat happens if it is not live on the date. Require the commercial consequence in writing. A supplier who carries nothing will say so here, in careful language.
ASKThe cost of a change, and who classifies it. Ask who decides whether something is a defect or a change request. If the supplier decides and the supplier bills for it, you have found your overrun.
ASKOne operationally similar engagement, with measured before-and-after numbers and a reference contact willing to take a twenty-minute call. Not a logo wall.
ASKThe boundary between deterministic and generated. Which components can a model influence, and which can it never touch? Require an architecture answer, not an assurance.
ASKWho owns the code, the data and the model artefacts at the end — and on what terms you exit.
OMITA named platform, model or vendor stack. You are buying an outcome. Naming the means guarantees you own the consequences of the choice.
OMITWeighting that rewards technical credentials over business integration. The most common scoring error in AI procurement: impressive technical demonstrations from suppliers who cannot change how the work is done.
OMITA long discovery phase before anything runs. Every month of specification is a month in which the business changes and the specification ages.
AI Washing

Four questions that separate a system from a wrapper.

Relabelled rule engines, thin wrappers over a third-party API, and in some cases human work presented as automation, all survive a normal RFP process comfortably. These four questions do not need a technical evaluator to be useful.

Q1What does this do that a decision tree could not? A real answer describes pattern recognition or language understanding on messy input. A washed answer restates the brochure in different words.
Q2Whose model is it, and what happens if your upstream provider disappears tomorrow? If the answer is a rebuild, you are buying a wrapper — and inheriting somebody else's roadmap and pricing power.
Q3What is deterministic and what is generated? Ask specifically whether the model can alter the ledger, entitlements, or the audit trail. “The model is very accurate” is not an answer to this question; it is an admission that no boundary exists.
Q4Show a live operation and its measured numbers. Not a demo environment, not a pilot deck — a running operation, with the before-and-after figures and the definition of what was counted.
How Xamun Answers

One problem. One price. Unlimited changes.

Xamun is built to be answerable to the questions above rather than to survive them. Three commitments, all verifiable before you sign anything.

RISKThe supplier carries it. One operation live end to end in 30 days. If it is not live, under Usage pricing the meter never starts and you owe nothing; under Licence the final tranche is never paid and you keep what was built. If the agreed number does not move afterwards, we fix it for three years without asking why it moved — not bug versus change, not data versus code.
PRICENever for effort. No rate card, no timesheets, and no change request has ever been invoiced. Changes are unlimited by construction: a supplier who does not bill for effort has nothing to charge for a change. This is the commitment that competitors cannot copy without changing how they earn.
ARCHTwo Minds — deterministic where it matters. A deterministic engine holds the rules, entitlements, calculations, compliance and audit trail; it runs first and has the final say. A governed language model handles extraction, drafting and explanation, and is never the last word. Both operate over one graph memory. Hallucination is a feature — just never in your ledger. This is what makes the guarantee underwritable rather than rhetorical.

Twenty-five years of delivery sit behind that: systems still in production after two decades, over $500M in transactions processed on a platform we built, clients in more than ten countries. A three-year commitment is only credible from a delivery engine old enough to have honoured one.

Questions

AI procurement, answered plainly.

Should we write an AI RFP before or after talking to vendors?

Talk to capable vendors before the RFP is written, not after it is issued. An RFP is a specification document: the moment you issue it, the solution is fixed and every respondent is competing to deliver the thing you already named. If the named thing is wrong — and research on AI programmes suggests it usually is, because the buyer is asked to specify a technology they have not yet operated — then the procurement runs perfectly and still produces a failure. Pre-RFP conversations are not a probity risk if you speak to several suppliers, document what you learn, and put the resulting problem statement into the RFP rather than any single vendor's answer.

Why do so many AI RFPs produce failed projects?

Because they specify means instead of outcomes. MIT's NANDA study of more than 300 AI deployments found roughly 95% produced no measurable P&L impact, and McKinsey's survey of about 1,993 organisations found only around 6% attribute more than 5% of EBIT to AI. The common thread is not model quality — it is that the work around the system was never redesigned. An RFP that lists a platform, a model and an integration scope buys exactly those three things and leaves the operation unchanged.

What should an AI RFP ask for instead of a technical specification?

Ask for a business outcome with a baseline, and make the supplier own it. Specify the operation to be fixed, the metric that must move (cycle time per case end to end, error rate against governing rules, throughput at constant headcount), the date by which it must be running, and who carries the cost if it is not. Let suppliers propose the architecture. Requiring a named technology stack in an AI RFP transfers the design risk to you while leaving the delivery risk with you as well.

How do we tell genuine AI from AI washing in RFP responses?

Ask four questions that a wrapper cannot answer well. First: what does the system do that a decision tree could not? Second: whose model is it, and what happens if the upstream API provider disappears tomorrow? Third: what is deterministic and what is generated — specifically, can the language model alter the ledger, the entitlements or the audit trail? Fourth: show a live operation and its measured before-and-after numbers, not a demo. Vendors relabelling rule-based automation, or wrapping a third-party API, answer these in marketing language rather than architecture.

What is a reasonable timeline to require in an AI RFP?

Require one operation running end to end in production within 30 days, and judge the supplier on that before committing to anything further. Long discovery phases are where AI programmes quietly die: every month of specification is a month in which the business changes and the spec ages. A supplier who cannot put a single bounded operation into production inside a month is describing a research project, not a delivery.

How should an AI RFP handle change requests and scope?

Treat the change-request line as a diagnostic. Suppliers who bill for effort — day rates, headcount-months, seats — earn more when scope grows, so they must price change requests, and the overrun is structural rather than accidental. Ask each respondent directly: what does a change cost, and who decides whether something is a change or a defect? Xamun's answer is one problem, one price, unlimited changes: no rate card, no timesheets, and no change request has ever been invoiced.

Can Xamun help before we issue the RFP?

Yes — that is the point at which help is worth most. Xamun Intelligence reads the business and maps where the value actually sits, so the RFP you issue describes the right problem with a measurable baseline attached. That work is useful even if Xamun never bids: you keep the problem statement, the baseline and the opportunity map. If Xamun does deliver, the engagement is judged on one live operation in 30 days, and if it is not live you owe nothing under Usage pricing.

Talk to us before the RFP, not after it.

Half a day. We read the business, map where the value actually sits, and leave you with a problem statement and a measured baseline you can put into any tender — whether or not Xamun ever bids. You keep the work either way.