← back to blog
StrategyOperations10 min read

Pilot to production: why most AI stalls — and the operating model that ships it

Only ~48% of AI projects reach production. The blocker isn't the model — it's the operating model. Here's the playbook that ships.

AMDIM · Operations

June 14, 2026

You've seen the demo work. You've probably also seen that same demo never reach production. The distance between those two moments is where most AI programmes quietly die — and in our experience, it's almost never a model problem.

The numbers bear that out, and they're sobering.

48%of AI projects reach production — and that takes ~8 monthsGartner · 2024
80%+of AI projects fail — about twice the rate of non-AI ITRAND · 2024
17→42%of firms abandoned most AI initiatives, 2024 to 2025S&P Global

If you own the operating model — as a VP of operations, a chief data officer, a transformation lead — this is your problem to solve. Which is actually the good news, because the blockers turn out to be organisational, not algorithmic. They're yours to move.

The gist: only about half of AI projects reach production, and the cause is the operating model, not the algorithm. A pilot is a notebook; a product is a pipeline. Ship with a hub-and-spoke model, cross-functional pods that own outcomes, a value gate on every initiative — and adoption designed in from day one.

How many AI projects actually reach production?

Two famous statistics dominate this conversation, and a credible operations leader should handle both with care.

The first — "87% of models never make it to production" — you should quietly retire. It traces back to a 2019 sponsored conference article, not a study, and it's been laundered through hundreds of vendor blogs since. Repeat it in a board meeting and the sharp person across the table will know you didn't check.

The second is real, but almost always misquoted. MIT's 2025 study found that 95% of enterprise generative-AI pilots delivered no measurable P&L returnnot that 95% "failed technically." The models mostly worked. The value never showed up. MIT calls the cause a "learning gap": a failure to weave AI into workflows, structures and culture. Explicitly not a model-quality problem.

Put the defensible numbers side by side and they tell one story. McKinsey found 88% of organisations now use AI in at least one function — but only about a third have begun to scale it. Adoption is easy. Scaling is where it stalls.

Across RAND, Gartner, McKinsey and MIT, the conclusion is the same: the technology mostly works. AI stalls on objectives, data, integration and adoption — and every one of those is an operating-model decision.

Why pilots stall

RAND's 2024 study produced the most rigorous list of causes we've seen, and it's striking how little of it is about the model. The top reason isn't data or infrastructure — it's that stakeholders misunderstand or miscommunicate the problem AI is meant to solve. After that come inadequate data, chasing technology for its own sake, missing production infrastructure, and the occasional problem that's simply too hard for today's AI.

McKinsey adds the organisational layer. Around 46% of COOs blame data or IT/OT system limits for stalled scaling. Roughly half point to culture. And the most damning finding of all: about 80% of organisations layer AI onto their existing processes without redesigning the workflow — even though workflow redesign is the single strongest predictor of bottom-line impact McKinsey could find. Automate a step inside a broken process and you get a faster broken process. The value evaporates somewhere on the way to production.

It's worth saying plainly, because it reframes the whole effort: AI stalls between pilot and production for the same reason transformation programmes have always stalled. The technology was the easy part, and the operating model got treated as an afterthought.

The operating model that ships

If the blockers are organisational, the fix is an operating model built to scale — not a heroic team improvising a fresh approach on every project. Three things separate the organisations that ship.

A hub-and-spoke structure. A pure centre of excellence hoards the scarce talent but becomes a bottleneck. A fully federated model moves fast but duplicates work and lets governance rot. McKinsey finds most successful organisations land on the hybrid: a central hub that owns standards, platform and governance, and spokes — embedded teams inside the business units — that execute locally with real autonomy.

exhibit

Hub-and-spoke AI operating model Hubstandards · platform · governance Operations spoke Finance spoke Service spoke Supply spoke
The hub sets the rails; the spokes own delivery in the business. Most organisations that scale AI end up here.

Pods that own the outcome. The handoff from "the business" to "IT" to "data science" is exactly where accountability — and value — leaks away. The organisations that ship put business, operations and technical people in a standing pod with end-to-end ownership of a domain, and they measure that pod on the outcome, not on whether it delivered the feature.

A value gate on everything. Tie each initiative to a specific value driver and KPI before it starts. McKinsey's rule is the right one: if you can't answer the value question, don't launch. It's also the cure for "pilot sprawl" — that portfolio of impressive demos nobody will fund to production, because nobody can say what any of them is worth.

A pilot is a notebook. A product is a pipeline.

Here's the technical reality every operations leader should internalise, because it explains the eight-month average and most of the abandonment. As Andrew Ng puts it, in production systems "the machine learning code is just a small piece of the puzzle." The model is the small part. The machinery around it is the actual product.

Google Cloud's widely-used MLOps maturity model makes the gap painfully concrete:

exhibit

MLOps maturity: level 0 to level 2 Level 0 — Manualnotebooks, hand-offs Level 1 — Pipelinecontinuous training Level 2 — CI/CD/CTautomated, reliable stalled pilots live at Level 0 → products live at Level 1–2
"Pilot to production" is mostly this climb — from a notebook to a pipeline that monitors and retrains itself.

A Level 0 system is all manual: the model lives in a notebook, every step is hand-run, results move around as files. That's where stalled pilots live. Level 1 automates the pipeline so the model retrains itself on fresh data — continuous training, the piece ordinary DevOps doesn't have, because models decay in a way software doesn't. Level 2 automates the whole loop: integration, delivery and training together, for fast, reliable updates.

And lurking inside all of this is the trap that catches even good teams: silent degradation. A normal deployment pipeline cheerfully reports "success" while a model's predictions quietly rot — because the upstream data shifted, or customer behaviour changed. Accuracy erodes invisibly until a business metric drops and someone asks why. Mature teams watch the input distributions, not just the outputs, and set drift alarms on the inputs — because input drift shows up before output failure and buys you time to react. None of that exists in a pilot. All of it is mandatory in a product.

The last mile is adoption — and it's the biggest gap

You can clear every technical hurdle and still fail, for the dullest reason imaginable: a deployed model nobody uses returns exactly nothing. This is the most under-budgeted line in the whole programme, and the data on it is genuinely startling.

Prosci's research found that projects with excellent change management hit their objectives 88% of the time — versus just 13% when change management is poor or absent. Roughly a sevenfold difference, driven not by the technology but by whether people actually adopt it.

The reframe we find most useful for operations leaders is Prosci's idea of people-dependent ROI: the share of a project's projected benefit that only materialises if people change how they work. So the question to put to any business case isn't "what's the ROI of change management?" It's "how much of this project's ROI depends on adoption — and what, exactly, are we doing to earn it?"

The symptom of getting this wrong is everywhere now: shadow AI. MIT found that while only about 40% of companies had bought an official AI subscription, employees at over 90% were already using personal AI tools — because the sanctioned ones didn't fit how they actually work. When the official tool doesn't land, people route around it, and your governance, security and value capture walk out the door with them.

Sequencing for time-to-value

How you sequence the work decides whether value shows up in months or never. The evidence points the same way every time.

Iterate; don't big-bang. McKinsey's Rewired research pushes organisations away from 12-to-18-month deliveries toward short, frequent releases reviewed with the business — value in weeks, with the chance to steer before you've spent the budget. Redesign the workflow before you roll out the tool, not after; it's the highest-leverage move and the one most often skipped. And start in the back office — the winning minority in MIT's data picked one well-defined pain point, partnered with specialists rather than building everything in-house, and went after high-frequency, measurable, fast-payback work before chasing the customer-facing showcase.

The point here is simpler: sequencing is a deliberate operating-model choice, not a project-management afterthought — map the first 90 days before you build, not after.

Where this goes wrong

Even with the right model, three failure modes keep recurring, and they're worth watching for.

The platform that outruns the use cases. Building an elaborate MLOps platform before you have two or three live use cases to justify it is its own kind of pilot purgatory. Let real products pull the platform into existence, not the other way round.

Governance as a brake instead of a rail. Controls that make the official path slower than shadow AI guarantee shadow AI. Governance has to be the path of least resistance, not a toll booth.

Change management bolted on at the end. Adoption designed in during week one is an operating discipline. Adoption "communicated" the week before launch is theatre — and that 88%-versus-13% gap is decided long before go-live.

The takeaway

If you keep one idea from this: "pilot to production" is an operating-model problem wearing a technology costume. The model works. What actually ships it is a hub-and-spoke structure, pods that own outcomes, a value gate on every initiative, the unglamorous pipeline-and-monitoring work that turns a notebook into a product, and adoption designed in from day one. Get those right and you move from the half that stall to the minority that scale — and the returns follow the operating model, not the other way round.

Frequently asked questions

What percentage of AI projects make it to production? Gartner's 2024 research found about 48% of AI projects move from prototype into production, taking around eight months on average. RAND reports that more than 80% of AI projects fail overall — about twice the rate of non-AI IT projects. (Ignore the popular "87% never reach production" figure; it traces to a 2019 sponsored article, not a study.)

Why do AI pilots fail to scale? The causes are organisational, not technical: unclear objectives, inadequate or poorly-governed data, no production infrastructure, and — above all — layering AI onto existing processes without redesigning the workflow. RAND ranks unclear or miscommunicated objectives as the number-one cause; McKinsey names workflow redesign as the strongest predictor of impact.

What's the difference between an AI pilot and a production system? A pilot is usually a manual, notebook-based proof of value (MLOps maturity Level 0). A production system is an automated pipeline with continuous training, monitoring for data and concept drift, governance and human oversight (Level 1–2). The gap between them — the industrialisation work — is where most programmes stall.

How important is change management to AI success? Decisive. Prosci's research shows projects with excellent change management meet their objectives 88% of the time, versus 13% for those with poor or absent change management. A deployed model nobody adopts returns nothing — much of a project's ROI is "people-dependent" on actual adoption, proficiency and use.

Stuck in pilot purgatory — and want to know exactly what's blocking the climb to production? That's the work we do: industrialising AI with MLOps and AIOps, and standing up the operating model to run it. The MLOps Maturity Scorecard will find your gaps in about ten minutes — or, if the money question is the one in the room, see how the economics stack up in our guide to funding AI that pays back.

/ go deeper

Put this to work on your numbers.