Pilot to production: what the minority who ship do differently
AI Consulting and Readiness

Pilot to production: what the minority who ship do differently

AI Consulting and Readiness
Pilot to production: what the minority who ship do differently

MIT’s 2025 research found 95 percent of enterprise AI pilots produced no measurable return. Here is what the minority who actually shipped to production did differently.

Pilot to production: what the minority who ship do differently
3 min read
Key takeaways
  • What does the research actually say happened to most pilots?
  • Why does a pilot die even when the technology works?
  • What does scoring a use case before building actually involve?

Answer: MIT’s 2025 enterprise AI research found roughly 95% of organizations reported no measurable return from their AI pilots. The minority who did ship to production shared 3 traits: they scored use cases on payback and risk before building, they funded the unglamorous back-office workflows instead of only the visible sales and marketing pilots, and they wrote a build-or-buy decision down instead of letting a pilot run indefinitely.

What does the research actually say happened to most pilots?

MIT Project NANDA’s 2025 State of AI in Business study (preliminary, not yet peer reviewed) found that despite billions in enterprise AI spend, approximately 95% of organizations reported no measurable return on their generative AI initiatives. The money concentrated in sales and marketing pilots — the visible, easy-to-greenlight category — while the highest measured returns sat in back-office automation, which received the least funding. That’s not a story about AI not working. It’s a story about budget going to the pilots that were easiest to approve, not the workflows that actually paid back.

Why does a pilot die even when the technology works?

Because “the technology works” and “this specific workflow is worth building” are different questions, and most pilots only answer the first one. A pilot that produces an impressive demo gets funded for a second phase almost automatically, regardless of whether anyone scored what it would actually save or whether the risk of it being wrong was ever quantified. Without that scoring step, a company ends up with a portfolio of pilots that all “seem promising” and no defensible way to decide which one deserves the build budget. Most stall not because the model failed, but because nobody could answer “compared to what, and worth how much” when the renewal conversation came up.

What does scoring a use case before building actually involve?

Every candidate workflow gets scored on payback and risk with the arithmetic shown, not a gut-feel ranking. Payback means an honest estimate of hours saved or revenue protected, using your actual data rather than a vendor’s benchmark. Risk means naming what happens if the system is wrong — a wrong answer in an internal drafting tool costs an edit; a wrong answer in a customer-facing billing decision costs a support ticket and possibly a customer. Scoring both before committing budget is what separates a company that funds 1 workflow and ships it from one that funds 5 workflows and ships none, because the fifth pilot never displaced the first when it turned out to be worse.

Why do the minority fund back-office workflows first?

Because that’s where MIT’s research found the returns actually concentrated, and because back-office workflows tend to have a clear, measurable “before”: a specific process a specific person repeats on a specific cadence. That clarity makes the payback arithmetic tractable in a way a novel customer-facing pilot rarely is. A sales-assist pilot is exciting to demo in a board meeting; an invoice-reconciliation automation is boring, measurable, and profitable, and the companies that ship don’t let the first quality outrank the second.

What does “write the decision down” actually mean in practice?

It means the pilot has a defined end date and a defined decision: build, buy, or kill, made by a named decision-maker in a scheduled review, not an indefinite “let’s keep iterating.” A pilot without a decision date doesn’t die, exactly — it just never converts, quietly consuming attention and a line in next year’s budget while nobody is accountable for calling it. The minority who ship treat the decision review as a deliverable, with a build-or-buy recommendation and a monthly run-cost estimate attached, the same way they’d treat any other deliverable with a deadline.

That’s the structure behind our own 2-week AI readiness sprint: score every candidate workflow on payback and risk with the arithmetic shown, then end in a build-or-buy decision review rather than a slide deck. It’s built specifically to close the gap MIT’s research surfaced — before the build budget is committed, which is the only point at which the decision is still cheap to get right.

Source: MIT Project NANDA, State of AI in Business, 2025 (preliminary, not peer reviewed).

Abdullah Shah
Written by

Abdullah Shah

Comments

Join the discussion. Be constructive, on-topic, and kind.

Leave a reply

Your email address will not be published.

Thanks. Your comment has been submitted for moderation.
Layout