Pilot vs production: the checklist that decides whether an AI project ever pays
A demo that worked once in a meeting room is not a production system, and the gap between the two is where most AI budgets disappear. Here is what actually has to change, and a checklist to run before you call something live.
"We ran a pilot" and "it's in production" get used almost interchangeably in meetings, and the gap between them is where most AI budget disappears. A pilot proves an idea can work. Production means it keeps working, on real data, at real volume, without someone watching it constantly. Confusing the two is how a business ends up with a demo it is quietly proud of and nothing running on a Tuesday morning.
What actually separates a pilot from production?
| Pilot | Production | |
|---|---|---|
| Data | A clean, chosen sample | The real inbox, the real files, the messy long tail |
| Users | A few willing volunteers | Everyone who needs it, including the sceptical ones |
| Ownership | Whoever built it | A named person inside the business, with time set aside for it |
| Watching it | Someone checks occasionally | Volume, errors and exceptions are logged and reviewed on a schedule |
| Failure | Shrugged off; "it's just a pilot" | A rollback plan that has actually been tested |
| Success measure | "It worked" | A baseline it is measured against, in pounds or hours |
Why so few pilots cross the line
MIT NANDA's widely cited 2025 study of enterprise AI found 60% of organisations had evaluated AI tools, 20% had reached a pilot, and only 5% had a custom tool running in production with a measurable return. Each stage loses most of what came before it. The same study found mid-market companies that did reach production moved fast once they committed, about 90 days from pilot to full rollout, against nine months or more for large enterprises slowed by procurement and governance. Speed is the smaller company's advantage; the discipline below is what it is spent on.
Gartner's own forecasting shows the problem getting worse, not better. In July 2024 it predicted 30% of generative AI projects would be abandoned after proof of concept by the end of 2025. By January 2026 it reported the real figure as at least 50%, citing poor data quality, unclear business value and escalating costs: three things a pilot, by design, never has to face. In the UK specifically, the ONS found AI use among businesses with ten or more staff rising from about 12% to 35% since late 2023, but the number of technologies in active use per adopter barely moved (1.4 to 1.6), which is a reasonable proxy for how many pilots are quietly staying pilots.
Six things production needs that a pilot usually skips
- A named owner, with time set aside. Not "IT will keep an eye on it." One person whose job includes checking exceptions and who gets asked first when something looks wrong.
- A rollback that has actually been tried. Not a plan on paper. Switch it off, or fall back to the manual process, once in a test run before anyone relies on it for real.
- Monitoring from day one. Volume, error rate, and how often a person has to override the system. If nobody is looking at these numbers, nobody will notice when they get worse.
- A test on real data, including the messy bits. The unusual invoice, the email with three attachments, the customer who writes in bullet points. A pilot that only ever saw clean examples has not been tested.
- A baseline, written down before launch. What did the process cost, in hours or pounds, before this started. Without it, "is this working" is a feeling, not an answer.
- A human approval step on anything that matters. Nothing reaches a customer, moves money, or changes a record without someone checking it first, at least until the system has earned the right to run unwatched (see human-in-the-loop automation).
None of these six are difficult on their own. What usually happens is that a pilot proves the AI part works, everyone is pleased, and the project quietly skips straight to "roll it out to everyone" without doing any of the above, because the exciting part already happened.
A staged rollout that actually works
The fix is not more caution, it is a narrower first step with a clear exit. Pick one slice of the real work, not a sample: one supplier's invoices, one team's enquiries, one branch's reporting. Run it in the real system, with the real data, alongside the existing manual process rather than instead of it, for a fixed period agreed in advance. Set the numbers that decide whether it continues before you start, not afterwards. If it clears the bar, widen it one slice at a time, repeating the monitoring and baseline step each time rather than assuming what worked for ten invoices a day will hold at a thousand. If it does not clear the bar, you have a stop date and a reason, not a slow fade into "we're still looking into it."
This is also why a pilot that never had a stop date is a warning sign on its own. If nobody can say when or how it would be judged to have failed, it was never really being tested, it was being hoped for. For the fuller list of where AI projects stall after a successful pilot, see why AI projects fail in mid-sized companies.
How does M22 handle this?
M22 Consultancy treats the pilot-to-production gap as the main risk in any AI project, not an afterthought. The AI Audit, a fixed fee from £1,500 agreed before day one, sets the baseline before anything is built, so there is a number to measure against from the start. Where M22 builds, the first version runs on a narrow, real slice of the work inside the client's own systems, with a named approval step on anything touching customers, money or records, and a monitoring view the client can read without asking M22 for a status update. The run-and-improve retainer then tracks volume, exceptions and the baseline every month, which is what turns "we ran a pilot" into a system still running a year later. Book a thirty-minute call to talk through where your own pilot stands, or see what this typically costs.
- MIT NANDA: The GenAI Divide, State of AI in Business 2025
- Gartner: Why Many Generative AI Projects Lack Value (January 2026)
- Gartner: Gartner Predicts 30% of Generative AI Projects Will Be Abandoned After Proof of Concept by End of 2025 (July 2024)
- ONS: Artificial intelligence in UK businesses, 2023 to 2026