Why AI Pilots Fail Before Anyone Writes Code
Ehtisham ul Haq
Founder of SeedInov. AI engineer building production-ready AI systems for businesses in 8 countries.

MIT's NANDA initiative reported in its 2025 State of AI in Business study that roughly 95% of enterprise generative-AI pilots produced no measurable effect on the P&L. Gartner expects more than 40% of agentic-AI projects to be cancelled before the end of 2027, principally for unclear business value. Those numbers get quoted constantly and almost always for the wrong purpose: as evidence that the technology is overhyped.
They are not evidence of that. In nearly every failure we have been asked to examine, the model worked. It classified the tickets, it extracted the invoice lines, it drafted the responses. The project died anyway, and it died for reasons that were fixed in place weeks before an engineer opened a terminal.
A pilot that cannot be judged is a pilot that will be cancelled. Most are unjudgeable by construction.
Failure one: no baseline was ever recorded
This is the single most common cause and the most avoidable. A team runs a pilot, the pilot appears to work, and then somebody asks whether it is actually better than what the humans were doing. Nobody measured what the humans were doing, so the question cannot be answered.
Without a baseline, every conversation about the system becomes a matter of impression. The sponsor thinks it helped. The team that uses it says it is roughly the same. Finance sees a bill. In that argument the bill wins, because it is the only number in the room.
What a baseline needs to contain, recorded before the pilot starts:
- Time per case. Measured across enough real cases to be representative, including the awkward ones people avoid demonstrating.
- Volume. How many cases per week. This is what converts a per-case saving into an annual number.
- Error and rework rate. How often the current process gets it wrong, and what a redo costs. This is the benchmark accuracy gets compared against, and it is almost always worse than people assume.
- Wait time. How long a case sits between handoffs. Frequently the largest saving available, and the one nobody thinks to measure.
If a pilot produced no number you could measure again six months later, it was a demonstration, not a pilot.
Failure two: nobody owned the metric
Pilots are usually sponsored by someone senior and run by someone technical. Neither of them owns the operational number the system is meant to move. So when the pilot ends and the question becomes who funds it, operates it, and defends it at budget time, the answer is nobody in particular.
This is why so many programmes stall specifically between pilot and production rather than at the start. The technology cleared its bar. The organisational question was never asked. The test is blunt: can you name the person whose performance review is affected by the number this system moves? If you cannot, the pilot has no owner and will not survive contact with a budget cycle.
Failure three: the data was never checked against the task
Teams assess data quality in the abstract, conclude that it is "messy but workable", and start building. The relevant question is much narrower: can the specific fields this specific use case needs support it?
The failures are specific and repetitive. A status field that is free text, so the model has to interpret twelve spellings of the same state. A customer identifier that means different things in two systems. Historical records that lack the outcome you want to predict, because nobody recorded it. Documents that exist only as photographs of printouts.
None of these show up in a data-quality score. All of them stop a use case. The productive version of the question is: name the tables, name the fields, and check whether the values in them can answer the question you intend to ask.
Failure four: the accuracy target was never agreed
Late in a project someone asks how accurate the system is. The answer is 91%, and the room divides into people who think that is impressive and people who think it is unacceptable. Both are arguing without a reference point.
Accuracy is meaningless as an absolute. The reference is the current process. If your team currently misclassifies 12% of tickets, a system at 6% with a confidence threshold routing uncertain cases to a human is a clear improvement, even though 6% sounds bad said out loud. Agreeing that comparison before the build, in writing, prevents a successful system being killed by a number that was never contextualised.
The other half of this is deciding where the confidence threshold sits. Above it the system acts; below it a person decides. That boundary is a business decision about the cost of being wrong, not a technical one, and it should be made by the people who bear that cost.
Failure five: the workflow was never changed
Licences get rolled out, training gets delivered, and nothing about anybody's day changes. People open the tool, find no obvious place for it in the sequence of things they actually do, and quietly stop.
Adoption failures are almost always process design failures wearing a training costume. If the automation lives in a separate application that people must remember to open, it will stop being used within a quarter regardless of how well it performs. The systems that survive are embedded in the tool where the work already happens, so using them is not a decision anybody has to make each morning.
Failure six: the redeployment question was avoided
If a workflow automation removes six hours a week from a team's workload, somebody in that team is going to ask what happens to those six hours, and probably not out loud. Organisations that have not answered this find that process knowledge becomes strangely hard to obtain during the build, exception cases are described vaguely, and edge cases surface late.
Nobody is being obstructive. People are being sensible about a question their employer declined to answer. Settling it in writing before the build, whatever the answer is, removes a failure mode that otherwise appears as unexplained technical delay.
Failure seven: the governance route did not exist
A pilot runs in a sandbox where nobody needs approval. Then it has to go into production, and it emerges that no one knows who signs off a model that makes decisions about customers, what has to be logged, or what happens when it is wrong.
In regulated environments this can stop a working system permanently. In Saudi Arabia the framework is layered: PDPL for personal data, SDAIA's AI Ethics Principles for the AI system itself, NCA controls for the platform underneath, and then sector rules from SAMA, the Ministry of Health, CST or others depending on who you answer to. Any one of them can rule out an architecture that was otherwise the obvious technical choice.
Discovering the constraint after choosing the model is how teams end up rebuilding a system that performed perfectly well. Establish it in scoping and model selection becomes a straightforward engineering question with several acceptable answers.
What a pilot that survives looks like
None of the fixes are technically difficult. They are all decisions taken before the build, and taken in writing.
- One workflow, measured first. Time per case, volume, error rate and wait time recorded before anything is built.
- A named owner. One person whose operational number the system is meant to move, who will be there at budget time.
- Data checked at field level. The specific tables and fields the use case needs, verified against the task rather than scored in general.
- An agreed accuracy target and confidence threshold. Set against the recorded baseline, with the human-in-the-loop boundary decided by whoever bears the cost of an error.
- Embedded in the existing tool. No new application anybody has to remember to open.
- The redeployment question answered. In writing, before the build, whatever the answer is.
- A governance route confirmed. Who approves production, what gets logged, which regulator applies, settled before the model is chosen.
Do those seven things and the pilot becomes judgeable. A judgeable pilot that fails is cheap and informative. An unjudgeable pilot that succeeds is indistinguishable from one that did not, which is precisely how 95% of them end up in the same bucket.
Where to start
Pick the process that costs your team the most hours and measure it for two weeks before deciding anything. That measurement is useful whether or not you ever build the automation: it tells you what the process actually costs, which is information most organisations do not have about their own operations.
If you want that done independently, our AI workflow audit is the two-week version with the baseline and the ranked opportunities written up. If the problem is broader than one process, the AI readiness assessment scores five dimensions separately and publishes the framework so you can run it yourself first. Both sit under AI consulting, and there is a Saudi-specific version covering SDAIA and PDPL.



