Blog
/
AI implementation

Why AI implementations die after the POC

Table of contents
Section oneSection two

Every failed AI implementation has the same first act. The proof of concept worked. The demo drew actual applause. Somebody said "this will save us thousands of hours," and meant it. Then months passed, the rollout date slid twice, the champion changed jobs, and the project entered the peculiar afterlife where nobody kills it and nobody ships it. Eventually a budget review does what no meeting had the nerve to do.

The scale of this pattern is now among the best-documented facts in enterprise software. MIT's Project NANDA found that 95 percent of enterprise generative AI pilots produced no measurable profit and loss impact. RAND, drawing on interviews with 65 experienced practitioners, notes that by some estimates more than 80 percent of AI projects fail, twice the failure rate of comparable IT projects. S&P Global Market Intelligence found the share of companies abandoning most of their AI initiatives jumped from 17 percent to 42 percent in a single year, and Gartner forecasts that over 40 percent of agentic AI projects will be canceled by the end of 2027.

Read those studies together and something stands out: none of them blames the models. RAND's root causes are organizational almost across the board, and MIT's divide separates companies not by which AI they bought but by whether it was adapted into real workflows. The technology cleared the bar. The implementations died anyway, and they died in a specific place: the gap between proof of concept and production.

This essay is about that gap, the five mechanisms that operate inside it, and the discipline visible in the small group that crosses it.

The POC proves the wrong thing

A proof of concept answers the question it was designed to answer: can the model do the task? Against sample data, in a clean environment, with the builder present, the answer is very often yes, and honestly so. The demo is not a lie. It is just an answer to a question that stops mattering the moment it is answered.

Production asks a different question: can this system do the task inside our environment, against our real systems, under our security constraints, every day, when nobody is watching? That question is not a harder version of the first one. It is a different subject. The first is about intelligence; the second is about integration and operations. Passing the first while failing the second is not a paradox, and the 95 percent are not fools. They simply funded an answer to question one and assumed it covered question two.

The gap between the questions has a job title now. Demand for forward deployed engineers, the people whose entire role is carrying working AI across that gap into specific customer environments, grew so fast that Indeed postings rose 5,230 percent above their January 2025 baseline by April 2026. The market diagnosed the cliff before most postmortems did.

The afterlife, month by month

The death is slow enough that nobody attends it, so it is worth narrating the composite timeline, because most readers will recognize their own calendar in it.

Month one is the victory lap: the demo circulates, the steering committee asks for a production plan, and an engineer produces the first honest integration estimate, which is larger than the entire pilot budget. Nobody panics; the number is parked for planning. Month two, the security questionnaire arrives, forty rows deep, and the row about credential storage gets answered aspirationally. Month three, the champion who sponsored the pilot is promoted, inherits a bigger portfolio, and hands the project to someone who was not in the room for the applause. Month four, the engineer who built the POC is pulled to a customer escalation, taking the only working knowledge of the prototype along. Month five, the quarterly plan arrives and the integration estimate from month one finally meets the roadmap it must displace; it loses politely, deferred one quarter. Month six, the deferral repeats, and the project stops appearing in status decks, which in enterprise software is what death looks like.

At no point did anyone decide to kill it. At no point did the model underperform. The implementation simply needed a lane, a custody answer, a vehicle, and an owner, and each month one of them failed to materialize on schedule. Multiply this calendar by the 42 percent of companies now abandoning most AI initiatives and you have the industry's quiet epidemic: not failed AI, but unfielded AI.

The five mechanisms of death

Watch enough implementations stall and the same five mechanisms appear, usually in combination.

1. The integration gap. The POC ran on exported CSVs and a friendly sandbox. Production means the customer's actual ERP, their actual permissions model, their actual data shapes, none of which resemble the sandbox. MIT's research draws the line precisely here: generic tools stall in enterprise use because they do not learn from or adapt to workflows, and the 5 percent that succeed run systems adapted to the specific workflow. Adaptation is integration work, integration work is custom code, and custom code for one environment is precisely the work most organizations have no lane for. So the POC waits in a queue behind the product roadmap, and queues are where momentum goes to die.

2. The security stall. The POC never met the security team, because pilots run on samples and goodwill. Production means real credentials to real systems, which means a review: where does this run, where do our API keys live, who can read them, what happens when someone leaves? If the honest answer is "in the engineer's env file," the review does not fail the project outright. It does something slower: it generates questions, the questions generate meetings, and the meetings generate the delay in which sponsors lose interest. Implementations rarely die of a rejected review; they die of a review that never quite concludes.

3. Deployment infrastructure that nobody budgeted. RAND names it directly: among the five root causes of AI project failure is underinvestment in the infrastructure needed to deploy and maintain models in production. The pilot budget bought model access and a demo. Nobody budgeted where the connective code would run, how it would be versioned, who would monitor it, or how the next change would ship. The project reaches the end of the POC with working intelligence and no vehicle, and building the vehicle from scratch costs more than the pilot did, which is an awkward thing to explain to the steering committee that thought the hard part was done.

4. The person who was the project leaves. POCs are built by individuals: an enthusiastic engineer, a consultant, a champion with weekend energy. Whatever they built lives where they built it, its logic lives in their head, and its credentials live in their password manager. When they change roles, leave, or simply get reassigned, the implementation does not fail loudly. It becomes unownable: running but frozen, because nobody can safely touch it. This is the bus-factor failure, and in AI implementations the bus factor of the pilot is almost always exactly one.

5. Leadership kills it by redirecting it. RAND's interviewees ranked this first: 84 percent of them cited leadership-driven causes as the primary reason AI projects fail, projects pointed at the wrong problem, success never defined, priorities shuffled mid-flight. After the POC is exactly when this bites, because the demo's applause invites scope: the assistant that answered order questions should now also do returns, and forecasting, and maybe the Berlin office. The project that proved one workflow is re-aimed at six, proves none of them, and the eventual cancellation gets blamed on the technology that did its job in act one.

Notice what the five have in common. Every one of them is about the scaffolding around the model: the integration lane, the credential custody, the deployment vehicle, the ownership continuity, the scope discipline. The model appears nowhere on the list. That is the finding hiding in all four studies, and it is cheerful news in disguise, because scaffolding, unlike intelligence, is a solved category of problem when someone owns it.

An AI POC does not fail in production. It fails to reach production, and the distance it fails to cross is measured in integrations, credentials, and owners, not in model quality.

Nobody's job, therefore nobody's fault

There is a structural reason the five mechanisms go unowned, visible the moment you ask whose job the gap was. Product engineering looks at the customer-specific integration and correctly says it is not roadmap work: it serves one environment, not the product. IT looks at it and correctly says it is not their application: they did not buy it and cannot support what they cannot see. The AI vendor looks at the customer's ERP and correctly says environments are the customer's side of the line. The consultant who built the pilot has an end date, and the end date came. Four parties, four correct answers, and a working system dead in the space between them.

This is why the fix is organizational before it is technical, and why the companies crossing the gap keep converging on the same structure: a named lane for environment-specific work, whether staffed in-house or through the consultant-versus-FDE decision, with its own operational backbone so the lane never has to borrow one from teams that correctly refuse to lend theirs. The gap stops killing projects the day it appears on an org chart with a budget line under it.

What the surviving 5 percent do

Invert the mechanisms and the survivors' playbook writes itself. They pick one workflow with a measurable outcome and refuse scope until it ships. They adapt the system to the workflow, which means they fund the custom integration work as the project rather than a footnote. They meet the security team before the pilot, not after, and arrive with a custody answer instead of an env file. They decide on day one where the connective code will live, how it versions, and who inherits it. And they treat the POC as the first version of production rather than a disposable demo.

The same disciplines read differently depending on which side of the deal you sit. For the buyer, they are risk management. For the vendor, they are a sales weapon: the software company whose POC ships on production rails converts technical wins while competitors are still scheduling their rebuild, and the difference shows up directly in sales-cycle length. Every mechanism above that kills a buyer's internal project also kills a vendor's deal, just with a commission attached.

That last discipline deserves its own sentence, because it is the cheapest of the five and the most often skipped: build the pilot on the path it would take to production, so that graduating is a promotion rather than a rebuild. Teams that build the POC on production rails can graduate it in days. Teams that build it on a laptop discover the POC was a beautiful sketch of a thing that now must be built twice.

The lane this creates, and the tooling it needs

Companies that internalize all this usually arrive at the same organizational answer: a lane for customer-and-environment-specific work, staffed by forward deployed engineers, with its own operational backbone so it never competes with the product roadmap. Whether to staff that lane in-house or rent it is its own decision, covered in our guide to AI implementation consultants versus FDEs, and the software for the whole stack is mapped in the AI implementation software guide.

The backbone half is the part we build. Archway is the platform for forward deployed engineering teams: the connective code that turns a pilot into a production system ships as bridges, serverless functions connecting the product to one customer's systems, deployed in minutes with no core-product change. Credentials sit in an AES-256 vault with zero-access custody, so the security review from mechanism two gets a one-sentence answer. Every version is retained and owned by the organization. When the person who built the pilot leaves, you reassign the bridges and leave a note; the running estate stays. That architecture retires mechanisms three and four. A POC built as a bridge does not graduate by being rebuilt; it graduates by being promoted, which is the entire difference between act one and act three.

The pilot contract: six lines to write before the next POC

Prevention compresses into a one-page agreement worth writing before any pilot starts, because every mechanism above is cheapest to kill before the applause. Six lines do it. The workflow: one process, named, with the person who owns it in the room. The number: the metric that defines success, with its current baseline measured before the demo. The rails: the pilot is built on the path production would use, or the rebuild cost is estimated in the same document. The custody: where credentials live during the pilot, an answer that would survive the security review reading it. The operator: the named person who runs this in month two, which is how you discover in week zero that no such person exists. The gate: the date and criteria on which this either gets a production budget or gets killed, pre-committed, so the afterlife has no room to form. Ten minutes of writing, and the five mechanisms lose their habitat: scope cannot drift from a named workflow, the security stall cannot ambush a custody answer that already exists, and the quiet deferral cannot outlive a pre-committed gate. Most organizations resist the contract not because it is hard but because it makes the pilot falsifiable, which is precisely its value.

Autopsy your own stalled pilot in five questions

If you have a pilot in the afterlife right now, five questions locate which mechanism has it. Who owns getting it into the customer's or the company's real systems, by name, this quarter? Where would its credentials live in production, and would that answer survive a security review read aloud? Where does the connective code run today, and what happens to it if its author is reassigned tomorrow? What is the one workflow and the one number that define success, and has either changed since the demo? And if the answer to any of these is a shrug: which team's budget does fixing the shrug come from? A stalled pilot with five answers is a scheduling problem. A stalled pilot with five shrugs is already dead, and knowing it early is worth a quarter of pretending.

The failure statistics at the top of this essay are usually quoted as gloom. They are better read as a map: 80 to 95 percent of the market is losing implementations to five mechanisms that have nothing to do with intelligence and everything to do with the last mile. The companies that fix the last mile get to keep the applause from act one. If your pilots keep dying between the demo and the deployment, see what Archway does; the first two bridges are free, and the next POC can be built on the rails it will retire on.

Sources and notes