Search "AI implementation roadmap" and you will be offered a transformation program: phases named Discover, Align, Scale, a maturity model, a steering committee. That document is not what the failure data describes, and it is not what this guide is. The companies drowning in pilots already have transformation programs. What they do not have is a sequence for getting one system into one real workflow, under real credentials, with a named owner, and keeping it there when the person who built it changes jobs.
The numbers are public and they keep repeating. MIT's Project NANDA found that 95 percent of enterprise generative AI pilots produced no measurable profit and loss impact, based on 150 interviews, a survey of 350 employees, and analysis of 300 public deployments. S&P Global Market Intelligence, as reported by CIO Dive, found the share of companies abandoning most of their AI initiatives jumped from 17 percent to 42 percent in a year, with the average organization scrapping 46 percent of proofs of concept before production. Gartner forecasts that over 40 percent of agentic AI projects will be canceled by the end of 2027. RAND, interviewing 65 experienced practitioners, notes that by some estimates more than 80 percent of AI projects fail, twice the rate of comparable IT projects, and that 84 percent of interviewees cited leadership-driven causes as the primary reason.
None of those studies blames the models. MIT's divide is whether systems were adapted into real workflows. RAND's root causes are organizational, including underinvestment in the infrastructure to deploy completed models. CIO Dive's S&P writeup names cost, data privacy, and security risks as the top obstacles, and cites an Informatica finding that two-thirds of enterprises cannot transition pilots into production even as they increase generative AI investment. The roadmap that matches this evidence is a last-mile roadmap: how to implement AI in a specific environment, not how to announce an AI transformation.
This guide is that sequence. After-POC failure modes belong to why AI implementations die after the POC; this page hops there and does not retell the autopsy. Consultant versus in-house belongs to AI implementation consultants. The boardroom sentence about governance belongs to AI transformation is a problem of governance. What remains is the work of implementing: pick, staff, build on rails, integrate, vault, operate, expand.
What this roadmap is not
It is not a use-case catalog. Lists of forty opportunities are how programs lose a year before a single workflow ships. RAND's advice from the same study is to choose enduring problems and commit a team to one of them for at least a year. If a candidate workflow is not worth a year, it is not worth a roadmap line.
It is not a model-selection exercise. The models are capable. The implementations die in integration, custody, ownership, and scope. Spending the first quarter comparing frontier APIs is how a program postpones the part that actually fails.
It is not a change-management novel. Training and adoption matter for internal tools. They do not substitute for a running system in the customer's or the company's real environment. MIT's 5 percent succeeded by adapting systems to workflows, which is implementation work, not a lunch-and-learn.
It is not a generic "how to implement AI" maturity staircase. If you searched that phrase, you are in the right place, but the staircase you were offered usually ends at "scale across the enterprise" without ever specifying where the connective code runs, who can read the credentials, or what happens when the builder leaves. Those three questions are the roadmap.
The last-mile diagnosis, before you draw phases
An AI implementation has two different products inside it, and most roadmaps fund only the first. Product one is intelligence: a model that can do a task in a demo. Product two is the last mile: the customer-specific or environment-specific code that connects that intelligence to the systems where work already happens, under the security constraints those systems impose, on a path that someone else can operate. Product one is cheap now. Product two is the job forward deployed engineers were hired to do, which is why Indeed postings for the role sat 5,230 percent above a January 2025 baseline by April 2026.
If you cannot answer the following four questions about your current AI work, you do not have an implementation. You have a demo with a budget line.
- Which single workflow, named, with a current baseline metric, is this for?
- Who owns getting it into the real systems, by name, this quarter?
- Where do the credentials live, and would that answer survive a security review read aloud?
- What happens to the running system if the person who built it is reassigned on Friday?
Question two is a staffing question; how to staff an FDE team is the headcount version. Question three and four are infrastructure questions; they are also the governance questions, which is the argument of the governance essay. This roadmap assumes you will take both seriously and then does the sequencing.
Step 1. Name one workflow and one number
Start smaller than the steering committee wants. One process, owned by a person who is in the room, with a metric that already exists. "Invoice coding in AP, currently 14 minutes a document, owned by Priya" is a workflow. "Finance transformation" is not. "Tier-one support deflection on the returns queue" is a workflow. "Customer experience AI" is not.
Write three lines before anyone opens a notebook:
- The workflow. The steps a human currently takes, including the ugly ones, the spreadsheet, the exception pile.
- The number. The baseline, measured before the demo, so success cannot be redefined after applause.
- The systems. The two or three systems of record the workflow already touches. If you cannot name them, you are not ready to implement. You are ready to interview.
Refuse additional workflows until the first one is in production, meaning a real user in a real environment, on a real credential path, with a named operator. RAND's year-long commitment is the discipline underneath this stubbornness. Scope that expands after the demo is mechanism five in the after-POC essay; kill it here, on the page, before it has a calendar.
A useful filter, stolen from the failure data: if the workflow does not require the system to adapt to this organization's data, permissions, and exception paths, you may not need an implementation at all. You may need the enterprise assistant layer, ChatGPT Enterprise or Claude for Work, which MIT found people actually adopt for individual tasks and which then stall when the work demands context. Do not run a last-mile roadmap at a chatbot problem. Do not run a chatbot rollout at a last-mile problem.
Step 2. Staff a lane, not a committee
The last mile is nobody's job in most companies, which is why it does not happen. Product engineering correctly refuses customer-specific or environment-specific work that will never generalize. IT correctly refuses to operate a system it cannot see. The vendor correctly draws a line at the customer's environment. The committee correctly meets monthly. Four correct answers, one dead implementation.
You need a named lane for environment-specific work, with a budget line, staffed by people whose job is to ship in someone else's systems. That lane has a job title now: forward deployed engineer. Whether you hire it or rent it is a real decision. Renting is the consultant guide: rates, engagement shapes, the four questions that separate an engagement from a PDF, and when to convert the embed into a hire. Hiring is the staffing guide: there is no public accounts-per-FDE benchmark, the sixth hire is usually a coverage hire, and sales engineers are not FDEs with a new title. Do not re-litigate those pages here. Make the decision, put a name on the lane, and give that name a deploy path that does not transit the product backlog.
If you staff the lane with a committee instead of a builder, you have staffed oversight. Oversight cannot vault a credential. If you staff it with a data science team that has never deployed into a customer's ERP, you have staffed product one. The roadmap needs product two.
A minimum viable lane, for a company implementing AI inside its own operations rather than at customers: one engineer who can ship full-stack work, one operator who owns the workflow in the business, and one security partner who will see the custody design before the pilot, not after. For a vendor implementing at customers, replace "operator" with the customer's champion and add the FDE. Either way, the lane has to be allowed to ship without waiting for the core product's release train. If every environment-specific commitment has to transit the product backlog, you have rebuilt the bottleneck the lane exists to remove.
Step 3. Build the first version on production rails
This is the step most roadmaps skip, and it is the one the POC essay exists to explain. A proof of concept that runs on exported CSVs, in a sandbox, with the builder present, answers a question that stops mattering the moment it is answered. Production asks whether the system can do the task in the real environment, against real systems, under real security constraints, when nobody is watching. Passing the first and failing the second is not a paradox. It is what 95 percent looks like from the inside.
The rule: the pilot is built on the path production would use, or the rebuild cost is estimated in the same document as the pilot budget. Teams that build the POC as a laptop demo discover they must build it twice. Teams that build it as the first version of the production system graduate it by promoting it. The afterlife of stalled pilots (champion leaves, security questionnaire never concludes, engineer reassigned, roadmap deferral) is narrated in the POC essay. This roadmap's job is to refuse that afterlife as a phase.
Practically, "on rails" means five properties on day one, not at go-live:
- There is one sanctioned path from working code to running code.
- The code runs somewhere owned, not on the laptop that wrote it.
- Credentials sit in custody the engineer can reference and cannot read.
- Every version is retained.
- Someone other than the author can be assigned the running system.
Those five verbs are FDE ops. They are also the cheapest insurance you can buy against the five death mechanisms. If your current pilot cannot satisfy them, you are not in step 3. You are in a demo, and you should not call it an implementation on a slide.
Step 4. Integrate into the real systems, on purpose
MIT's researchers drew the line here: generic tools stall in enterprise use because they do not learn from or adapt to workflows, and the 5 percent that succeed run systems adapted to the specific workflow. Adaptation is integration. Integration is custom code. Custom code for one environment is the work most organizations have no lane for, which is why it waits behind the product roadmap until the sponsor loses interest.
Name the integrations as the project, not as a footnote. For each system of record from step 1, write down: what we read, what we write, whose credentials, what failure looks like to the user, and who gets paged. Then build those connections as first-class artifacts, one environment at a time. This is customer-specific integration work, even when the "customer" is your own finance team. The shape does not change because the logo is internal.
Two sequencing rules that save quarters.
Narrow data, then widen. Ground the system on the documents and records this workflow needs, ship, and widen later. Implementations that fail often spend two quarters building a perfect knowledge base for a workflow that never goes live. The data layer is real; it is also a famous way to postpone the last mile.
Do not wait for a unified API to save you. Standards and connectors shrink protocol work. They do not build the connector to this organization's peculiar ERP configuration, host it, hold its credentials, or keep it running after the author changes jobs. Budget the last mile as if the standards already existed, because even when they do, the operational remainder is still the job.
If the integration estimate from the engineer is larger than the entire pilot budget, believe the engineer. That number is the implementation. The pilot budget that ignored it was the demo budget. CIO Dive's S&P story is, in part, what happens when that surprise arrives after applause: cost becomes the reason the project is abandoned, which is another way of saying the last mile was never priced.
Step 5. Pass security as architecture, not as a questionnaire at the end
The security stall is mechanism two of after-POC death: the review does not fail the project outright, it generates questions, the questions generate meetings, and the meetings generate the delay in which sponsors lose interest. S&P's respondents, per CIO Dive, already cite data privacy and security risks among the top obstacles. You will not talk your way around that with a policy binder.
Arrive at the review with answers that are properties of the system:
- Where it runs.
- Where the credentials live, and the fact that engineers cannot read the values.
- Who can deploy a change, and how you would revoke that.
- What the running version is, and how you would roll it back.
- Who inherits the system if the author is gone.
Those answers are governance as deployment, which is the essay to read if your transformation office is currently funding committees. This roadmap only needs one operational consequence: meet security before the pilot, not after, and do not start a POC whose custody story is an env file. An AES-256 vault with zero-access custody is what that answer looks like when it is infrastructure rather than a promise. The questionnaire then has somewhere to land, which is how reviews conclude.
If your organization has an AI review board, make it a customer of this infrastructure rather than a substitute for it. The board is good at "should we automate this decision." It is bad at "what is running where." Wire it to the inventory and it becomes useful. Ask it to certify a laptop demo and it becomes delay.
Step 6. Operate it as production, including the day someone leaves
A system that cannot be changed by a second person is not in production. It is a working prototype with users. RAND's infrastructure finding is about this moment: completed models that cannot be deployed and maintained, underinvestment presenting as project failure. The operating bar is boring on purpose.
- The named operator from step 2 can ship a small, safe change without the original builder on the call.
- Logs, versions, and credential paths are findable without the original builder's memory.
- A handoff is a reassignment, not a two-week pairing marathon. The FDE departure playbook is the full procedure when the person is actually leaving; do not wait for a resignation to discover you needed it.
- Evaluation is continuous against real cases, not a one-time demo metric. If you cannot show how often the system is right, someone else will decide the answer is "not often enough," and that is how projects die in budget review even after they "shipped."
This is also where tribal knowledge either gets captured as artifacts or becomes the bus factor. On many AI implementations the bus factor of the pilot is exactly one. Operating as production means driving that number to at least two, on purpose, before you expand.
Step 7. Expand only from a live system
The temptation, once the first workflow works, is to point the same system at six more. That is how leadership kills implementations by redirecting them, RAND's most-cited failure pattern. Expansion is earned, and it is earned by a live system with a measured number, not by a promising demo.
The expansion test:
- The first workflow is in production by the definition in step 6.
- The number moved, or you have an honest postmortem of why it did not and a change already shipped.
- The second workflow shares systems, credentials paths, or patterns with the first, so you are compounding, not starting a second hero project.
- The lane that shipped the first still has coverage. Expanding onto an already overloaded FDE or consultant is how the first system freezes the week the second begins.
Reuse is an infrastructure property. You cannot reuse what you cannot find, version, or reassign. Pattern extraction, the Palantir loop in which field work teaches the product, only works if the field work is visible. That is the argument for treating last-mile code as an estate rather than a collection of personal scripts, and it is why expansion without an ops layer produces forty one-off integrations and no compounding.
A 12-month roadmap that respects RAND's "at least a year" looks like this: months 1-3, one workflow on rails, in production. Months 4-6, operate, measure, fix, and run a handoff drill. Months 7-12, a second workflow that reuses the lane, the custody design, and whatever patterns survived contact with production. Anything labeled "scale across the enterprise" before month 12 is a slide, not a plan.
A 90-day version, for people who do not have a year on paper
Executives will still ask for ninety days. Give them ninety days that cannot lie.
Days 1-30. Pick the workflow, measure the baseline, name the lane, meet security with a custody design, and inventory whatever is already running in shadow form. If this month produces a use-case workshop and no named owner, you are in the afterlife already.
Days 31-60. Build the first version on the production path. One integration into a real system, not a sandbox. Credentials in custody. A second person watches a deploy. If this month produces a demo on sample data, you have spent 60 days on product one.
Days 61-90. Put a real user on it, measure the number, run the "builder is out" test, and either fund the next two quarters against the live system or kill the project in writing. A pre-committed kill gate is how you avoid the six-month deferral that the POC essay treats as death.
Ninety days will not finish a transformation. It will tell you whether you have an implementation function. That is a better use of a quarter than another strategy deck. If the honest result is "we need a consultant for one quarter while we hire," that is a valid output, and the consultant page is the briefing for that buy. If the honest result is "we need two FDEs and a deploy path," that is a valid output, and the staffing page is the headcount math.
Budget the last mile, not the model
Model API spend is the line everyone can see. The last mile is the line that decides whether the model spend produces a P&L number. Price it as three things finance already understands.
Labor. The lane, hired or rented. FDE pay in the broad market is Indeed's $170,000 to over $200,000; lab total compensation runs higher, as the jobs list compiles. Consultants annualize higher still, which is the price of optionality, not a scandal. Understaffing the lane to protect a model budget is how you fund the 95 percent.
Time. RAND's year. S&P's 46 percent of POCs scrapped before production. A budget that assumes the hard part is the demo will be surprised by the integration estimate, and surprise is what abandonment looks like from the finance side.
Infrastructure. Where the connective code runs, where credentials live, how versions are retained, how handoff works. This is a small line against any of the labor numbers above, and it is the line that converts "we did a pilot" into "we can show an auditor what is running." The stack, layer by layer, is in the AI implementation software guide. Do not buy a transformation suite and hope the last mile is included.
If you need a single sentence for the budget meeting: we are not funding intelligence, we are funding the last mile that makes intelligence touch this business, and we will know in ninety days whether that mile exists.
Where the last-mile code runs
The honest close is also the product fact. The last mile is code: small, environment-specific functions that connect a product or a model to one organization's systems. That code has to be written, hosted, credentialed, versioned, and handed off. Most companies currently do those five things on laptops, personal cloud accounts, and goodwill. That is why the implementations die, and it is why the hiring wave exists.
Archway is where that code runs when you do not want it on a laptop. A bridge is a serverless function, written in JavaScript, that connects your product to one customer's (or one environment's) systems. Deploys take minutes, with no core-product change. Credentials go in an AES-256 vault with zero-access custody, which is the security answer from step 5. Every version is retained and owned by the organization. Handoff is a same-org reassignment, a note, and the versions, which is step 6 without a scramble. Two bridges are free, then $45 per bridge per month. There is AI-powered coding help in the editor; it is a feature on top of the ops layer, not the product. There is no deal tracker and no separate continuity module.
We do not run your transformation office. We do not replace your model vendor, your review board, or your FDE. We are the runtime for the last-mile work those people produce, so that a POC built as a bridge can graduate by being promoted, and so that the estate is still there when the builder is not. If you want the software map rather than the runtime pitch, use the implementation software guide. If you want the champion kit for the meeting that funds the lane, use the business case.
How to implement AI, reduced to a sentence you can put on a roadmap slide: pick one workflow, staff a lane, build on the rails production will use, integrate on purpose, vault the secrets, operate it as something a second person can touch, and expand only from there. The 95 percent skipped the middle of that sentence. The 5 percent did not. If your current roadmap still reads like a transformation program, rewrite it around the last mile, then go build. See what Archway does. The first two bridges are free, and they are enough to put the next workflow on the path it will retire on.
Sources and notes
-
MIT Project NANDA, "The GenAI Divide: State of AI in Business 2025," as reported by Fortune (August 18, 2025): 95 percent of enterprise generative AI pilots showed no measurable P&L impact; research based on 150 interviews, a survey of 350 employees, and analysis of 300 public AI deployments; generic tools stall in enterprise use because they do not adapt to workflows. https://fortune.com/2025/08/18/mit-report-95-percent-generative-ai-pilots-at-companies-failing-cfo/
-
S&P Global Market Intelligence, Voice of the Enterprise: AI and Machine Learning, as reported by CIO Dive (March 14, 2025): share of companies abandoning most AI initiatives rose from 17 percent to 42 percent; average organization scrapped 46 percent of AI proofs of concept before production; survey of more than 1,000 respondents in North America and Europe; cost, data privacy, and security risks cited as top obstacles; Informatica finding, cited in the same article, that two-thirds of enterprises cannot transition pilots into production. https://www.ciodive.com/news/AI-project-fail-data-SPGlobal/742590/
-
Gartner press release, 25 June 2025: over 40 percent of agentic AI projects forecast to be canceled by the end of 2027. https://www.gartner.com/en/newsroom/press-releases/2025-06-25-gartner-predicts-over-40-percent-of-agentic-ai-projects-will-be-canceled-by-end-of-2027
-
RAND, "The Root Causes of Failure for Artificial Intelligence Projects and How They Can Succeed" (RRA2680-1): notes that by some estimates more than 80 percent of AI projects fail, twice the rate of non-AI IT projects; 84 percent of 65 interviewees cited leadership-driven causes as the primary reason; underinvestment in infrastructure to deploy completed models among the five root causes; recommendation to commit a team to a specific problem for at least a year. https://www.rand.org/pubs/research_reports/RRA2680-1.html
-
Business Insider (Cadie Thompson and Lakshmi Varanasi, May 2026): Indeed posting growth for forward deployed engineers, 5,230 percent above a January 2025 baseline by April 2026, and the Indeed pay range. https://www.businessinsider.com/forward-deployed-engineer-jobs-in-demand-2026-5
-
Archway product claims are first-party and match the published product facts at https://www.tryarchway.ai/llms.txt.
