Why Enterprise AI Stalls After the Pilot Stage: Lessons from an IT Leader
In this interview with The Left Shift, Abhradeep Chatterjee, Associate Director at NTT DATA Services, argues that enterprises are asking the wrong question. Instead of focusing on whether AI models work, leaders should ask whether AI can transform how the business operates.
Despite launching 100s of Proof of Concepts (PoC), enterprises still struggle to scale AI. According to McKinsey's 2025 State of AI survey, around 88% of organisations now use AI in at least one business function, yet only about one-third have begun scaling AI across the enterprise.
Another study suggests that 95% of generative AI pilots fail to deliver measurable business impact, while Gartner expects more than 40% of agentic AI projects to be cancelled by the end of 2027 as organisations struggle with governance, integration, ownership, and proving ROI beyond early demonstrations.
While AI experimentation has become commonplace, most enterprises are still struggling to turn promising pilots into AI systems that deliver real business value at scale. So why do enterprises continue launching PoCs but struggle to scale them? Is the problem the technology or the organisation behind it?
In this interview with The Left Shift, Abhradeep Chatterjee, Associate Director at NTT DATA Services, argues that enterprises are asking the wrong question. Instead of focusing on whether AI models work, leaders should ask whether AI can transform how the business operates.
He explains why so many organisations remain trapped in "POC purgatory," why workflow redesign matters more than model accuracy, and what CIOs, CTOs, and business leaders must do to move enterprise AI from experimentation to measurable business value.
Edited Excerpts
Many enterprises have dozens of AI proof of concepts, but only a handful ever make it into production. From your experience, what is the biggest reason for this disconnect?
The biggest reason is that many AI PoCs are designed to prove that the technology works, but not that the enterprise is ready to absorb it.
A pilot can succeed in a controlled environment with clean data, limited users, and low operational risk. Production, however, is very different. Once AI enters a real enterprise workflow, it must operate across legacy systems, fragmented data, security controls, compliance requirements, role-based access, unclear ownership, and measurable business expectations.
In my experience, the disconnect is rarely just about model performance. Many AI pilots fail to scale because they are not connected to an operating model. They generate insights, but those insights do not always change a process, improve a decision, reduce effort, or create a measurable business outcome.
I have seen this in real enterprise environments where the AI capability performed well during the pilot and generated useful recommendations. But when the organisation evaluated production deployment, the real issues appeared:
- Data was spread across multiple systems,
- Approval paths were unclear,
- Support ownership was not defined, and
- Operational teams were not fully prepared to consume the AI output.
The challenge was not whether the model could work. The challenge was whether the enterprise was ready to operationalise it. The real question is not, “Can AI produce a useful answer?” The real question is, “Can AI improve how the enterprise works?”
Do you think enterprises are still treating AI as a technology experiment rather than a business transformation initiative? What mindset shift is required at the leadership level?
Yes, many enterprises are still treating AI as a technology experiment. They launch tools, pilots, and use cases, but the deeper transformation work is often missing.
The leadership mindset needs to shift from “Where can we use AI?” to “Which business outcomes do we need to improve, and how can AI help us redesign the process?”
That is a very different conversation. AI should not be treated as a layer added on top of existing inefficiencies. If a workflow is broken, AI may make the broken workflow faster, but it will not necessarily make it better.
Leaders need to view AI as part of operating model redesign. That includes how work is routed, how decisions are made, how risks are governed, how employees adopt new tools, and how outcomes are measured.
The organisations that succeed with AI will not be the ones with the most pilots. They will be the ones that connect AI to strategy, process change, governance, and measurable execution.
How much of the scaling challenge comes down to legacy infrastructure, fragmented data, and integration issues versus organisational factors like governance and change management?
Both matter, but organisational factors are often underestimated. Legacy infrastructure, fragmented data, and integration challenges are real. Many large enterprises operate across decades of technology investments. Data sits in different systems. Business logic lives in applications, spreadsheets, emails, knowledge bases, and people’s heads. AI cannot deliver reliable outcomes if it does not have access to the right evidence, context, and system integration.
But even when the technical foundation exists, AI can still fail to scale if governance and change management are weak. Who owns the AI output? Who approves high-impact actions? How do employees trust the recommendation? What happens when AI is wrong? How are decisions audited? How does the organisation train people to work differently?
I have also seen situations where the infrastructure was not the biggest blocker. The data and systems were largely available, but adoption remained limited because teams were unsure when to trust AI recommendations, who would be accountable for decisions, and how exceptions should be handled. In these situations, the technology may be ready, but the governance model is not.
In many cases, the technology problem is visible, but the operating model problem is the real blocker. Successful scaling requires both: a modern data and integration foundation, and a governance model that makes AI usable, trusted, and accountable inside real business workflows.
We’ve seen companies announce hundreds of AI use cases, yet employees often continue working the same way. What distinguishes an AI project that delivers measurable business value from one that remains a successful demo?
The difference is workflow adoption. A successful demo shows that AI can perform a task. A successful enterprise AI project changes how work gets done.
In one operational support environment, AI was introduced to help analyse incidents and surface probable causes. The technical results were impressive, but the real value appeared only after those recommendations were embedded into the incident-resolution workflow.
Once teams began using the insights during investigations, decision-making became faster and manual effort was reduced. That experience reinforced an important lesson: value comes from workflow adoption, not from the model alone.
For example, an AI tool that summarises an incident, a customer complaint, or a policy document may look impressive. But if that summary does not help an employee make a faster decision, complete a workflow, reduce manual effort, improve response quality, or serve a customer better, the business value will be limited.
AI projects that deliver measurable value usually have five characteristics:
- They are tied to a clear business problem.
- They are embedded directly into the workflow.
- They have both executive ownership and operational ownership.
- They include governance, controls, and human oversight.
- They are measured against business KPIs, not only technical metrics.
The most important test is simple: after AI is deployed, does the work actually change? If employees continue working the same way, the project is probably still a demo, even if the technology is impressive.
As generative AI adoption accelerates, many organisations are facing rising inference costs, security concerns, and compliance requirements. How should enterprises build an AI strategy that is both scalable and economically sustainable?
Enterprises need to move from experimentation-first AI to architecture-first AI. In the early stage, it is natural for teams to experiment with many models, tools, and use cases. But at scale, that approach becomes expensive and risky.
Inference costs can grow quickly. Data exposure risks increase. Compliance requirements become more complex. Without discipline, AI can become another layer of uncontrolled technology spend.
I have seen organisations initially apply large language models to a wide range of tasks, only to later realise that the cost profile was difficult to justify for some use cases. After reassessing the actual business need, some workloads could be handled through smaller models, automation workflows, traditional analytics, or rules-based approaches. The business outcome remained meaningful, but the operating cost and risk profile became much more manageable.
A sustainable AI strategy should begin with use-case prioritisation. Not every process needs generative AI. Some problems can be solved with automation, analytics, rules-based workflows, traditional machine learning, or smaller models. Enterprises should match the technology to the value, sensitivity, and risk profile of the use case.
Second, organisations need strong data governance and access controls. AI should only use the data it is authorised to use, and sensitive information must be protected by design.
Third, they need cost governance. Leaders should understand the economics of inference, model selection, token usage, usage patterns, compute consumption, and return on investment.
Fourth, they should build reusable AI platforms instead of isolated one-off solutions. Common capabilities such as identity, audit logging, prompt governance, retrieval, monitoring, evaluation, human approval, and compliance controls should be standardised.
Scalable AI is not just about deploying more models. It is about building an AI operating foundation that is secure, governed, cost-aware, reusable, and aligned to business value.
In your view, what role should CIOs, CTOs, and business unit leaders each play in moving AI initiatives from isolated pilots to enterprise-wide deployment? Where do you see ownership breaking down most often?
AI scaling requires shared ownership, but the roles need to be clear. The CIO should focus on enterprise integration, governance, data access, cybersecurity, operational readiness, and scalability.
The CTO should focus on architecture, platform strategy, engineering standards, model integration, reliability, observability, and technical sustainability.
Business unit leaders should own the business problem, workflow adoption, process redesign, employee enablement, and value realisation. Where ownership often breaks down is between the pilot and production stages.
A team may build an impressive prototype, but no one clearly owns the transition into production. IT may see it as a business experiment. The business may see it as a technology project. Risk and compliance may get involved too late. Operations may not be prepared to support it. Employees may not be trained to use it.
That is how good pilots enter what I often call “POC purgatory.” In one case, a prototype gained strong stakeholder interest and generated positive feedback during demonstrations. However, ownership of production support, risk review, operational monitoring, and long-term adoption had not been assigned.
As a result, the initiative struggled to move forward despite clear technical success. Experiences like this show why ownership alignment must begin at the start of an AI programme, not after the pilot succeeds.
To avoid this, every AI initiative should have three owners from the beginning: a business outcome owner, a technology or platform owner, and an operational adoption owner. Without all three, scaling becomes difficult.
How can organisations define the right success metrics for AI beyond accuracy or model performance? What business KPIs should determine whether an AI initiative deserves to be scaled?
Model accuracy is important, but it is not enough. In enterprise settings, the best AI metrics are tied to business and operational outcomes. The question should be: did AI improve speed, quality, cost, resilience, employee productivity, or customer experience?
For operations, useful KPIs may include reduction in mean time to resolution, faster time to diagnosis, fewer escalations, improved first-time resolution, better routing accuracy, lower manual effort, and improved service availability.
In my own experience, some of the strongest indicators of AI success were not model-centric metrics. They were reductions in troubleshooting time, faster access to operational knowledge, fewer manual investigations, improved resolution quality, and better consistency in decision-making. Those outcomes were far more meaningful to business stakeholders than model accuracy alone.
For customer-facing use cases, metrics may include faster response time, higher customer satisfaction, reduced rework, improved resolution quality, lower cost per interaction, and better consistency across channels.
For knowledge-work use cases, metrics may include cycle-time reduction, document-processing speed, decision quality, compliance accuracy, reduced handoffs, and improved employee productivity.
There should also be risk and trust metrics: how often AI recommendations are accepted, how often they are overridden, how frequently they produce errors, whether outputs are explainable, and whether decisions are auditable.
An AI initiative deserves to be scaled only if it shows measurable improvement in a meaningful business workflow. If the only metric is that the model performs well in a test environment, it is too early to scale.
Looking ahead over the next two to three years, do you expect enterprises to continue running numerous AI experiments, or will we see a shift toward fewer, more strategic AI deployments? What advice would you give organisations that are still stuck in “POC purgatory”?
I believe we will see a shift from many disconnected experiments to fewer, more strategic AI deployments. The first wave of generative AI adoption was about exploration. Companies wanted to understand what was possible.
The next wave will be about execution discipline. Boards and executive teams will ask harder questions: What value did this create? What risk did it introduce? What did it cost? Did it scale? Did employees actually use it? Did it improve the business?
Over the past few years, I have observed a clear pattern across enterprise AI initiatives. The most successful efforts were rarely the ones that received the most attention during the experimentation phase. They were the ones connected to a specific business process, supported by executive sponsorship, adopted by operational teams, and continuously measured against business outcomes.
Organisations stuck in POC purgatory should take three steps.
- Reduce the number of pilots and focus on the few use cases closest to measurable business value.
- Redesign the workflow around AI instead of placing AI beside the workflow.
- Build the governance, data, integration, security, and change-management foundation required for production.
My advice is simple: stop measuring AI progress by the number of experiments launched. Measure it by the number of business processes improved.
The next phase of enterprise AI will not be defined by model innovation alone. It will be defined by execution. The winners will not be the companies with the most demos. They will be the companies that turn AI into a reliable, governed, and measurable operating capability.
Abhradeep Chatterjee is an Associate Director at NTT DATA, based in Minnesota, USA. He specialises in enterprise AI transformation, intelligent operations, AIOps, automation, and operational resilience. With more than 15 years of industry experience, his work focuses on helping large organisations move from reactive operations and isolated AI pilots toward predictive, governed, and production-scale AI capabilities. He is also an active researcher, IEEE Senior Member, international conference speaker, reviewer, and judge in the fields of AI and enterprise technology.
Comments ()