Many AI pilots do not move beyond the testing stage because implementation complexity is underestimated. A clearer path forward starts with identifying stakeholders, understanding where complexity moves, and defining metrics for the value created.

Why do so many AI pilots fail to scale?
Organizations experiment with AI, test a few use cases, and see promising results. Then the initiative stays at the pilot stage:
- It does not become part of the operating model.
- It does not create a measurable business impact at scale.
Deloitte reported that only 25 percent of respondents had moved 40% or more of their AI experiments into production.1
McKinsey described a similar situation: nearly two thirds of organizations had yet to scale AI beyond a few pilots.2
In my view, one of the main reasons is that senior stakeholders often underestimate the complexity of the problem:
An AI initiative is presented as a solution that will remove complexity. In practice, it moves complexity somewhere else.
A Simple Example: An AI-Powered Chatbot
Consider a chatbot on a website. The chatbot will answer customer questions automatically and reduce the workload of the support team. A large percentage of requests will be handled without human involvement.
From the perspective of the customer, this works well. A user receives an answer immediately and does not need to contact support.
The complexity is removed from the customer journey.

At the same time, part of that complexity is transferred to the support team and to the AI team that maintains the system.
What Changes for End Users?
For end users, the main question is: did the chatbot solve the problem?
A starting metric is the first-contact resolution rate. It shows how many requests are resolved without escalation to a support specialist.
What Changes for the Support Team?
The support team receives the requests that the chatbot cannot resolve. Sometimes, the chatbot escalates too late. The customer has already spent time trying to solve the problem and is frustrated by the time a specialist joins the conversation.
Sometimes, the chatbot escalates too early. It transfers the case before collecting enough information. The support specialist receives an incomplete request and needs to ask the same questions again.
This can be quantified by:
- Escalation Rate. Measure the percentage of requests transferred to the support team.
- Late-Escalation Rate. Track cases where the customer experienced avoidable friction before escalation.
- Missing Context. Measure how often specialists need to request information that could have been collected earlier.
- Support Workload. Monitor changes in the volume and difficulty of support requests.
The Long-Term Effect on the Support Team
There is also a longer-term question:
- How should support specialists be trained?
Before the chatbot was introduced, junior specialists learned by answering simple questions. Over time, they developed the experience required for more difficult cases.
Once routine requests are automated, the support team receives a higher concentration of complex cases. The old training model no longer works.
The organization needs a new way to build expertise.

What Changes for the AI Team?
In production, the chatbot becomes a system that needs continuous maintenance.
New customer questions appear. The knowledge base needs updates. Existing answers need to remain accurate after each change. Safety protocols need to be reviewed.
We can quantify this by:
- Time to Add a New Use Case. Measure how long it takes to support a new customer scenario.
- Regression Testing. Check whether new changes affect previously supported use cases.
- Knowledge-Base Maintenance. Track the effort required to keep information current.
- Safety-Protocol Coverage. Confirm that sensitive topics have clear rules and escalation paths.
Questions about refunds, finance, account access, or order changes complicate the case even more… Should the chatbot answer the question directly? Should it provide general guidance? Should it transfer the request to a specialist?
Someone needs to define these rules, test them, and update them when the business context changes.
Do We Really Need Metrics for AI Initiatives?
A reasonable question is whether all this measurement effort is necessary. My answer is “Yes!” The discussion itself creates value. It forces the team to become more specific:
- What exactly do we mean by a successful chatbot?
- What kind of escalation is acceptable?
- How much maintenance work can the AI team handle?
- How will we train support specialists?
Even when the exact values for the KPIs are not available, those questions improve the quality of the implementation.
McKinsey analyzed 12 adoption and scaling practices for generative AI. Each practice had a positive correlation with EBIT impact. Among the practices included in the study, tracking well-defined KPIs for generative AI solutions had the strong impact on the bottom line.3
Continuous tracking of the metrics related to each AI pilot requires discipline. The BSC Designer platform can significantly facilitate this process by ensuring consistency in performance measurement and aligning KPIs with the required business context.
Start With Complexity Metrics
Each AI project needs its own set of indicators. Still, there are a few common areas that are worth reviewing in almost every implementation. The first is complexity.
Complexity is difficult to quantify – look at the areas where the organization spends time, effort, and budget. Look for repeated communication, manual reviews, delayed decisions, and exceptions.
Good candidates are:
- Exception-Handling Time. Measure the effort required to resolve cases that the AI system cannot handle.
- Repeated Communication. Track additional interactions caused by missing context or unclear answers.
- Manual Reviews. Measure the volume of cases that require human validation.
- Maintenance Effort. Track the resources required to update and test the system.
- Control Costs. Monitor the effort required to manage safety and compliance risks.
These metrics show where complexity has moved after automation.
Measure Trust Empirically
The second area is trust. Stakeholders will ask whether the AI output is reliable. This question becomes difficult when the final result depends on the input data, the model, the knowledge base, the controls, and the escalation process.
Trust can be measured empirically. The team can test the system with known data, identify typical failure modes, and check what happens if we combine various systems that break in different way.
Moving From Pilot to Implementation
To move from AI pilots to full-scale implementation, review existing AI initiatives through the lenses of stakeholders, complexity, and trust.
A structured workshop can help the team answer a few practical questions:
- Stakeholders. Who is affected by the AI initiative?
- Expected Value. What improvement should each stakeholder experience?
- Transferred Complexity. Where does the new workload appear?
- Capabilities. What systems, resources, and skills are required?
- Risks. What can go wrong during implementation?
- Metrics. How will the team monitor progress and outcomes?
The Strategy Execution Canvas can be used as a starting point for this discussion. It helps connect stakeholder needs with goals, capabilities, risks, assumptions, and metrics.
- The State of AI in the Enterprise, Deloitte, 2026 ↩
- Are Your People Ready for AI at Scale?, Alex Camp, Drew Goldstein, Laura Pineault, Holly Price, and Nicolette Rainone, McKinsey & Company, 2026 ↩
- The State of AI: How Organizations Are Rewiring to Capture Value, Alex Singla, Alexander Sukharevsky, Lareina Yee, and Michael Chui, with Bryce Hall, McKinsey & Company, 2025 ↩
Alexis Savkin is a Strategy Architect and founder of BSC Designer, a strategy execution software platform with the Balanced Scorecard at its core. He helps organizations translate strategy into measurable objectives, KPIs, and initiatives. Alexis is the creator of the Strategy Execution Canvas, the author of 100+ articles on strategy and performance measurement, and a regular speaker.
