How to Measure the ROI of AI
A practical framework for measuring AI ROI from individual use cases through an enterprise portfolio. The article distinguishes adoption and technical health from business value, explains how to translate time saved and operational improvements into economic outcomes, and argues for baselines, attribution, full costs, and stop-or-scale decisions.
· Blaize Stewart · #ai-roi #kpi #measurement #ai-adoption #portfolio-management #productivity #value-realization
AI programs often start in the wrong place. They lead with AI, not with value, and the project fizzles. What makes a difference though is when technology changes an a value outcome that matters to the organization. That changes the first question for any AI use case. The useful question is what measurable business outcome should improve. If there is no meaningful outcome to improve, the business case for AI is already weak. This sounds obvious, yet current research suggests it is still uncommon. McKinsey's 2025 global survey found that tracking well-defined KPIs for generative AI solutions had the strongest relationship with bottom-line impact among the 12 adoption and scaling practices it tested. Fewer than one in five respondents said their organizations were actually doing it. McKinsey's 2026 work on AI measurement makes the same point more explicitly. Technical performance, adoption, operational KPIs, strategic outcomes, and financial impact have to connect in a chain. A model can perform well and still create little business value if the process, product, or customer outcome does not move.
The simplest use cases are often the easiest to measure. If an employee spends ten hours each week preparing a report and a redesigned AI-enabled workflow reduces the work to one hour, the time reduction is observable. The economic value still needs another step. Nine hours saved is only a realized return when something happens with that capacity. The organization may complete more work with the same staff, avoid additional hiring, reduce overtime, shorten a revenue-producing cycle, or eliminate labor expense. Time saved is an operational KPI. The business outcome is what the organization does with the saved time. Productivity estimates can be misleading when they rely on perception. METR ran a randomized controlled trial in 2025 with experienced open-source developers working in repositories they knew well. The developers expected AI tools to make them 24 percent faster. After the study, they still believed AI had made them about 20 percent faster. The measured result went in the other direction. Tasks took 19 percent longer when the developers were allowed to use the AI tools. METR is careful about the limits of that study, and later tools may perform differently, but the experiment illustrates a durable measurement problem. Satisfaction and perceived productivity can be useful signals, while objective task and business metrics are needed to establish value. Other field research shows what strong measurement looks like when the KPI is well chosen. Brynjolfsson, Li, and Raymond studied more than 5,000 customer support agents using a generative AI assistant. Productivity was measured as issues resolved per hour rather than as self-reported usefulness. The published study found an average productivity increase of about 15 percent, with larger gains among less experienced and lower-skilled workers. The researchers also examined quality and customer experience rather than assuming that faster work was automatically better. That is the kind of measurement that can support an economic case because the intervention is tied to a real operating metric.
The same logic applies at the workflow level. For example, a claims, finance, service, or other operational process may contain manual research, summarization, routing, and handoffs that consume hours without changing the final outcome. AI can accelerate selected steps, but the KPI should remain the workflow outcome. Measure elapsed cycle time, labor hours per completed case, throughput, rework, and quality before and after the change. If the process becomes materially faster without degrading quality, the value can be translated into released capacity, avoided cost, or higher throughput.
Revenue-oriented use cases require the same discipline. A customer-facing AI feature may look impressive without changing customer behavior. The useful metrics are conversion, average order value, retention, sales-cycle duration, product usage, or another outcome that connects to the economics of the product. Attribution can come from A/B testing, staged rollout, matched cohorts, or another design that creates a credible comparison. If the feature does not improve a metric that matters, the fact that it uses AI adds no economic value. Some use cases are harder to express directly in dollars, but they are still measurable. The KPI needs to describe the value the organization actually wants rather than the activity produced by the model.
A backend recommendation system is a good example. AI might rank products, next-best actions, content, or offers for a customer and present alternatives that are more relevant to that person. The feature has value only if those recommendations change customer behavior or improve the economics of the product. Conversion rate, average order value, click-through, retention, or revenue per session can provide the measurable outcome, ideally through an A/B test or controlled rollout. If the recommendations become more sophisticated but customers do not respond differently, the feature has not demonstrated ROI.
Adoption metrics belong in the measurement system, too, especially for broad tools such as Microsoft Copilot, Claude, or internal assistants that do not impact products or specific use cases directly. Daily active users, eligible users, feature usage, task penetration, acceptance rates, and employee satisfaction can reveal whether the tool has become part of normal work. They are leading indicators. High adoption can explain why a useful capability is or is not producing results, but adoption itself is not ROI. Ten thousand active users can still represent a poor investment if the relevant operating and financial measures remain unchanged. The same distinction applies to technical metrics. Accuracy, latency, hallucination rates, token consumption, acceptance rates, and model cost matter because they determine whether the system can operate safely and economically. They should be treated as the health measures of the AI component. The business case sits above them. A model with excellent latency and low token cost can still be attached to a use case that nobody needs.
A credible ROI calculation takes in the the full cost. License fees and token usage are only part of the investment. Integration, data preparation, evaluation, security, governance, human review, change management, training, monitoring, support, and ongoing maintenance all consume resources. McKinsey's 2026 measurement framework explicitly places total cost of ownership alongside revenue uplift, cost-to-serve reduction, and margin improvement. The same discipline used for another capital or technology investment should apply to AI.
At the use-case level, the calculation can remain simple.
ROI = (realized benefit - total cost) / total cost
The hard part is usually the realized benefit. A projected value is a hypothesis. A measured change with credible attribution is evidence. That is why Deloitte's 2026 guidance for CFOs emphasizes establishing pre-deployment baselines, locking success criteria before development, and isolating the AI contribution from other changes happening at the same time. Without a baseline, almost any post-implementation improvement can be attributed to the technology after the fact.
Across an enterprise, this can become a portfolio discipline rather than a collection of unrelated success stories. Each AI use case can occupy one row in a value ledger. That row carries the business owner, the KPI being changed, its baseline, the expected improvement, the measured result, the confidence in attribution, the realized economic benefit, the total cost, and the resulting return. Small automations may contribute modest value. A redesigned core process may contribute much more. A customer-facing feature may create revenue while another initiative primarily releases capacity. The unit of measurement differs by use case, but each eventually has to translate into an economic outcome that can be defended. The portfolio view also makes prioritization easier. A technically sophisticated project with weak measured value should compete poorly for funding against a simple automation that reliably removes thousands of hours of work. A highly adopted copilot may deserve continued investment if objective measures show improved throughput or quality. Measurement provides a reason to scale one and stop the other.
When measured wholistically, the enterprise-level AI number then becomes much less mysterious. The total value of AI adoption is the aggregation of the returns from individual use cases, adjusted for shared platform costs and any overlap in attributed benefits. The organization can still report adoption, satisfaction, capability growth, and qualitative stories. Those measures provide useful context and explain how the change is happening. The financial case comes from the use cases that can show what changed, how much it changed, what the change is worth, and what it cost to produce.
AI should therefore enter a portfolio only after the value hypothesis is clear enough to measure. Every candidate needs a specific business outcome, an established baseline, a target KPI, an attribution method, and a credible path from operational change to economic value. The implementation can then prove or disprove that hypothesis. If the metric does not move, the organization has learned something useful and can stop spending. If it moves enough to justify the full cost, the use case has earned the right to scale. That creates a much healthier AI program. Technology can catalyze experiments and expose opportunities, but the investment case must remain anchored in the business. The question stops being how much AI the organization has deployed and becomes how much measurable value those deployments have produced.
A minimal AI value dossier
Before approving or scaling an AI use case, capture these items on one page:
- Use case: Describe the person, workflow, and problem. State what AI will change, and keep the scope narrow enough to test.
- Business value: Name the outcome the organization expects, such as revenue, cost, capacity, cycle time, quality, or risk, and explain why it matters.
- Measure: Choose one primary business KPI and any guardrail metrics. Record the unit, owner, measurement window, and how the KPI translates into dollars or another defensible measure of value.
- Baseline and control: Record current performance before introducing AI. Compare the AI-enabled workflow with a control, such as the current process, a holdout group, a matched cohort, or an equivalent historical period.
- Test case: Define the participants, tasks, sample, duration, and success threshold before the pilot begins. Keep the conditions comparable and track quality alongside speed or throughput.
- Result and ROI: Measure the difference between the test and control, translate only the attributable improvement into realized benefit, include the full cost, and calculate
ROI = (realized benefit - total cost) / total cost. Use the result to decide whether to scale, change, or stop the use case.
Further reading
McKinsey The State of AI and How Organizations Are Rewiring to Capture Value
McKinsey From Promise to Impact
Stanford Generative AI at Work
METR Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity