Enterprise AI has entered a new phase of maturity. Most large companies can now point to chatbot deployments, copilots or AI-powered workflows. Proving those investments have actually delivered measurable business value is much harder.
For the past two years, the focus has been on deployment: how many copilots have been rolled out, which model is most capable, how many customer conversations are automated. Increasingly, those questions are giving way to a more important one. Is AI actually creating measurable business value? Boards and executive teams increasingly want evidence that AI is improving customer outcomes, productivity and financial performance, not simply increasing automation. For customer experience teams, that means chatbot containment and automation rates are no longer enough on their own.
OpenAI's scorecard for the AI age
Last week, OpenAI CFO Sarah Friar set out an argument that reframes how AI spend should be judged. Rather than tracking software adoption metrics such as seats or logins, she proposes measuring "Useful Intelligence per Dollar", asking whether the value of the work AI completes grows faster than the cost of producing it. Friar sets out four questions businesses should ask: how much useful work gets done, what a successful task actually costs, how dependable the AI is, and whether each dollar buys more work as usage scales.
More important than the metric itself is what it says about where enterprise AI is heading. The timing is significant. As AI moves from experimental budgets into mainstream operational spending, finance leaders are demanding the same accountability expected of any other major technology investment. Whether or not OpenAI's scorecard becomes widely adopted, the direction of travel is clear. AI is increasingly being evaluated by what it achieves, not simply by how widely it is deployed.
McKinsey: outcomes over activity
OpenAI is not alone in reaching this conclusion. In April, McKinsey researchers made a similar case, with their global AI survey finding that nearly eight in ten companies use generative AI in at least one function, yet 60 percent have still not seen enterprise-wide EBIT impact.
McKinsey argues the gap reflects a common mistake: deploying AI without redesigning the underlying workflows it is meant to improve. The firm proposes a five-layer measurement framework that runs from technical performance and user adoption, through operational KPIs and strategic outcomes, to financial impact. Crucially, it argues the biggest gains come not from deploying more tools but from redesigning workflows around AI. Earlier this month, a separate McKinsey analysis uncovered that only 21 percent of companies have fundamentally redesigned their operating models around AI, yet top performers who attribute at least 5 percent of EBIT to AI were three times more likely to have done so.
OpenAI approaches the problem from a vendor's economics; McKinsey from organisational transformation. Both arrive at the same conclusion, however, that AI should be judged by business outcomes, not deployment metrics.
Why this matters for customer experience
The implications for CX are significant. Customer service teams have traditionally measured AI using metrics such as chatbot containment, automation rates, cost per conversation and average handling time. Those numbers still have value, but they only tell part of the story. They tell you what the AI is doing. They don't tell you whether it's making customers' lives easier or improving business performance.
If you apply McKinsey's thinking to customer experience, a different set of questions starts to matter. Are more issues being resolved at the first attempt? Has customer effort fallen? Are agents spending less time on repetitive work? Has the overall cost-to-serve improved? Those are the measures that connect AI performance to business performance.
On the surface, containment and resolution can look like similar ideas. In practice, they measure very different things. One counts what AI handled, the other counts whether the customer's problem actually went away.
This distinction matters most where interactions span multiple channels and require handoffs between AI and human agents. In those cases, measuring "useful work" usually means measuring a resolution, not a single automated exchange.
So what does measuring useful work actually look like in customer service? Consider a chatbot that contains 80 percent of enquiries. On paper, that looks like a success. But if those customers call back, abandon purchases or need an agent to sort things out later, containment alone says very little about business value. First contact resolution, repeat contact rates and customer satisfaction give a much clearer picture of whether AI is genuinely improving the experience.
Evidence across the industry
Recent product launches illustrate the same broader shift. Medallia's latest AI capabilities focus on measurable frontline productivity gains as the basis for the next phase of CX transformation, while Sprinklr is positioning AI as a means of turning voice-of-customer data into real-time operational action rather than reporting. In both cases, the emphasis is less on the AI itself than on the business outcomes it enables. This suggests CX vendors are increasingly judged on what work gets completed, not simply on how much AI has been deployed.
How CX leaders should measure AI success
This is where things become interesting for anyone running a CX function today. Drawing on both pieces of research, leaders can begin by asking a narrower set of questions. Instead of asking how many conversations AI automated or what the containment rate was, try asking:
● Which customer journeys have measurably improved because of AI?
● Where has AI reduced customer effort?
● Which workflows have genuinely become faster, rather than merely busier?
● Which tasks have been eliminated entirely, rather than shifted elsewhere?
● Has first contact resolution improved, and is customer satisfaction moving with it?
● Are agents spending more time on higher-value work?
● Has the overall cost-to-serve improved enough to justify continued investment?
These questions do the same job as McKinsey's five layers. They tie technical deployment to a business result that finance, not just IT, can recognise.
The takeaway
For much of the generative AI era, success was measured by access to the latest models and the speed of rollout. For many organisations, AI is moving from the innovation budget to the operating budget. Once that happens, finance starts asking different questions. The most successful companies are likely to be the ones that can demonstrate, with measurable improvements in customer experience and business performance, that AI is delivering value. Adoption is the starting point. Measurement is becoming the real differentiator.

