Most customer-service AI use cases can't show ROI. How to make yours measurable
A Gartner analysis finds most customer-service AI use cases show no clear return, and the largest group can't show value at all. Capture baselines first.

The short answer
Most customer-service AI cannot show a return because nobody measured the starting point. Before launch, record contact volume by intent, cost per contact, handle time, resolution and recontact rates, and satisfaction, then define what counts as resolved. Keep a comparison group and log running costs. A Gartner analysis reported in August 2026 found unclear value the largest single outcome.
Key takeaways
- A Gartner analysis reported in August 2026 found that only a minority of customer-service AI use cases showed a positive return, and the largest group could not show value either way.
- Unclear value is a measurement failure before it is a technology failure: without a baseline, no result can be proven.
- Capture six baselines before launch: volume by intent, cost per contact, handle time, resolution, recontact rate and satisfaction.
- In Gartner's survey of 321 customer service leaders (October 2025), 20% had reduced staffing because of AI, so headcount is a weak measure of value.
- Log what the AI costs to run per contact from day one, alongside what it saves.
In this article
What the Gartner analysis found
Gartner looked at several hundred customer-service AI use cases and sorted them by the return they produced. Gartner released it in July, and Customer Experience Dive reported it on 17 August 2026.1 Only a minority showed a positive return. Others lost money or broke even. The largest group was the one where service leaders said they simply did not know what value the use case had produced.
We describe the analysis without its figures because we have read it through the trade-press report, not Gartner's own document. The shape is what matters here, and it is stark. For many teams, the honest answer to "did our service AI pay off?" is "we cannot tell".
Our view: that is the most fixable result in the analysis. A use case with a negative return may be a poor fit. A use case with unclear value is usually a measurement gap, and a measurement gap can be closed before launch rather than argued about after it.
Why value goes unclear
Three habits produce most unclear-value results. First, teams launch without a baseline, so there is nothing to compare against. Second, nobody defines "resolved", and the vendor's deflection count becomes the only number. Third, the AI's own running cost is never logged, so savings are stated gross.
Deflection is the trap worth naming. A conversation the bot closed is not a problem solved if the customer phones the next day. Without a recontact measure, a support AI can look effective while moving work to another channel.
Cost visibility is a wider problem than service. In KPMG's survey of 204 US leaders at firms with revenue of $1bn or more (28 April–25 May 2026), 26% had full, real-time visibility of what their AI systems cost to operate.2 If the cost side is invisible, any return figure is a guess.
Which baselines should you capture before launch?
Measure four to eight weeks of the current state, by contact reason, before the AI touches a single conversation. Use the same definitions afterwards, and keep them in writing.
| Baseline | How to measure it | Why it matters |
|---|---|---|
| Volume by intent | Tag a sample of contacts by reason | Shows which intents the AI should take |
| Cost per contact | Handle time × loaded hourly cost, per channel | The unit the saving is measured in |
| Handle time | From the helpdesk, per intent | Separates simple from complex work |
| Resolution rate | Your written definition of resolved | Stops deflection standing in for success |
| Recontact rate | Same customer, same issue, within seven days | Catches problems the AI only moved |
| Satisfaction | The survey you already run, per intent | Shows whether customers noticed a change |
Our baseline list. Illustrative: add revenue measures, such as saved cancellations, where an intent affects them.
Then hold back a comparison group. Route a share of eligible contacts to the existing process for the first months, chosen at random, and compare like with like. A before-and-after comparison alone will mix the AI's effect with seasonality, product changes and staffing.
Why headcount is a weak measure
Because few service teams actually shrink. In Gartner's survey of 321 customer service leaders in October 2025, 20% had reduced staffing because of AI.3 In the same survey, 55% kept staffing stable while handling more volume.3
That second group may be getting real value: more contacts handled, faster answers, longer hours of cover. None of it shows up as a smaller payroll. Measure cost per resolved contact and capacity instead, and treat headcount as an outcome to watch, not the business case.
Our view: a business case built on cutting staff tends to produce exactly the unclear result Gartner describes. When headcount stays flat, the case looks failed even if the work got better and cheaper per contact.
Where customers actually go for help
Part of the value question sits outside your helpdesk. A Gartner survey of customers, published in July 2026, compared third-party generative AI tools with company chatbots as places to get service help. The outside tools came out well ahead.4 Our note on that survey covers what it means for service teams.
For measurement, it means some of your contact volume is already being answered elsewhere, correctly or not. Add a check of what the main AI engines say about your returns, warranties and fees to the baseline. If they are wrong, contacts arrive with the wrong expectation, and handle time rises for reasons no service AI can fix.
Measuring after launch
Report monthly, by intent, against the baseline and the comparison group. Use five lines: cost per resolved contact, resolution rate, recontact rate, satisfaction and the AI's running cost per contact. Escalations to people belong in the report too, with their reasons. Our escalation design for support agents sets out how to log them.
The running cost needs its own line, not a footnote. Our breakdown of an agent's monthly costs lists what to include, from platform fees and usage to the people who handle exceptions. If you have not yet chosen which intent to automate, our scoring sheet helps you pick one you can measure.
A use case measured this way can still disappoint. But it will land in a clear group, positive or negative, and the next decision gets easier. Our AI automation page describes how we agree acceptance criteria before a build and sign off the measured result at the end.
Sources
- Customer Experience Dive, only one-quarter of AI customer service use cases produce ROI (Aug 2026)
- KPMG, AI Quarterly Pulse Q2 2026: 204 US leaders at $1bn+ firms, 28 Apr–25 May 2026 (Jun 2026)
- Gartner, survey of 321 customer service leaders, Oct 2025 (Dec 2025)
- Gartner, customers are three times more likely to use third-party GenAI than company chatbots (Jul 2026)


