AI in customer support: real resolution rates and ROI
AI customer support ROI is real, measurable, and routinely oversold, and you need to hold both facts at once. On high-volume support with documented answers, an AI layer resolves a meaningful share of tickets at a fraction of the cost per contact. The disappointments come almost entirely from teams that bought the vendor headline number and skipped designing the escalation path.
The most famous number in the field illustrates both halves. Klarna announced in early 2024 that its assistant was handling two-thirds of customer service chats in its first month, the workload of 700 full-time agents, with a projected 40 million dollar profit improvement. Fifteen months later the same company was publicly recruiting human agents again, its CEO conceding that the cost focus had gone too far and that customers must always be able to reach a person. Both parts are true, and together they are the whole subject in miniature: the deflection economics work, and the quality constraint bites.
Deflection is not resolution
The industry vocabulary does a lot of quiet work. Deflection means the conversation ended without reaching a human. Resolution means the problem was actually solved. A bot can deflect 90 percent of conversations while resolving far fewer, because the customer who gives up and emails tomorrow, or churns, still counts as deflected. When a vendor quotes a rate, your first question is which one.
The metrics that keep you honest: resolution confirmed by the customer or by the absence of a reopen, reopen rate on AI-handled tickets, and CSAT measured separately on the AI-handled share rather than blended into the team average, where damage hides.
What realistic rates look like
The figures vendors publish in 2026 cluster in the same band: around two-thirds of conversations resolved for mature deployments, 70 to 75 percent counting as strong, anything above 80 best in class, with first-year figures typically landing below that. These are vendor numbers, so put the deflection question to every one of them. B2B tends to run well under B2C, because the questions are account-specific and the volumes thinner.
The more useful number is one you can compute yourself this week: your achievable rate is roughly the share of tickets that are known questions with documented answers or simple account actions. Pull last quarter’s tickets and count. If 30 percent of your volume consists of genuinely novel judgment calls, no product on the market will resolve 80 percent of it, whatever the demo showed.
The three patterns that pay
Full self-service works on known questions, answered from your own help content with citations, so the customer can verify and the system rests on an architecture that controls wrong answers rather than on a disclaimer. Drafting for agents is the underrated second pattern: the AI writes, a human reviews and sends. It captures much of the time saving in exactly the places where autonomy would be reckless. Triage is the third: classification, prioritization, and a context summary waiting for the agent before they open the ticket. Unglamorous and reliably positive. Strong programs run all three at once; weak ones bet everything on the first. The same three patterns are now reaching the phone channel, where voice agents got good enough to matter, with the same escalation caveats.
The escalation design that protects CSAT
The Klarna lesson: the handover is the product. Four rules we hold in every system we build. A human is always reachable, and the path is visible rather than buried behind three loops of "did this answer help". The bot hands over with full context, so the customer never repeats themselves; forced repetition is the fastest known way to turn a mediocre interaction into an angry one. Emotion, billing disputes, cancellations and anything regulated route to a human immediately, without a self-service detour. And the system says it does not know instead of improvising, because one confidently wrong answer about a refund costs more goodwill than fifty escalations.
The ROI arithmetic worth doing
Compare cost per resolved ticket, all-in, before and after: platform and model costs, the humans who still review, the content maintenance the bot depends on, and the new tickets that bad answers create. Set that against the loaded cost of a human-resolved contact, then check the quality column: reopen rate and CSAT on the AI-handled share. Support automation earns its place when volume is high and answers are documented, which is the same test we apply to any first AI use case, and its return deserves a baseline and honest measurement like anything else that touches customers.
Run the pilot against your own ticket history, write the escalation rules before go-live, and report resolution rather than deflection to the board. Those three habits protect more CSAT than any model upgrade will.