Is your AI support agent actually working? Measuring ROI after launch

Most teams measure ROI of AI customer support in week one. That reading is noise. Five baselines to take before go-live, and the three gates it must clear.

Is your AI support agent actually working? Measuring ROI after launch
Created time
Sep 3, 2026 03:35 PM
Title length (<60)
Author
Last optimised
Ecomm?
Image
how-to-measure-roi-of-ai-customer-support-header.png
Publish date
Aug 13, 2026
Video
Slug
how-to-measure-roi-of-ai-customer-support
Featured
Type
Article
Ready to Publish
Ready to Publish
💡
A high deflection number sitting on top of poor verified resolution is support debt, and you repay it in reopened tickets and lost customers.
The standard ROI sum is your cost per ticket before, your cost per ticket after, and the difference between them. It builds a solid business case, and in the weeks after go-live it returns mostly noise wearing a percentage sign. Measuring it after launch means writing down five baseline numbers before the agent handles its first ticket, taking readings at day 7, day 30 and day 90, and running each one through three gates: volume, stability, quality.
My take is that your first ROI reading is usually wrong, for three reasons that all sit upstream of the formula. You take the reading too early, you have nothing firm to hold it up against, and nobody has gone back to see if those tickets are still closed.
I'm Mike, co-founder of My AskAI. We help 200+ ecommerce and SaaS businesses run AI customer service inside Zendesk, Intercom, Freshdesk, Freshchat, Gorgias and HubSpot, and our agents have now resolved over 1,000,000 tickets. Across the whole customer base we sit at a 72%+ resolution rate on a rolling 30-day basis.
Post-launch reviews run to a pattern: someone shows a number, someone else asks whether it is real, and the room has no way to settle it. What settles it is a test you can run on the number before you argue about it.
RecruitCRM, one of our customers, went from around 35% AI resolution at go-live to 68%, and a weekly review habit is what moved it.

Why is your first AI support ROI number wrong?

TL;DR: Because it was read before the AI had resolved enough tickets to be more than noise, against a baseline nobody wrote down, and with no check on whether those tickets stayed shut.
The published advice on this is easy to find, and it is good advice. Gorgias sets out three steps and puts baselines first, naming AHT, FCR, CSAT, churn and cost per contact. Balto's step one is to capture current performance before you change anything, and Zendesk defines customer service ROI as a formula for weighing what you made against what you spent (the same order we use on an onboarding call).
Every one of them answers the question you had before you bought. None of them answers the one you have the week after go-live: can the inputs you just fed the formula be trusted yet?
The returns out there are uneven. Trade press reported a Gartner analysis of 432 AI use cases in customer service: a quarter produce a return, and a quarter produce a negative one. So the reading you take is not confirming a win, it is telling you which quarter you landed in, and we want you to know which one before finance asks.
Julie Geller of Info-Tech Research Group named the mechanism in that piece:
"Delaying contact with a human agent is not the same as resolving the customer's problem." Julie Geller, principal research director at Info-Tech Research Group, speaking to CX Dive.
A second reason we run into constantly: teams put real effort in up front, reach a rate they are content with, and stop, because the next steps mean connecting APIs and doing something more involved. The number on their dashboard is a floor, and most of the climb is still available above it.
Before you act on an ROI reading, you need a way to decide whether it is trustworthy at all. Three gates do that job.
Two-panel not-equal contrast card: Week 1 shows a raw rate with no history behind it, with reopens and recontacts not yet counted; Day 90 shows a verified rate, net of tickets that reopened, as a number the gates have actually cleared.
Two-panel not-equal contrast card: Week 1 shows a raw rate with no history behind it, with reopens and recontacts not yet counted; Day 90 shows a verified rate, net of tickets that reopened, as a number the gates have actually cleared.

The Three Gates: when is an AI support ROI number safe to act on?

TL;DR: A reading has to clear three gates before it means anything: enough volume, a settled rate of change, and quality inside threshold. The third is the one nearly every post-launch report leaves out.
We run the Three Gates in order, because they answer different questions. The first two ask whether you have a measurement at all; the third asks what the measurement cost you.
Gate
What it asks
What a fail means
What you do
1. Volume
Have enough conversations been resolved for the rate to be a measurement?
One ticket still moves the rate 4.2 points over 24 resolutions
Wait, and read quality signals in the meantime
2. Stability
Has the rate of change flattened out?
You are reading a trend as though it were a level
Wait, and never annualize a climbing rate
3. Quality
Are verified resolution, reopens, recontacts and escalation accuracy inside your own baseline?
The savings are being funded by tickets that did not stay shut
Stop, then narrow the scope or drop back to copilot

Gate 1: volume

Vendor documentation defines when a resolution counts and stops there. The floor at which an automated resolution rate becomes a measurement is left to you.
Treat this gate against your own volume (a rate computed over 24 resolutions moves 4.2 points on a single ticket).
Fewer tickets a week need longer on the calendar. Gate 1 counts events, and with fewer tickets a week they arrive slowly.
For the band to read your own rate against, our AI resolution rate benchmark study puts the median AI handling rate at 70% across 195 rated deployments spanning 38 vendors, with the middle half between 56% and 80%.
Distribution chart of AI resolution rates across 195 rated deployments spanning 38 vendors (median taken across all labels, however each vendor names its own rate), cropped to the 40-100% range where deployments actually sit, with a red band from P25 (56%) to P75 (80%) around the median marker at 70%.
Distribution chart of AI resolution rates across 195 rated deployments spanning 38 vendors (median taken across all labels, however each vendor names its own rate), cropped to the 40-100% range where deployments actually sit, with a red band from P25 (56%) to P75 (80%) around the median marker at 70%.
Take them with a grain of salt. They are an aggregate across the field, and they are directional, since no two deployments run the same ticket mix. The set is also self-selected, so it carries a ceiling.
The label also moves the number, with deployments reporting "resolution" at a 72.5% median, "automation" at 61% and "containment" at 58.2%.

Gate 2: stability

A resolution rate that is still climbing is a trend, and annualizing a trend overstates the return. We treat the rate of change as its own number, because it tells you whether you are still rising or have flattened out.
In our own deployments, early gains come from knowledge and then taper off; the bigger jumps arrive later, when live customer data and completed tasks get connected.
When a rate flattens at week six, find out whether the knowledge work has just finished or the system has stopped improving.

Gate 3: quality

Gate 3 is the only one of the three that can force a pause: Gates 1 and 2 say wait, Gate 3 says stop. We run it on four numbers.
Four-row table for Gate 3's signals: verified resolution and escalation accuracy carry no published threshold ('No'); reopen rate is an average, not a ceiling, read against MetricHQ's 3.1%; recontact rate is a window, not a ceiling, using a 72-hour window here. A second column gives how each number is actually obtained: pick a definition and hold it, reopened over solved, pick your own window, and sampled by hand.
Four-row table for Gate 3's signals: verified resolution and escalation accuracy carry no published threshold ('No'); reopen rate is an average, not a ceiling, read against MetricHQ's 3.1%; recontact rate is a window, not a ceiling, using a 72-hour window here. A second column gives how each number is actually obtained: pick a definition and hold it, reopened over solved, pick your own window, and sampled by hand.
  • Verified resolution. This gate turns on a disagreement: vendors do not agree on what "resolved" means. Verified resolution is also the number we get asked to explain most often. Intercom's Fin documentation is explicit about it:
"A resolution is a type of outcome that is counted when, following Fin's last answer in a conversation, the customer either confirms the answer was satisfactory (confirmed resolution), or exits the conversation without requesting further assistance (assumed resolution)." Intercom's Fin AI agent outcomes documentation.
Zendesk counts an email or web-form automated resolution after 72 hours of inactivity. It only counts one where the AI gave a generative reply, the end user gave positive feedback or none at all, no human agent responded, and "The AI evaluation process confirmed that the AI agent's response was relevant."
Both of them count silence as success. Intercom names an inactivity timeout as the mechanism that closes the conversation, without publishing a duration for it.
We count a conversation resolved when it was not escalated to a human. That only holds up because we make escalation easy.
  • Reopen rate. MetricHQ defines it as reopened tickets divided by tickets solved:
"Ticket Reopen Rate (RR) is the percentage of resolved support tickets that customers reopen, signalling the issue was not fully resolved." MetricHQ's ticket reopen rate definition.
The same page reports an average, with no limit attached, citing an Endsight survey of 260 companies that found 3.1% of tickets reopened. It also calls reopen rate a lagging indicator of resolution quality: it tells you something fell short, and the why comes out of sampling.
  • Recontact rate. Kustomer's glossary defines it as the share of customers who contact you more than once inside a defined window about the same unresolved issue, and points out that it is the direct inverse of first contact resolution. At 70% FCR, roughly 30% of customers are coming back.
Kustomer puts the typical measurement window at 7 to 30 days. We picked 72 hours instead, to mirror the clock Zendesk uses to declare an email resolution: if that much silence is good enough to count a win, a contact inside the same span should be good enough to cancel it.
  • Escalation accuracy. Nobody publishes a threshold for this one, so it gets sampled by hand: you read closed conversations and mark every one where the customer was reaching for a person and did not get handed to one. Intercom's own documentation names the failure mode: a customer who leaves after Fin's answer without requesting further help counts as an assumed resolution and gets charged as one. The tells sit in the last few messages: the same question asked twice in different words, an explicit ask for a human, or a closing line that reads as giving up rather than as satisfied (yes, including a polite one).

What a point of resolution is actually worth

The gates rate how trustworthy a number is, and in our data a trustworthy number and a valuable one come apart quickly.
A point of resolution won from a help center answer and a point won through a completed task each score one point on the dashboard, and the completed task saves far more handling time. We look at the time-weighted value of each point, and at the ladder underneath it: knowledge first, then live customer data, then actions the agent completes end to end.
Gate 3 needs evidence. Most of it is already sitting in your helpdesk.
In My AskAI, Self-Learning drafts new knowledge articles by comparing the AI's reply to the human agent's actual reply on handed-over tickets. The questions the AI could not answer come back to you as handed-over tickets. You can also ask Echo, the AI agent inside our dashboard, why a given conversation went the way it did and which knowledge it used.

How do you measure the ROI of AI customer support after launch?

TL;DR: Write down five numbers before you switch it on, then take one reading at day 7, one at day 30 and one at day 90, knowing which of the three gates each reading is allowed to clear.
Start with the baseline, and keep it to five numbers that each feed a gate later:
Four-step measurement sequence: define the five baseline numbers, record them before the AI goes live, wait through the early smoke-test window, then verify the reading against the Three Gates.
Four-step measurement sequence: define the five baseline numbers, record them before the AI goes live, wait through the early smoke-test window, then verify the reading against the Three Gates.
  1. Cost per contact, plus the headcount and hours behind it. This is what your cost per resolution gets compared against.
  1. Ticket mix by topic. Without it you cannot tell a resolution-rate move from a change in what customers are asking.
  1. CSAT on the tickets the AI is about to take, measured on those tickets alone.
  1. Reopen rate.
  1. Recontact rate, on whichever window you have picked.
We ask for reopen rate and recontact rate on day one, before the AI touches a single ticket.
A reopen rate with no pre-AI value cannot be read as a regression afterwards, so the day the AI goes live is the day that number becomes unrecoverable.
Then take three readings, and be strict about what each one is allowed to prove.
Day
What it can tell you
Which gate it may clear
Day 0
Nothing yet. This is the baseline everything else is read against
None
Day 7
Whether anything is going badly wrong
Gate 3, partially
Day 30
Whether the rate has become a measurement
Gate 1, usually
Day 90
Whether the rate has settled
Gate 2, usually
We treat day 7 as a smoke test that carries quality signals only. Even those signals are partial, because the vendors' own resolution clocks are still running and your day-7 window is still filling as you read it.
Day 30 is normally the first reading with enough events behind it to clear Gate 1. Day 90 is normally the first that can clear Gate 2, and in our experience these systems keep moving well past the quarter, because the big jumps land when data and tasks get connected.
Video preview
AI Customer Support Analytics
Our 30/60/90 day plan covers the month-by-month version, with what to do in month one, two and three.
Your baseline numbers are already in your helpdesk before you switch anything on. Pull the last 90 days out of your Zendesk or Intercom reporting and save it somewhere outside the tool, because you will want it long after the reporting window has rolled. And if you are starting without much written down, Train on Historic Tickets drafts starter articles from your last 5,000 historic tickets (more on request), so a thin help center is not a reason to skip the baseline.

When should you pause or roll back an AI support agent?

TL;DR: When Gate 3 fails, with reopens and recontacts climbing while deflection climbs, you pause, because that combination is support debt being taken on.
One condition triggers a pause. Your resolution or deflection rate is going up, and your reopen and recontact rates are going up with it.
That pairing is the first thing I look for when someone sends me a dashboard. Each reopened ticket is a resolution that did not hold. Each repeat contact is a dissatisfied customer whose problem survived the first attempt, at a second lot of cost.
Deflection climbing on top of that means the same work is arriving twice.
On thresholds, take only what the sources publish. MetricHQ gives an average of 3.1% reopened, and Kustomer gives a window of 7 to 30 days and the FCR inverse; neither publishes a ceiling.
The trigger is your own baseline moving: set the threshold against the number you recorded at day zero, and review it weekly (we do ours on a Monday, with the previous week's sample in front of us).
A pause and a roll back are different moves:
  • A pause narrows the scope. Keep the agent replying on the topics where Gate 3 is clean, and stop it on the ones where it is not. Nothing gets turned off.
  • A roll back turns off direct replies. The agent drops to copilot, drafting each reply for a human to send.
Teams tend to reach straight for the roll back, because it is the move they know. In our own product both are settings you change in an afternoon. Every action is a choice between running autonomously and proposing a step for a human to approve.
You can also trigger the agent on selected helpdesk workflows only, which is what RecruitCRM did. The copilot itself is the AI Copilot Chrome Extension, which drafts inside whichever helpdesk your team already works in.
Deflection on its own is the number I trust least. It does not show whether the customer's issue was resolved, and pushed hard it gets dangerous: you could run a perfect 100% deflection rate by making it impossible to speak to anyone.
The vendor definitions have the same hole. An assumed resolution is counted when a customer leaves, so a frustrated exit and a satisfied one score the same unless escalation fires first.

What does this look like in real rollouts?

TL;DR: RecruitCRM went from about 35% AI resolution at go-live to 68%, and the thing that moved it was a weekly QA review, with no model change involved.

RecruitCRM: 68% AI resolution, up from around 35%

RecruitCRM are an all-in-one SaaS platform for recruitment agencies, combining an applicant tracking system with a CRM, and they run our agent inside Intercom.
They handle about 1,088 conversations a month, roughly 740 of which are never touched by a human, saving 62 hours a month at five minutes a ticket. AI CSAT sits at 75%, and the AI has resolved around 5,700 tickets in the first year.
They ran a disciplined weekly QA review: working through the questions the AI could not answer, adding custom answers, tightening guidance.
They also connected live user data through our API and turned direct replies on from day one. Nothing about the underlying model changed between 35% and 68%.
Two facing metric-scale figures for RecruitCRM: ~35% AI resolution at go-live, before any structured review cadence, next to 68% AI resolution after a weekly QA review: a 33-point lift, from fixed knowledge gaps, added custom answers and tightened guidance.
Two facing metric-scale figures for RecruitCRM: ~35% AI resolution at go-live, before any structured review cadence, next to 68% AI resolution after a weekly QA review: a 33-point lift, from fixed knowledge gaps, added custom answers and tightened guidance.

A high-volume platform: where Gate 1 clears in days

One of our highest-volume deployments is a prop-trading platform on Intercom. They run about 105,000 tickets a month at 73% AI resolution and 68% AI CSAT, saving around 5,650 hours a month.
At that volume a single week produces a larger sample than a deployment taking fewer tickets manages in a quarter, so Gate 1 clears in days.

Sofar Sounds: 26%, and correct

Sofar Sounds run intimate live-music gigs in unconventional venues, and they are one of our Zendesk customers. Their AI resolution rate is 26%, with 85% AI CSAT.
About 195 tickets a month are resolved by the AI and about 555 are escalated to a human with full context, out of roughly 750, saving around 16 hours a month. The low rate is deliberate.
Their Handover and Escalation Guidance rules send most tickets to a person on purpose, and each one arrives with the context our agent has already gathered.
Sofar Sounds were aiming for 26%, in a deployment designed to hand most tickets to a person.
A My AskAI stat card reading "Customer service tickets resolved by our AI agents", showing 1,170,303, with a line beneath it stating that the AI resolution rate for the last 30 days was 74.9%, framed on the brand cream canvas with a white matte and two red sparkle accents.
A My AskAI stat card reading "Customer service tickets resolved by our AI agents", showing 1,170,303, with a line beneath it stating that the AI resolution rate for the last 30 days was 74.9%, framed on the brand cream canvas with a white matte and two red sparkle accents.

What should you do this week?

TL;DR: Write the baseline down today, even if you went live three months ago, because your helpdesk reporting still holds the pre-launch period.
  1. Write down the five baseline numbers, today. If you went live months ago, reconstruct them: pull the pre-go-live period out of your helpdesk reporting and save it outside the tool. Around 30 minutes. Every later number then reads as a change against something.
  1. Add reopen rate and recontact rate to the same sheet. Pick your recontact window and write it down next to the number, so nobody re-derives it differently in a month (ours is 72 hours). About an hour the first time, then five minutes a week.
  1. Label your last ROI reading with the gate it cleared. Go back to whatever number you last showed your finance lead and mark it Gate 1, Gate 2 or neither. 15 minutes, and it usually explains any argument you have had about that number since.
  1. Put a 45-minute weekly QA review in the calendar. Pull ten conversations the AI closed last week, and pick them deliberately: a few from your highest-volume topic, a few from wherever your reopens landed, and a few at random (the random ones are where the surprises live). Read each one against three questions: was the answer correct, did the conversation stay shut, and should it have gone to a person instead. Then route every miss to one of three fixes: a gap that needs a knowledge article, an answer the AI keeps getting subtly wrong that needs a custom answer written for it, or a handover that should have fired and needs a guidance rule. That is the loop RecruitCRM ran between 35% and 68%, so move it in the calendar if you must, but never drop it.
  1. Agree your pause rule before you need one. Write the sentence that will trigger it, with your own numbers in it, and get your team to agree it in your next stand-up. 20 minutes, and it turns a stressful judgment call into a decision you already made.

How do I get AI to run the Three Gates on my numbers?

Paste your own figures into the prompt below (your helpdesk reporting holds all of them) and it will tell you which gates your reading clears and which it does not. It cannot judge answer quality for you, so Gate 3 comes back as a sampling instruction rather than a verdict.
A My AskAI internal note in Zendesk headed "Want to make these AI replies even better?", listing four linked actions: inspect this conversation to see what knowledge was used; add guidance to control reply tone and style and how scenarios are handled; create custom answers to fill knowledge gaps (draft with AI); and continue drafting AI replies with the copilot extension. Framed on the brand cream canvas with a white matte and two red sparkle accents.
A My AskAI internal note in Zendesk headed "Want to make these AI replies even better?", listing four linked actions: inspect this conversation to see what knowledge was used; add guidance to control reply tone and style and how scenarios are handled; create custom answers to fill knowledge gaps (draft with AI); and continue drafting AI replies with the copilot extension. Framed on the brand cream canvas with a white matte and two red sparkle accents.
You are helping me decide whether an AI customer support ROI reading is safe to act on.

My numbers:
- Days since go-live: [number]
- Conversations resolved by the AI so far: [number]
- Current AI resolution or deflection rate: [number]%
- The same rate four weeks ago and two weeks ago: [number]% and [number]%
- Reopen rate, pre-AI baseline and now: [number]% and [number]%
- Recontact rate, pre-AI baseline and now: [number]% and [number]%, over a [number]-hour window
- Cost per contact before: [amount]. Cost per resolution now: [amount]

Answer each gate separately.
Gate 1 (volume): with that many resolutions behind the rate, how many points does one more ticket move it? Say whether the rate is a measurement yet.
Gate 2 (stability): has the rate of change flattened, or is the rate still climbing? Say whether annualising this reading would overstate the return.
Gate 3 (quality): are reopens and recontacts inside my pre-AI baseline? Flag a pause if my resolution or deflection rate is rising while reopens or recontacts rise with it.

Finish with one line: which gates this reading clears, and what it is therefore allowed to prove.
Where a number I gave you is missing or cannot settle a gate, write "unverified, sample it by hand" rather than estimating.

When does this framework not apply?

TL;DR: Low-volume support, deliberately narrow deployments and seasonal businesses all break the gates, and in a genuine incident you skip straight to the roll back.
The gates assume enough tickets to average out, and a ticket mix that stays roughly comparable on both sides of the switch. Plenty of real rollouts break one of those assumptions, and when they do, the framework stops applying.
Four cards in a 2x2 grid plus a full-width red card, showing exceptions to the Three Gates framework and what each one breaks: low-volume support (how long Gate 1 takes), a low rate on purpose (how the reading should be judged), narrow deployments (which population the reading covers), seasonal businesses (the baseline's like-for-like comparison), and, in red, a genuine incident (past the gates entirely, straight to the roll back).
Four cards in a 2x2 grid plus a full-width red card, showing exceptions to the Three Gates framework and what each one breaks: low-volume support (how long Gate 1 takes), a low rate on purpose (how the reading should be judged), narrow deployments (which population the reading covers), seasonal businesses (the baseline's like-for-like comparison), and, in red, a genuine incident (past the gates entirely, straight to the roll back).
Five come up repeatedly (two of them are on our own customer list):
  • Low-volume support. On a few hundred conversations a month, one ticket moves the rate by a full point. Gate 1 still needs the same number of events, so with fewer tickets a week it takes longer. Monthly reporting will lie to you either way.
  • A low rate that is the right outcome. Sofar Sounds at 26% is the on-file example. It should be read against the intent the deployment was built for.
  • Deliberately narrow deployments. An agent scoped to two topics should be measured against those two topics, and the other twenty are out of scope by design.
  • Seasonal businesses. A baseline taken in a trough and a reading taken in a peak are different populations (our ecommerce customers hit this every peak season). Either compare like periods year on year, or accept that the comparison is directional.
  • A genuine incident. Wrong answers going out at volume is an incident. Roll back first, diagnose afterwards.
The case that bothers me most clears everything. A team gets through Gates 1 and 2 comfortably and is still sitting on a knowledge-only floor that has stopped improving. The gates will certify that number as trustworthy, because it is.

The takeaway

TL;DR: Run every AI support ROI reading through the Three Gates before you act on it, and write your baseline down today so the gates have something to measure against.
A high deflection number with poor verified resolution is support debt. It gets repaid in reopened tickets, repeat contacts and customers who stop bothering, and it shows up on the dashboard as a win for months before it shows up in the numbers you care about.
The Three Gates are how we tell the difference: volume, stability, quality. Run every reading through all three, and give Gate 3 the same weight you give the rate itself.
The single action is the baseline: five numbers, written down and saved outside your helpdesk, today. If you are still building the business case, our guide to calculating AI customer service ROI carries the formula and the worked examples.
The RecruitCRM case study has the full rollout, including the weekly review that moved the number.

FAQs

How long should I wait before judging whether my AI support agent is working?
Day 7 supports quality signals only, and even those are partial while the vendors' own resolution clocks are still running. Zendesk waits 72 hours of inactivity before counting an email automated resolution, and Intercom counts a silent exit as an assumed resolution without publishing the duration behind it.
Day 30 is normally the first reading with enough events to be a measurement, and day 90 the first where the rate has settled enough to annualize. In our own deployments the rate keeps moving well past the quarter, especially once live data and tasks get connected.
What should I baseline before turning on an AI support agent?
The standard list is AHT, FCR, CSAT, churn and cost per contact, which Gorgias sets out well. Add reopen rate and recontact rate to it, because neither can be read as a regression later without a pre-AI value to compare against.
Five numbers is enough: cost per contact, ticket mix by topic, CSAT on the tickets the AI will take, reopen rate and recontact rate. Save them somewhere outside your helpdesk, since reporting windows roll.
Baseline number
What it is for later
Cost per contact
What cost per resolution gets compared against
Ticket mix by topic
Separates a rate move from a mix change
CSAT on the tickets the AI is about to take
Isolates the tickets the agent takes over
Reopen rate
No pre-AI value, no regression signal
Recontact rate
Same, plus the window you measured it over
What resolution rate should I expect from AI customer support?
Our AI resolution rate benchmark study puts the median AI handling rate at 70% across 195 rated deployments spanning 38 vendors, with the 25th percentile at 56% and the 75th at 80%. Across our own customer base we sit at 72%+.
Read those as an aggregate across the field; treat them as directional, since no two deployments run the same ticket mix; and remember that a self-selected set carries a ceiling. Watch the label too, because it moves the number.
Label the deployment reports
Median rate
Resolution
72.5%
Automation
61%
Containment
58.2%
How much can AI reduce customer support costs?
Nobody publishes a percentage you can safely adopt as a target. Fin publishes average returns of $3.50 for every $1 spent as a vendor claim (their number, their methodology). Set against that, Gartner's analysis of 432 use cases found only a quarter produce a return and another quarter produce a negative one.
Fin's own cost-savings page argues down from its industry's marketing, putting a realistic net reduction across a whole support organization at 20 to 35% in year one after infrastructure costs and the long tail of complex tickets.
The number to compute is your own. Zendesk defines cost per resolution as total support costs divided by verified resolutions, excluding abandoned conversations, unresolved tickets, duplicate contacts and anything that reopens inside your recontact window. It also warns that a falling cost per resolution alongside a rising repeat-contact rate may mean the work has moved somewhere else (the same pairing that trips Gate 3).
How do I make sure the AI isn't giving customers wrong answers?
Sampling plus escalation design, because no dashboard catches this on its own. The vendor documentation shows why: an assumed resolution counts a customer who left, so the metric scores a silent failure as a win unless something else catches it.
Our own resolution count rests on the same logic, which is why we make escalation easy and why we ask people to sample ten closed conversations a week. Echo will tell you why a specific answer landed the way it did and which knowledge it drew on.
How do you measure ROI with AI?
Compute the return the way you would for any support investment, then run the answer through the gates before you act on it. Our guide to calculating AI customer service ROI covers the formula, the cost-per-resolved-ticket model and the worked examples at 1,000 and 10,000 tickets.
The gates then tell you whether the number that comes out is safe to spend against.
What is the best KPI for measuring customer satisfaction?
There isn't one, and the metric owners say so. Qualtrics names NPS, CSAT and customer effort score as three metrics that complement each other, with CES coming out of CEB research finding that reducing customer effort predicts loyalty more strongly than delighting customers.
Zendesk puts a CSAT of 75 to 85% in the "good" band and above 90% as exemplary (while stating outright that the number is subjective, and that the trend over time is the useful signal). On tickets the AI handles, pair CSAT with a behavioral measure nobody has to opt into. Reopen rate and recontact rate count every customer who came back, including the ones who never answer a CSAT survey.

Start using AI customer service in your business today

Create AI customer service agent

Written by

Mike Heap
Mike Heap

Mike is an experienced Product Manager who focuses on all the “non-development” areas of My AskAI, from finance and customer success to product design, copywriting, testing and more.

Related posts

How to calculate AI customer service ROI (formula + worked example)

How to calculate AI customer service ROI (formula + worked example)

Most AI customer service ROI math counts deflected tickets and sticker prices. Here's the formula that uses cost per resolved ticket, with a worked example.

AI Customer Service KPIs That Actually Matter (and the 5 to Stop Leading With)

AI Customer Service KPIs That Actually Matter (and the 5 to Stop Leading With)

Most AI support dashboards lead with deflection rate, a proxy. Here are the AI customer service KPIs that actually predict ROI, and 5 to stop tracking.

The First 90 Days of AI Customer Service: A 30/60/90 Plan

The First 90 Days of AI Customer Service: A 30/60/90 Plan

AI customer service goes live in minutes; a real resolution rate takes a quarter. Here's the 30/60/90-day plan: what to do, measure and expect each month.

What Is a Good AI Resolution Rate? Benchmarks From 195 Real Deployments

What Is a Good AI Resolution Rate? Benchmarks From 195 Real Deployments

Everyone asks "what's a good AI resolution rate?" and gets a hand-waved number. We pulled real data from 195 deployments across 38 vendors. Here's the truth.

Containment vs deflection vs resolution: three metrics, decoded

Containment vs deflection vs resolution: three metrics, decoded

Containment, deflection, and resolution aren't the same metric. Here's the decoder: what each measures, the formulas, and the one number to report on.

What is deflection rate? The formula, benchmarks, and what it misses

What is deflection rate? The formula, benchmarks, and what it misses

Deflection rate is the % of support contacts handled before they reach a human. Here's the formula, what counts, real benchmarks, and why it isn't resolution.

The 7 Most Common AI Customer Service Mistakes (and How to Avoid Them)

The 7 Most Common AI Customer Service Mistakes (and How to Avoid Them)

The most common AI customer service mistakes trace back to one: treating it as set-and-forget. Here's the operator's fix for each of the 7, with real numbers.

The Hidden Costs of AI Customer Service (and How to Find Them Before You Sign)

The Hidden Costs of AI Customer Service (and How to Find Them Before You Sign)

The hidden costs of AI customer service never make the pricing page: setup, integrations, escalation, upgrade fees, oversight. Price them before you sign.

How RecruitCRM achieves 68% AI resolution, saving 62 hrs each month

How RecruitCRM achieves 68% AI resolution, saving 62 hrs each month

RecruitCRM rejected Fin AI at $0.99/resolution, chose My AskAI instead. Started at 35%, now at 68% — through disciplined weekly QA reviews.