Mike is an experienced Product Manager who focuses on all the “non-development” areas of My AskAI, from finance and customer success to product design, copywriting, testing and more.
If your weekly report has ten customer service metrics on it, you're tracking a dashboard; the support underneath is going unmeasured.
Every guide to customer service metrics hands you roughly the same list of ten numbers, and every helpdesk puts most of them on the default dashboard. So most weekly support reports look the same: a wall of metrics, every number given the same weight.
Customer service metrics only answer three questions: did we solve it, how fast, and what did it cost. You need one scoreboard number per question; the rest are diagnostics you check when a scoreboard number moves. Treating all ten as equals is how a team ends up proud of a fast first reply while customers churn.
I'm Mike, co-founder of My AskAI. We help 200+ ecommerce and SaaS businesses run AI customer service inside the helpdesks they already use (Zendesk, Intercom, Freshdesk, Gorgias and HubSpot), and we publish resolution benchmarks from 195 real AI deployments. I spend a lot of my week looking at other teams' support dashboards, and the numbers on them rarely line up with the ones that predict a good quarter.
Why does every customer service metrics list look the same (and what does it get wrong)?
⚡
TL;DR: The standard 10-metric list is legitimate but unweighted. It hands a team ten numbers with no hierarchy, so teams default to reporting the easiest ones (the speed metrics) while the outcome question goes unmeasured.
The consensus sources aren't wrong about the metrics themselves. Qualtrics' top-10 guide splits experience measures from operational ones and gives real per-channel response targets: 24 hours for email, 60 minutes for social, three minutes for phone, instant for live chat. Zendesk's 18-metric list leads with a stat I'd keep on file: 73% of business leaders report a direct link between their customer service and business performance.
Gainsight at least tiers its ten by program maturity, telling beginners to start with CSAT and NPS. That's already more guidance than I found anywhere else on page one.
None of them gives you a hierarchy. Qualtrics avoids prescribing which metrics to prioritize, Zendesk doesn't say which of its 18 to report, and Intercom's guide has clean formulas for every metric but stops short of ranking them. The closest any of them gets is one guide's advice to pick 3 to 5 metrics tied to a single goal, with no rule for the picking, and a rule for the picking is exactly what I'd want first.
The search results give the game away, too. Sitting among the ranking guides is a thread on r/CustomerSuccess asking what metric tells you support is working, beyond the standard ones. Practitioners are asking because the listicles ranking above the thread don't answer it.
When nothing tells a team which number to report, the easiest number wins, and the easiest numbers are the speed metrics. Customers do rank speed highly, so a fast first reply time reads like proof the operation works. I don't blame teams for defaulting there.
A team in that spot reports first reply time and ticket volume every week, and both trend well. CSAT drifts down for a quarter and nothing on the dashboard says so, because the dashboard measures speed and activity while the outcome question (did we solve it?) has no number at all.
The Three-Question Scoreboard: the only 3 customer service metrics that matter weekly
⚡
TL;DR: Every customer service metric answers one of three questions: did we solve it, how fast, and at what cost. Pick one scoreboard metric per question (resolution rate paired with CSAT, total time to resolution, cost per resolution) and demote the rest to diagnostics.
Every customer service metric you've ever been handed answers one of three questions:
Did we solve it? (the outcome)
How fast? (the speed)
At what cost? (the load)
The Three-Question Scoreboard is the rule we use to organize them: one scoreboard metric per question, everything else demoted. A scoreboard metric is a number you report weekly and act on. A diagnostic metric is a number you open only when a scoreboard number moves.
Hub-and-spoke diagram showing every customer service metric folding into three questions: did we solve it, how fast, and at what cost, each with its scoreboard metric.
Question
Scoreboard metric
Diagnostics underneath it
Did we solve it?
Resolution rate, paired with CSAT
FCR, CES, reopen rate
How fast?
Total time to resolution (median, per channel)
First reply time, average reply time, time to first human response on escalations
At what cost?
Cost per resolution
Ticket volume, backlog, tickets per agent, occupancy
Q1: Did we solve it? (the outcome)
The scoreboard here is a pair, reported together every week. Resolution rate is the number I'd protect first (it means an issue got solved rather than deferred or abandoned), and the CSAT next to it keeps resolution from being gamed.
The formula is resolved conversations divided by total conversations. The trap is the numerator: vendors define "resolved" differently, and the definition you write down decides what the number tells you.
For AI-handled conversations we have first-party data. Across 195 rated AI deployments covering ~55 vendors, the median AI-handled rate is 70%; a quarter of deployments sit below 56%, a quarter above 80%, and the full range runs from 15% to 98.3%. The label alone moves the number: deployments reporting "resolution rate" show a 72.5% median while those reporting "automation rate" show 61%, so two vendors quoting different headline numbers are usually counting different things rather than performing differently.
Stat callout showing the 70% median AI-handled resolution rate, the 56 to 80 percent middle half, and the 72.5% vs 61% medians for resolution versus automation labels.
Three caveats travel with that dataset: we publish it as aggregate bands, never head-to-head vendor duels, definitions vary enough that comparisons are rough at best, and published rates skew toward best cases. I posted the headline findings on LinkedIn when we published the dataset.
Resolution rate is only as good as its definition. Ours counts a conversation as resolved when the customer didn't need a human.
We keep escalation easy (a customer can ask for a person in any words at any point, and the AI hands over when it can't answer, spots frustration, or hits a topic you've marked for humans), so a conversation that stays with the AI is the customer telling you they got what they needed. Across our customer base the rolling 30-day figure is 72%.
For the CSAT half of the pair, Retently's benchmarks put 65% to 80% as the dominant band across industries, with 70 to 90 a good place to be. Their 2026 industry values run from 83 for consulting and 81 for financial services down to 77 for ecommerce and retail, with B2B software in the high 70s (worth checking your own industry's band before you judge your number against consulting's). The formula: satisfied responses divided by total responses, times 100.
When the pair moves, open the diagnostics. First contact resolution benchmarks at a 70% average on SQM Group's data, with 70% to 79% a good rate, 80%+ world-class, and only 5% of call centers reaching world-class.
Human-rep FCR at ~70% and the AI-handled median at ~70% are different numerators that happen to share a number, and the mix-up comes up on plenty of our onboarding calls: one counts first-touch fixes by an agent, the other counts the share of a queue an AI handled end to end.
Customer effort score is the other outcome diagnostic, and it predicts loyalty hard: Gartner research found 96% of customers with high-effort service interactions become more disloyal, against 9% for low-effort ones. Reopen rate is the anti-gaming check: a "resolved" ticket that reopens wasn't solved. When CSAT dips under a steady resolution rate, reopens are where I'd look first.
The failure mode for Q1 is reporting resolution without CSAT. A fast, wrong answer counts as resolved in most vendors' math and looks fine on the dashboard until the CSAT next to it says otherwise.
Q2: How fast? (the speed)
The scoreboard is total time to resolution: the clock starts when the ticket opens and only stops when the issue is done. A fast first reply doesn't pause it. Report it as a median, per channel.
Jitbit's classic 1,000-company dataset put the median resolution time at 82 hours, or 3 days and 10 hours (it's a 2017 analysis, and the 2026 update to the article notes the figures now read on the generous side). MetricNet's ROI of Support report has the quartiles for IT service desks: top-quartile desks resolve in 0.8 hours against 5.0 hours for the bottom quartile, and desktop support runs 2.9 against 12.3.
The per-channel numbers sit much tighter than the overall medians. Live chat is judged in seconds and email in hours, so a single blended speed target punishes one channel and flatters the other (we keep a channel-banded benchmark table in our time to first response glossary).
The diagnostics under the speed scoreboard: first reply time (you'll also see it as time to first response), average reply time, and time to first human response on escalations. That last one is the number that still measures human responsiveness once an AI answers first.
First reply time is the most-gamed number in support. An auto-acknowledgment lands in seconds and helps nobody, and once an AI answers first (ours included), raw first reply time collapses to seconds across the board and stops measuring anything real. Keep the speed scoreboard on the full clock, and keep a separate diagnostic clock on how fast a human picks up escalations.
Q3: At what cost? (the load)
The scoreboard is cost per resolution: total support cost divided by issues resolved (the number a budget conversation runs on). Cost per ticket is the diagnostic underneath it: dividing by tickets pays the team for activity, whether or not anything got solved.
The best published cost ladder is IT service desk benchmarking, so for a CX queue I'd treat it as rough context only. MetricNet's ROI of Support report puts the average fully loaded cost per ticket at $22 for the service desk, $69 for desktop support and $104 for level-3 IT support. The full escalation ladder runs from $0 for incident prevention and $2 for self-help up to $221 for field support and $599 for vendor support.
Horizontal bar chart of MetricNet's cost-per-ticket escalation ladder from $0 incident prevention up to $599 vendor support.
The same report found 21% of tickets resolved by desktop support could have been resolved a level down; that's what a mis-escalation costs.
Channel economics from the same source: a chat resolution costs 76% as much as a voice one, and email runs about 81%. I can't give you a clean benchmark for what your cost per resolution should be; nobody publishes a credible cross-industry ladder for ecommerce or SaaS queues, so the IT ladder is context and your own trend line is the number to manage.
The diagnostics under the cost scoreboard: ticket volume, backlog, tickets per agent, and occupancy (the industry standard is 83.3%, with the recommended band between 85 and 90%).
The cost failure mode is celebrating falling volume that's customers giving up. You could run a "perfect" 100% deflection rate by making it impossible to speak to anyone, and a deflection rate above 90% without matching resolution and CSAT numbers usually means the tool is blocking customers on their way to a human.
As I put it on LinkedIn, any vendor can hit 90% resolution if they block the humans. Falling volume is a question to answer before anyone reports it as a win.
Which customer service metrics should you stop reporting weekly?
⚡
TL;DR: Stop reporting NPS, raw ticket volume, first reply time (once an AI answers first), occupancy, and all-channel averages every week. Keep measuring them; each one either lags too far, games too easily, or answers a question nobody asked this week.
These five are diagnostics wearing scoreboard clothes: stop reporting them weekly, keep measuring them.
NPS. A company-wide, lagging signal, and no lever a support team can pull. MeasuringU's walk through the academic record found the original analysis correlated NPS to growth that had already happened, and replications (including a study in the Journal of the Academy of Marketing Science) found satisfaction metrics performed comparably; their verdict lands at "qualified", short of "discredited". Either way, a support team can't move NPS week to week; it belongs in the exec deck.
Raw ticket volume. Falling volume can mean fewer problems, seasonality, or customers giving up, and rising volume can mean growth. The number has no direction on its own; keep it as a cost diagnostic.
First reply time, once an AI answers first. AI replies land in seconds, so the raw number collapses to seconds and stops telling you anything. Swytch told us their time-to-first-response metric no longer made sense to track once our AI agents went live; their case study records reply times for automated inquiries dropping to zero and overall response times going from days to minutes. The replacement diagnostic is time to first human response on escalations.
Occupancy. A staffing and capacity metric: right for workforce planning, useless as a weekly "is support working" number. The industry standard is 83.3%, and chasing it higher burns agents out.
All-channel averages. Averages hide the tail; a handful of weekend tickets inflates a mean. Jitbit publishes its dataset in medians for exactly this reason. Medians and percentiles, per channel, is how I'd report every speed number.
This demotion list is also what practitioners keep asking for. The r/CustomerSuccess thread is support leaders asking what to look at beyond the standard dashboard numbers; the answer is three numbers and a cut list.
Two-column graphic mapping five metrics to stop reporting weekly, NPS, raw ticket volume, first reply time, occupancy and all-channel averages, to where each belongs instead.
What does this look like in real rollouts?
⚡
TL;DR: Named rollouts, scoreboard first: YesLMS runs 76% resolution with 88% CSAT, Honeygain 90% with 78% across ~3,400 tickets a month, GiveCard 95% with 90%. When a scoreboard number moved, the diagnostics underneath it held the answer.
Here are four of our rollouts, read through the scoreboard. All four run on Zendesk, and every figure comes from the published case studies.
YesLMS: 76% resolution, 88% CSAT
YesLMS is an edtech learning platform handling ~300 Zendesk tickets a month. The outcome pair reads healthy on both sides: 76% AI resolution with 88% CSAT, worth roughly 17 hours a month back to the team.
The lever behind the pair was Self-Learning, our feature that drafts new knowledge articles from the tickets the AI hands to humans. About 200 ticket responses in a 30-day window came from those auto-drafted articles.
Honeygain: 90% resolution, 78% CSAT, ~3,400 tickets a month
Honeygain is a consumer passive-income app, which means a high-friction queue: account bans, withheld payouts, users who arrive annoyed. The AI resolves 90% of it (~3,060 of ~3,400 monthly tickets) at 78% CSAT, saving ~507 hours a month, with ~600 tickets a month answered by Self-Learning's auto-drafted knowledge alone.
In a ban-appeal queue, reporting resolution alone would have shown the 90% and hidden the 78%.
GiveCard: 95% resolution, 90% CSAT
GiveCard runs fintech disbursement support: ~255 tickets a month at 95% AI resolution and 90% CSAT, around 20 hours a month saved.
At that queue size the cost question reads differently: the denominator is small, so I'd watch hours saved over cost per resolution.
Swytch: the retired speed metric
Swytch sells e-bike conversion kits and resolves 4,050+ tickets a month entirely by AI, at an 81% deflection rate. Before the rollout they told us reply times were creeping up and customers were waiting too long for answers to very simple questions.
Afterward, reply times for automated inquiries dropped to zero and overall response times went from days to minutes. At that point first reply time stopped being a metric worth tracking, and they retired it; in their words:
"The difference was immediate. Customers now get instant answers to routine queries, and our team can dedicate more time to solving complex problems."
Set against the 195-deployment distribution, these four sit around the top quarter of the field (the 75th-percentile mark is 80%). The dataset is presented as distribution bands, so read it as context for your own target. More named rollouts live on our case studies page.
How do you wire the scoreboard into your weekly review?
⚡
TL;DR: Five actions, each two hours or less: pick your three scoreboard metrics, define "resolved" in writing, move speed to medians, set diagnostic tripwires, and review the same three numbers every week.
Pick your three scoreboard metrics. One per question; the defaults from this post are resolution rate paired with CSAT, total time to resolution, and cost per resolution. About an hour, and the output is a one-line scoreboard definition your team can read.
Define "resolved" in writing. Vendors don't agree on what the word means: Intercom's Fin counts an end-to-end resolution or a configured Procedure ending in handoff, Zendesk's Autoresolve fires after a 72-hour quiet period, Gorgias Automate counts a ticket the customer didn't reply to, and Sierra counts agreed outcomes including saved cancellations. Ours is one line: a conversation is resolved when the customer didn't need a human. Write yours the same way so anyone can audit the numerator. Budget two hours.
Move speed metrics to medians, per channel. Most helpdesk reporting can do this in about 30 minutes. The payoff: the speed number stops jumping every time one weekend ticket lands.
Set diagnostic tripwires. Give every diagnostic a threshold that triggers a look: reopen rate above X, backlog older than Y, escalation response time over Z. About an hour. Diagnostics leave the weekly report but can still raise a hand.
Review the scoreboard weekly, same three numbers every week. A 15-minute standing agenda, and the boring-but-effective option: trend lines only mean something when the metric underneath them holds still, so resist swapping metrics between weeks. (If you're running our AI agents, Insights does the grouping for you: it sorts every conversation into topics and scores 100% of conversations for AI CSAT, so the topics needing attention surface on their own.)
How do I get AI to sort my metrics into scoreboard and diagnostics?
The sorting itself is mechanical once the framework is in front of you, so it's a good job to hand an LLM. Paste this into ChatGPT or Claude along with your current weekly report and it applies the Three-Question Scoreboard to your numbers.
AI Customer Support Analytics
One limit to know before you run it: the model can only sort the metrics you paste in, and it can't audit what your vendor counts as "resolved". That part stays with you (it's action 2 above).
I'm going to give you my support team's weekly metrics report. Sort every metric on it using the Three-Question Scoreboard:
1. Every customer service metric answers one of three questions: Did we solve it? How fast? At what cost?
2. Each question gets exactly ONE scoreboard metric, reported weekly. The defaults: resolution rate paired with CSAT (outcome), total time to resolution as a median per channel (speed), and cost per resolution (cost).
3. Everything else is a diagnostic: it comes off the weekly report and gets a tripwire threshold that triggers a look when crossed.
My weekly report: [paste your metrics report, with the numbers]
My monthly ticket volume and channels: [e.g. 2,000 tickets/mo, email + live chat]
How my helpdesk or AI vendor defines "resolved", if you know it: [paste the definition, or write "unknown"]
Produce:
1. A table sorting every metric I gave you: the metric, which of the three questions it answers, scoreboard or diagnostic, and a one-sentence reason.
2. The gaps: any of the three questions with no metric currently answering it.
3. A suggested tripwire threshold for each diagnostic, based on the numbers I pasted.
4. For anything you can't judge from what I gave you, write "check this yourself" instead of guessing - especially whether "resolved" is being counted honestly.
When is the standard 10-metric dashboard the right call?
⚡
TL;DR: Three legitimate exceptions: teams in diagnosis mode after a known problem, contact centers with metrics written into their SLA contracts, and teams under ~200 tickets a month who read every ticket anyway.
The full dashboard has three legitimate homes. In all three, the big consensus lists from Qualtrics and Zendesk are the right tool.
The first is diagnosis mode. After a known problem (a CSAT drop, a helpdesk migration, a new product line), the full diagnostic wall is the right view for a set stretch. The mistake is leaving the wall up after the problem closes.
The second is SLA-contracted contact centers and BPOs, where some metrics are written into the contract. The traditional service level is to answer 80% of calls in 20 seconds, many centers now target 90/15, and average handle time benchmarks at a little over six minutes. Missing a contracted level can carry financial penalties for a BPO, so I'd never argue a contracted metric off a weekly report.
The third is small queues. Under roughly 200 tickets a month you read every ticket anyway, so the operator is the diagnostic layer. Dashboard tooling adds little until volume grows past what one person can hold in their head.
The takeaway
⚡
TL;DR: Three questions, one scoreboard metric each; everything else is a diagnostic you open when a scoreboard number moves. Write your three down and take the rest off the weekly report.
Customer service metrics answer three questions: did we solve it, how fast, and at what cost. The Three-Question Scoreboard is one metric per question (resolution rate paired with CSAT, total time to resolution, cost per resolution), with everything else demoted to a diagnostic you open when a scoreboard number moves.
If you take one action from this post, write your three down and take the rest off the weekly report.
We'd start with the outcome pair, because it's the piece that catches a gamed number. When your AI rollout needs a scoreboard of its own, the AI-specific version of this argument sits in our AI customer service KPIs post, and the data behind the 70% median is in our resolution-rate benchmarks report.
FAQs
What is customer service metrics?
Customer service metrics are the numbers a support team uses to measure whether customer issues get solved, how fast, and at what cost. Most guides split them into operational metrics such as first reply time, resolution time and ticket volume, and experience metrics such as CSAT, CES and NPS. I'd split them differently: three scoreboard metrics you report weekly, and diagnostics you open when a scoreboard number moves.
What are the 4 metrics of customer service?
The classic four are CSAT, NPS, CES and first contact resolution (FCR). All four sit on the outcome side; a working scoreboard also needs a speed number and a cost number.
Metric
What it measures
CSAT
Satisfaction with a specific interaction
NPS
Relationship-level loyalty
CES
How easy you were to deal with
FCR
Share of issues fixed on first touch
What are the most important customer service metrics?
Resolution rate paired with CSAT: together they say the issue ended and ended well, and resolution is the number I'd protect first because it implies an issue got solved. After the pair come total time to resolution, measured as a median per channel, and cost per resolution. Everything else is a diagnostic.
What are some examples of customer service performance metrics?
Scoreboard examples with benchmarks: resolution rate (the AI-handled median across 195 deployments is 70%), CSAT (65% to 80% is the dominant band across industries), total time to resolution (82 hours median in Jitbit's classic dataset), and cost per resolution ($22 per ticket is the IT service desk average). Diagnostic examples: first contact resolution (70% average), customer effort score, reopen rate, backlog, and occupancy (83.3% industry standard).
How do you measure customer service metrics?
Start with the formulas: CSAT is satisfied responses divided by total responses, times 100; FCR is issues resolved on first contact divided by total issues. Resolution rate is resolved conversations divided by total conversations, counting "resolved" however your written definition says; vendors differ, so pin yours down. Cost per ticket, as MetricNet defines it, is the total monthly operating expense of a service desk divided by the monthly ticket volume.
What metrics measure customer satisfaction?
CSAT is the transactional one: how satisfied was the customer with this specific interaction. CES measures how easy you were to deal with; Gartner research found 96% of customers with high-effort service interactions become more disloyal, against 9% for low-effort ones. NPS measures the whole relationship, which is why I'd keep it off the weekly support report; it lags too far behind to steer by.
What resolution rate should I expect from AI customer support?
Across 195 rated AI deployments, the median AI-handled rate is around 70%; a quarter of deployments sit below 56% and a quarter above 80%. The label moves the number (deployments quoting "resolution rate" show a 72.5% median against 61% for "automation rate"), and quoted vendor numbers are self-selected best cases. The only way to know yours is to test on your own tickets; the full data is in our resolution-rate benchmarks report.
Is there a customer service metrics template I can copy?
Here's the scoreboard as a copyable template:
Question
Scoreboard metric
Formula
Benchmark
Diagnostics
Tripwire example
Did we solve it?
Resolution rate + CSAT
Resolved conversations ÷ total (write your definition down); satisfied responses ÷ total × 100
~70% AI-handled median; CSAT 65-80%
FCR, CES, reopen rate
Reopen rate rises two weeks running
How fast?
Total time to resolution
Opened to done, median per channel
82 hrs median (Jitbit, dated); top-quartile IT desks resolve in 0.8 hrs
First reply time, average reply time, time to first human response
Escalation response time over four hours
At what cost?
Cost per resolution
Total support cost ÷ issues resolved
$22/ticket IT service desk average
Volume, backlog, tickets per agent, occupancy
Backlog over a week old starts growing
Fill in your own numbers, then review the same three scoreboard rows every week.
Mike is an experienced Product Manager who focuses on all the “non-development” areas of My AskAI, from finance and customer success to product design, copywriting, testing and more.