How to Run an AI Customer Service Pilot: The 30-Day Plan
An AI customer service pilot judged on one day-30 resolution rate can fail the wrong vendor. Take three readings and decide on where the rate is heading.
Mike is an experienced Product Manager who focuses on all the “non-development” areas of My AskAI, from finance and customer success to product design, copywriting, testing and more.
A 30-day pilot cannot show how good an AI support agent will become, but it can show how fast it is improving and why it still hands tickets to your team. Read the resolution rate at around day 10, day 20 and day 30, and base the buying decision on the day-30 rate projected one step forward.
You have a trial running, a day-30 meeting in the calendar, and a number you are supposed to bring to it. Everyone in that meeting expects one figure: the resolution rate, set against a threshold somebody wrote down before anyone had touched the tool. The standard playbook says run it for longer, because a single-category pilot commonly takes eight to twelve weeks, but the business gave you thirty days. Most of the trials we start look exactly like that.
Thirty days is enough, and the trouble is the number of times you plan to read the rate. One reading, taken at the very end, mostly tells you how much of your documentation was ready in month one. That is a fact about your content and your calendar.
My view is that the day-30 number taken alone is the least informative figure the whole exercise produces. You want a pilot to tell you the direction the resolution rate is moving in across the month.
We tell every customer that day one is the worst it will ever be, and the rollouts we watch bear that out. The same agent that reads 35% in week two settles in the high sixties, while the one that started at 24% on an incumbent tool never moved.
So take three readings instead of one. I have spent the last two years watching this number move for support teams on Zendesk, Intercom, Freshdesk, Gorgias and HubSpot. The teams that pick well are the ones who set the pilot up to show them where it is heading. That is the whole method, and we run our trials the same way.
Why does a day-30 resolution rate pick the wrong vendor?
⚡
TL;DR: A single end-of-pilot reading mostly measures how much knowledge you loaded in month one, and the published spread is so wide that the number on its own cannot separate a good agent from a badly prepared one.
Start with the strongest version of the standard advice, because the discipline in it is real: "Agree, in writing, the result that would make you pay", as a guide to running an AI tool trial puts it. The same guide wants that result set before anyone has seen the tool, because two weeks of use is enough to shift everyone's judgment. That is right about human nature, and a pilot without a pre-agreed threshold does end in opinions. The part I disagree with is what the threshold gets applied to: one measurement, taken at the end.
The first problem is the published spread, and I raise it on calls before anyone sets a threshold. Where the field actually lands is a median of 70%, with a mean of 66.6%, a 25th percentile of 56% and a 75th percentile of 80%. The full range is 15% to 98.3%, across 195 rated deployments and 38 vendors.
A histogram showing where AI resolution rates land across 195 deployments and 38 vendors, with markers at the 25th percentile (56%), median (70%), and 75th percentile (80%).
That is not a spread you can pick a vendor out of with one number. A pilot reading 62% could be a weak agent that has already stopped moving, or a strong one that has been live for eleven days. I cannot tell those two apart from the figure, and neither can you.
The second problem is that the word in front of the number moves it more than the capability does. In that same set, deployments whose vendors report a "resolution" rate have a median of 72.5% across 108 readings. The ones reporting an "automation" rate show 61% across 52, so roughly twelve points of the difference come from the label alone.
Three caveats apply every time those figures come up, including when I use them:
The field as a whole - they describe the market in aggregate and never a single named competitor.
Directional only - vendors count a resolution differently, some the moment the AI stops replying and some only once the customer confirms the fix, so two rates rarely rest on the same rules.
Skewed high - vendors publish their good deployments and leave the disappointing ones alone, ours included.
The third problem catches the most careful buyers: a rate can be held down on purpose, and held up the same way. Any vendor can reach 90% by blocking the path to a human, and any customer can drop to 25% by routing most conversations to a person. The raw rate cannot tell a design choice from a failure, which is why I have written about this at greater length on LinkedIn. It is the most common way a good tool gets failed by a pilot.
The last problem is coverage. Most pilots judge quality on the conversations a customer bothered to rate, typically the ones included in a CSAT survey. The sourced position on response rates is that no universal good number exists, with plenty of healthy feedback programs down in the single digits or low double digits. We score AI CSAT across 100% of conversations in My AskAI's Insights, with no survey sample to wait for, so you can read the quality signal at day 20, well before the month is up.
What should you actually measure in a 30-day AI customer service pilot?
⚡
TL;DR: Take a resolution rate at about day 10, see how far it has climbed by about day 20, take the day-30 figure, and let only that last number, pushed forward by the same climb, decide whether you buy.
The first two readings tell you what to fix, and only the third one decides. The standard method folds all three into a single number and then asks that number to make a decision it cannot make. Here is what you record at each point, and what each reading is for.
Reading
When
What you record
What it tells you
Reading 1
around day 10
the resolution rate since go-live, on live tickets the AI is allowed to answer
how much usable content you had ready before the clock started
Reading 2
around day 20
the same rate, plus how much it has risen since day 10
whether the agent gets better when you teach it, and how quickly
Reading 3
day 30
the day-30 rate, projected forward by one more step of the same rise
whether to buy: the only reading you compare with your pass mark
The closest thing I have found to this is the enterprise pilot run in fixed phases, and two published versions are worth a look. One enterprise pilot guide sets conversational AI pilots at eight to twelve weeks, and says pilot performance "should be documented by week", which is a genuine weekly record. A published pilot scorecard evaluates after 30 measured operating days, against a baseline taken before the AI went live.
A three-step process showing the three readings of an AI pilot: at around day 10 the resolution rate shows how ready your content was, at around day 20 the rise since day 10 shows whether the agent improves when taught, and at day 30 the rate projected one step forward is the only number compared with your pass mark.
Both of them stop at a single decision point. Neither reads the AI's number more than once inside the pilot, and neither makes the buying decision on where that number is heading. Their weekly record exists to hand a starting figure to the live rollout, while the pass-or-fail call still falls on the last reading. I wanted the decision to rest on the direction instead.
Reading 1: the resolution rate at around day 10
The first reading is simply the resolution rate the agent produces on live tickets. Take it at around day 10, once the agent has seen enough real conversations to be answering from your material. At that point it mostly reflects how much usable content you had ready before the clock started. Write it down, then resist every instinct to compare it to anything (at this point it has no external meaning whatsoever).
What the day-10 rate is good for is showing you which content is missing, before you have spent the month. A low first reading is a to-do list for your content team. In our own data the gap between where a rollout starts and where it settles ranges from four to sixty points, so the day-10 number cannot be allowed anywhere near a decision.
You can take this first reading without sending a single AI answer to a customer, and on a first pilot that is usually the sensible route. Internal Note Replies, our shadow mode, drafts a reply on every ticket as a private internal note, so your team reads and compares the draft while continuing to answer the customer themselves. You get a day-10 rate built on real tickets without a customer ever seeing a draft, which usually makes the trial far easier to get approved internally.
An email support conversation with the subject line 'Zendesk Ticket Notification 26457' in which the My AskAI agent's draft reply is tagged Internal, answering the customer's pricing, API error and password questions as a private note.
Reading 2: how much the rate has risen by day 20
Between the first and second readings you spend one deliberate week improving the content the agent learns from, and nothing else changes: no second vendor, no new escalation policy, no new ticket types for the AI to answer. Hold everything else still, and the change you read at day 20 belongs to the teaching you did, which makes the movement a fact about the agent. I have watched a mid-pilot scope change wipe out a perfectly good rise.
The rise between day 10 and day 20 is the most useful number in the pilot. It tells you whether the agent turns the material you feed it into resolutions, and how fast. An agent that starts low and climbs quickly is a much better buy than one that starts in the middle and barely moves, and a single end-of-month reading cannot see the difference.
AI Customer Support Analytics
That week of work is bounded, and worth scoping properly before you start. If your team's view is that the documentation needs work first, pair the week with one or both of these:
Train on your past tickets - the agent learns from 5,000 of your past conversations by default, and more on request, which usually covers more ground than a week of writing articles would.
At day 30 you take the third reading, and then you do one piece of arithmetic on it (the only arithmetic in the whole method). Add the rise you saw between day 20 and day 30 on top of the day-30 rate: if the day-20 rate was 55% and the day-30 rate is 68%, the projected rate is 81%. It is an estimate of direction, so project it forward by exactly one step and no further.
The hard stop at one step is there because three points is the least any projection can be built on. As a standard forecasting textbook puts it, "The only theoretical limit is that we need more observations than there are parameters in our forecasting model. However, in practice, we usually need substantially more observations than that." Stretching a short trend forward is favored in the short term for its simplicity, and it stops being defensible the moment you push it out a quarter. Treat the projected rate as a direction, with a band around it.
Then take the pass mark you agreed before the pilot began, and compare the projected rate with it. The published spread gives you sensible bands to set that mark against.
Projected day-30 rate
Decision
Why
below 56%
no
the projection falls under the 25th percentile of the published field, with a month of teaching already spent
56% to just under 70%
undecided, check why tickets were handed over
the projection is inside the field but below the median, so the reason the agent handed tickets over decides it
70% or above
yes
the projection is at or above the published median across 195 rated deployments
above 80%
yes, with a check on the scope
the projection is above the 75th percentile, which more often means the range of tickets is narrow than that the agent is exceptional
Record why every escalation happened
The undecided middle band is where most of the pilots we see actually finish, and no further rate will resolve it. Instead, put a one-word label on every conversation the AI handed to a human, on the day it happens, while the reason is still visible. Four labels cover almost everything.
Label
What happened
Where the fix lives
Knowledge
the answer exists somewhere in your business but was not in the material the agent was trained on
new help center content, or training the agent on your past tickets
Data
the answer needed a live customer record, such as order status or subscription state
a live data connection, through the User Data API or the Shopify connector
Task
the agent had the right answer but could not carry out the action the customer wanted
actions and tools, configured to run autonomously or to propose and hand over
Out of scope by design
your escalation rules sent it to a human on purpose
nothing, so exclude it from the judgment instead of counting it as a failure
The mix of those four labels is what settles an undecided result. A pilot at 63% with three quarters of its handovers labeled Knowledge is a buy, and in my experience that is the most common pattern for a first deployment. Knowledge handovers are the cheapest kind to fix: you add the missing content to a connected knowledge source, the agent syncs it within 24 hours, and you have already proved the agent turns new material into resolutions. A 63% built mostly from Task handovers points at an integration project, which is a different kind of work on a much longer timeline.
Labeling is only cheap if the reason is visible at the moment of escalation. In My AskAI you ask Echo, our in-dashboard assistant, why the agent answered the way it did and which source it used. Whatever tool you are piloting, check before day one that somebody on your team can see the reason without filing a support request.
An internal note from the AI Agent headed 'Want to make these AI replies even better?', listing four links: inspect this conversation to see what knowledge was used, add guidance, create custom answers, and continue drafting with the copilot extension.
How many tickets do you need before a reading counts?
None of this is worth anything if you read it on forty tickets, and a narrow scope is how I have seen pilots get there. At a rate near 70%, the margin of error at 95% confidence follows straight from the standard formula, and the arithmetic is unforgiving at small volumes.
In-scope tickets in the reading
Margin of error at 95%
50
±12.7%
100
±9.0%
200
±6.4%
300
±5.2%
400
±4.5%
500
±4.0%
1,000
±2.8%
Find your volume in the left column, then read what it costs you. A 300-ticket reading of 68% has a band of roughly 63% to 73%, which straddles the 70% median, so it cannot tell you whether you passed (and anyone who decides on it anyway is reading noise).
There is also a practical side that the arithmetic cannot give you, and we ask about it before anyone starts. Each reading should meet two conditions as well as the ticket count:
A full business cycle - the reading covers at least one peak day, or you have measured a quiet Tuesday and called it a month.
Every ticket type in the test - each type the AI is allowed to answer appears at least once.
If your volume cannot meet all of that inside thirty days, narrow the scope until it does.
How do I get AI to run the three readings for me?
Every calculation above uses numbers you already hold, so it is work you can hand to an AI assistant and read back in a couple of minutes. Fill in the bracketed parts of the prompt with four sets of figures from your helpdesk:
Ticket types - the ones the AI is allowed to answer.
Three rates - the resolution rate at each reading.
Three ticket counts - how many tickets each reading covered.
Four label counts - how many escalations fell under each of the four labels.
Paste the prompt below with those filled in, and it returns the rise, the projected day-30 rate, the margin of error, the decision and a short list of fixes.
You are helping me run a 30-day AI customer service pilot in which I take three readings of the AI's resolution rate.
My figures:
- Ticket types the AI is allowed to answer: [list them]
- Resolution rate around day 10: [rate]% on [n] tickets
- Resolution rate around day 20: [rate]% on [n] tickets
- Resolution rate at day 30: [rate]% on [n] tickets
- Escalations by reason: Knowledge [n], Data [n], Task [n], Out of scope by design [n]
Do this, in order:
1. Report how much the rate rose from day 10 to day 20, and say whether the agent
is turning the material it was taught into resolutions.
2. Project the day-30 rate forward: add the rise from day 20 to day 30 on top of the
day-30 rate. Project by exactly one step and no further.
3. Give the margin of error at 95% confidence on my smallest ticket count, and say
whether the band straddles 70%.
4. If the band straddles 70%, report that the reading cannot settle the decision, name
what would settle it, and stop. Otherwise compare the projected rate with these bands:
below 56% is a no, 56% to just under 70% is undecided, 70% or above is a yes, and
above 80% is a yes with a check that the ticket types are not drawn too narrowly.
5. Where the result is undecided, settle it on the mix of escalation reasons, excluding
Out of scope by design: a majority of Knowledge escalations points at content I can
write, a majority of Task escalations points at an integration project that will
take much longer.
Rules for you:
- If one of my figures is missing, or the counts do not add up, say so and stop. Write
"unverified, go and check" and never fill the gap with an assumption.
- Show your working for the projected rate and for the margin of error.
- Do not judge answer quality from these numbers. A person must read the conversations;
no arithmetic can stand in for it.
Output: a short table of the three readings and the projected rate, then a one-line
decision, then the two or three things I should fix fi
What do the readings look like in real rollouts?
⚡
TL;DR: Across three rollouts on Zendesk and Intercom, the starting rate and the settled rate were between four and sixty points apart, which is the gap a single day-30 reading would have been asked to cover.
That argument only holds if the rate at go-live and the settled rate really do differ, so here are three of our own rollouts.
A before-and-after comparison titled 'Where three pilots started and settled': RecruitCRM started at around 35% at go-live and settled at 68%; TravelJoy started at 76% and settled at 80%; Edel Optics started at 20 to 30% and settled at 75 to 79%.
Rollout
Helpdesk
Starting rate
Where it settled
RecruitCRM
Intercom
around 35% at go-live
68%
TravelJoy
Zendesk
76% on day one
80%
Edel Optics
Zendesk
20% to 30% before live data
75% to 79%
RecruitCRM: 68% AI resolution, up from around 35% at go-live
RecruitCRM run their support on Intercom, and handle roughly 1,088 conversations a month through the AI. Their go-live rate was around 35%, and a pass mark set at the field median and applied to the raw day-30 rate would have failed them somewhere in the low fifties. They now run at 68% with 75% AI CSAT, saving about 62 hours a month, and the AI resolved roughly 5,700 tickets in its first year.
The rise between the first two readings would have called this correctly, because a rate climbing out of the mid-thirties belongs to an agent that is learning from what it is taught. It is the first number we look at on any pilot.
TravelJoy: 80% AI resolution, up from 24% on the incumbent
TravelJoy handle around 2,500 to 2,700 tickets a month across Zendesk Ticketing and Messaging, saving 193 hours a month at 77% AI CSAT across the last twelve months. Their starting rate was unusually high, at 76% on day one, because most of the content work had been done before we arrived.
What makes TravelJoy the useful case is the comparison underneath it, because the AI agent they ran before us was reading 24%. A single day-30 reading would have passed the new agent here, and it would also have passed it on day one, which tells you the reading was doing no work.
Edel Optics: 75% to 79% AI resolution, up from 20% to 30%
Edel Optics handle 4,000 or more tickets a month on Zendesk, and now resolve 75% to 79% with AI, at 92% AI CSAT across 4,067 rated tickets. The AI has resolved more than 18,000 tickets since launch, saving around 150 hours a month. Before they connected live customer data, the same agent on the same content was reading 20% to 30%.
That jump is the clearest case we have of the reason behind the escalations deciding the result. Almost all of them came from the agent needing a live customer record it could not see, and adding the User Data API lifted resolution from the 20% to 30% band into the high seventies overnight.
A pilot produces one more thing, and I have seen teams throw it away. The three readings, the mix of escalation reasons and the ticket counts are exactly the inputs a return-on-investment calculation needs, so you are not starting the business case from scratch in week five. Deflected volume, hours saved and the escalation share all fall out of the same record. More rollout detail is in our published case studies.
What should you do this week?
⚡
TL;DR: Fix the scope, put three dated readings in the calendar with an owner on each, book one week of content work, and label every escalation as it happens: about two hours of setup in total.
Four actions, in order, and none of them needs a conversation with us or any other vendor first. The times below assume one person with admin access to the helpdesk, using AI to do the work.
Fix the scope before the clock starts - pick the ticket types the AI is allowed to answer, and write down the ones it is not, then check that the in-scope volume reaches the ticket counts in the table above over thirty days. Testing a few ticket types properly beats testing everything at once, because a wide scope leaves too few tickets of each type to read. The outcome is a written in-scope list and an expected ticket count per reading, and it takes about an hour.
Put the three readings in the calendar - dated entries for day 10, day 20 and day 30, each with a named person and a single field to fill in. This is what keeps the readings happening through a busy month (an undated reading is a reading nobody takes). The outcome is three calendar entries and one shared record, and it takes about fifteen minutes.
Book the week of content work - protect the time between the day-10 and day-20 readings now, because in our rollouts that week is where most of the improvement comes from. Plan half a day to a day for content, and one to three hours per data connection if you need the agent to read live records. The outcome is a booked block with a named owner, and booking it takes about ten minutes.
Label every escalation on the day it happens - one word per escalation, from the four labels I set out above, recorded by whoever picks the conversation up. The outcome is a record of why tickets were handed over that you can read at day 30 without reconstructing it, and it costs about ten minutes a day.
None of this requires a purchase to start: our trial at My AskAI runs for 30 days with every feature unlocked, unlimited tickets and no credit card, which is deliberately long enough to take all three readings inside it. If you are piloting somebody else's tool, check the trial length against your reading dates before you begin, because a fourteen-day window can only ever give you the first reading.
When does a 30-day pilot not settle the question?
⚡
TL;DR: Three readings do not help if the rate is suppressed by your escalation rules, if the volume is too thin to read, or if the thing that is not improving is your organization.
This method has real limits, and there are cases where I would not trust its result at all.
Start with the case that costs buyers the most: a resolution rate can be low entirely by design, and the pilot will read that as failure. One high-volume marketplace we work with on Zendesk records 21% AI resolution, because their escalation rules deliberately force a handover on around 77% of conversations. They still handle roughly 6,100 tickets a month, save about 105 hours, and hold 66% AI CSAT.
A live-events company we support reads 26% for the same reason, with most escalations driven by their handover and escalation guidance. They hold 85% AI CSAT, and pass around 555 of their 750 monthly conversations to a human with full context. Both are working as intended, and both would fail a pass mark applied to the raw rate.
Volume is the next limit, and the one I check first. Below roughly 165 in-scope conversations per reading, the margin of error is wider than the gap between passing and failing, so the readings cannot settle anything. Narrowing the scope raises the volume per type, which is usually the better trade.
A rate that barely moves can also be a fact about your organization. A team too busy to do the content work produces exactly the same stalled rate as a weak agent, and thirty days is often not long enough for me to tell them apart. The same goes for a team that leaves a long gap between the decision and the work.
The Task label matters for a reason the resolution rate will not show you. The AI may answer correctly but fail to update an order, issue a refund or write a CRM record, and only the escalation label will surface it.
Structurally, I am arguing against the orthodox position here. One framework for defining pilot success states it cleanly: the pilot's success criterion is checked once, at the end, and continuous monitoring only starts in production.
One last thing, on the failure statistics. You will find percentages for how many AI pilots fail quoted everywhere, and the ones in circulation disagree with each other and, in one case, with the study underneath the citation. I treat any such number as decoration until I can follow it to a method. The failure modes themselves are well described, even where the headline figures are not.
The by-design case is worth one practical note from our side, because it is fixable. Handover and escalation guidance in My AskAI is your configuration, so a low reading is a reason to check the rules before you blame the agent, and the conversations that do escalate arrive with the full context already attached.
The takeaway
⚡
TL;DR: Take three readings, base the decision on the day-30 rate projected forward, and check you have enough tickets before you let any of it decide anything.
A 30-day pilot can answer the buying question, but only if you stop asking it for a snapshot. The day-30 resolution rate on its own mostly reflects how much knowledge you loaded in month one. A field with a published median of 70%, and a range from 15% to 98.3%, cannot be navigated with a single figure.
Three readings is the smallest fix I know. Read the rate at day 10 to see where your content stands and again at day 20 to see whether the agent learns, then let only the day-30 rate, projected forward, make the decision. The reasons behind each escalation settle the close calls, and the ticket count decides whether any of it can be trusted at all.
If you do one thing this week, put the three dates in the calendar with a name against each. Everything else here is cheap once the readings exist. I have never managed to rebuild a day-10 reading after the fact.
When the pilot says yes, the question changes from whether the agent works to how you run it every day. We treat that as a longer horizon and a different plan: the operating cadence for the first ninety days.
FAQs
How long should an AI support pilot run before I decide?
Thirty days is enough if it produces three readings on enough tickets, and it is not enough if it produces one. The standard advice runs longer, commonly eight to twelve weeks for a single-category pilot including scoping, integration, testing and the live run. Structured trial plans, on the other hand, run as short as 14 or 28 days. The length falls out of how many readings you need and how much in-scope volume you have, so work backwards from those two things. We set our trial at 30 days for exactly that reason.
What should I measure during an AI customer service pilot?
Four things: the resolution rate at around day 10, how much it has risen by around day 20, the day-30 rate projected one step forward, and a one-word label on every escalation saying why it happened. The first two tell you what to fix, and only the third decides whether to buy. The escalation labels are what I use to turn an ambiguous rate into a decision, because missing content and a missing integration point at completely different amounts of work.
What resolution rate should an AI pilot hit to pass?
Judge the day-30 rate projected one step forward: below 56% is a no, 70% or above is a yes, and anything in between is decided by why the agent handed tickets over. Above 80% is still a pass, but check that the scope has not been drawn too narrowly. Those thresholds come from the published spread of rated deployments, which describes the field as a whole, is directional at best, and skews high because vendors publish their better deployments.
How big a ticket sample do I need to trust a pilot result?
The arithmetic I use is simple: around 400 in-scope tickets gets you to roughly ±5% at 95% confidence, and about 1,080 gets you to ±3%. Below roughly 165 the band is wider than the distance between passing and failing, so the reading cannot settle anything. Each reading should also cover a full business cycle with at least one peak day, and include every in-scope ticket type at least once.
What resolution rate should I expect from AI customer support?
The median is 70% across the published field, and our rate at My AskAI is 72%. Read both of those, and every figure in the spread below, as directional, because vendors count a resolution differently and publish their better deployments.
Measure
Across 195 rated deployments, 38 vendors
Median
70%
Mean
66.6%
25th percentile
56%
75th percentile
80%
Full range
15% to 98.3%
How to read it
The field as a whole, never head to head; directional, because the counting rules differ; skewed high, because vendors publish their wins
Which AI customer support tools have a free trial I can test before committing?
Trial terms vary more than the marketing does, so check the length against your reading dates before you sign up. A fourteen-day trial only leaves time for the first reading, so ask whether it can be extended or plan to pay for the rest of the month. Check too whether the trial caps tickets or features, because a capped trial cannot reach the volume the readings need. We set My AskAI's trial at 30 days with every feature unlocked, unlimited tickets and no credit card, which is long enough to take all three readings inside the window.
How to evaluate AI customer service tools: what should I look for?
Evaluate on the things a pilot can actually show you: whether the agent converts new material into resolutions, why it hands tickets over, and whether it can reach your live customer data and act on it. Price on total cost at the volume you actually handle. The fuller checklist, including the security and integration questions that fall outside a 30-day window, is a separate exercise worth doing before you start the trial.
Mike is an experienced Product Manager who focuses on all the “non-development” areas of My AskAI, from finance and customer success to product design, copywriting, testing and more.