The AI customer service performance gap is bigger than buyers think
Run an AI customer service quality comparison before you buy: some big-name vendors still fail on hallucinations and threaded email. Three tests catch them.
Mike is an experienced Product Manager who focuses on all the “non-development” areas of My AskAI, from finance and customer success to product design, copywriting, testing and more.
The performance gap between AI customer service vendors is far bigger than buyers assume. Big and well-funded doesn't mean a vendor has solved the basics. The conventional view is that the household names cleared hallucinations and threaded email years ago, so the brand on the box is a safe proxy for quality. Yet one customer watched a big-name product invent advice about caring for a dog, and another told us their helpdesk's AI couldn't reply past the first email in a thread. Across the 195 deployments we track, published AI-handling rates span sub-40% to 90%+. Test the basics before you trust the logo.
What everyone assumes
⚡
TL;DR: The market treats brand strength as a stand-in for answer quality, and the published buying guides encourage it. Accuracy almost never makes the comparison list.
I understand why buyers lean on brand as a proxy for quality. Funding, enterprise logos and analyst coverage usually track engineering depth, and nobody expects a category leader to fail a first-week test. We assumed the same thing ourselves.
The standard buying advice points the same way. One enterprise buyer's guide, to pick one example, tells companies to reduce risk by weighing vendor maturity, analyst recognition and large enterprise deployments. Zendesk's own AI agents explainer concedes that AI systems vary in quality, then moves on to its own pitch without giving you a way to check.
And buyers behave accordingly. The comparison questions we see people ask (of Google, of ChatGPT, of us on calls) are about features, price and integrations. Whether the answers will be right almost never comes up.
Why that's wrong
⚡
TL;DR: One customer's Gorgias AI Agent invented dog-care advice and another's Freddy AI dropped every email thread after one reply, both big, well-funded products. Published rates across the 195 rated deployments stretch from 15% to 98.3%.
We run an AI customer service product, so we see what happens when customers arrive from other tools.
One customer had tried Gorgias's AI Agent (the product formerly sold as Automate) before coming to us. The usual failure we see is a docs gap, where the customer asks something the help center never covered and the AI has nothing to work with.
This was a different class of failure. It answered off-topic and made things up, including advice on how to care for a dog.
And Gorgias is a big, well-funded business, with a $710M valuation at its 2022 Series C and roughly $104M raised to date. This is the kind of company the brand-as-proxy shortcut says you can trust with the basics.
Another customer told us their Freshdesk Freddy AI Agent couldn't respond past the first email in a thread. It answered the opening message, then escalated the moment a reply landed. We haven't run that test ourselves, so treat it as one customer's report.
Producing good answers in production is still a skill set a company has to learn or hire in, and most vendors only started learning when generative AI landed. Fin (formerly Intercom) is the clearest exception. Intercom had a team of data scientists and ML engineers primed before the boom, and it shows in the product.
The Best Helpdesk AI Agents of 2026 (Intercom Fin vs Zendesk AI vs HubSpot Breeze vs Gorgias AI)
The same gap shows up in the data. We maintain a corpus of published AI-handling rates covering 195 rated deployments across roughly 55 vendors.
The median sits at 70%, the mean at 66.6%, and individual published rates run from 15% to 98.3%. A quarter of deployments sit below 56%, and a quarter above 80%.
Spectrum chart of published AI-handling rates across 195 deployments: lowest 15%, lower quartile 56%, median 70%, upper quartile 80%, highest 98.3%.
Vendors define their metrics differently ("resolution", "automation" and "deflection" count different events, and the same work can read 12 points apart depending on the label).
Stat callout showing median published AI-handling rates by metric label: resolution 72.5% across 108 rows, deflection 70% across 17 rows, automation 61% across 52 rows.
Published figures are also marketing wins, so take them with a grain of salt and expect the true field average to sit below that median. The full breakdown is in our AI resolution rate benchmarks.
I shared the first cut of these numbers on LinkedIn. We hold ourselves to a high bar on hallucinations and multi-reply email threads, a bar we wrongly assumed the big names had cleared years ago.
Where this doesn't hold
⚡
TL;DR: Company size predicts little: Fin is excellent, and plenty of small vendors would fail the same checks. A vendor can also fix a failure mode within a quarter, so test at purchase time.
The gap doesn't run along company-size lines. Fin is proof that some big names are excellent, and some small vendors would flunk the basics. In our benchmark corpus, company size barely moves the published rates; industry and metric definition predict far more.
An anecdote is also a snapshot. A vendor can close a specific failure mode in a quarter. Either way, test at purchase time instead of trusting the brand or anyone's anecdote, including ours.
And for some buyers, platform breadth wins on the merits. If you're all-in on one suite, running one fewer tool can be worth more than a marginal edge in answer quality. I'd make that call knowing the size of the edge.
What this means for you
⚡
TL;DR: Run the trial against your own tickets: one off-topic question, one question your docs can't answer, and a reply to the agent's first email. Failing any of the three disqualifies the vendor, whatever the brand.
Stop using brand as a proxy and run the basics yourself, on a trial, against your own tickets. Three tests catch a hallucinating or thread-dropping agent before you sign:
Ask something off-topic, even a bit adversarial. A good agent stays in its lane; a primitive one improvises (that's how you end up with dog-care advice).
Ask something your docs don't answer. You want it to say it doesn't know, because an invented answer only gets more expensive once real customers see it.
Reply to its first email answer. Watch whether it holds the thread or hands straight to a human, the failure the Freddy customer described to us.
Run all three against your own ticket history rather than the vendor's demo, which is their best case.
We publish our own numbers so you can hold us to the same three tests. Our resolution rate is 72%, measured on a rolling 30-day window across our whole customer base, and the benchmark corpus it sits in is public. Our 30-day free trial has every feature unlocked and unlimited tickets, with no card needed.
A screenshot of My AskAI AI integrated into the Zendesk Ticket and Messaging platforms, directly replying to customer questions.
How do I build the three tests from my own tickets?
An LLM can draft the test set for you. Paste this into ChatGPT or Claude with the placeholders filled in. Judging the agent's live answers is still your job, on your own trial.
I'm evaluating an AI customer service agent on a trial. Build me a test set using the three-test protocol below, grounded in my real business.
My business: [one line on what you sell and who your customers are]
My help center: [URL, or paste your main articles]
Sample tickets: [paste 10-20 recent customer messages, anonymised]
Build:
1. Five off-topic or mildly adversarial questions a real customer of mine might plausibly send, designed to tempt the agent out of its lane.
2. Five questions that sound like my tickets but that my help center does not answer, so a good agent should admit it doesn't know. For each, note which missing article would be needed to answer it.
3. One realistic multi-reply email scenario: an opening customer email, then two follow-up replies that each add a complication, so I can check whether the agent holds the thread or escalates.
Format the output as a checklist I can work through on the trial, with a pass/fail line under each item describing what a good agent's response looks like. Where you can't tell whether my docs cover a question, mark it "check your docs" instead of guessing.
FAQs
Aren't the big AI customer service vendors safer to buy?
Sometimes. Fin (formerly Intercom) is a real exception, built on an ML team that predates the generative-AI boom.
But published AI-handling rates span 15% to 98.3% across the 195 deployments we track, and both failures we heard about came from big, well-funded products. A big brand proves the vendor has resources. Run the three tests to check the answers, mine included.
How do I test an AI support vendor before buying?
Run a trial against your own ticket history and score three things: whether an off-topic question keeps the agent in its lane, whether it admits the gap when your docs don't hold the answer, and how it handles a multi-reply email thread. Skip judging the canned demo. It's the vendor at their most rehearsed, and I've never watched one flunk its own demo.
Do all AI customer service agents hallucinate?
The risk exists everywhere, because it comes with the underlying technology. The rate is a vendor-quality variable. In our experience, grounding, retrieval and guardrail engineering are what move it, and that engineering is the skill set some vendors haven't built yet.
Assume the risk is present, then measure it. Ask the agent something your docs don't cover and watch what it does.
Mike is an experienced Product Manager who focuses on all the “non-development” areas of My AskAI, from finance and customer success to product design, copywriting, testing and more.