Why "what model do you use?" is the wrong question to ask an AI support vendor
A real AI customer service model is 10+ orchestrated, A/B-tested LLM calls. A vendor who answers "what model do you use?" in one word is waving a red flag.
Mike is an experienced Product Manager who focuses on all the “non-development” areas of My AskAI, from finance and customer success to product design, copywriting, testing and more.
If an AI support vendor can answer "what model do you use?" in one word, treat it as a red flag instead of a spec. A real AI customer service agent fires 10+ orchestrated LLM calls between question and answer, each doing a different job, each under a continuous A/B test. A vendor who can name a single model is describing a single-prompt setup, and that's where the bad answers come from.
What everyone assumes
⚡
TL;DR: The standing assumption is that the LLM behind an AI support agent is the one spec worth comparing. The question stands in for what a buyer needs to know: whether the vendor is current, capable and safe.
It's a reasonable question, and I understand the instinct. Models do differ, every frontier release gets its hype cycle, and when three vendors all promise the same resolution rates, "which LLM is this built on?" feels like the one spec you can compare.
A whole content industry feeds it: search "best LLM for customer service" and you'll find a stack of rankings matching models to use cases.
And if you run a support team, you need something concrete to take to your Director. Telling them the vendor runs the latest model sounds like due diligence. The question is a proxy for "are you current, capable and safe?", which is what I'd want to know too, and those answers sit a layer below the model name.
Why that's wrong
⚡
TL;DR: Choosing the model is work the vendor should own, with cross-customer visibility on what each release does to resolution and accuracy. A production agent also runs upwards of 10 tuned, A/B-tested calls per answer; Fin publishes seven purpose-built models working the same way.
It's wrong on two levels.
The first is that picking the model is the vendor's job, and I'd argue you should want it to be. Models change all the time. Even a small update to the one you're on can shift behavior overnight.
Only the vendor can see, across all of their customers, how each model moves resolution rate, accuracy and escalation. That visibility is a big part of what you're paying us for.
The second is the premise that an AI support agent is one model with a prompt in front of it. Between question and answer, our agent fires upwards of 10 orchestrated LLM calls, each doing a different job: one rewrites and summarizes the conversation, one detects the language, one applies your guidance and tone, one decides whether a task or process should run.
Control Your AI Agent's Answers
We tune one hard to follow instructions and avoid made-up facts, because it writes the reply your customer reads. The language-detection call gets a smaller, faster model; spotting that a message is in German is a simple job.
Process flow showing five of the ten-plus orchestrated LLM calls a real AI support agent fires between a customer question and the answer: rewrite and summarise, detect language, apply guidance and tone, decide if a task runs, write the reply.
We run continuous A/B tests on each of those steps with a suite of models: OpenAI for most of them, with Google and Anthropic on the tasks where they do better. We keep reasoning models out of the live path; they're too slow for a customer waiting in chat. The current mix is in our LLM docs.
Fin (formerly Intercom), the one to beat in this category, describes its own agent the same way on its CX Models page:
"The Fin CX Model Suite is a system of seven purpose-built AI models, each handling a discrete stage of resolving a customer support query: detecting language, summarizing the issue, retrieving knowledge, reranking results, generating the final answer, parsing feedback, and routing escalations."
To answer "what model do you use?" in full, a vendor would have to walk you through that entire structure. A vendor who can answer in one word is telling you the setup is question in, one prompt, answer out, and that setup won't produce a high-quality reply.
Stat callout comparing one model named by a single-prompt setup, the seven purpose-built models Fin publishes, and the ten-plus orchestrated LLM calls behind one My AskAI answer.
TL;DR: Three model questions still earn a place on a vendor call: whether customer data trains models, whether a degrading model can be swapped, and how regressions get caught after an update. Specific answers to all three point to a real multi-model setup.
The worry underneath the question is fair. If you suspect a vendor is running a cheap, dated model to pad their margin, that failure is real; the question is a clumsy way to catch it. Some model questions still belong on the call:
Is our data used to train models? A yes should end the call. We never train models on your data, and we hold SOC 2 Type II certification and GDPR compliance (the live trust report has the detail).
Can you switch models when one degrades?
How do you catch a regression after a model update?
A vendor running a real multi-model setup will answer all three in specifics. In my experience, a one-model vendor can't.
What this means for you
⚡
TL;DR: Take five questions into a demo: resolution rate on your kind of tickets, multi-reply email handling, who catches a model regression first, pipeline A/B-test cadence, and whether customer data feeds training. Then verify on your own tickets with a side-by-side trial.
When a model update degrades answers, who notices first, you or us?
How often do you A/B test the steps in your pipeline?
Is our data training anything?
On the third one, our answer is us. We never stop benchmarking the mix, and your team can ask Echo, the assistant in our dashboard, why the agent gave any answer and which knowledge source it used.
Then test the basics on your own tickets. The gap between vendors is wider than the marketing suggests.
One customer told us their previous AI, Gorgias Automate, once answered a support question with made-up advice about caring for a dog. Another told us Freshdesk Freddy would answer the first email in a thread, then escalate the moment the customer replied.
These are big, well-funded products. Running AI support well is still a skill set a company has to learn or hire in.
The cheapest way to find out is our Internal Notes mode. It drafts a reply on every ticket as a private note, your customers see nothing, and you compare it side by side with whatever you run today. The 30-day free trial has every feature unlocked, unlimited tickets, and needs no card.
A screenshot of the My AskAI agent within Intercom replying in notes mode to a user question.
FAQs
Should I ask an AI support vendor what model they use?
You can, but I'd read the answer as a signal about the setup behind it. A vendor who explains how they orchestrate, test and swap models is showing you a real pipeline. A one-word answer describes a single prompt.
Their resolution rate on tickets like yours tells you far more.
Does it matter if an AI support agent uses GPT-5 or Claude?
Less than you'd think. No model is best at every step, so a production agent (ours included) mixes small fast models for jobs like language detection with larger ones tuned for the final reply.
Models also change under you; only the vendor sees how each release moves resolution and accuracy across its whole customer base. Pin the vendor down on who owns the swap when a model degrades.
Why do AI support agents use more than one model?
Because each job in the pipeline gets its own tuning. Language detection needs raw speed, and the customer-facing answer needs instruction-following plus a strong resistance to made-up facts.
Fin (formerly Intercom) publishes that its agent runs a system of seven purpose-built models. Ours fires 10+ A/B-tested calls on every answer.
What questions should I ask an AI customer service vendor instead?
Ask for the scores on the doors first: their resolution rate on tickets like yours. Then ask how they handle multi-reply threads, who notices a model regression first, how often they A/B test their pipeline steps, and whether your data trains anything. Then run a trial on your own tickets and check the basics yourself.
Mike is an experienced Product Manager who focuses on all the “non-development” areas of My AskAI, from finance and customer success to product design, copywriting, testing and more.