Most AI support programs miss their resolution-rate target because the knowledge behind the AI is thin: docs that lag the last release, edge cases nobody's written up, and answers that only ever made it into a Slack thread. Qualtrics' 2026 report found AI customer service fails at 4x the rate of other AI use cases, and nearly 1 in 5 consumers who tried AI for support saw no benefit at all. Forrester's 2024 CX Index adds useful context: overall CX effectiveness fell to 64%, its lowest recorded level in three years running. Brainfish customer Smokeball runs an 83% self-serve rate with an AI agent grounded in its real product knowledge, running alongside Zendesk. CareMaster took self-serve from 30% to 76%, a 2.5x lift, with no team growth. What explains the gap is simple: how much real, current knowledge each AI actually has to work with.
How is AI resolution rate calculated?
The formula itself is simple: conversations the AI resolves without escalation, divided by total conversations the AI handles, times 100. A team running 10,000 AI-touched conversations a month with 9,200 closed hands-free lands at a 92% resolution rate.
The math is the easy part. What actually counts as "resolved" is where vendors start to disagree. Some vendors count a conversation as resolved the moment the AI stops replying, whether the answer was correct or not. Others only count it resolved once the customer confirms the issue is closed. If you're comparing two vendors' numbers, ask both how they define and verify "resolved" first.
What's a good AI resolution rate?
Honestly, there isn't a single industry benchmark worth pointing to, because resolution rate is only comparable when the definition of "resolved" and the ticket mix match. Worth noting: Brainfish reports these two as self-serve rate, not strictly resolution rate. Smokeball runs at 83%, and CareMaster went from 30% to 76%, a 2.5x lift. Both come from an AI agent grounded in real product knowledge.
If a vendor hands you an impressive-looking number, ask the same two questions you'd ask about your own: is it measured against real production traffic, and does "resolved" mean verified-correct, or just that the AI stopped talking?
How is AI resolution rate different from deflection rate, containment rate, and FCR?
These four terms get used interchangeably in vendor decks, and they really shouldn't be.
| Metric | What it measures | Formula | Where it can mislead |
|---|---|---|---|
| AI resolution rate | % of AI-handled conversations closed with no human handoff | AI-resolved ÷ AI-handled × 100 | Can count "AI stopped replying" as resolved unless the vendor verifies the answer |
| Deflection rate | % of total support volume that never reaches a human agent | Deflected ÷ total incoming × 100 | Can count a customer giving up as a win |
| Containment rate | % of conversations that stay inside the AI channel without transferring | Contained ÷ total in-channel × 100 | Says nothing about whether the issue was actually solved |
| First contact resolution (FCR) | % of tickets closed in a single interaction, human or AI | Resolved on first contact ÷ total tickets × 100 | Doesn't distinguish AI-resolved from human-resolved |
Of the four, the resolution rate is the narrowest. It's the only one of the four that has to prove two things at once: that the problem actually got solved, and that the AI solved it alone.
Why do AI resolution rates vary so widely between tools?
Because the AI is only ever as good as what it's allowed to read. Intercom's own implementation guide for Fin tells customers to run weekly content reviews, keep articles narrowly scoped, and maintain consistent terminology just to keep answers accurate. That's the incumbent, in its own documentation, admitting the knowledge feeding the AI decides the resolution rate.
Point an AI agent at knowledge that's thin, stale, or scattered across five tools, and it will guess or punt to a human the moment a question gets even slightly hard. Swap in structured, always-current knowledge for that same model, and the resolution rate climbs on its own, proof the knowledge was doing the heavy lifting the whole time. That's the specific layer Brainfish runs: an AI agent grounded in a team's real docs, tickets, and transcripts, deployed alongside the helpdesk they already have.
How do you improve AI resolution rate?
Three things actually move the number, roughly in order of impact.
Start with the knowledge itself. Prompt tweaks rarely move the number much; fresh, well-structured source material does. Pull from every place product truth lives: docs, tickets, transcripts, release notes, instead of a quarterly content sprint that's stale by the time it ships. Then structure that content for retrieval, not just reading: deduplicated, consistently worded, mapped back to a source. A well-written article a person enjoys isn't automatically something an AI can query cleanly. Finally, close the loop on what the AI can't answer, and turn that gap into new content, or the same failure repeats every week. Brainfish automates all three for teams running Zendesk, Intercom, or Salesforce: continuous ingestion, structured retrieval, and a closed loop on unanswered questions.
What should you look for in a tool that reports AI resolution rate?
Not every "resolution rate" is measured the same way. Before you trust one, ask the vendor how "resolved" is defined and verified, whether the number comes from real production traffic or a curated demo set, and whether it separately tracks deflection and containment, or quietly does the work of all three.
Brainfish CEO Daniel Kimber has made the same case about deflection, and it holds just as well for resolution rate:
"90% ticket deflection!" sounds impressive, until you realize it's missing the point entirely. In our race to reduce support tickets with AI, we've forgotten what actually makes customer support valuable in B2B, helping users get important work done.
Daniel Kimber, CEO, Brainfish, in Beyond Deflection