Resolution rate is the number every AI chat vendor puts on the first slide. On its own it tells you almost nothing about whether customers are being helped, and it can rise for reasons you would hate.
Ask most AI chat platforms how they define a resolved conversation and the answer is some version of this: the customer engaged with the bot and did not request a human agent.
Read that again, because it covers several situations that are not resolutions at all.
The customer got a correct, useful answer and left happy. That is a resolution. The customer got a vague answer, assumed the company could not help, and gave up. That counts too. The customer got the wrong answer, believed it, and will contact you again in four days when the consequence arrives. Also counted. The customer got frustrated, closed the window, and went to write a public review instead. Counted as a resolution.
A rate built on the absence of an escalation request measures how successfully the bot avoided a handoff. Whether the customer's problem went away is a separate question that most dashboards never ask.
Silent abandonment. The customer disengages without escalating. The bot records success. You can find these by looking at conversation length and end state: short sessions that end without a clear answer being given, especially on intents the bot is not confident on, are abandonments wearing a resolution badge.
Confident wrong answers. The most expensive failure. The bot answers a policy or status question incorrectly, in a tone that suggests certainty, and the customer acts on it. The cost shows up later as a dispute, a refund, or a complaint that now takes three times as long to handle because the customer was told something different by your own system.
Handoff churn. The bot collects information, fails, and passes the conversation to a human who cannot see what was already asked. The customer repeats themselves. Handle time on the human side goes up, and the automation has added cost rather than removed it. This one is easy to miss because both systems report their own numbers as fine.
Four measurements turn resolution rate from a vanity number into something you can manage.
Satisfaction on automated conversations only. Not the blended score for the whole department, which the human team's performance will mask. Score the contacts the bot handled alone, and compare them against the equivalent intents when a human handles them. If the automated score is materially worse on the same intent, the automation is trading service for cost.
Repeat contact within seven days. Take every conversation marked resolved by automation and check whether that customer came back within a week on a related issue. This single number exposes both silent abandonment and confident wrong answers. A high repeat rate means the resolutions are not real, and your true automation rate is much lower than reported.
Handoff rate by intent, not in aggregate. An overall handoff rate hides the shape of the problem. Broken down by intent, it tells you exactly which topics the bot should own, which it should stop attempting, and where the next improvement is worth the effort.
Attribution in negative feedback. Read the verbatim comments on poor satisfaction scores and classify how many mention the automated experience. It is one of the fastest signals available, it needs no integration work, and it is usually the number that changes the conversation internally. When a meaningful share of your worst feedback names the bot, expanding its scope is the wrong next move.
Once you can see quality by intent, the decisions become obvious rather than political.
Expand where the bot is beating the human baseline. Usually order status, delivery timing, returns initiation, and simple account changes. Here the automated answer is faster, available overnight, and consistent. Customers prefer it, and the satisfaction data will say so.
Restrict where emotion or money is involved. Complaints, damaged goods, refund disputes, anything involving a failed delivery of something time-sensitive. These need a human early, and the best thing automation can do is recognise them fast and route them well rather than attempt a resolution.
Fix the handoff before you widen the scope. If a conversation escalates, everything the bot collected has to arrive with it, visible to the agent, with no repetition asked of the customer. Getting this right typically improves both cost and satisfaction more than any expansion of bot coverage.
The uncomfortable version of this analysis is that the honest automation rate is often well below the reported one. That is a better position to start from. A bot that genuinely resolves a quarter of contacts well is worth more than one that claims half and generates rework, disputes, and public reviews you then pay to answer.
None of this is an argument against automating customer service. Automation is the single largest lever available in most retail support operations, and the intents it handles well it handles better than people do, at any hour, without a queue.
The argument is against measuring it with one number that cannot fail. Put a quality metric beside every efficiency metric, break it down by intent, and read what customers write. You will end up automating less than the vendor promised and getting more from it, with a cost curve that keeps improving instead of quietly rebounding two quarters later.
We analyse automated conversations against repeat contacts, satisfaction by intent, and handoff quality, then tell you where to expand and where to stop.