Ai
Networks
Insights

How to measure chatbot containment rate honestly

How to measure chatbot containment rate honestly

Chatbot containment rate — the share of conversations resolved without a human — is the number most often quoted to justify a deployment, and the easiest one to quietly inflate.

It is not a dishonest metric. It is just a metric with several defensible definitions, and the most flattering one tends to be the one that gets reported.

Abandonment is not containment

The single biggest distortion: counting a conversation as contained because the customer left without reaching an agent.

Someone who gave up and phoned instead has not been contained. Someone who gave up and went to a competitor certainly has not. Yet by a naive definition — no human involved — both look like successes, and the number improves precisely as the experience gets worse.

Any honest definition has to separate three outcomes: resolved, escalated, and abandoned. Only the first is containment.

Check what happened next

A conversation that ends without escalation but is followed by the same customer calling within twenty-four hours was not resolved. It was deferred into a more expensive channel, and the containment number took credit for it.

This is the most useful single improvement to the measurement: join bot sessions to subsequent contacts across every channel, over a sensible window. The gap between raw containment and containment-with-no-follow-up is usually large, and it is the number worth managing.

Decide what counts as a conversation

Definitions that quietly inflate the denominator or shrink it:

  • Counting sessions where the user sent nothing. They are not conversations, and including them as contained inflates the rate substantially on high-traffic pages.
  • Counting a “what are your opening hours” exchange the same as a billing dispute. Both are one conversation; they are not equivalent work.
  • Excluding out-of-hours sessions, where there was no human available to escalate to. Containment is not meaningful when the alternative did not exist.

None of these are wrong to include, as long as they are stated. The failure is an unqualified percentage.

Segment by intent, or the average will mislead you

An aggregate containment rate blends trivially easy intents with genuinely hard ones, and the blend moves with traffic mix rather than with performance.

A bot can look like it improved from 55% to 68% because a marketing campaign drove a wave of simple status queries. Nothing about the system changed. Reported by intent, the picture is honest and actionable: this intent is well handled, this one escalates constantly and should be looked at.

Report it with three companions

Containment alone tells you how often the human was avoided, not whether that was good. Alongside it:

Resolution quality. A sample of contained conversations, read by a person, scored on whether the answer was actually correct. Automation fails quietly — plausible and wrong looks exactly like plausible and right in the logs.

Escalation experience. How long the customer spent with the bot before reaching a human, and whether context was passed. A high containment rate bought by making escalation difficult is a cost shifted onto customers.

Customer effort. How many turns to resolution. A contained conversation that took fourteen exchanges is not a good outcome.

Set the target against the right baseline

Vendor benchmarks are close to meaningless, because containment depends almost entirely on intent mix. A bot handling password resets and one handling insurance claims are not comparable at any number.

The baseline that matters is your own: what proportion of contacts were previously resolved by a first-line agent reading a script. That is the work genuinely available to automate, and it is usually a smaller share than the headline suggests.

Our Call Center Agent escalates with context rather than optimising for a containment number. See AI chatbots and virtual agents, or read integrating AI into human workflows.

Common questions

What is a good chatbot containment rate?

There is no portable answer, because containment depends almost entirely on intent mix. A bot handling password resets and one handling insurance claims are not comparable. Benchmark against your own baseline: what first-line agents previously resolved from a script.

How is containment rate calculated honestly?

Separate resolved, escalated and abandoned, and only count the first. Then join bot sessions to subsequent contacts across every channel — a conversation followed by a phone call within twenty-four hours was deferred, not resolved.

What should we report alongside containment?

Resolution quality from a human-read sample, escalation experience including whether context was passed, and customer effort measured in turns to resolution. Containment alone tells you the human was avoided, not whether that was good.

More from Insights

Keep reading.

Want this applied to your business?

Tell us what you are building. We respond within one business day, in English or Arabic.

Book an appointment

Let us find a time.

Tell us what you need and when suits you. We reply within one business day, in English or Arabic — across the United States, Saudi Arabia and South Africa.