Ir al contenido
Comparta un flujo. Dring AI le llama en unos dos minutos y califica la necesidad. Solicitar llamada de la IA
Esta página está disponible en inglés por ahora. Ver la página en inglés
Call center metrics

Containment vs resolution rate in voice AI: what should you measure?

A call that does not transfer can still fail. Enterprise teams need a metric dictionary that describes what happened after the caller said goodbye.

OPERATING PLAYBOOKREVIEWABLE FLOW
Resolution over containment
01
IntentUnderstand what the caller needs
02
OutcomeMeasure whether the job actually finished
03
ReviewFix the gap by workflow, not by averages
FROM SIGNALShorter queues are useful only when the work is doneTO OWNED OUTCOME

For an operations or CX leader, voice AI reporting should make three decisions easier: what to automate, what to route and what to fix next. That requires more than one automation percentage. A call can end without a transfer and still leave work unfinished; a call can transfer and still be a good customer experience if the right person receives the right context.

Start with four different questions

Containment is an exit state

Containment usually means the call ended inside the automated flow without a live-agent transfer. It describes where the conversation ended, not whether the caller got what they needed. A caller may hang up after a confusing answer, abandon a slow authentication step or accept a promised callback that nobody owns. Each event can look contained unless the outcome model records what happened next.

Resolution is a completed outcome

Resolution asks whether the contact reason was completed and accepted. For an order-status call, that may mean a current status was retrieved, stated accurately and understood. For a cancellation request, it may require a confirmed system update, not just an explanation of the policy. Define the evidence before launch: a verified write-back, an explicit caller confirmation or an owned next action with a completed downstream state can count, depending on the workflow. A polite closing, a long call or a confident tone cannot prove resolution by itself.

Transfer quality and repeat contact expose the gap

Transfer rate tells you how often a call moved to a person. Transfer quality tells you whether that move worked: the caller reached the appropriate queue, the transfer succeeded, the relevant details were passed through and the person did not have to restart the conversation. A transfer with no context is technically complete but operationally expensive. Also measure what happened after the transfer: acceptance by the receiving queue, time to human action, correction work and whether the promised next step was completed.

Repeat contact measures whether the same customer, household, case or transaction returns with the same reason within a defined window. Use a case-based denominator where possible, and choose a window that matches the workflow. A delivery exception may recur quickly; a billing dispute may need a longer window. Repeat contact is not proof that the first call failed, but it is a strong review signal when paired with the original transcript, system outcome and reason for the later contact.

Write a metric dictionary before you build the dashboard

Put the definitions in a shared document that operations, support, quality and whoever owns the integration can all approve. A rate is only meaningful when its population, numerator, denominator, evidence and exclusions are visible. Use the following as a starting point, then adapt the evidence rule to the workflow.

MetricNumeratorDenominator and evidence
Containment rateEligible calls that ended in the automated flow without a live transfer.All eligible calls entering the measured workflow. Expose abandonment, technical failure and unknown outcome separately; do not silently count them as successes.
Resolution rateCalls with a verified terminal business outcome, such as a matching write-back, accepted answer or completed owned action.All eligible in-scope calls, unless a second diagnostic rate is clearly labeled. A closing phrase, short duration or no transfer is not evidence.
Transfer qualityReviewed transfers with the correct destination, reason, context and usable next action.Reviewed transfers, with the sample size shown. Also report the share of calls that needed a transfer and the receiving queue's acceptance or completion result.
Repeat contact rateUnique customers or cases with another related contact inside the agreed window.Unique customers or cases in the original cohort. Document the identity match, related-intent rule and observation window, and mark recent cohorts as not yet fully observable.
Task completion rateCallbacks, tickets or reconciliation tasks closed with the required outcome by the due time.Tasks due in the reporting window. Pending, rejected, overdue and unowned tasks stay visible instead of entering the success numerator.
Unknown outcome rateCalls for which available signals cannot support resolved, unresolved or transferred classification.All eligible calls. Treat missing instrumentation as a data-quality issue with an owner, not as a neutral success state.

Keep at least two views of important metrics: a headline rate for the defined eligible population and a diagnostic rate for a narrower stage, such as answered calls or valid action requests. Never compare those rates without naming the denominator. Dring's analytics view can be useful here because outcome, confidence, reason and next action can be examined together rather than collapsed into one score.

Find false containment in the outcome chain

False containment is a call that appears automated and complete but creates no satisfactory customer or business outcome. Common patterns include a caller hanging up during a confusing explanation, an answer that is accurate but does not address the actual reason for contact, an integration returning a stale record, a requested change being spoken but not written, or a callback promise without a named owner and due time. A caller who stops asking questions is not necessarily a caller who received help.

To detect it, model the conversation as a chain rather than a single terminal event: contact reason identified, identity or eligibility checked, information retrieved, decision or action taken, result confirmed, record updated and next action owned. Capture a status for each stage. A failure before the final evidence should be classified by cause, such as abandonment, misunderstanding, policy boundary, tool failure, data mismatch, transfer failure or missing follow-up. This gives a supervisor something to fix and keeps a high containment rate from hiding operational work.

Pair automated labels with a review sample. Reviewers should be able to compare the caller's stated need, the agent's answer, the downstream record and any later contact. When the agent says "resolved" but the system record is unchanged, count the workflow as unresolved or unknown according to the agreed rule and open a data or process issue. Do not overwrite the event to make the dashboard tidy.

Establish a baseline and segment it

Before changing prompts, tools or routing, capture the current process. Define the baseline cohort first: the intents included, the channels and hours covered, the identity or case key used for matching, the start and end dates, and the exclusions. Separate attempted calls, connected calls and calls that entered the automated workflow. Tests, spam, duplicates and known instrumentation checks can be excluded when documented. Outages and queue closures should remain visible as operational conditions, even if they are reported in a separate availability view.

Use a lookback long enough to cover normal variation across workdays, busy periods, holidays, languages and the people or systems that receive escalations. Record the existing disposition, time to completion, transfer destination, manual work, repeat contact, correction effort and downstream write-back. If the current team does not have reliable resolution data, mark that limitation and estimate the missingness. A weak baseline is still more useful than a precise-looking guess, but a pilot should improve the instrumentation before it claims a business improvement.

Segment the baseline and every pilot report by intent, complexity, language, time of day, customer type, authentication path, integration path, queue and handoff reason. Add a segment for caller-requested human help and for calls that fall outside the approved scope. A blended containment rate can rise while an important intent gets worse. For example, a support flow may answer a simple policy question accurately while struggling with an exception that needs an order lookup. Those should be separate operating decisions, not one average.

Keep the cohort definition stable while comparing baseline and pilot, or explain the change. Show counts beside percentages and suppress confident interpretation when a segment has too little observed volume for a useful comparison. Repeat contact and task completion need a lag: the latest calls may not have had time to generate a second contact or reach a due date. Label those cohorts as immature rather than treating them as zero.

Build a dashboard for decisions, not applause

An executive dashboard should answer four questions in order: how much traffic is in scope, what outcome did the caller receive, what work did the operation inherit and where is risk increasing? Put volume, eligible population and data coverage near the top. Show containment beside resolution, false-containment review findings and unknown outcomes. Then show transfer quality, repeat contact, time to human action, correction effort and task completion. A single blended score can be useful for a local decision, but it should never replace these component measures.

Give leaders a trend and a drill-down. The trend should compare the same definitions by cohort and show the pilot or release boundary. The drill-down should move from intent and segment to reason code, call, system event and owner. Display the count behind each rate, the reviewed sample size for quality metrics, and the number of records still waiting for a downstream outcome. Use the dashboard to choose an action: update a source, change a tool guard, retrain a queue, narrow scope, add a test or pause traffic.

Keep a small exception view visible: unresolved calls, false containment, failed write-backs, unowned callbacks, incorrect destinations, repeated authentication, opt-outs and critical policy or data-integrity errors. These are not merely negative annotations. They are the queue of work that connects measurement to operating improvement.

Make human handoff part of the product

Decide the handoff contract before the agent takes traffic. The agent should know when to transfer, which queue owns the case, what information must be collected, whether the caller has authenticated and what the caller should expect next. Send a concise summary containing the reason, relevant identifiers, actions already taken, uncertainty and promised follow-up. Tell the caller what is happening, preserve their place where the operation allows it and provide a callback path when no suitable person is available.

Measure handoff quality as an outcome, not just as a transport event. A receiving employee should be able to tell whether the transfer is appropriate, whether the summary is accurate, whether the caller has to repeat discovery and whether the next action is clear. Track failed transfers, wrong queues, dropped context, repeat authentication, time to first human action and downstream completion. The appropriate denominator depends on the question: use transfers for context quality, calls requiring escalation for routing coverage and tasks due for callback completion. State which one you chose.

Test the failure states as carefully as the happy path: the target queue is closed, the transfer fails, an integration is unavailable, the caller refuses authentication, the caller requests a person immediately or the agent is unsure. A safe fallback may be a clearly owned callback task, a limited answer with a next step or a direct escalation. The platform workflow should make those states visible to the team that owns the work.

Use staged implementation with explicit pilot gates

1. Scope one measurable workflow

Choose a narrow, repeatable reason for contact with a known source of truth and a clear human boundary. Write the resolution evidence, exclusions, caller choice, handoff destination and fallback in plain language. Name the operations owner, quality reviewer, data owner and person who can pause the route.

2. Observe before optimising

Run the baseline review, listen to representative calls and label failure reasons. Agree on the dictionary and review rubric before looking at a new headline number. Reconcile conversation outcomes with the system record so the team knows which signals are reliable.

3. Launch with constrained coverage

Start with the intents, languages and hours the team can support. Keep unknown and unresolved outcomes visible, and give supervisors a way to inspect individual calls and downstream events. Make the pilot boundary easy to reverse and tell the receiving queues how to report a bad handoff.

4. Set continue, expand and pause gates

Set the gates before live traffic, with local thresholds rather than retrofitting them to the result. Continue only when critical policy, identity, consent, data-integrity and handoff checks are passing and the exception queue is owned. Expand only when resolution, transfer quality, task completion and repeat-contact signals are acceptable for each important segment across a mature observation window. Pause or roll back when a critical error appears, write-backs cannot be reconciled, the receiving queue is unavailable, unresolved work is unowned or an agreed guardrail deteriorates. Keep the decision and evidence in a release record.

5. Operate an improvement loop

Every review should end with a decision: change the prompt, add a tool guard, update policy, adjust routing, improve the handoff summary, fix instrumentation or leave the behavior alone. Dring's Agent Factory improvement loop connects reviewed calls to regression tests, controlled changes and post-release monitoring. For useful release feedback, keep the original call, failure label, expected outcome, affected segment, changed layer, test result and post-release observation together. An unexpected repeat contact or correction should become a new scenario or regression case when it represents a repeatable failure.

Release feedback should also check for trade-offs. A change that reduces transfers may increase repeat contact; a stricter identity step may improve data integrity while increasing abandonment; a new fallback may reduce unknown outcomes while increasing queue work. Compare the intended metric with the quality and workload measures that could move in the opposite direction. The quality harness should check the change against known edge cases before more callers encounter it, and the release owner should record whether the live result confirms, limits or rejects the hypothesis.

Assign governance to the metric

Give each metric a named owner and a named data source. Operations owns the workflow definition and decision; quality owns the review rubric and sampled evidence; the integration or systems owner owns event completeness and write-back reconciliation; the queue owner owns handoff and callback completion. A sponsor can decide scope, but should not be the only person who can explain a denominator.

Version the dictionary when the workflow, event schema, identity rule or evidence standard changes. Keep a short decision log with the release date, scope, segments reviewed, exceptions, gate result and follow-up owner. Review headline metrics on the operating cadence and review definitions when people start interpreting a label differently. Governance is working when a leader can ask why a rate moved and the team can trace the answer to calls, system outcomes and a recorded change.

Executive review checklist

  • Is each metric defined with an owner, numerator, denominator, evidence rule and version?
  • Can the system distinguish containment, verified resolution, false containment, transfer quality, repeat contact, task completion and unknown outcome?
  • Have you recorded a baseline with a stable cohort, explicit exclusions, counts and data-quality limits?
  • Are the important results segmented by intent, complexity, language, time, customer type, integration and handoff reason?
  • Does every escalation have a destination, context payload, fallback, due time and human owner?
  • Are unresolved, abandoned, identity, policy, integration-failure and data-write paths included in the test set?
  • Were continue, expand and pause gates agreed before the pilot, with a reversible route?
  • Does each release turn reviewed call evidence into a tested change, a release record and post-release feedback?

Measure what the caller actually got

Request a callback to define resolution for your first workflow.