Ir al contenido
Comparta un flujo. Dring AI le llama en unos dos minutos y califica la necesidad. Solicitar llamada de la IA
Esta página está disponible en inglés por ahora. Ver la página en inglés
Finance and operations

Voice AI ROI: how to model cost per resolved call

An automation rate can look impressive while the team still does the expensive work. Cost per resolved call gives finance and operations a shared way to evaluate the result.

OPERATING PLAYBOOKREVIEWABLE FLOW
Cost per resolved call
01
BaselineCount the full cost of today's workflow
02
CompareSeparate answered from actually resolved
03
MeasureTrack the improvement by intent
FROM SIGNALA business case built around completed workTO OWNED OUTCOME

Why cost per resolved call is the useful denominator

Cost per minute answers a narrow question: what does it cost to keep a conversation running? It helps with capacity planning, but rewards shorter calls even when the caller repeats the request or a colleague finishes the work. AI can lower talk time while increasing transfers, review queues and callbacks.

Containment has a different weakness. It says whether a call ended without a human, not whether the customer got the right result. A bot can contain a call with an incomplete answer, wrong information or a delayed repeat contact. Cost per resolved call asks the useful question: what did it cost to produce a verified outcome?

Use "resolved" for an outcome checked in the business system: a correctly updated account, completed payment arrangement, qualified handoff with context, or support request closed under policy. Allow a good human handoff when automation is not the right endpoint, but do not count silence, abandonment or an untracked transfer as success.

The metric is a unit-economics lens, not a complete business case. It does not by itself prove that customers prefer the experience, that revenue increased, that risk fell, or that a team can be resized. It also does not compare unlike work fairly: a simple status call and a regulated exception may both end in a CRM disposition while requiring different controls. Pair the metric with service quality, customer effort, compliance review, capacity released and the outcome the workflow is meant to create.

Keep two denominators visible. Handled calls describes the workload that arrived. Verified resolved calls describes the workload that met the agreed outcome and quality gate. Reporting only the latter can make a small resolved subset look efficient, so show the resolution rate beside the cost metric and explain every excluded call.

Define the cost components before calculating savings

List every resource required to move a call from arrival to resolution. Keep definitions stable across the current process and pilot.

Labor and follow-up

Labor cost is the fully loaded cost of the people doing the work, not only talk-time wages. Multiply productive minutes by an hourly rate including benefits, payroll costs and relevant overhead. Count queue handling, account reading, authentication, after-call work, quality review and supervisor intervention. If roles have materially different rates, keep them as separate rows instead of hiding a specialist handoff inside one blended rate. Follow-up cost covers the email, callback, ticket update or internal task created afterward, including specialist research.

Use observed minutes when they exist, and label estimates when they do not. A useful split is: caller-facing minutes, post-call minutes, review minutes and exception minutes. A voice AI interaction can reduce the first line while increasing the other three. For internal labor, use scheduled productive capacity or an agreed loaded rate consistently; do not treat a vacant hour as zero cost unless finance explicitly wants a short-run marginal view.

Platform, telephony and transfer

Platform cost includes voice AI runtime, orchestration, logging, storage, monitoring and connected tools used per interaction. Telephony cost includes inbound or outbound minutes, numbers, carrier fees and recording charges. Transfer cost is the extra call leg plus the receiving human's time, including correction of missing context. A transfer that saves no human effort is not free.

Make transfer treatment auditable. Record the transfer rate, the receiving role, warm-transfer minutes, hold time, re-authentication and any failed transfer that returns to a queue. If the AI passes a summary and verified fields, credit the human time actually avoided rather than assuming the handoff is equivalent to a fully automated resolution. If the human repeats the entire discovery step, count that work in full.

Repeat contact, unresolved work and implementation

Repeat-contact cost is the expected cost of another call, message or ticket when the first interaction did not stick. Link repeats where possible; use a documented estimate when identity matching is imperfect. Unresolved work is remaining effort on abandoned, incorrect or incomplete calls, including queue work with no owner. Do not hide it by excluding those calls from the denominator.

Choose a repeat window before looking at results, such as the same intent within a defined number of days, and apply it to both baseline and pilot. Count the follow-on interaction's labor, platform and telephony costs, not just its talk time. For unresolved calls, estimate expected completion work from sampled cases or track the task to closure. A conservative model can carry unresolved work as a separate reserve until the outcome is known; that prevents early pilot data from looking artificially cheap.

Implementation cost includes discovery, integration, testing, policy review, training, launch support and maintenance. Allocate it over a stated period or expected resolved-call volume. A monthly model can show it separately; an annual model can amortize it. Make the choice visible.

Build a transparent baseline and worksheet

Start with a representative four- to six-week period. Export call volume, intent, outcome, handling time, transfer, repeat contact and follow-up fields, then reconcile totals with the phone system and CRM. Write the model so another person can reproduce it:

Cost per resolved call = (labor + platform + telephony + transfer + repeat contact + follow-up + unresolved work + implementation allocation) / verified resolved calls.

For the baseline, platform and implementation may be zero or may represent existing tools. For AI, replace each line with observed or estimated values. Keep rows for volume, resolved calls, human minutes, AI minutes, transfer minutes, repeats, follow-up minutes and quality failures. Add notes for the labor rate, repeat-contact rule and allocation period.

Worksheet inputBaseline and pilot valueSource or treatment
Handled callsCount by intent and periodPhone system; reconcile with CRM arrivals
Verified resolved callsCount after outcome and quality gateCRM disposition plus sample or control check
Labor minutesTalk, queue, after-call, review and exceptionsWorkforce data, time study or sampled estimate
AI and telephony costUsage, recording and carrier chargesInvoices or usage export; allocate shared tools
Transfer and follow-upReceiving minutes, callbacks, tickets and researchTransfer logs and linked tasks; include rework
Repeat and unresolved reserveExpected downstream interactions or open workPredefined repeat window and sampled completion cost
Quality remediationCorrections, escalations or policy reviewQuality sample and incident records
Implementation allocationPeriod or resolved-call allocationApproved finance assumption; show separately

For each row, record the unit, owner, date range, source system and confidence level. This turns the formula into a reviewable worksheet rather than a single spreadsheet output. Show the numerator in currency and minutes, and show the denominator, resolution rate and quality pass rate beside it. When an input is missing, mark it as unknown and run a sensitivity case; do not silently substitute a favorable value.

Compare AI with the same formula, not a headline automation percentage. Show low, base and high cases for uncertain inputs such as review rate, extra transfer handling and repeat-contact probability. A useful model lets a finance partner change one assumption and see which conclusion moves.

Before approving a result, finance should be able to answer: Which calls are included? What event makes a call resolved? Which costs are incremental now, and which are allocated for a full-cost view? Are baseline and pilot using the same labor rate, repeat window and quality rule? Can the team reproduce the totals from source exports? These questions are more important than adding decimal places to the final ratio.

Price resolution quality and human handoff honestly

A resolved call must pass a quality gate. Sample outcomes against policy, verify required fields reached the CRM, and check that the caller's next action was clear. Track correctness, completeness, compliance and customer effort separately. A time-saving call that creates a policy exception carries remediation cost.

Human handoff is part of the design, not a failure state. Record intent, transfer reason, confidence, summary quality and time to completion. A good handoff reduces human work to verification; a poor one makes the human start over. Model them differently. The quality workflow keeps review criteria explicit, while analytics separates outcomes, reasons and next actions for audit.

There are two honest ways to reflect quality. First, include the expected cost of correction, escalation and downstream failure in the numerator. Second, keep a stricter quality-adjusted view in which calls failing the gate are removed from verified resolved calls until corrected. Label the views clearly and avoid double counting the same remediation. A high apparent resolution rate with a weak quality sample should be reported as provisional, not rounded into a win.

Set handoff acceptance criteria before launch: the receiving team, reason for transfer, caller identity or authentication state, relevant account fields, promised next action and ownership. The receiving team should be able to tell whether it accepted the case, completed it, returned it or created follow-up. This makes handoff a measurable operating path and gives product teams a specific failure mode to improve.

Segment the model before trusting the average

An overall average can conceal the calls that matter most. Segment by intent: billing, scheduling, status checks, lead qualification and exceptions have different systems, risk and handoff rates. Segment by channel, and by time when after-hours, peak or seasonal calls have different staffing costs.

Keep separate views for human-only and high-risk intents. Watch mix shift: a pilot full of simple status requests is not comparable with a baseline full of complex account issues. Report volume share with each segment's resolution rate, quality result, human minutes and cost per resolved call.

Use a consistent segment key in both the phone data and CRM. At minimum, preserve intent, customer or case type, channel, time band, language where relevant, transfer destination and risk tier. Report segments large enough to be interpretable, and flag small samples rather than making a precise claim from them. For an executive roll-up, show the weighted total and the segment table that explains it; never let a favorable mix shift disappear inside the average.

Design a pilot that earns confidence

Choose one workflow with a clear outcome, stable policy and enough volume to observe repeats. Freeze definitions and the worksheet before launch, capture the same fields in baseline and pilot, and assign an owner for unresolved work. Prefer a concurrent comparison; otherwise use matched weeks adjusted for intent and time of day.

Review in stages: verify that transfers, follow-ups and CRM outcomes are recorded; sample resolved and unresolved calls with support and quality owners; then compare low, base and high cases with observed ranges. Confidence is earned when outcome definition, call mix, quality sample and cost inputs are complete. A small pilot can provide direction without proving a precise annual return.

Pilot gates

Gate 1, data readiness: pause the comparison if call IDs cannot be joined to outcomes, if transfer events are missing, or if the baseline uses a different resolution definition. Gate 2, customer and quality safety: hold expansion when required disclosures, authentication, escalation or CRM fields fail review. Gate 3, economics: advance only when the base case remains understandable after transfer, repeat, follow-up and unresolved reserves are included. Gate 4, operating readiness: confirm the receiving team, exception owner, monitoring cadence and rollback path before adding volume.

Governance should have named owners: operations owns the outcome definition, finance owns cost assumptions, quality or compliance owns the gate, data owns the joins, and the workflow owner owns changes. Keep a dated decision log with the version, scope, assumptions, sample notes and approval. Review weekly during the pilot and at a fixed monthly or quarterly cadence after launch. A metric without an owner becomes a presentation number instead of a control.

Connect CRM outcomes to continuous improvement

The CRM is where a conversation becomes an operational result. Map the voice AI outcome to a real status, disposition, appointment, case, payment or task. A transcript alone is not resolution. Measure whether the record was accurate and usable by the next person, and include correction work when it was not.

Design the outcome event before designing the dashboard. Store the interaction ID, intent, disposition, resolution state, transfer reason, follow-up owner, next-action date and quality result. Where policy permits, link repeat contacts to the original case so the model can distinguish a new request from a failure to finish the first one. Reconcile closed CRM outcomes against open tasks and queue records; an apparently resolved call with an orphaned task is unresolved work in another location.

Improvements in an Agent Factory should change model inputs, not just a version label. Lower repeat-contact and review assumptions only after a prompt, tool guardrail or routing change shows better first-pass correctness. If transfer minutes fall but follow-up rises, update both lines. Track each change by intent, deployment date and quality result.

Use the improvement loop at the level of a concrete change: identify the failure pattern, adjust the instruction, tool permission, retrieval source, validation rule or routing, then compare the same segment against the prior version. Keep a holdout or review sample where feasible. Promote a change only when the outcome and quality evidence support it, and record whether the improvement saves human minutes, reduces repeats, prevents rework or simply shifts effort to another queue.

Know what this metric cannot prove

Cost per resolved call should not be used to claim a guaranteed payback period, a precise annual saving, universal customer satisfaction, lower compliance risk or permission to remove human coverage. It is also not a substitute for contribution margin, retention, collections performance, service-level attainment or a safety review. A lower ratio can result from cheaper labor assumptions, easier call mix, deferred follow-up or a looser definition of resolution.

Use it to compare like-for-like workflows and to expose where effort moves. Then pair it with the decision metric for the workflow: accurate account action, completed appointment, compliant payment arrangement, closed case or another verified outcome. State the time horizon and whether the view is marginal, fully loaded or inclusive of implementation. That discipline keeps a useful operational measure from becoming an unsupported ROI headline.

Voice AI ROI checklist

  • Write a verified definition of resolution for each intent.
  • Use fully loaded labor minutes, including after-call work and supervision.
  • Separate platform, telephony, transfer, repeat-contact and follow-up costs.
  • Assign a cost to unresolved work and state how implementation is allocated.
  • Keep handled calls, verified resolved calls, resolution rate and quality pass rate visible together.
  • Compare cost per resolved call with quality, human minutes and customer effort.
  • Segment results by intent, channel, time and human-only coverage.
  • Log CRM outcomes, handoff quality and correction work.
  • Run low, base and high cases and document the assumptions a reviewer can change.
  • Set pilot gates, owners, review cadence and a rollback path.
  • Recalculate when an Agent Factory improvement changes routing, tooling or quality.
  • Do not use the ratio alone to claim savings, payback, customer preference or lower risk.

Build a defensible ROI model

Tell us the workflow, volume and current handling effort. We will help define the first measurement window.