Ir al contenido
Comparta un flujo. Dring AI le llama en unos dos minutos y califica la necesidad. Solicitar llamada de la IA
Esta página está disponible en inglés por ahora. Ver la página en inglés
Quality

Voice AI quality assurance before production and after launch

Before an agent takes a real call, it is exercised against realistic conversations, including difficult turns, policy boundaries, tool failures and handoff conditions. Each selected locale and workflow is validated before production. After launch, live outcomes are reviewed against the release and the workflow's agreed quality measures.

300Kabout customer conversations/month across inbound+outbound
10K+agent conversation minutes/day
72%about resolution across current production workloads
62language technical coverage, with 10 launch-priority languages
Pipeline

What happens before an agent launches

The same pipeline runs for every agent, regardless of how simple the use case looks.

  1. 01

    Persona generation

    Test personas include angry callers, wrong-order scenarios, policy traps, silence and interruptions.

  2. 02

    Automated conversations

    The agent runs each persona through a full simulated conversation.

  3. 03

    Independent evaluation

    Automated evaluators score each conversation; disagreements and edge cases are audited by humans.

  4. 04

    Regression check

    Results are compared against the previous release before anything ships.

  5. 05

    Handover path checks

    Escalation to a human is verified explicitly, not assumed.

  6. 06

    Staged rollout

    Live traffic increases in controlled stages, with scores watched at each step and the previous release available for rollback.

Judging

Why independent evaluation

Automated evaluation can drift or miss a nuance. Independent checks and human review of disagreements make the quality signal more useful than a single score alone.

This process is checked against human QA review, with disagreements investigated and the evaluation set refined when the workflow exposes a new edge case.

Test call #1042Reviewed
Review AResolved
Review BResolved
AgreementMatch
Human auditNot required
A/B testing in production

The better opener wins on outcomes, not opinions.

Which opening line, tone or voice works is decided on live calls with a controlled share of traffic. Variants are compared on the outcome the campaign exists for, the winner is promoted in stages, and every promoted change still passes the regression suite. The visual below is an illustrative example; each real experiment uses an agreed sample, outcome metric and guardrails.

Illustrative test · opener variantsPhase 1 · equal traffic
Variant A"Hello, I am calling from the logistics team about your shipment request."
traffic33%
green outcomes0%
0 calls
Variant B"Hi, you asked us for a freight quote last month. Is now a good moment for two questions?"
traffic33%
green outcomes0%
0 calls
Variant C"Good afternoon. This is a short call about international shipping rates."
traffic34%
green outcomes0%
0 calls
Minimum sample not reached. Keep splitting traffic.

What gets tested

Opening line, tone, voice, pacing, the order of questions, the way an objection is handled. One variable per variant, so the result is attributable.

How the winner is chosen

On the campaign's own outcome metric, the green share of the outcome status, with a minimum sample per variant before any decision. Sentiment and handover rate act as guardrails: a variant that wins on outcomes but loses on reception does not ship.

How it goes live

The winner is promoted in controlled stages, the other variants are retired, and the change is recorded so the next comparison starts from a known baseline.

Production

In production, every call is scored

Testing does not stop at launch. Every live call is scored the same way, continuously.

Resolution

Whether the call actually resolved what the customer called for.

Sentiment

How the customer's tone moved over the course of the call.

Policy adherence

Whether the agent stayed within the rules it was given.

Latency

Response time throughout the conversation.

Each call also produces a structured end-of-call summary. Drift alerts flag agents whose scores move outside the expected range, and weekly sweeps run across all live agents to catch it early.

Guardrails

What an agent will never do

Refuse out of scope

An agent declines requests outside what it was built to handle.

Never invent policy

It does not guess at a policy it was not given.

Escalate sensitive topics

Sensitive requests are routed to a human rather than handled by the agent.

Always offer handover

A human option is always available to the caller.

Measurement

What we measure

MetricDefinitionWhy it matters
ResolutionWhether the customer's reason for calling was actually addressed.The core measure of whether the agent did its job.
SentimentHow the customer's tone changed from the start of the call to the end.Catches calls that resolved on paper but frustrated the customer.
Policy adherenceWhether the agent's actions stayed within the rules configured for it.Keeps the agent inside the boundaries the business set.
LatencyResponse time from the customer finishing speaking to the agent replying.Slow responses feel like a bad connection, even when the answer is right.
Handover accuracyWhether the agent escalated to a human when it should have.A missed escalation is worse than an unnecessary one.
Regression deltaScore change against the previous release on the same test set.Catches a release that quietly makes the agent worse.
FAQ

Testing questions

Does testing happen before every release, not just launch?+

Yes. Regression sweeps run weekly against the previous release, and every change goes through the same staged rollout.

What happens when evaluations disagree?+

Disagreements are routed to a human reviewer for audit before the result is finalized.

Can we see the test scenarios for our agent?+

Yes. Test personas and scenarios are built around your specific policies and edge cases and are visible during onboarding.

What happens when a live agent's scores drift?+

Drift alerts flag it automatically, and it is reviewed as part of the weekly sweep across all live agents.

Ver todas las preguntas

See the testing pipeline on your own use case

Walk through what a measured launch would look like for your agent, from the first workflow brief to the post-call review loop. A consented form submission can be followed by an AI qualification callback in about two minutes, with timing confirmed for the workflow. Read the Agent Factory overview, the voice AI operations snapshot and the 62-language customer service guide before you scope the next step.