跳转到正文
分享一个工作流。Dring AI 会在约两分钟内致电并梳理需求。 申请 AI 回呼
此页面目前仅提供英文版本。 查看英文页面
Quality operations

Call center quality sampling for voice AI

The goal is not to listen to calls at random. It is to create enough structured review to find risk, explain change and improve the next release.

OPERATING PLAYBOOKREVIEWABLE FLOW
Operational guide
01
SignalUnderstand the request
02
RunApply the right rule
03
OutcomeWrite back the next action
FROM SIGNALA useful conversation with a visible ownerTO OWNED OUTCOME

Traditional call quality programs often sample a small percentage of interactions because a human cannot listen to everything. Voice AI changes the economics of analysis, but it does not remove the need for judgement. A transcript can show what was said, yet a reviewer may still need to decide whether the customer was understood, whether the policy was applied correctly and whether the next action was genuinely completed.

A practical quality program combines broad automated checks with targeted human review. It samples high-risk conversations, new workflows, new languages, unusual outcomes and the calls that the system itself flags as uncertain. Dring's quality and testing workflows are designed to connect those reviews to the Agent Factory release cycle.

Sample by risk, not only by volume

A purely random sample is useful for a general health check, but it can miss rare, consequential failures. Add risk-based strata: sensitive payment requests, identity uncertainty, complaints, clinical boundaries, human handoffs, low-confidence transcriptions, failed tool calls and negative sentiment signals. Also review a deliberate sample of successful calls so the team knows which behaviours to preserve.

Include newness as a risk factor. A new language, voice, knowledge article, integration or prompt deserves closer review during its first release window. The same applies to a workflow that expands into a new market. Sampling should follow change as well as demand.

Write a rubric that two reviewers can use

A quality rubric should describe observable behaviour. “Sounded helpful” is hard to calibrate. “Confirmed the customer's requested delivery date, did not invent a carrier status and created a callback when the record was unavailable” is reviewable. Separate dimensions so a call can be strong in one area and weak in another.

  • Intent and context: Did the agent understand what the caller needed?
  • Knowledge and policy: Was the response accurate and within scope?
  • Conversation: Did it listen, handle interruption and avoid unnecessary repetition?
  • Action: Were tools used correctly and only when authorised?
  • Handoff: Did the right queue receive a complete, honest context packet?
  • Outcome: Did the customer reach the agreed next step?

Calibrate reviewers with the same examples. Discuss disagreements and update the rubric when a new failure category appears. Keep an audit trail of the score, reviewer, version and evidence used.

Review the whole call journey

Quality begins before the first sentence and ends after the record is written. Check the greeting and disclosure, number or identity path, speech recognition, knowledge retrieval, tool call, transfer, transcript, summary and CRM outcome. A beautiful answer followed by a wrong write-back is a failed workflow. A correct handoff with no owner is also incomplete.

Dring's shared context model is useful here because a customer may move from voice to WhatsApp or SMS. The reviewer should be able to see whether the case remained coherent across the channel change.

Balance automated checks and human judgement

Automated checks are good at consistency. They can test whether a required disclosure appeared, a field was written, a prohibited action was avoided or a call ended with the correct disposition. Human reviewers are better at nuance: whether a phrase sounded dismissive, whether the caller's correction was truly heard or whether an escalation arrived at the right emotional moment.

Use disagreement as information. If automated and human scores diverge frequently on one language or workflow, the rubric or model may need attention. Do not silently average the two into a number that hides the underlying question.

Make language and voice part of the sample

A 62-language technical capability inventory is not a claim of uniform quality. A global metric can remain stable while one language develops a pronunciation, translation or handoff problem. The public launch-priority set is ten languages, and every requested locale and workflow needs actual-path validation before production. Sample by language, channel and workflow. Listen for proper names, numbers, dates, product vocabulary, polite forms and code-switching. Review whether the same policy is being expressed naturally, not whether the sentence is a literal translation.

The platform supports a common agent brain across voice, WhatsApp, SMS and email, but common knowledge does not remove language-specific QA. Keep the outcome model shared and the test examples locally meaningful.

Feed findings into a release process

Every material finding should have a category and an owner. A pronunciation problem may belong to terminology. A wrong answer may belong to knowledge. A failed lookup may belong to integration. A late transfer may belong to policy or conversation strategy. Add the verified reproduction to the regression set and state the guardrail that must remain green.

The Agent Factory can turn this process into a weekly or monthly cadence: review calls, prioritise defects, create the candidate, run simulations, release in stages and compare production evidence. The value is not the dashboard itself. It is the shorter path from a caller's friction to a controlled improvement.

Do not optimise for a perfect average

Averages can encourage the wrong behaviour. If the team chases a high containment score, agents may avoid handoff when it is needed. If it chases a short call, it may skip clarification. If it chases a high sentiment score, it may use pleasant language while leaving the task unfinished. Keep outcome, quality and guardrails visible together.

Report both the rate and the sample size. Mark the period, workflow, language and release version. When a metric changes, show the underlying category movement. A manager should be able to see whether improvement came from better resolution, fewer risky calls, a mix change or a real behaviour change.

A quality sampling plan

  • Maintain a random health sample plus risk-based samples.
  • Oversample new workflows, languages, integrations and release candidates.
  • Use a rubric with separate intent, policy, action, handoff and outcome scores.
  • Combine automated checks with calibrated human review.
  • Link every important defect to a regression case and an owner.
  • Publish period, denominator, release version and sample size with the metric.

Quality sampling is how a voice AI operation stays honest after the launch announcement. It gives supervisors a way to see what customers experience, gives engineers reproducible failures and gives the next release a standard it must earn.

Further reading

Review the calls that shape the next release

We will help turn your current QA sample into a risk-based operating scorecard.