Zum Inhalt springen
Teilen Sie einen Workflow. Dring AI ruft in etwa zwei Minuten an und qualifiziert den Bedarf. KI-Rückruf anfordern
Diese Seite ist derzeit nur auf Englisch verfügbar. Zur englischen Seite
Speech experience

Voice AI pronunciation dictionaries for names, brands and products

Natural voice starts with being able to say the words that matter to the business, the customer and the next action.

OPERATING PLAYBOOKREVIEWABLE FLOW
Operational guide
01
SignalUnderstand the request
02
RunApply the right rule
03
OutcomeWrite back the next action
FROM SIGNALA useful conversation with a visible ownerTO OWNED OUTCOME

People notice pronunciation before they notice architecture. A voice agent that mispronounces a customer's name, a clinic procedure, a product family or a city can sound careless even when the underlying answer is correct. In a sales call, the wrong product name can weaken trust. In a logistics call, a location error can change the operational instruction. In a support call, a caller may spend the next minute correcting the system.

A pronunciation dictionary gives the agent a managed way to recognise and say important terms. It is more than a list of phonetic spellings. It should connect written forms, spoken variants, language, context, confidence and the workflow where a mistake would matter. Dring's agent layer and multilingual platform keep speech quality tied to the conversation's actual job.

Choose the terms that deserve attention

Do not begin by collecting every name in the CRM. Prioritise terms that affect identity, routing, product choice, financial value, safety or customer trust. Include recurring brand names, plan tiers, model numbers, partner names, street names, regional places, technical abbreviations and words that callers frequently correct.

Ask operators which terms make customers pause. Review transcripts or quality notes for corrections such as “I said…” or “that's pronounced…”. Combine frequency with consequence. A rare medical term or financial product may deserve more attention than a common greeting because the cost of getting it wrong is higher.

Model recognition and speech separately

The agent must recognise what the caller says and pronounce what it says back. Those are related but distinct problems. A speech recogniser may need a phrase set or custom class to favour a rare term. The text-to-speech layer may need a pronunciation instruction or lexicon entry to produce the natural local form. Test both paths.

Keep the canonical written value for CRM and analytics. Store a spoken alias or phonetic representation separately. This prevents a phonetic workaround from polluting the customer record. If the term has multiple valid pronunciations by region, associate each with the caller's language or selected locale rather than choosing one globally.

Context decides the correct word

A name can be ambiguous without context. A product code may sound like a date, a street number or a person. A phrase may be a brand in one workflow and a generic noun in another. The dictionary should therefore include context tags and example phrases. Recognition should use the surrounding words and the expected entity type rather than a single isolated keyword.

When context remains unclear, ask a short confirmation. “Did you mean the Atlas plan or the Atlas account?” is better than silently selecting one. The agent should confirm before taking a consequential action and should preserve the caller's correction for future review.

Handle numbers, dates and spelling carefully

Many voice errors are not words but representations. Customers say model numbers with pauses, dates in local order, amounts with decimal conventions and email addresses in fragments. Define how the agent asks, repeats and normalises these values. A CRM record may need a standard numeric form while the spoken confirmation should sound natural in the customer's language.

Test leading zeros, repeated digits, letters that sound alike and names with diacritics. For a shipping or appointment workflow, confirm the value back before writing it. For a sales call, distinguish a budget range from a final commitment. A pronunciation dictionary cannot solve every ambiguity; it should make the safe confirmation path easier.

Review language-specific naturalness

A pronunciation that sounds correct in one language may sound artificial in another. Use native or highly fluent reviewers for important terms and ask them to judge the complete sentence, not only the isolated word. Include formality, stress, rhythm and whether the agent sounds like the kind of operator the customer expects.

Dring's 62-language technical capability inventory spans voice, WhatsApp, SMS and email. Shared knowledge and tone should remain stable, but spoken expression may need local adaptation. Ten languages are public launch priorities, and every requested locale/workflow is validated on the actual path before production. Keep a shared concept record and language-specific pronunciation tests. The quality layer can show whether a change improved one locale while affecting another.

Connect pronunciation to the Agent Factory

Pronunciation issues should not be fixed only in production. Add representative terms to the simulation set and score recognition, response pronunciation, action selection and record accuracy. When a live call reveals a new term, classify it, redact unnecessary personal data and add a controlled regression case.

The Agent Factory can then package the dictionary update with the right release candidate. The team can see what changed, test it across the affected workflows and release it gradually. This is safer than adding a broad instruction that changes every conversation.

Use a small language quality panel

One reviewer rarely hears every problem. For important terms, use a small panel that combines a native or highly fluent speaker, an operator who knows the workflow and a quality owner who can inspect the transcript and record. Ask them to judge the full exchange: whether the term was recognised, whether the agent said it naturally, whether the correction was handled politely and whether the final action remained correct.

Keep the review set intentionally small and repeatable. Five to ten carefully chosen examples can reveal a regression faster than a large unlabelled call sample. Include a normal pronunciation, a regional variant, a noisy version, a correction and an ambiguous phrase. Store the decision with the language, model or instruction version and affected workflow. That makes it possible to improve one locale without accidentally changing the experience everywhere else.

A pronunciation dictionary record

  • Canonical written term and entity type.
  • Language, region and accepted spoken variants.
  • Preferred agent pronunciation and context examples.
  • Common misrecognitions and the safe confirmation phrase.
  • Workflow, tool or policy impact if the term is wrong.
  • Owner, review date, test cases and release version.

Pronunciation is part of the customer experience and part of data quality. When the agent recognises and speaks the terms that matter, customers correct it less, handoffs contain better context and the operation feels considered rather than generic.

The strongest dictionaries are also living release assets. A term enters because a caller used it, an operator flagged it or a quality review found a risk. It leaves when the product or policy changes. Keeping that history means the team can explain why a pronunciation was added, compare the affected calls and roll back a change without guessing which instruction caused the drift.

Further reading

Say the words your customers use

Bring your difficult names, product terms and local pronunciations to a focused voice review.