Naar inhoud
Deel één workflow. Dring AI belt binnen ongeveer twee minuten en kwalificeert de behoefte. Vraag een AI-terugbelgesprek aan
Deze pagina is voorlopig alleen in het Engels beschikbaar. Naar de Engelse pagina
Multilingual voice AI · Language quality

Multilingual voice AI language support: what 62 languages really means

Dring has a confirmed 62-language technical capability inventory across voice, WhatsApp, SMS and email. The public launch-priority set is smaller: ten languages and locale tracks selected for market relevance and testability.

By Dring AI · Published 2026-09-05 · Reviewed 2026-09-05

OPERATING PLAYBOOKREVIEWABLE FLOW
Operational guide
01
SignalUnderstand the request
02
RunApply the right rule
03
OutcomeWrite back the next action
FROM SIGNALA useful conversation with a visible ownerTO OWNED OUTCOME

A language count is useful only when it is honest about what it counts. Dring's confirmed technical capability inventory covers 62 languages across voice, WhatsApp, SMS and email. That inventory describes the speech and messaging configurations that can be assembled across the product stack. It does not mean that every language has the same voice, the same dialect coverage, the same telephone performance or the same level of production evidence.

For public launch priorities, Dring recommends a deliberately smaller set of ten: English with UK, Australian and Singaporean tracks; Turkish; German; French; Dutch; Spanish; UAE Arabic; Italian; European Portuguese; and Polish. Every requested locale and workflow is validated on the actual telephony path before production. The result is a clearer promise: broad technical reach, with a narrower quality-assured set that earns a public claim through testing.

This distinction matters for a customer support leader choosing a voice agent, but it also matters for sales qualification, appointment booking, logistics, healthcare coordination, finance support, HR workflows, surveys and reactivation. The agent is not successful because it can pronounce a greeting. It is successful when a caller can complete the intended job, the system records the right outcome and a person receives the right context when the workflow reaches its boundary.

Capability is not the same as production quality

Provider documentation normally answers a compatibility question: can a model accept or generate a language code? A production buyer needs a harder answer. Can the selected voice remain intelligible through an 8 kHz telephone route? Can the speech recognizer capture a company name, an order number and a local address? Can the agent understand an interruption, use a tool correctly and hand a difficult call to a human without losing the case?

These are different layers. A language may be listed by a text-to-speech provider but have no Dring-approved voice for a particular dialect. A speech-to-text model may list a language while a specific carrier, codec, noise profile or speaking rate creates errors. A fluent transcript may still lead to a wrong tool argument. A natural voice may still read a number incorrectly. The only defensible product decision is therefore tied to a tested combination of language, locale, workflow, STT model, TTS model, prompt, glossary, telephony route and fallback.

The public language page should make that evidence legible. It should show whether a language is a priority launch language, available on request, or experimental; which locale was tested; when it was last reviewed; which workflows are in scope; and what happens when recognition or language routing is uncertain.

The audio architecture: three separate jobs

A multilingual Dring call is best understood as a chain:

telephone audio -> speech-to-text -> GPT-5.6 text and reasoning brain -> text-to-speech -> telephone audio

The chosen GPT-5.6 text/reasoning model is not STT/TTS. OpenAI's current GPT-5.6 model documentation lists text input and output, image input, and audio as not supported. GPT-5.6 can be the reasoning and tool-use brain after a transcript arrives. It can select a workflow, apply policy, decide whether to ask a clarifying question, call an approved tool and produce the response text. It does not itself hear the phone signal or synthesize the spoken answer.

OpenAI documents audio-capable models separately, including gpt-realtime-2.1 and gpt-audio-1.5. Those models are not interchangeable with the GPT-5.6 text/reasoning claim. Dring's website should name the actual STT and TTS providers used for each voice path and should not imply that GPT-5.6 supplies the microphone, telephone recognition or voice.

This separation also makes testing more useful. Speech recognition errors belong to the STT and audio path. Pronunciation and naturalness belong to the TTS voice and text-normalization path. Wrong decisions, policy violations and incorrect tool calls belong to the reasoning and workflow path. A combined call test must measure all three without hiding one layer behind an overall success percentage. See the Dring platform for the connected workflow surface.

Provider comparison for multilingual telephone work

LayerDocumented capabilityWhat Dring must still prove
ElevenLabs TTSElevenLabs documents Eleven v3 at 70+ languages, with its current language list showing 74. It lists Multilingual v2 at 29 languages and Flash v2.5 at 32. The v2 list includes English, German, French, Spanish, Dutch, Turkish, Bulgarian, Romanian, Arabic, Italian, Portuguese and Polish.Choose a locale-appropriate voice, test it through the telephone codec, check naturalness with native listeners and maintain pronunciation rules for names, places, products, abbreviations and numbers.
ElevenLabs pronunciationElevenLabs documents IPA and CMU pronunciation dictionaries. Its current guidance says phoneme tags work with eleven_flash_v2 and eleven_v3; aliases can be used where phoneme tags are not accepted.Confirm that a dictionary rule works with the selected model, version the rules, and regression-test every critical term after changing a voice or normalization strategy.
Deepgram STTDeepgram documents Nova-3 as a broad language-specific catalog, including Turkish, Arabic locale tags, Bulgarian, Polish and Romanian. Nova-2 remains documented as an option for languages not yet available on Nova-3.Use the exact language or locale route, measure recognition on telephone audio, and verify extracted names, numbers, entities and tool arguments rather than relying on a general language label.
Deepgram Fluxflux-general-multi is documented for ten languages: English, Spanish, French, German, Hindi, Russian, Portuguese, Japanese, Italian and Dutch. It supports language hints and turn-aware, interruption-aware streaming.Do not imply that Flux multilingual covers Turkish, Arabic or Bulgarian. For those languages, prove the Nova-3 path and the complete latency, turn-taking and fallback behavior separately.
Deepgram language toolsDeepgram documents language detection and keyterm prompting. Keyterms are plain terms or phrases for Nova-3 and Flux, with up to 100 terms described in the current guidance.Use detection when the language is genuinely unknown, use hints when it is known, and use keyterms to improve a controlled vocabulary. Neither feature proves native naturalness or safe task completion.
OpenAI GPT-5.6The GPT-5.6 model page lists text reasoning and tool capabilities but audio not supported. OpenAI's audio guide describes separate audio models and separate audio APIs.Keep GPT-5.6's role explicit in architecture, logs and website copy. Evaluate the STT, reasoning/tool and TTS stages as separate failure surfaces.

These are provider-documented capabilities, not Dring performance scores. For example, ElevenLabs documents 8 kHz mu-law and A-law as telephony-optimized output formats, but that does not establish that any chosen voice will sound natural after encoding. Deepgram documents language support and language controls, but that does not establish the error profile of a Turkish, UAE Arabic or Polish workflow on a particular carrier. The quality and testing workflow is where a compatibility claim becomes a release decision.

The ten-language public priority set

The priority set is a market and evidence decision, not a ranking of human languages. It combines Dring's named acquisition markets with documented provider coverage and a realistic ability to recruit native reviewers, assemble terminology and test telephone calls.

Language trackMarket reasonProvider and locale riskAcceptance focus
English
en-GB, en-AU, en-SG
UK, Australia and Singapore are named acquisition markets. English also supports broader European buying teams.English is not one accent. Names, postcodes, currencies, date formats and business vocabulary differ. Deepgram documents broad English accent coverage, while ElevenLabs documents US, UK and Australian English in Multilingual v2.Run separate UK, Australian and Singaporean caller panels. Test postcodes, dates, currencies, names, interruptions and tool readback for each track.
Turkish
tr-TR
Turkey is the first named acquisition market and requires native product language, not only translated pages.Deepgram documents tr and tr-TR for Nova-3 and Turkish for Nova-2; ElevenLabs lists Turkish. Turkish is not in Flux multilingual's ten-language list. Suffixes, rapid speech and local names need focused review.Use a native Turkish panel, a Turkish glossary and local numbers, dates, addresses and company names. Test an explicit Nova-3 route and human correction of critical prompts.
German
de-DE
Germany is a named acquisition market and a major European expansion path.Deepgram documents German on Flux multilingual and Nova-3, including de-CH as a separate locale. Germany, Austria and Switzerland must not be treated as one voice promise. Compounds and formal address are high-risk.Use Germany-native reviewers. Test compound nouns, umlauts, long numbers, formal address, handoff context and interruption recovery without implying Swiss or Austrian coverage.
French
fr-FR
France is a named acquisition market and French opens a wider international and European path.Deepgram documents French and fr-CA; ElevenLabs documents France and Canada in Multilingual v2. Liaison, politeness and number phrasing need a France-specific review.Test France-native speech, amounts, dates, names, liaison-heavy prompts, confirmation readback and the correct voice locale.
Dutch
nl-NL
The Netherlands is a named market and Dutch is a useful regional hub language.Deepgram documents Dutch on Flux and Nova-3; ElevenLabs and Deepgram Aura-2 document Dutch. Netherlands Dutch and Flemish are separate QA tracks.Test Dutch names, addresses, compounds, product terms and tool arguments. Keep Flemish out of the public claim until it has its own panel.
Spanish
es-ES
Spain is a named market. A Spain-first track is a controlled extension into wider Spanish-speaking demand.Deepgram documents Spanish and es-419 on Flux and Nova-3. ElevenLabs and Aura-2 document multiple regional varieties. Peninsular and Latin American Spanish must not be silently merged.Use Spain-native reviewers, test accent contrast, names, addresses, dates, currencies and polite turn-taking. Do not label a Latin American voice as es-ES without evidence.
UAE Arabic
ar-AE
The UAE is a named acquisition market and Arabic is strategically important for Gulf conversations.Deepgram documents ar-AE and other Arabic tags on Nova-3; ElevenLabs lists Arabic including UAE in Multilingual v2. Arabic is not in Flux multilingual. Dialect, formality, script normalization and names require ownership.Use UAE-native reviewers. Test Gulf commercial terms, Arabic names, numbers, Modern Standard Arabic versus dialect prompts and explicit Nova-3 routing.
Italian
it-IT
Italy is a practical broader-Europe extension with a strong documented model intersection.Deepgram documents Italian on Flux and Nova-3; Aura-2 and ElevenLabs list Italian. Regional accents, doubled consonants and entity names can change perceived competence or extracted data.Test Italy-native speech, names, places, amounts, dates, barge-in and confirmation of extracted values.
European Portuguese
pt-PT
Portugal is a natural broader-Europe extension and gives Dring a clear distinction from Brazilian Portuguese.Deepgram documents Portuguese on Flux and Nova-3; ElevenLabs lists Portuguese. pt-PT and pt-BR differ in rhythm, vocabulary and pronunciation.Test Portugal-native callers, currency, dates, names and numbers. Use a separate Brazilian track rather than letting it pass as European Portuguese.
Polish
pl-PL
Poland is a defensible broader-Europe priority with a large terminology and name surface.Deepgram documents Polish on Nova-3 and Nova-2; ElevenLabs lists Polish. Polish is not in Flux multilingual. Consonant clusters, inflection and declined names need regression coverage.Test Polish names, addresses, inflected forms, numbers, readback and exact tool arguments on the language-specific Nova-3 path.

Where Bulgarian belongs

Bulgarian is explicitly validated-on-request, not dismissed and not quietly counted as public priority. ElevenLabs documents Bulgarian in Multilingual v2. Deepgram documents Bulgarian in both Nova-3 and Nova-2, while Bulgarian is outside Flux multilingual's ten-language set. That makes Bulgarian a technically credible request path with a real model option, but not enough evidence for a blanket public quality claim.

A Bulgarian customer pilot should have a bg-BG voice decision, native reviewer panel, Bulgarian company and place names, number and address tests, keyterm list, human handoff, fallback language and a dated release record. If the card passes the same telephone gate as Tier 1, Bulgarian can be promoted without changing the broader 62-language capability claim.

Language tiers and release policy

Public priority: English with three locale tracks, Turkish, German, French, Dutch, Spanish, UAE Arabic, Italian, European Portuguese and Polish. These are the only languages recommended for prominent quality-oriented launch copy, and only for the workflows and locales that have passed the gate.

Validated-on-request: Bulgarian, Romanian, Czech, Greek, Hungarian, Croatian, Slovak, Ukrainian, Danish, Swedish, Finnish and Norwegian are sensible named-customer candidates because relevant provider paths are documented. They need a language card, native reviewers and a complete telephone test before a live promise. The same rule applies to any public-priority language when a customer requests a new dialect, carrier or workflow.

Experimental or holdback: long-tail languages that rely mainly on broad Eleven v3 TTS coverage, lack a tested Deepgram real-time STT path for the workflow, lack an approved telephone voice or lack native review. This can include untested Arabic variants, Chinese variants, Japanese, Korean, Hindi, Russian, Vietnamese, Indonesian, Malay, Filipino and Tamil. Experimental means an engineering or customer pilot is possible; it does not mean equivalent public quality.

Workflow-specific acceptance method

Do not test a language with a translated greeting and a few polite questions. Freeze the complete call contract first: intended outcome, in-scope intents, prohibited actions, authoritative systems, tools, human queue, consent or disclosure, business hours, fallback language and critical terminology. A support flow, a lead qualification flow and a healthcare coordination flow need different test cases even when they use the same language.

Build a language card containing the locale, STT model, TTS model, voice ID, prompt version, pronunciation dictionary version, keyterms, telephone carrier and codec, fallback path, test date and owners. Use consented or synthetic examples that reflect the actual market. Include customer names, company names, addresses, product names, abbreviations, local currencies, dates, reference IDs and terms that the workflow must capture exactly.

Run the test set in three passes. First, test STT and TTS independently: can the transcript preserve the terms, and can a native listener understand the generated response? Second, test the reasoning and tool layer with fixed transcripts: does GPT-5.6 apply the policy, ask the right question and choose the correct tool arguments? Third, test the complete telephone loop with interruptions, noise, silence, barge-in, caller correction, tool delay and handoff. A good isolated score cannot rescue a failed end-to-end call.

The Agent Factory should turn reviewed calls into regression cases. A pronunciation correction, model change, telephony change, keyterm update or prompt edit should be tested against the existing cases before traffic expands. Keep the prior configuration available and record the reviewer decision, not just a pass percentage.

The measurable launch gate

The thresholds below are proposed Dring release criteria, not provider guarantees and not claimed Dring scores. Run at least 120 scripted caller turns per language, three native speakers per locale track, 30 business terms, 30 names or structured identifiers, 20 tool scenarios and interruption cases. Use at least three independent native listeners for voice review and a fourth adjudicator for critical disagreement.

  • WER and keyterm recall: target WER at or below 12% on the telephone set and keyterm recall at or above 95% for the approved vocabulary. These are diagnostics, not the decision by themselves. Keyterm prompting can improve a known vocabulary while general conversation remains weak.
  • Human MOS and naturalness: mean at least 4.0 out of 5, with no locale panel mean below 3.7. Review clarity, pacing, pronunciation, confidence, politeness and whether the voice sounds like the promised locale.
  • Task completion: at least 90% of scripted calls reach the intended business outcome without human repair. Consent, opt-out, safety and required escalation paths must be 100% correct.
  • Tool accuracy: at least 95% correct tool selection and arguments overall, and 98% exactness for customer-visible or irreversible values. A fluent wrong booking or wrong CRM write is a release failure.
  • Interruption handling: at least 90% of interruption cases resume the correct turn without speaking over the caller, dropping context or accepting an unspoken instruction.
  • Names and numbers: at least 98% exact capture and readback for the critical set of phone numbers, dates, amounts, names, addresses and identifiers. A low WER with one wrong phone number is still unacceptable.
  • Language routing: at least 98% correct language selection on expected-language cases, with an explicit caller correction path and deterministic human fallback when confidence is low.
  • Audio integrity: no repeated clipping, dropped audio, codec regression or turn timeout in the accepted call set. Record p50 and p95 response latency against the approved workflow SLA.

Fail the release for any repeated unauthorized disclosure, wrong critical number, false completion, unsafe tool action, missed opt-out, lost handoff context or unowned exception. Do not average a critical error away with fluent routine calls. A language can pass one narrow workflow and remain validated-on-request for general product claims.

Code-switching, detection and dialect caveats

Code-switching is common in business calls. A Turkish caller may use an English product name. A Dutch caller may mention a global platform. An Arabic caller may use a technical acronym. Test the actual mixture instead of assuming that a language detector or multilingual model will always choose the intended route.

When the expected language is known, use a language-specific route or a language hint where the provider documents it. Deepgram documents language hints for Flux multilingual and language parameters for Nova models. When the language is genuinely unknown, detection can identify the dominant language, but the system needs a confirmation and correction experience. After detection, locking to the chosen language may be more predictable than allowing every subsequent turn to drift, while genuinely bilingual workflows may need multiple hints. The decision belongs in the call design and test set.

Keyterm prompting is valuable for Dring product names, customer systems, sectors and workflow vocabulary. It should be maintained as a controlled glossary, not as a list of every word in a language. Pronunciation dictionaries and aliases are useful on the TTS side, but they do not repair an incorrect intent, a missing policy or a bad tool mapping. Keep source spelling, spoken form, meaning, forbidden alternatives and locale in one versioned record.

Dialect claims need the same discipline. `en-GB`, `en-AU` and `en-SG` are acceptance tracks, not decorative language tags. `de-DE` is not automatically `de-CH`. `fr-FR` is not automatically `fr-CA`. `es-ES` is not automatically `es-419`. `pt-PT` is not automatically `pt-BR`. `ar-AE` is not a promise for every Arabic-speaking market. A locale tag should narrow the promise to the voices, terminology and reviewer evidence that actually exist.

Use cases beyond a receptionist

Multilingual voice AI earns its place when language continuity removes work from a real operation. A customer support agent can identify an intent, retrieve an approved order or account status, answer within policy and create a human handoff with the original language and context. A sales agent can qualify a lead, capture timing and volume, handle an objection and write a structured brief without forcing the next seller to repeat the discovery call.

In appointment booking, the agent can ask for the right service, search approved availability, confirm the selected slot and escalate exceptions. In logistics, it can capture a shipment or load reference, report the latest approved checkpoint, open an exception and pass the correct details to dispatch. In healthcare coordination, it can collect administrative context and arrange an approved next step while sending clinical, urgent or sensitive questions to the right human queue.

Finance teams can use a bounded agent for routine status questions and case intake while routing identity, fraud, payment and dispute decisions to authorized staff. HR teams can use it for structured screening or scheduling with a reviewed rubric and human decision boundary. Surveys can gather responses and classify themes, while complaints, distress and urgent matters receive a human route. WhatsApp, SMS and email can carry the same case context after the voice interaction, but the content and consent rules for each channel still need their own review.

The common pattern is not "one voice for everything." It is one operating model with language-specific expression: authoritative data, approved actions, observable outcomes and a handoff that preserves what the customer said. Explore the voice AI operations benchmark for the difference between a broad capability statement and an evidence-backed operating measure.

How to state the claim on the website

Keep two inventories. The first is the confirmed 62-language technical capability inventory across the configured voice, WhatsApp, SMS and email stack. The second is the smaller quality-assured priority inventory, which currently has ten recommended language tracks but only becomes a production claim per locale and workflow after testing.

A defensible website statement is:

Dring has a confirmed 62-language technical capability inventory across voice, WhatsApp, SMS and email. Our public launch-priority set is smaller: English with UK, Australian and Singaporean tracks, Turkish, German, French, Dutch, Spanish, UAE Arabic, Italian, European Portuguese and Polish. Every requested locale and workflow is validated on the actual telephone path before production. Other languages, including Bulgarian, are available on request after workflow-specific validation.

Do not claim native-level or equal quality for all 62 languages, imply that every model supports every language or imply that one language label covers every dialect. Do not imply that GPT-5.6 is an audio model. Do not publish a universal Dring WER, MOS or completion score unless the dataset, denominator, model versions, telephone route and review date are public and accurate. Link readers to the callback form for a language and workflow assessment rather than promising an untested locale.

Frequently asked questions

Does 62-language capability mean all 62 languages are production-ready?

No. It is a technical capability inventory across voice, WhatsApp, SMS and email. Production readiness is narrower and depends on the locale, workflow, models, voice, terminology, telephony path, native review and fallback. The public priority set is ten language tracks, with each requested production path validated separately.

What is the role of GPT-5.6 in a voice agent?

GPT-5.6 can act as the text and reasoning brain after speech has been transcribed. It can apply workflow rules, produce text, select approved tools and prepare a structured outcome. OpenAI's current documentation lists audio as not supported for GPT-5.6, so STT and TTS remain separate audio components.

Why is Bulgarian validated-on-request?

Bulgarian is explicitly supported in the documented ElevenLabs and Deepgram model catalogs, so it is a credible request path. It is outside the named acquisition markets and Flux multilingual's ten-language set, and it needs its own telephone, terminology, native-listener and workflow tests before a public quality claim.

Are WER and keyterm recall enough to choose a language?

No. They help diagnose recognition, but a production call also needs naturalness, locale-appropriate accent, task completion, tool accuracy, interruption recovery, exact names and numbers, language routing, audio integrity and a safe human handoff. A single aggregate score can hide a critical failure.

Can one language label cover every dialect?

No. Use locale tracks such as `en-GB`, `en-AU`, `en-SG`, `de-DE`, `fr-FR`, `es-ES`, `pt-PT` and `ar-AE` when that is the tested promise. Add a new dialect only after its voice, terminology, caller panel and workflow gate are complete.

Official technical sources

Provider catalogs and model behavior can change. Re-check the linked official documentation, the selected API model IDs and the actual Dring test record before expanding a language claim.

Validate the language your workflow needs

Bring the locale, terminology and first business outcome. We will map the telephone test path and the human fallback.