Skip to main content
Text to speech in DialNexa is what the caller actually hears. It turns the model’s reply into audio through a selected voice, language, voice model, speed, stability, volume, and provider path. A voice can make a correct answer feel helpful, rushed, unclear, or strangely formal. DialNexa voice selector showing voice filters, sample playback, Nexa voice ID copy button, and row language selection. DialNexa voice settings popover showing voice model, speed, stability, and volume controls for a selected voice.
The caller does not hear your provider architecture. They hear a voice saying their name, amount, date, and next step. Test those words.

Choose DialNexa Text To Speech And Voice Providers

For provider background, use the ElevenLabs integration catalog page. Cartesia voices are selected from the DialNexa voice selector.
Choose ElevenLabs when the voice personality matters and you want to audition a wider library. In the current dashboard path, ElevenLabs agent versions are standardized on Flash v2.5 (eleven_flash_v2_5) where supported, so treat that as the main model to test.

How The Voice Selector Works

The selector is designed for large voice libraries.

Voice Settings In The Popover

Do not copy provider documentation numbers into DialNexa sliders. Use the UI values and test calls. The dashboard maps provider ranges before saving.

Audio Cache And Repeated Speech

Audio Cache stores synthesized audio for repeated phrases. It works best when the generated text, voice provider, voice id, voice settings, and output format repeat.
Audio Cache loves repetition. If every sentence is personalized confetti, cache will politely sit there doing very little.
DialNexa does not store incomplete interrupted segments as reusable cache entries. A segment must finish cleanly and match the text that was sent for synthesis before it can become future cached audio.

Streaming Speech Continuity

For streamed voices, DialNexa buffers and plays synthesized segments in order so callers hear the response as a coherent sentence. This is especially important for fast Cartesia paths and multi-part ElevenLabs output, where chunks can finish at different times. Hindi punctuation is also treated as a sentence boundary for supported TTS segmentation, so Hindi and Hinglish responses can flush at natural pause points instead of waiting for only English punctuation.

Where Voice Quality Shows Up Outside The Call

Voice quality is not only a caller comfort issue. It changes whether downstream work is trusted.

Voice Review Checklist

1

Test the first sentence

The welcome line sets trust. Check pace, pronunciation, greeting tone, and whether the voice fits the use case.
2

Test difficult words

Include brand terms, product names, locality names, acronyms, medicine names, plan names, and agent names.
3

Test numbers and dates

Amounts, due dates, order IDs, phone numbers, and appointment slots reveal speech issues quickly.
4

Test interruption recovery

Interrupt the agent during the greeting and check how naturally it resumes.
5

Review recording and transcript together

The transcript shows content. The recording shows delivery.

Supported Voices And Models

Review voice fields and model fields.

Speech Settings

Enable Audio Cache and tune speech behavior.

Multilingual And Hinglish Calls

Match voice, language, and transcriber.

Audio Cache Monitoring

Read cache evidence on the call detail page.