Voice
Voice Agents use TTS (Text to Speech), which generates audio that LLMs generate during the course of a conversation. This is the audio that the end user having the conversation listens to.
Talkr platform supports ElevenLabs, OpenAI, Google, Azure Speech, Deepgram, Cartesia, Smallest AI, MiniMax, Sarvam, Rime, Inworld, Camb.ai, xAI, and LMNT. There are some voices from the providers that we ship by default. You can refer to the providers API documentation to select a voice ID that's most relevant for your language requirement.
If your agent needs to speak more than one language in the same call, see Voice & Language Profiles — Talkr can switch TTS voice mid-call as the detected language changes.
Voice providers receive the data needed to synthesize speech, such as generated text, selected voice, model settings, and request metadata. Review the provider's data processing, retention, model training, and regional hosting policies before using sensitive data.
For locally deployed or self-hosted TTS models, Talkr also supports Speaches, an OpenAI API-compatible server for speech generation.
If you don't find your favourite voice, you can always add the voice ID manually.
