Overview
TheTextToSpeech class provides an async interface to the Unmute TTS server. It streams text input and receives synthesized audio in real-time with precise timing control.
Class Definition
Constructor
str
default:"TTS_SERVER"
URL of the TTS server instance
Recorder | None
default:"None"
Optional recorder instance for logging TTS events
Callable[[], float] | None
default:"None"
Optional callback function to get current time (for synchronization)
str | None
default:"None"
Voice identifier. Can be a preset voice name or “custom:” prefixed for custom voice embeddings
Properties
voice
received_samples
received_samples_yielded
Core Methods
send
str | TTSClientMessage
required
Text string or structured message to synthesize. Strings are automatically preprocessed to remove unpronounceable characters.
- Empty strings are ignored
- String messages are preprocessed by
prepare_text_for_tts() TTSClientTextMessagebypasses preprocessing
start_up
MissingServiceAtCapacity: If TTS server is at capacityAssertionError: If connection setup fails
- Sends custom voice embeddings if voice starts with “custom:”
- Waits for
TTSReadyMessagebefore considering startup complete
shutdown
- Active sessions count
- Total audio duration
- Generation duration
state
Literal["not_created", "connecting", "connected", "closing", "closed"]
Async Iterator
aiter
TTSAudioMessage: Synthesized audio chunksTTSTextMessage: Text alignment with timing information
- Audio is buffered and released with
AUDIO_BUFFER_SECdelay (approx. 160ms) - Text messages are synchronized with audio playback timing
Message Types
Client Messages (sent to server)
TTSClientTextMessage
TTSClientVoiceMessage
TTSClientEosMessage
Server Messages (received from server)
TTSAudioMessage
TTSTextMessage
TTSErrorMessage
TTSReadyMessage
Helper Functions
prepare_text_for_tts
- Strips leading/trailing whitespace
- Removes unpronounceable characters:
*,_,` - Normalizes curly quotes to straight quotes
- Removes spaces around colons
Example Usage
Basic Synthesis
Streaming Synthesis
Custom Voice
Configuration
TtsStreamingQuery
Metrics
The class automatically tracks:TTS_SESSIONS: Total TTS sessionsTTS_ACTIVE_SESSIONS: Active synthesis sessionsTTS_SENT_FRAMES: Text chunks sentTTS_RECV_FRAMES: Audio chunks receivedTTS_RECV_WORDS: Words with timing info receivedTTS_TTFT: Time to first token (audio)TTS_AUDIO_DURATION: Total audio generatedTTS_GEN_DURATION: Total generation time
Notes
- Audio output is 24kHz PCM float32 format
- Audio buffering is approximately 160ms (4 frames × 40ms)
- Text preprocessing improves pronunciation quality
- Custom voices require pre-cached embeddings
- Messages are encoded using MessagePack format
- Connection automatically closes when iteration completes