Skip to main content

Overview

The TextToSpeech class provides an async interface to the Unmute TTS server. It streams text input and receives synthesized audio in real-time with precise timing control.

Class Definition

Constructor

Initializes the text-to-speech client.
str
default:"TTS_SERVER"
URL of the TTS server instance
Recorder | None
default:"None"
Optional recorder instance for logging TTS events
Callable[[], float] | None
default:"None"
Optional callback function to get current time (for synchronization)
str | None
default:"None"
Voice identifier. Can be a preset voice name or “custom:” prefixed for custom voice embeddings

Properties

voice

The currently configured voice identifier.

received_samples

Total number of audio samples received from the TTS server.

received_samples_yielded

Number of audio samples that have been yielded to the consumer (after buffering).

Core Methods

send

Sends text or a message to the TTS server for synthesis.
str | TTSClientMessage
required
Text string or structured message to synthesize. Strings are automatically preprocessed to remove unpronounceable characters.
Notes:
  • Empty strings are ignored
  • String messages are preprocessed by prepare_text_for_tts()
  • TTSClientTextMessage bypasses preprocessing

start_up

Establishes WebSocket connection to the TTS server and configures the voice. Raises:
  • MissingServiceAtCapacity: If TTS server is at capacity
  • AssertionError: If connection setup fails
Notes:
  • Sends custom voice embeddings if voice starts with “custom:”
  • Waits for TTSReadyMessage before considering startup complete

shutdown

Closes the WebSocket connection and records session metrics. Metrics recorded:
  • Active sessions count
  • Total audio duration
  • Generation duration

state

Returns the current WebSocket connection state. Returns: Literal["not_created", "connecting", "connected", "closing", "closed"]

Async Iterator

aiter

Iterates over synthesized audio and text alignment messages from the TTS server. Yields:
  • TTSAudioMessage: Synthesized audio chunks
  • TTSTextMessage: Text alignment with timing information
Notes:
  • Audio is buffered and released with AUDIO_BUFFER_SEC delay (approx. 160ms)
  • Text messages are synchronized with audio playback timing
Example:

Message Types

Client Messages (sent to server)

TTSClientTextMessage

Text to synthesize.

TTSClientVoiceMessage

Custom voice embeddings.

TTSClientEosMessage

End of stream signal indicating no more text will be sent.

Server Messages (received from server)

TTSAudioMessage

Synthesized audio chunk in PCM float32 format at 24kHz.

TTSTextMessage

Text alignment information with timing.

TTSErrorMessage

Error message from the server.

TTSReadyMessage

Server ready signal.

Helper Functions

prepare_text_for_tts

Preprocesses text for better TTS pronunciation. Transformations:
  • Strips leading/trailing whitespace
  • Removes unpronounceable characters: *, _, `
  • Normalizes curly quotes to straight quotes
  • Removes spaces around colons
Example:

Example Usage

Basic Synthesis

Streaming Synthesis

Custom Voice

Configuration

TtsStreamingQuery

Query parameters sent to the TTS server during connection.

Metrics

The class automatically tracks:
  • TTS_SESSIONS: Total TTS sessions
  • TTS_ACTIVE_SESSIONS: Active synthesis sessions
  • TTS_SENT_FRAMES: Text chunks sent
  • TTS_RECV_FRAMES: Audio chunks received
  • TTS_RECV_WORDS: Words with timing info received
  • TTS_TTFT: Time to first token (audio)
  • TTS_AUDIO_DURATION: Total audio generated
  • TTS_GEN_DURATION: Total generation time

Notes

  • Audio output is 24kHz PCM float32 format
  • Audio buffering is approximately 160ms (4 frames × 40ms)
  • Text preprocessing improves pronunciation quality
  • Custom voices require pre-cached embeddings
  • Messages are encoded using MessagePack format
  • Connection automatically closes when iteration completes