Skip to main content

Overview

The SpeechToText class provides an async interface to the Unmute STT server via WebSocket. It streams audio data and receives real-time transcription along with Voice Activity Detection (VAD) pause predictions.

Class Definition

Constructor

Initializes the speech-to-text client.
str
default:"STT_SERVER"
URL of the STT server instance
float
default:"STT_DELAY_SEC"
Processing delay in seconds for the STT pipeline

Properties

pause_prediction

Exponential moving average of pause prediction scores (0-1 range). Higher values indicate more confidence that the user has paused speaking. Configured with:
  • attack_time: 0.01 seconds
  • release_time: 0.01 seconds
  • initial_value: 1.0

sent_samples

Total number of audio samples sent to the STT server.

current_time

Current processing time in seconds, accounting for the STT delay.

received_words

Count of words received from the STT server.

Core Methods

send_audio

Sends audio data to the STT server for transcription.
np.ndarray
required
1D numpy array of audio samples (float32 format)
Raises:
  • ValueError: If audio is not a 1D array
Notes:
  • Automatically converts audio to float32 if needed
  • Updates sent_samples counter
  • Increments metrics for monitoring

send_marker

Sends a marker message to the STT server for synchronization.
int
required
Unique marker identifier

start_up

Establishes WebSocket connection to the STT server and waits for ready signal. Raises:
  • MissingServiceAtCapacity: If STT server is at capacity
  • RuntimeError: If unexpected message type received during startup

shutdown

Closes the WebSocket connection and records session metrics. Metrics recorded:
  • Session duration
  • Total audio duration
  • Number of words transcribed

state

Returns the current WebSocket connection state. Returns: Literal["not_created", "connecting", "connected", "closing", "closed"]

Async Iterator

aiter

Iterates over messages received from the STT server. Yields:
  • STTWordMessage: Transcribed word with timing information
  • STTMarkerMessage: Synchronization marker
Example:

Message Types

STTWordMessage

Represents a transcribed word or phrase.

STTMarkerMessage

Synchronization marker echoed back from the server.

STTStepMessage

Processing step update with pause prediction scores.

STTErrorMessage

Error message from the server.

STTReadyMessage

Server ready signal.

Example Usage

Advanced Usage: Pause Detection

Metrics

The class automatically tracks the following metrics:
  • STT_ACTIVE_SESSIONS: Active transcription sessions
  • STT_SENT_FRAMES: Audio frames sent to server
  • STT_RECV_FRAMES: Processing steps received
  • STT_RECV_WORDS: Words transcribed
  • STT_TTFT: Time to first token (transcription)
  • STT_SESSION_DURATION: Total session duration
  • STT_AUDIO_DURATION: Total audio processed
  • STT_NUM_WORDS: Total words per session

Notes

  • Audio is expected at 24kHz sample rate
  • The STT pipeline has an inherent delay (configurable via delay_sec)
  • Pause predictions are smoothed using exponential moving average
  • First 12 processing steps are ignored for pause prediction to avoid initial noise
  • Connection is automatically closed when iteration completes