Skip to main content

Overview

Unmute uses a WebSocket-based protocol inspired by the OpenAI Realtime API for real-time voice conversations. The protocol handles:
  • Real-time audio streaming (bidirectional)
  • Voice conversation transcription
  • Session configuration
  • Error handling and debugging

Connection Setup

Endpoint

string
required
/v1/realtime
string
required
realtime (WebSocket subprotocol)
number
  • Development: 8000
  • Production: Routed through Traefik (HTTP port 80, HTTPS port 443)

Establishing Connection

The WebSocket connection is established using the realtime subprotocol. This subprotocol identifier is required for the client to connect properly. Frontend implementation (frontend/src/app/Unmute.tsx:91-97):
Backend implementation (unmute/main_websocket.py:310-314):

Message Structure

All messages are JSON-encoded with a common structure defined in unmute/openai_realtime_api_events.py. Every message inherits from BaseEvent which provides:
string
required
The event type identifier (e.g., "session.update", "response.audio.delta")
string
Unique identifier for the event (format: event_<21_random_chars>)

Client → Server Messages

Audio Input Streaming

Type: input_audio_buffer.append Streams real-time audio data from the microphone to the backend.
string
required
"input_audio_buffer.append"
string
required
Base64-encoded Opus audio data
Audio Format:
  • Codec: Opus
  • Sample Rate: 24kHz
  • Channels: Mono
  • Encoding: Base64-encoded bytes
Frontend Example (frontend/src/app/Unmute.tsx:100-110):
Backend Processing (unmute/main_websocket.py:460-477):

Session Configuration

Type: session.update Configures the voice character and conversation instructions. The backend will not start processing until it receives this message.
string
required
"session.update"
object
required
Session configuration object
object
Conversation instructions (Unmute extension)
string
Voice identifier for TTS
boolean
required
Whether to allow recording of the conversation
Frontend Example (frontend/src/app/Unmute.tsx:232-240):

Server → Client Messages

Audio Response Streaming

Type: response.audio.delta Streams generated speech audio to the frontend.
string
"response.audio.delta"
string
Unique event identifier
string
Base64-encoded Opus audio data chunk
Frontend Handling (frontend/src/app/Unmute.tsx:164-175):

Speech Transcription

Type: conversation.item.input_audio_transcription.delta Real-time transcription of user speech.
string
"conversation.item.input_audio_transcription.delta"
string
Transcribed text chunk
number
Start time of the transcription (Unmute extension)
Frontend Handling (frontend/src/app/Unmute.tsx:186-193):

Text Response Streaming

Type: response.text.delta Streams generated text responses for display or debugging.
string
"response.text.delta"
string
Text chunk from the LLM response
Frontend Handling (frontend/src/app/Unmute.tsx:194-202):

Speech Detection Events

Types:
  • input_audio_buffer.speech_started
  • input_audio_buffer.speech_stopped
Indicate when the user starts or stops speaking based on Voice Activity Detection (VAD).
string
"input_audio_buffer.speech_started" or "input_audio_buffer.speech_stopped"
These events are currently reported but not actively used in the Unmute frontend for UI feedback.

Response Status Updates

Type: response.created Indicates when the assistant starts generating a response.
string
"response.created"
object
Response metadata object
string
"realtime.response"
string
One of: "in_progress", "completed", "cancelled", "failed", "incomplete"
string
Voice identifier being used
array
Array of chat history objects

Error Handling

Type: error Communicates errors and warnings to the client.
string
"error"
object
Error details object
string
Error type (e.g., "warning", "fatal", "invalid_request_error")
string
Error code (optional)
string
Human-readable error message
string
Parameter that caused the error (optional)
object
Additional error details (Unmute extension)
Frontend Handling (frontend/src/app/Unmute.tsx:178-185):

Unmute-Specific Events

Additional Outputs

Type: unmute.additional_outputs Provides debugging information and additional outputs.
string
"unmute.additional_outputs"
any
Debug dictionary or additional output data

Text Delta Ready

Type: unmute.response.text.delta.ready Indicates that a text delta is ready to be sent.
string
"unmute.response.text.delta.ready"
string
Text delta content

Audio Delta Ready

Type: unmute.response.audio.delta.ready Indicates that audio samples are ready.
string
"unmute.response.audio.delta.ready"
number
Number of audio samples ready

VAD Interruption

Type: unmute.interrupted_by_vad Indicates that the VAD interrupted the response generation.
string
"unmute.interrupted_by_vad"

Connection Lifecycle

  1. Health Check: Frontend checks /v1/health endpoint before connecting
  2. WebSocket Connection: Establish connection with realtime protocol
  3. Session Setup: Send session.update with voice and instructions
    • Backend will not process audio until this is received
  4. Audio Streaming: Bidirectional real-time audio communication
    • Client sends input_audio_buffer.append messages
    • Server sends response.audio.delta messages
    • Transcription and text deltas flow concurrently
  5. Graceful Shutdown: Handle disconnection and cleanup
    • Frontend stops audio processing
    • Backend cleans up resources via UnmuteHandler.cleanup()

Implementation Details

Backend Message Loop

The backend uses two concurrent loops (unmute/main_websocket.py:391-403):
  • receive_loop: Receives messages from the WebSocket, processes audio, handles session updates
  • emit_loop: Sends messages to the WebSocket from the emit queue and handler
  • quest_manager: Manages processing quests and tasks

Audio Encoding/Decoding

Frontend:
  • Uses opus-recorder library for recording microphone input
  • Encodes to Opus at 24kHz sample rate
  • Uses Web Audio API decoder for playback
Backend:
  • Uses sphn.OpusStreamReader for decoding incoming audio
  • Uses sphn.OpusStreamWriter for encoding outgoing audio
  • Processes audio at 24kHz sample rate

Error Handling

The protocol includes comprehensive error handling:
  • Invalid JSON: Returns invalid_request_error with details
  • Validation Errors: Returns invalid_request_error with validation details
  • Service Unavailable: Returns fatal error and closes connection
  • Warnings: Logged but don’t disrupt the connection