> ## Documentation Index
> Fetch the complete documentation index at: https://mintlify.com/kyutai-labs/unmute/llms.txt
> Use this file to discover all available pages before exploring further.

# WebSocket Protocol

> Real-time bidirectional communication protocol for audio streaming and conversation management

## Overview

Unmute uses a WebSocket-based protocol inspired by the [OpenAI Realtime API](https://platform.openai.com/docs/api-reference/realtime) for real-time voice conversations. The protocol handles:

* Real-time audio streaming (bidirectional)
* Voice conversation transcription
* Session configuration
* Error handling and debugging

## Connection Setup

### Endpoint

<ParamField path="URL" type="string" required>
  `/v1/realtime`
</ParamField>

<ParamField path="Protocol" type="string" required>
  `realtime` (WebSocket subprotocol)
</ParamField>

<ParamField path="Port" type="number">
  * Development: `8000`
  * Production: Routed through Traefik (HTTP port 80, HTTPS port 443)
</ParamField>

### Establishing Connection

The WebSocket connection is established using the `realtime` subprotocol. This subprotocol identifier is required for the client to connect properly.

**Frontend implementation** (`frontend/src/app/Unmute.tsx:91-97`):

```typescript theme={null}
const { sendMessage, lastMessage, readyState } = useWebSocket(
  webSocketUrl || null,
  {
    protocols: ["realtime"],
  },
  shouldConnect,
);
```

**Backend implementation** (`unmute/main_websocket.py:310-314`):

```python theme={null}
# The `subprotocol` argument is important because the client specifies what
# protocol(s) it supports and OpenAI uses "realtime" as the value.
await websocket.accept(subprotocol="realtime")
```

## Message Structure

All messages are JSON-encoded with a common structure defined in `unmute/openai_realtime_api_events.py`. Every message inherits from `BaseEvent` which provides:

<ResponseField name="type" type="string" required>
  The event type identifier (e.g., `"session.update"`, `"response.audio.delta"`)
</ResponseField>

<ResponseField name="event_id" type="string">
  Unique identifier for the event (format: `event_<21_random_chars>`)
</ResponseField>

## Client → Server Messages

### Audio Input Streaming

**Type:** `input_audio_buffer.append`

Streams real-time audio data from the microphone to the backend.

<ParamField body="type" type="string" required>
  `"input_audio_buffer.append"`
</ParamField>

<ParamField body="audio" type="string" required>
  Base64-encoded Opus audio data
</ParamField>

**Audio Format:**

* **Codec:** Opus
* **Sample Rate:** 24kHz
* **Channels:** Mono
* **Encoding:** Base64-encoded bytes

**Frontend Example** (`frontend/src/app/Unmute.tsx:100-110`):

```typescript theme={null}
const onOpusRecorded = useCallback(
  (opus: Uint8Array) => {
    sendMessage(
      JSON.stringify({
        type: "input_audio_buffer.append",
        audio: base64EncodeOpus(opus),
      }),
    );
  },
  [sendMessage],
);
```

**Backend Processing** (`unmute/main_websocket.py:460-477`):

```python theme={null}
if isinstance(message, ora.InputAudioBufferAppend):
    opus_bytes = base64.b64decode(message.audio)
    if wait_for_first_opus:
        # Check for first packet bit
        if opus_bytes[5] & 2:
            wait_for_first_opus = False
        else:
            continue
    pcm = await asyncio.to_thread(opus_reader.append_bytes, opus_bytes)
    
    if pcm.size:
        await handler.receive((SAMPLE_RATE, pcm[np.newaxis, :]))
```

### Session Configuration

**Type:** `session.update`

Configures the voice character and conversation instructions. The backend will not start processing until it receives this message.

<ParamField body="type" type="string" required>
  `"session.update"`
</ParamField>

<ParamField body="session" type="object" required>
  Session configuration object
</ParamField>

<ParamField body="session.instructions" type="object">
  Conversation instructions (Unmute extension)
</ParamField>

<ParamField body="session.voice" type="string">
  Voice identifier for TTS
</ParamField>

<ParamField body="session.allow_recording" type="boolean" required>
  Whether to allow recording of the conversation
</ParamField>

**Frontend Example** (`frontend/src/app/Unmute.tsx:232-240`):

```typescript theme={null}
sendMessage(
  JSON.stringify({
    type: "session.update",
    session: {
      instructions: unmuteConfig.instructions,
      voice: unmuteConfig.voice,
      allow_recording: recordingConsent,
    },
  }),
);
```

## Server → Client Messages

### Audio Response Streaming

**Type:** `response.audio.delta`

Streams generated speech audio to the frontend.

<ResponseField name="type" type="string">
  `"response.audio.delta"`
</ResponseField>

<ResponseField name="event_id" type="string">
  Unique event identifier
</ResponseField>

<ResponseField name="delta" type="string">
  Base64-encoded Opus audio data chunk
</ResponseField>

**Frontend Handling** (`frontend/src/app/Unmute.tsx:164-175`):

```typescript theme={null}
if (data.type === "response.audio.delta") {
  const opus = base64DecodeOpus(data.delta);
  const ap = audioProcessor.current;
  if (!ap) return;

  ap.decoder.postMessage(
    {
      command: "decode",
      pages: opus,
    },
    [opus.buffer],
  );
}
```

### Speech Transcription

**Type:** `conversation.item.input_audio_transcription.delta`

Real-time transcription of user speech.

<ResponseField name="type" type="string">
  `"conversation.item.input_audio_transcription.delta"`
</ResponseField>

<ResponseField name="delta" type="string">
  Transcribed text chunk
</ResponseField>

<ResponseField name="start_time" type="number">
  Start time of the transcription (Unmute extension)
</ResponseField>

**Frontend Handling** (`frontend/src/app/Unmute.tsx:186-193`):

```typescript theme={null}
else if (data.type === "conversation.item.input_audio_transcription.delta") {
  // Transcription of the user speech
  setRawChatHistory((prev) => [
    ...prev,
    { role: "user", content: data.delta },
  ]);
}
```

### Text Response Streaming

**Type:** `response.text.delta`

Streams generated text responses for display or debugging.

<ResponseField name="type" type="string">
  `"response.text.delta"`
</ResponseField>

<ResponseField name="delta" type="string">
  Text chunk from the LLM response
</ResponseField>

**Frontend Handling** (`frontend/src/app/Unmute.tsx:194-202`):

```typescript theme={null}
else if (data.type === "response.text.delta") {
  setRawChatHistory((prev) => [
    ...prev,
    // The TTS doesn't include spaces in its messages, so add a leading space
    { role: "assistant", content: " " + data.delta },
  ]);
}
```

### Speech Detection Events

**Types:**

* `input_audio_buffer.speech_started`
* `input_audio_buffer.speech_stopped`

Indicate when the user starts or stops speaking based on Voice Activity Detection (VAD).

<ResponseField name="type" type="string">
  `"input_audio_buffer.speech_started"` or `"input_audio_buffer.speech_stopped"`
</ResponseField>

<Note>
  These events are currently reported but not actively used in the Unmute frontend for UI feedback.
</Note>

### Response Status Updates

**Type:** `response.created`

Indicates when the assistant starts generating a response.

<ResponseField name="type" type="string">
  `"response.created"`
</ResponseField>

<ResponseField name="response" type="object">
  Response metadata object
</ResponseField>

<ResponseField name="response.object" type="string">
  `"realtime.response"`
</ResponseField>

<ResponseField name="response.status" type="string">
  One of: `"in_progress"`, `"completed"`, `"cancelled"`, `"failed"`, `"incomplete"`
</ResponseField>

<ResponseField name="response.voice" type="string">
  Voice identifier being used
</ResponseField>

<ResponseField name="response.chat_history" type="array">
  Array of chat history objects
</ResponseField>

### Error Handling

**Type:** `error`

Communicates errors and warnings to the client.

<ResponseField name="type" type="string">
  `"error"`
</ResponseField>

<ResponseField name="error" type="object">
  Error details object
</ResponseField>

<ResponseField name="error.type" type="string">
  Error type (e.g., `"warning"`, `"fatal"`, `"invalid_request_error"`)
</ResponseField>

<ResponseField name="error.code" type="string">
  Error code (optional)
</ResponseField>

<ResponseField name="error.message" type="string">
  Human-readable error message
</ResponseField>

<ResponseField name="error.param" type="string">
  Parameter that caused the error (optional)
</ResponseField>

<ResponseField name="error.details" type="object">
  Additional error details (Unmute extension)
</ResponseField>

**Frontend Handling** (`frontend/src/app/Unmute.tsx:178-185`):

```typescript theme={null}
else if (data.type === "error") {
  if (data.error.type === "warning") {
    console.warn(`Warning from server: ${data.error.message}`, data);
  } else {
    console.error(`Error from server: ${data.error.message}`, data);
    setErrors((prev) => [...prev, makeErrorItem(data.error.message)]);
  }
}
```

### Unmute-Specific Events

#### Additional Outputs

**Type:** `unmute.additional_outputs`

Provides debugging information and additional outputs.

<ResponseField name="type" type="string">
  `"unmute.additional_outputs"`
</ResponseField>

<ResponseField name="args" type="any">
  Debug dictionary or additional output data
</ResponseField>

#### Text Delta Ready

**Type:** `unmute.response.text.delta.ready`

Indicates that a text delta is ready to be sent.

<ResponseField name="type" type="string">
  `"unmute.response.text.delta.ready"`
</ResponseField>

<ResponseField name="delta" type="string">
  Text delta content
</ResponseField>

#### Audio Delta Ready

**Type:** `unmute.response.audio.delta.ready`

Indicates that audio samples are ready.

<ResponseField name="type" type="string">
  `"unmute.response.audio.delta.ready"`
</ResponseField>

<ResponseField name="number_of_samples" type="number">
  Number of audio samples ready
</ResponseField>

#### VAD Interruption

**Type:** `unmute.interrupted_by_vad`

Indicates that the VAD interrupted the response generation.

<ResponseField name="type" type="string">
  `"unmute.interrupted_by_vad"`
</ResponseField>

## Connection Lifecycle

1. **Health Check**: Frontend checks `/v1/health` endpoint before connecting
   ```typescript theme={null}
   const response = await fetch(`${backendServerUrl}/v1/health`);
   const data = await response.json();
   // Check data.ok, data.tts_up, data.stt_up, data.llm_up
   ```

2. **WebSocket Connection**: Establish connection with `realtime` protocol

3. **Session Setup**: Send `session.update` with voice and instructions
   * Backend will not process audio until this is received

4. **Audio Streaming**: Bidirectional real-time audio communication
   * Client sends `input_audio_buffer.append` messages
   * Server sends `response.audio.delta` messages
   * Transcription and text deltas flow concurrently

5. **Graceful Shutdown**: Handle disconnection and cleanup
   * Frontend stops audio processing
   * Backend cleans up resources via `UnmuteHandler.cleanup()`

## Implementation Details

### Backend Message Loop

The backend uses two concurrent loops (`unmute/main_websocket.py:391-403`):

```python theme={null}
async with asyncio.TaskGroup() as tg:
    tg.create_task(
        receive_loop(websocket, handler, emit_queue), name="receive_loop()"
    )
    tg.create_task(
        emit_loop(websocket, handler, emit_queue), name="emit_loop()"
    )
    tg.create_task(handler.quest_manager.wait(), name="quest_manager.wait()")
```

* **receive\_loop**: Receives messages from the WebSocket, processes audio, handles session updates
* **emit\_loop**: Sends messages to the WebSocket from the emit queue and handler
* **quest\_manager**: Manages processing quests and tasks

### Audio Encoding/Decoding

**Frontend:**

* Uses `opus-recorder` library for recording microphone input
* Encodes to Opus at 24kHz sample rate
* Uses Web Audio API decoder for playback

**Backend:**

* Uses `sphn.OpusStreamReader` for decoding incoming audio
* Uses `sphn.OpusStreamWriter` for encoding outgoing audio
* Processes audio at 24kHz sample rate

## Error Handling

The protocol includes comprehensive error handling:

* **Invalid JSON**: Returns `invalid_request_error` with details
* **Validation Errors**: Returns `invalid_request_error` with validation details
* **Service Unavailable**: Returns `fatal` error and closes connection
* **Warnings**: Logged but don't disrupt the connection

## Related Documentation

* [WebRTC](/architecture/webrtc) - Audio processing and streaming details
* [System Architecture](/architecture/overview) - Overall system design


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.