> ## Documentation Index
> Fetch the complete documentation index at: https://mintlify.com/kyutai-labs/unmute/llms.txt
> Use this file to discover all available pages before exploring further.

# Data Flow & Timing

> Detailed data flow and timing characteristics of the Unmute system

## Conversation Flow

A typical conversation in Unmute follows this sequence:

```mermaid theme={null}
sequenceDiagram
    participant User
    participant Frontend
    participant Backend
    participant STT
    participant LLM
    participant TTS
    
    User->>Frontend: Speak into microphone
    Frontend->>Backend: input_audio_buffer.append (Opus)
    Backend->>STT: Audio PCM frames
    STT->>Backend: Word deltas + VAD scores
    Backend->>Frontend: conversation.item.input_audio_transcription.delta
    
    Note over Backend,STT: Pause detected (VAD > 0.6)
    Backend->>STT: Flush with zeros
    
    Backend->>LLM: Chat completion request
    LLM->>Backend: Text deltas (streaming)
    Backend->>Frontend: response.text.delta
    Backend->>TTS: Text words
    TTS->>Backend: Audio PCM + Text timing
    Backend->>Frontend: response.audio.delta (Opus)
    Frontend->>User: Play audio
```

## Detailed Message Flow

### 1. Connection Setup

```mermaid theme={null}
sequenceDiagram
    participant Frontend
    participant Backend
    participant STT
    participant TTS
    
    Frontend->>Backend: GET /v1/health
    Backend-->>Frontend: {tts_up, stt_up, llm_up}
    
    Frontend->>Backend: WebSocket /v1/realtime (subprotocol: realtime)
    Backend-->>Frontend: Connection accepted
    
    Frontend->>Backend: session.update {instructions, voice}
    Backend->>STT: Connect WebSocket
    STT-->>Backend: Ready message
    Backend-->>Frontend: session.updated
```

**Timing**: Connection setup typically takes 100-300ms

### 2. User Speaking

```mermaid theme={null}
sequenceDiagram
    participant User
    participant Frontend
    participant Backend
    participant STT
    
    loop Every 20ms
        User->>Frontend: Microphone audio
        Frontend->>Frontend: Encode to Opus
        Frontend->>Backend: input_audio_buffer.append (base64)
        Backend->>Backend: Decode Opus to PCM
        Backend->>STT: Audio {type: Audio, pcm: [...]}
    end
    
    loop Every 20ms
        STT->>Backend: Step {prs: [pause_scores]}
        STT->>Backend: Word {text, start_time}
        Backend->>Frontend: input_audio_transcription.delta
    end
    
    Note over Backend: pause_prediction > 0.6
    Backend->>Frontend: input_audio_buffer.speech_stopped
```

**Key Timing**:

* Audio frame: 20ms (480 samples @ 24kHz)
* STT delay: \~2.5s (configurable via `STT_DELAY_SEC`)
* Time to first word: \~50-100ms after first audio frame

### 3. LLM Response Generation

```mermaid theme={null}
sequenceDiagram
    participant Backend
    participant LLM
    participant TTS
    participant Frontend
    
    Backend->>Backend: Flush STT with zeros
    Note over Backend: Wait for STT delay (~2.5s)
    
    Backend->>TTS: Connect WebSocket
    TTS-->>Backend: Ready message
    
    Backend->>LLM: POST /v1/chat/completions (stream=true)
    Backend->>Frontend: response.created
    
    LLM->>Backend: First token
    Note over Backend: VLLM TTFT metric recorded
    Backend->>Backend: Rechunk to word boundaries
    Backend->>Frontend: response.text.delta.ready
    Backend->>TTS: Text {text: "word"}
    
    loop For each word
        LLM->>Backend: Token delta
        Backend->>Backend: Accumulate to word
        Backend->>TTS: Text {text: "word"}
        Backend->>Frontend: response.text.delta.ready
    end
    
    Backend->>TTS: Eos
    Backend->>Frontend: response.text.done
```

**Key Timing**:

* Time to first token (TTFT): 100-500ms (depends on model and GPU)
* LLM streaming: \~50-200ms per word
* Temperature: 0.7 (first message), 0.3 (subsequent)

### 4. TTS Audio Generation

```mermaid theme={null}
sequenceDiagram
    participant TTS
    participant Backend
    participant Frontend
    
    loop For each word from LLM
        Note over TTS: Generate audio for word
        TTS->>Backend: Text {text, start_s, stop_s}
        TTS->>Backend: Audio {pcm: [samples]}
        
        Backend->>Backend: Queue with timestamp
        Note over Backend: Wait for correct playback time
        Backend->>Backend: Encode to Opus
        Backend->>Frontend: response.audio.delta (base64)
    end
    
    TTS-->>Backend: Connection closed
    Backend->>Frontend: response.audio.done
```

**Key Timing**:

* TTS TTFT: 200-750ms (depends on GPU and setup)
* Audio buffer: 80ms ahead of real-time (`AUDIO_BUFFER_SEC = 4 * 20ms`)
* Text/audio synchronization: Text released at `start_s` timestamp

### 5. User Interruption

```mermaid theme={null}
sequenceDiagram
    participant User
    participant Frontend
    participant Backend
    participant STT
    participant LLM
    participant TTS
    
    Note over Backend: Bot is speaking
    User->>Frontend: Start speaking
    Frontend->>Backend: input_audio_buffer.append
    Backend->>STT: Audio frames
    
    alt STT detects word
        STT->>Backend: Word {text: "..."}
        Backend->>Backend: interrupt_bot()
    else VAD detects speech
        Note over Backend: pause_prediction < 0.4
        Backend->>Backend: interrupt_bot()
    end
    
    Backend->>Backend: Cancel LLM task
    Backend->>Backend: Cancel TTS task
    Backend->>Backend: Clear output queue
    Backend->>Backend: Add INTERRUPTION_CHAR (—)
    Backend->>Frontend: unmute.interrupted_by_vad
    
    Note over Backend: Transition to user_speaking
```

**Key Features**:

* Interruption cooldown: First 3s (VAD-based only)
* VAD threshold: pause\_prediction \< 0.4
* Interruption character: em-dash (—) stripped from LLM context

## Conversation States

The backend manages conversation flow through three states:

```mermaid theme={null}
stateDiagram-v2
    [*] --> waiting_for_user: Initial connection
    
    waiting_for_user --> user_speaking: STT word received
    user_speaking --> waiting_for_user: Pause detected (VAD > 0.6)
    waiting_for_user --> bot_speaking: Generate response
    
    bot_speaking --> user_speaking: Interruption
    bot_speaking --> waiting_for_user: TTS complete
    
    waiting_for_user --> [*]: Long silence (7s)
    bot_speaking --> [*]: Bot says "Bye!"
```

**State Logic** (`unmute/llm/chatbot.py:21`):

* `waiting_for_user`: Last message is empty user message
* `user_speaking`: Last message is non-empty user message
* `bot_speaking`: Last message is assistant message

## Timing Characteristics

### Latency Breakdown (Typical)

| Stage | Duration | Notes |
| - | - | - |
| User finishes speaking | 0ms | Baseline |
| Pause detection | 100-500ms | VAD threshold crossing |
| STT flush | 2500ms | Configurable delay |
| LLM TTFT | 100-500ms | First word generated |
| TTS TTFT | 200-750ms | First audio chunk |
| **Total to first audio** | **\~3-4s** | End-to-end latency |

### Throughput

* **Backend**: 4 concurrent sessions per instance (configurable via `MAX_CLIENTS`)
* **STT**: Multiple concurrent streams (capacity-based)
* **TTS**: Multiple concurrent streams (capacity-based)
* **LLM**: Batch processing via VLLM

### Memory Usage

* **LLM**: 6.1 GB VRAM (Llama 3.2 1B)
* **STT**: 2.5 GB VRAM
* **TTS**: 5.3 GB VRAM
* **Total**: \~14 GB VRAM (single-GPU setup)

## Real-Time Constraints

### Audio Frames

Every 20ms:

1. Frontend captures 480 samples
2. Encodes to Opus
3. Sends over WebSocket
4. Backend decodes and forwards to STT
5. STT processes and returns word/VAD

### Output Synchronization

The backend carefully manages audio release timing:

* TTS generates audio faster than real-time (RTF > 1.0)
* Audio queued with timestamps (`RealtimeQueue`)
* Released exactly at `start_s - AUDIO_BUFFER_SEC`
* Prevents stuttering and maintains sync

### Frame Time Constant

```python theme={null}
FRAME_TIME_SEC = 0.02  # 20ms
SAMPLE_RATE = 24000    # 24kHz
SAMPLES_PER_FRAME = 480  # 20ms * 24kHz
```

Defined in `unmute/kyutai_constants.py:19-21`

## Error Handling

### Connection Failures

* Retry with exponential backoff (50ms → 75ms → 112ms...)
* Max retries: 5
* User notification after exhaustion

### Service Capacity

* STT/TTS return `Error` message when at capacity
* Backend catches `MissingServiceAtCapacity` exception
* Returns error to frontend with retry suggestion

### WebSocket Disconnects

* Clean shutdown via `CloseStream`
* Metrics recorded before disconnect
* Resources released (STT, TTS, LLM connections)

## Metrics Collection

Prometheus metrics tracked throughout the flow:

**Session Metrics**:

* `unmute_sessions_total`
* `unmute_active_sessions`
* `unmute_session_duration_seconds`

**STT Metrics**:

* `unmute_stt_ttft_seconds` (time to first token)
* `unmute_stt_sent_frames_total`
* `unmute_stt_recv_words_total`

**LLM Metrics**:

* `unmute_vllm_ttft_seconds`
* `unmute_vllm_request_length_words`
* `unmute_vllm_reply_length_words`

**TTS Metrics**:

* `unmute_tts_ttft_seconds`
* `unmute_tts_audio_duration_seconds`
* `unmute_tts_gen_duration_seconds`

See `unmute/metrics.py` for complete list.


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.