> ## Documentation Index
> Fetch the complete documentation index at: https://mintlify.com/kyutai-labs/unmute/llms.txt
> Use this file to discover all available pages before exploring further.

# Welcome to Unmute

> Transform text LLMs into voice-enabled conversational AI with real-time speech-to-text and text-to-speech

<img className="block dark:hidden" src="https://mintlify.s3.us-west-1.amazonaws.com/kyutai-labs-unmute/images/hero-light.svg" alt="Unmute Hero Light" />

<img className="hidden dark:block" src="https://mintlify.s3.us-west-1.amazonaws.com/kyutai-labs-unmute/images/hero-dark.svg" alt="Unmute Hero Dark" />

## What is Unmute?

Unmute is a complete system that allows text LLMs to listen and speak by wrapping them in Kyutai's state-of-the-art speech models. Experience natural voice conversations with any text LLM you like.

<CardGroup cols={2}>
  <Card title="Low Latency" icon="bolt" iconType="duotone">
    Optimized STT and TTS models deliver \~450ms response time in production
  </Card>

  <Card title="Any LLM" icon="brain" iconType="duotone">
    Works with any OpenAI-compatible LLM server - vLLM, Ollama, OpenAI, Mistral
  </Card>

  <Card title="Real-time Streaming" icon="wave-pulse" iconType="duotone">
    Stream audio and text bidirectionally over WebSocket connections
  </Card>

  <Card title="Custom Voices" icon="microphone-lines" iconType="duotone">
    Clone voices from audio samples and customize character personalities
  </Card>
</CardGroup>

## How It Works

Unmute orchestrates multiple AI services to create seamless voice conversations:

```mermaid theme={null}
graph LR
    UB[User Browser]
    UB --> B(Backend)
    UB --> F(Frontend)
    B --> STT(Speech-to-text)
    B --> LLM(LLM)
    B --> TTS(Text-to-speech)
```

<Steps>
  <Step title="User Speaks">
    Audio is captured in the browser and streamed to the backend via WebSocket
  </Step>

  <Step title="Speech to Text">
    Kyutai STT transcribes speech in real-time with 6-token delay for low latency
  </Step>

  <Step title="LLM Responds">
    Your chosen text LLM generates a response based on the conversation history
  </Step>

  <Step title="Text to Speech">
    Kyutai TTS converts the response to speech with customizable voices
  </Step>

  <Step title="Audio Playback">
    Generated audio streams back to the browser for immediate playback
  </Step>
</Steps>

## Key Components

<AccordionGroup>
  <Accordion title="Backend (Python/FastAPI)" icon="server">
    The core orchestration layer that manages:

    * WebSocket connections with OpenAI Realtime API compatibility
    * Conversation state and chat history
    * Service coordination between STT, LLM, and TTS
    * Voice management and cloning

    Located in `unmute/main_websocket.py`
  </Accordion>

  <Accordion title="Speech-to-Text (Kyutai STT 1B)" icon="ear">
    Low-latency speech recognition:

    * Model: Kyutai STT 1B (English/French)
    * Memory: \~2.5GB VRAM
    * Latency: 6-token delay for real-time transcription
    * Architecture: Transformer with 16 layers, 2048 d\_model
  </Accordion>

  <Accordion title="Text-to-Speech (Kyutai TTS 1.6B)" icon="volume">
    Natural voice synthesis:

    * Model: Kyutai TTS 1.6B (English/French)
    * Memory: \~5.3GB VRAM
    * Features: Voice cloning from audio samples
    * Voices: 100+ community-donated voices available
  </Accordion>

  <Accordion title="LLM (Configurable)" icon="brain">
    Any OpenAI-compatible text model:

    * Default: Llama 3.2 1B Instruct (16GB config)
    * Recommended: Mistral Small 3.2 24B, Gemma 3 12B
    * Memory: 6.1GB VRAM minimum (model dependent)
    * Hosted via vLLM, Ollama, or external APIs
  </Accordion>

  <Accordion title="Frontend (Next.js)" icon="browser">
    Modern web interface:

    * Real-time audio capture and playback
    * WebSocket communication
    * Subtitles and debug mode (press 'S' and 'D')
    * Character selection and voice customization
  </Accordion>
</AccordionGroup>

## Deployment Options

<CardGroup cols={3}>
  <Card title="Docker Compose" icon="docker" href="/quickstart">
    **Recommended** - Single GPU, single machine, very easy setup
  </Card>

  <Card title="Dockerless" icon="terminal" href="/deployment/dockerless">
    Manual service management for 1-3 GPUs across 1-5 machines
  </Card>

  <Card title="Docker Swarm" icon="server" href="/deployment/swarm">
    Production scaling for 1-100 GPUs (used by unmute.sh)
  </Card>
</CardGroup>

## Model Architecture

### STT Model Configuration

The speech-to-text model uses a transformer architecture optimized for streaming:

```toml theme={null}
[modules.asr.model.transformer]
d_model = 2048
num_heads = 16
num_layers = 16
dim_feedforward = 8192
causal = true
max_seq_len = 40960
asr_delay_in_tokens = 6  # Low latency configuration
```

### TTS Model Configuration

The text-to-speech model supports voice cloning and multiple languages:

```toml theme={null}
[modules.tts_py.py]
cfg_coef = 2.0          # Classifier-free guidance coefficient
n_q = 24                 # Number of quantization levels
padding_between = 1      # Token padding for prosody
```

## Performance Metrics

<Note>
  On unmute.sh with separate GPUs for each service:

  * **TTS Latency**: \~450ms (vs. \~750ms on single L40S GPU)
  * **Max Concurrent Users**: 4 per backend instance (GIL constraint)
  * **Model Memory**: 16GB VRAM total (STT: 2.5GB, TTS: 5.3GB, LLM: 6.1GB+)
</Note>

## Try It Now

Experience Unmute live at [unmute.sh](https://unmute.sh) or deploy your own instance:

<CardGroup cols={2}>
  <Card title="Quick Start" icon="rocket" href="/quickstart">
    Get Unmute running locally in 5 minutes with Docker Compose
  </Card>

  <Card title="Requirements" icon="list-check" href="/requirements">
    Check hardware, software, and configuration prerequisites
  </Card>
</CardGroup>

## Research & Development

<CardGroup cols={2}>
  <Card title="Research Paper" icon="file-lines" href="https://arxiv.org/pdf/2509.08753">
    Read the academic paper on delayed streams modeling
  </Card>

  <Card title="Kyutai Models" icon="github" href="https://github.com/kyutai-labs/delayed-streams-modeling">
    Use Kyutai STT or TTS independently in your projects
  </Card>
</CardGroup>

<Warning>
  Unmute requires:

  * **GPU**: CUDA-capable with 16GB+ VRAM
  * **Architecture**: x86\_64 only (no aarch64 support)
  * **OS**: Linux or Windows with WSL (no native Windows or macOS)
</Warning>

## Community

Unmute includes 100+ voices donated by the community through the [Unmute Voice Donation Project](https://unmute.sh/voice-donation) (June 2025 - February 2026). These voices are available for use with Kyutai TTS and other open-source TTS models.

Browse available voices in the [voice repository](https://huggingface.co/kyutai/tts-voices).


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.