> ## Documentation Index
> Fetch the complete documentation index at: https://mintlify.com/kyutai-labs/unmute/llms.txt
> Use this file to discover all available pages before exploring further.

# Requirements

> Hardware, software, and configuration prerequisites for running Unmute

## Hardware Requirements

### GPU Requirements

<Warning>
  Unmute requires a CUDA-capable NVIDIA GPU. CPU-only deployment is not supported.
</Warning>

<Tabs>
  <Tab title="Single GPU (Minimum)">
    **VRAM**: 16GB minimum

    **Example GPUs**:

    * NVIDIA RTX 4090 (24GB)
    * NVIDIA RTX 3090 (24GB)
    * NVIDIA L40S (48GB)
    * NVIDIA A100 (40GB/80GB)
    * NVIDIA RTX A6000 (48GB)

    **Memory Breakdown**:

    * **STT**: 2.5GB VRAM
    * **TTS**: 5.3GB VRAM
    * **LLM**: 6.1GB+ VRAM (model dependent)
    * **Overhead**: \~2GB for CUDA and buffers

    <Info>
      With 16GB VRAM, use Llama 3.2 1B with `--gpu-memory-utilization=0.4` and `--max-model-len=1536`
    </Info>
  </Tab>

  <Tab title="Multi-GPU (Recommended)">
    **Configuration**: 3 GPUs (one per service)

    **Benefits**:

    * **Reduced Latency**: TTS latency drops from \~750ms to \~450ms
    * **Better Isolation**: Services don't compete for GPU memory
    * **Higher Throughput**: Serve more concurrent users

    **GPU Allocation**:

    * **GPU 0**: STT (2.5GB VRAM)
    * **GPU 1**: TTS (5.3GB VRAM)
    * **GPU 2**: LLM (6.1GB+ VRAM)

    <Tip>
      You can use smaller GPUs for STT/TTS (e.g., RTX 3060 12GB) and reserve a larger GPU for the LLM
    </Tip>
  </Tab>

  <Tab title="Production (Swarm)">
    **Scale**: 1 to \~100 GPUs across multiple machines

    **Architecture**:

    * Load balancer (Traefik)
    * Multiple backend replicas (4 users per instance)
    * Shared STT/TTS/LLM pools
    * Distributed metrics and monitoring

    **Example unmute.sh Setup**:

    * Multiple L40S GPUs (48GB each)
    * Separate GPUs for STT, TTS, and LLM
    * Redis for service discovery
    * Prometheus + Grafana monitoring

    See [Docker Swarm documentation](/deployment/swarm) for details.
  </Tab>
</Tabs>

### Architecture Requirements

<Card title="x86_64 Only" icon="microchip">
  Unmute is built for **x86\_64 (AMD64)** architecture.

  **Not Supported**:

  * ARM64 (aarch64) - No support planned
  * Apple Silicon (M1/M2/M3) - No support planned

  This is due to dependencies on CUDA and compiled Rust binaries.
</Card>

### System Memory

<Info>
  **Recommended RAM**: 16GB+ system memory

  While models run on GPU, the host needs memory for:

  * Docker containers and Python processes
  * Model loading and initialization
  * Audio buffering and WebSocket connections
</Info>

## Software Requirements

### Operating System

<Tabs>
  <Tab title="Linux (Recommended)">
    **Supported**:

    * Ubuntu 20.04+
    * Debian 11+
    * Fedora 36+
    * Arch Linux
    * Any modern Linux distribution with Docker support

    **Why Linux?**

    * Best NVIDIA driver support
    * Native Docker integration
    * Used by Kyutai for development and production
  </Tab>

  <Tab title="Windows (WSL)">
    **Requirements**:

    * Windows 10 version 2004+ (Build 19041+) or Windows 11
    * WSL 2 ([installation guide](https://ubuntu.com/desktop/wsl))
    * NVIDIA GPU drivers for Windows
    * CUDA support in WSL

    **Setup Steps**:

    ```bash theme={null}
    # Install WSL 2 with Ubuntu
    wsl --install -d Ubuntu

    # Install Docker Desktop for Windows with WSL 2 backend
    # Download from docker.com

    # Verify GPU access in WSL
    wsl
    nvidia-smi
    ```

    <Warning>
      **Native Windows Not Supported**: Unmute cannot run directly on Windows without WSL ([issue #84](https://github.com/kyutai-labs/unmute/issues/84))
    </Warning>
  </Tab>

  <Tab title="macOS (Not Supported)">
    **Status**: Not supported ([issue #74](https://github.com/kyutai-labs/unmute/issues/74))

    **Reasons**:

    * No NVIDIA GPU support on Apple Silicon
    * CUDA unavailable on macOS
    * Rust bindings incompatible with Metal/MPS

    **Alternatives**:

    * Deploy on a Linux cloud instance
    * Use a remote Linux machine via SSH
    * Contribute to add Metal/MPS support (community effort)
  </Tab>
</Tabs>

### Docker (Docker Compose)

<Steps>
  <Step title="Install Docker">
    Follow the [official Docker installation guide](https://docs.docker.com/engine/install/) for your platform.

    **Linux** (Ubuntu/Debian):

    ```bash theme={null}
    curl -fsSL https://get.docker.com -o get-docker.sh
    sudo sh get-docker.sh
    sudo usermod -aG docker $USER
    ```

    **Windows**: Install [Docker Desktop](https://www.docker.com/products/docker-desktop) with WSL 2 backend
  </Step>

  <Step title="Install Docker Compose">
    Docker Compose is included with Docker Desktop. On Linux:

    ```bash theme={null}
    sudo apt-get install docker-compose-plugin
    ```

    Verify installation:

    ```bash theme={null}
    docker compose version
    ```
  </Step>

  <Step title="Install NVIDIA Container Toolkit">
    Required for GPU access from Docker containers:

    ```bash theme={null}
    # Ubuntu/Debian
    curl -fsSL https://nvidia.github.io/libnvidia-container/gpgkey | \
      sudo gpg --dearmor -o /usr/share/keyrings/nvidia-container-toolkit-keyring.gpg

    curl -s -L https://nvidia.github.io/libnvidia-container/stable/deb/nvidia-container-toolkit.list | \
      sed 's#deb https://#deb [signed-by=/usr/share/keyrings/nvidia-container-toolkit-keyring.gpg] https://#g' | \
      sudo tee /etc/apt/sources.list.d/nvidia-container-toolkit.list

    sudo apt-get update
    sudo apt-get install -y nvidia-container-toolkit
    sudo systemctl restart docker
    ```

    See [NVIDIA's official guide](https://docs.nvidia.com/datacenter/cloud-native/container-toolkit/latest/install-guide.html) for other distributions.
  </Step>

  <Step title="Verify GPU Access">
    Test that Docker can access your GPU:

    ```bash theme={null}
    sudo docker run --rm --runtime=nvidia --gpus all ubuntu nvidia-smi
    ```

    <Accordion title="Expected output">
      ```
      +-----------------------------------------------------------------------------+
      | NVIDIA-SMI 535.129.03   Driver Version: 535.129.03   CUDA Version: 12.2     |
      |-------------------------------+----------------------+----------------------+
      | GPU  Name        Persistence-M| Bus-Id        Disp.A | Volatile Uncorr. ECC |
      | Fan  Temp  Perf  Pwr:Usage/Cap|         Memory-Usage | GPU-Util  Compute M. |
      |===============================+======================+======================|
      |   0  NVIDIA L40S         Off  | 00000000:01:00.0 Off |                    0 |
      | N/A   32C    P0    35W / 350W |      0MiB / 46068MiB |      0%      Default |
      +-------------------------------+----------------------+----------------------+
      ```
    </Accordion>
  </Step>
</Steps>

### Dockerless Setup (Alternative)

<Accordion title="Software Requirements for Dockerless Deployment" icon="terminal">
  If you prefer to run services without Docker:

  <Note>
    This is more complex and requires manual dependency management. Docker Compose is recommended.
  </Note>

  **Required Tools**:

  * **uv**: Python package manager
    ```bash theme={null}
    curl -LsSf https://astral.sh/uv/install.sh | sh
    ```

  * **cargo**: Rust toolchain (for STT/TTS servers)
    ```bash theme={null}
    curl https://sh.rustup.rs -sSf | sh
    ```

  * **pnpm**: Node package manager (for frontend)
    ```bash theme={null}
    curl -fsSL https://get.pnpm.io/install.sh | sh -
    ```

  * **CUDA 12.1**: For Rust processes
    * Install via conda or from [NVIDIA website](https://developer.nvidia.com/cuda-downloads)

  **Start Services**:

  ```bash theme={null}
  # In separate terminals or tmux sessions
  ./dockerless/start_frontend.sh  # Port 3000
  ./dockerless/start_backend.sh   # Port 8000
  ./dockerless/start_llm.sh       # Needs 6.1GB VRAM
  ./dockerless/start_stt.sh       # Needs 2.5GB VRAM
  ./dockerless/start_tts.sh       # Needs 5.3GB VRAM
  ```

  Access at **[http://localhost:3000](http://localhost:3000)**
</Accordion>

## Configuration Requirements

### Hugging Face Access

<Card title="Model Access Token" icon="key">
  Unmute downloads models from Hugging Face Hub:

  **Required**:

  1. [Hugging Face account](https://huggingface.co/join)
  2. Accept licenses for models you'll use:
     * [Llama 3.2 1B Instruct](https://huggingface.co/meta-llama/Llama-3.2-1B-Instruct) (default)
     * [Mistral Small 3.2 24B](https://huggingface.co/mistralai/Mistral-Small-3.2-24B-Instruct-2506) (recommended)
     * [Gemma 3 12B](https://huggingface.co/google/gemma-3-12b-it) (alternative)
  3. [Generate access token](https://huggingface.co/settings/tokens) with read access
  4. Set environment variable:
     ```bash theme={null}
     export HUGGING_FACE_HUB_TOKEN=hf_your_token_here
     ```

  <Warning>
    **Security**: Never use tokens with write access in production deployments
  </Warning>
</Card>

### Network Requirements

<CardGroup cols={2}>
  <Card title="Ports" icon="network-wired">
    **Docker Compose**:

    * Port 80: Traefik (HTTP traffic)

    **Dockerless**:

    * Port 3000: Frontend
    * Port 8000: Backend WebSocket

    **Optional**:

    * Port 9090: Prometheus metrics
    * Port 3001: Grafana dashboards
  </Card>

  <Card title="Bandwidth" icon="gauge-high">
    **Per User**:

    * Audio upstream: \~16 KB/s
    * Audio downstream: \~16 KB/s
    * Total: \~32 KB/s bidirectional

    **For 10 concurrent users**:

    * \~320 KB/s (\~2.5 Mbps)
  </Card>
</CardGroup>

### Browser Requirements

<Card title="WebRTC & WebSocket Support" icon="browser">
  **Recommended Browsers**:

  * Chrome 90+
  * Firefox 88+
  * Edge 90+
  * Safari 14+ (requires HTTPS)

  **Required Features**:

  * WebSocket support
  * WebRTC (for optional WebRTC mode)
  * Microphone access (requires HTTPS or localhost)
  * Web Audio API

  <Info>
    Modern browsers require **HTTPS or localhost** for microphone access. Use SSH port forwarding for remote access over HTTP.
  </Info>
</Card>

## Model Requirements

### Default Models

<Tabs>
  <Tab title="Speech-to-Text">
    **Model**: Kyutai STT 1B (English/French)

    **Specifications**:

    * **Size**: \~2GB download
    * **VRAM**: 2.5GB
    * **Languages**: English, French
    * **Architecture**: Transformer (16 layers, 2048 d\_model)
    * **Latency**: 6-token delay (\~200ms)

    **Configuration** (`stt.toml`):

    ```toml theme={null}
    lm_model_file = "hf://kyutai/stt-1b-en_fr-candle/model.safetensors"
    text_tokenizer_file = "hf://kyutai/stt-1b-en_fr-candle/tokenizer_en_fr_audio_8000.model"
    audio_tokenizer_file = "hf://kyutai/stt-1b-en_fr-candle/mimi-pytorch-e351c8d8@125.safetensors"
    batch_size = 1
    temperature = 0.25
    ```
  </Tab>

  <Tab title="Text-to-Speech">
    **Model**: Kyutai TTS 1.6B (English/French)

    **Specifications**:

    * **Size**: \~3GB download
    * **VRAM**: 5.3GB
    * **Languages**: English, French
    * **Voices**: 100+ from community donations
    * **Latency**: \~450ms (multi-GPU), \~750ms (single GPU)

    **Configuration** (`tts.toml`):

    ```toml theme={null}
    text_tokenizer_file = "hf://kyutai/tts-1.6b-en_fr/tokenizer_spm_8k_en_fr_audio.model"
    voice_folder = "hf-snapshot://kyutai/tts-voices/**/*.safetensors"
    batch_size = 2
    cfg_coef = 2.0
    n_q = 24
    ```
  </Tab>

  <Tab title="Language Model">
    **Default**: Llama 3.2 1B Instruct

    **Specifications**:

    * **Size**: \~2.5GB download
    * **VRAM**: 6.1GB (with 16GB config)
    * **Context**: 1536 tokens (configurable)
    * **Format**: bfloat16

    **Alternatives**:

    | Model | VRAM | Quality | Context |
    | - | - | - | - |
    | Llama 3.2 1B | 6GB | Good | 1536 |
    | Gemma 3 12B | 20GB | Better | 4096 |
    | Mistral Small 24B | 30GB | Best | 8192 |

    **Configuration** (`docker-compose.yml`):

    ```yaml theme={null}
    llm:
      command:
        [
          "--model=meta-llama/Llama-3.2-1B-Instruct",
          "--max-model-len=1536",
          "--dtype=bfloat16",
          "--gpu-memory-utilization=0.4",
        ]
    ```
  </Tab>
</Tabs>

## Performance Targets

<CardGroup cols={2}>
  <Card title="Latency" icon="stopwatch">
    **Single GPU** (L40S):

    * STT: \~200ms
    * LLM: \~500ms (model dependent)
    * TTS: \~750ms
    * **Total**: \~1450ms

    **Multi-GPU**:

    * STT: \~200ms
    * LLM: \~500ms
    * TTS: \~450ms
    * **Total**: \~1150ms
  </Card>

  <Card title="Throughput" icon="users">
    **Per Backend Instance**:

    * Max concurrent users: 4
    * Limited by Python GIL

    **Scaling Strategy**:

    * Run multiple backend replicas
    * Each replica handles 4 users
    * Load balance with Traefik
    * Example: 10 replicas = 40 users
  </Card>
</CardGroup>

## Optional Components

<AccordionGroup>
  <Accordion title="NewsAPI (for News Character)" icon="newspaper">
    The "Dev (news)" character requires a NewsAPI key:

    1. Sign up at [newsapi.org](https://newsapi.org/)
    2. Get your free API key
    3. Add to environment:
       ```bash theme={null}
       export NEWSAPI_API_KEY=your_key_here
       ```

    Without this, the news character won't have current topics.
  </Accordion>

  <Accordion title="Redis (for Service Discovery)" icon="database">
    Optional for Docker Swarm deployments:

    * Used for service registration and health checks
    * Required for multi-node setups
    * Not needed for single-machine Docker Compose
  </Accordion>

  <Accordion title="Prometheus + Grafana (Monitoring)" icon="chart-line">
    Optional monitoring stack:

    * **Prometheus**: Metrics collection
    * **Grafana**: Visualization dashboards
    * Pre-configured for Unmute metrics
    * Included in Docker Swarm setup

    See `services/prometheus/` and `services/grafana/`
  </Accordion>
</AccordionGroup>

## Deployment Comparison

<Note>
  Choose the deployment method that matches your resources and use case:
</Note>

| Method | GPUs | Machines | Difficulty | Kyutai Support | Best For |
| - | - | - | - | - | - |
| **Docker Compose** | 1+ | 1 | Very Easy | ✅ Full | Development, testing, single-user |
| **Dockerless** | 1-3 | 1-5 | Easy | ✅ Full | Custom setups, debugging |
| **Docker Swarm** | 1-100 | 1-100 | Medium | ❌ None | Production, scaling, unmute.sh |

<Card title="Start with Docker Compose" icon="rocket">
  We strongly recommend starting with Docker Compose:

  * Fastest setup (5-10 minutes)
  * Fully supported by Kyutai team
  * Easy to troubleshoot
  * Perfect for learning and development

  Switch to other methods only when you need:

  * Fine-grained control (Dockerless)
  * Multi-machine scaling (Docker Swarm)
</Card>

## Ready to Start?

<CardGroup cols={2}>
  <Card title="Quick Start Guide" icon="rocket" href="/quickstart">
    Follow step-by-step instructions to get Unmute running
  </Card>

  <Card title="Join the Community" icon="github" href="https://github.com/kyutai-labs/unmute">
    Star the repo, report issues, and contribute
  </Card>
</CardGroup>


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.