> ## Documentation Index
> Fetch the complete documentation index at: https://mintlify.com/kyutai-labs/unmute/llms.txt
> Use this file to discover all available pages before exploring further.

# GPU Setup

> Configure NVIDIA GPU access for running Unmute's AI models

Unmute requires GPU acceleration to run the speech-to-text, text-to-speech, and LLM models with acceptable latency. This guide covers GPU requirements and configuration.

## Hardware Requirements

### Minimum Requirements

* **GPU**: NVIDIA GPU with CUDA support
* **VRAM**: At least 16 GB
* **Architecture**: x86\_64 (aarch64 is not supported)

### Memory Usage by Service

When running all services, approximate VRAM usage:

| Service | VRAM Required |
| - | - |
| Speech-to-text (STT) | 2.5 GB |
| Text-to-speech (TTS) | 5.3 GB |
| LLM (Llama 3.2 1B) | 6.1 GB |
| **Total** | **\~14 GB** |

<Note>
  The default `docker-compose.yml` uses Llama 3.2 1B which fits in 16GB VRAM. If using larger models like Mistral Small 3.2 24B, you'll need more VRAM.
</Note>

## Operating System Support

<Warning>
  **Windows**: Native Windows is not supported ([#84](https://github.com/kyutai-labs/unmute/issues/84)). Use [WSL (Windows Subsystem for Linux)](https://ubuntu.com/desktop/wsl) instead.

  **macOS**: Not supported ([#74](https://github.com/kyutai-labs/unmute/issues/74)). macOS does not have NVIDIA GPU support.
</Warning>

Supported platforms:

* Linux (native)
* Windows with WSL 2

## Docker Setup

### Install NVIDIA Container Toolkit

The NVIDIA Container Toolkit allows Docker containers to access your GPU.

<Steps>
  <Step title="Install the Container Toolkit">
    Follow the official [NVIDIA Container Toolkit installation guide](https://docs.nvidia.com/datacenter/cloud-native/container-toolkit/latest/install-guide.html).

    For Ubuntu/Debian:

    ```bash theme={null}
    distribution=$(. /etc/os-release;echo $ID$VERSION_ID)
    curl -s -L https://nvidia.github.io/nvidia-docker/gpgkey | sudo apt-key add -
    curl -s -L https://nvidia.github.io/nvidia-docker/$distribution/nvidia-docker.list | \
      sudo tee /etc/apt/sources.list.d/nvidia-docker.list

    sudo apt-get update
    sudo apt-get install -y nvidia-container-toolkit
    sudo systemctl restart docker
    ```
  </Step>

  <Step title="Verify the Installation">
    Test that Docker can access your GPU:

    ```bash theme={null}
    sudo docker run --rm --runtime=nvidia --gpus all ubuntu nvidia-smi
    ```

    You should see output showing your GPU(s), similar to:

    ```
    +-----------------------------------------------------------------------------+
    | NVIDIA-SMI 525.147.05   Driver Version: 525.147.05   CUDA Version: 12.0     |
    |-------------------------------+----------------------+----------------------+
    | GPU  Name        Persistence-M| Bus-Id        Disp.A | Volatile Uncorr. ECC |
    | Fan  Temp  Perf  Pwr:Usage/Cap|         Memory-Usage | GPU-Util  Compute M. |
    |                               |                      |               MIG M. |
    |===============================+======================+======================|
    |   0  NVIDIA L40S         Off  | 00000000:00:05.0 Off |                    0 |
    | N/A   28C    P0    32W / 350W |      0MiB / 46068MiB |      0%      Default |
    ```
  </Step>
</Steps>

### Configure GPU Access in Docker Compose

The `docker-compose.yml` file configures GPU access for the AI services:

```yaml theme={null}
services:
  tts:
    deploy:
      resources:
        reservations:
          devices:
            - driver: nvidia
              count: all  # Use all available GPUs
              capabilities: [gpu]
```

## Multi-GPU Configuration

Running services on separate GPUs significantly improves latency. On [unmute.sh](https://unmute.sh), TTS latency decreases from \~750ms (single L40S GPU) to \~450ms (multi-GPU setup).

### Single GPU Setup (Default)

By default, all services share GPU(s) using `count: all`:

```yaml theme={null}
tts:
  deploy:
    resources:
      reservations:
        devices:
          - driver: nvidia
            count: all
            capabilities: [gpu]
```

### Dedicated GPU per Service

If you have 3+ GPUs, assign one GPU to each service for optimal performance:

<Steps>
  <Step title="Check Available GPUs">
    List your GPUs:

    ```bash theme={null}
    nvidia-smi -L
    ```

    Example output:

    ```
    GPU 0: NVIDIA L40S
    GPU 1: NVIDIA L40S
    GPU 2: NVIDIA L40S
    ```
  </Step>

  <Step title="Update docker-compose.yml">
    Modify the `stt`, `tts`, and `llm` services to use dedicated GPUs:

    ```yaml theme={null}
    stt:
      deploy:
        resources:
          reservations:
            devices:
              - driver: nvidia
                count: 1  # Changed from 'all' to '1'
                capabilities: [gpu]

    tts:
      deploy:
        resources:
          reservations:
            devices:
              - driver: nvidia
                count: 1
                capabilities: [gpu]

    llm:
      deploy:
        resources:
          reservations:
            devices:
              - driver: nvidia
                count: 1
                capabilities: [gpu]
    ```
  </Step>

  <Step title="Restart Services">
    Apply the changes:

    ```bash theme={null}
    docker compose down
    docker compose up --build
    ```
  </Step>
</Steps>

<Note>
  Docker will automatically distribute services across available GPUs when using `count: 1`. You don't need to manually specify device IDs.
</Note>

## Memory Optimization

If you're running out of GPU memory, adjust these settings in `docker-compose.yml`:

### LLM Memory Settings

```yaml theme={null}
llm:
  command:
    - "--model=meta-llama/Llama-3.2-1B-Instruct"
    # Reduce max context length to save memory
    - "--max-model-len=1536"  # Lower this value
    - "--dtype=bfloat16"
    # Reduce GPU memory usage percentage
    - "--gpu-memory-utilization=0.4"  # Lower this value (e.g., 0.3)
```

<ParamField path="--max-model-len" type="integer" default="1536">
  Maximum context length for the LLM. Lower values use less memory but support shorter conversations.
</ParamField>

<ParamField path="--gpu-memory-utilization" type="float" default="0.4">
  Percentage of GPU memory to allocate (0.0-1.0). Lower values leave more memory for other services.
</ParamField>

### Switch to a Smaller Model

Use a smaller LLM model:

```yaml theme={null}
llm:
  command:
    - "--model=meta-llama/Llama-3.2-1B-Instruct"  # Smaller model
```

Available models (by size):

* `meta-llama/Llama-3.2-1B-Instruct` - \~6 GB VRAM
* `google/gemma-3-1b-it` - \~6 GB VRAM (note: slower on vLLM)
* `google/gemma-3-12b-it` - \~12 GB VRAM
* `mistralai/Mistral-Small-3.2-24B-Instruct-2506` - \~24 GB VRAM

## Dockerless Setup

For dockerless deployment, ensure CUDA 12.1+ is installed:

<Steps>
  <Step title="Install CUDA">
    Install CUDA 12.1 or later:

    * Via conda: `conda install cuda -c nvidia/label/cuda-12.1.0`
    * Or download from [NVIDIA's website](https://developer.nvidia.com/cuda-downloads)
  </Step>

  <Step title="Verify Installation">
    ```bash theme={null}
    nvcc --version
    nvidia-smi
    ```
  </Step>

  <Step title="Run Services">
    The dockerless scripts automatically detect and use available GPUs:

    ```bash theme={null}
    ./dockerless/start_stt.sh   # Uses 2.5GB VRAM
    ./dockerless/start_tts.sh   # Uses 5.3GB VRAM
    ./dockerless/start_llm.sh   # Uses 6.1GB VRAM
    ```
  </Step>
</Steps>

## Troubleshooting

<AccordionGroup>
  <Accordion title="Docker can't access GPU">
    If `nvidia-smi` works but Docker can't access the GPU:

    1. Verify NVIDIA Container Toolkit is installed
    2. Restart Docker: `sudo systemctl restart docker`
    3. Check Docker runtime: `docker info | grep -i runtime`
    4. Try the verification command again:
       ```bash theme={null}
       sudo docker run --rm --runtime=nvidia --gpus all ubuntu nvidia-smi
       ```
  </Accordion>

  <Accordion title="Out of memory errors">
    If services crash with OOM errors:

    1. Check VRAM usage: `nvidia-smi`
    2. Reduce `--gpu-memory-utilization` for the LLM
    3. Lower `--max-model-len` for shorter conversations
    4. Use a smaller LLM model
    5. Stop other GPU-intensive applications
  </Accordion>

  <Accordion title="WSL GPU issues">
    For Windows WSL users:

    1. Ensure you're using WSL 2 (not WSL 1)
    2. Update to the latest NVIDIA driver for Windows
    3. Install NVIDIA CUDA on WSL following [Microsoft's guide](https://docs.microsoft.com/en-us/windows/ai/directml/gpu-cuda-in-wsl)
    4. Don't install NVIDIA drivers inside WSL - use the Windows driver
  </Accordion>
</AccordionGroup>


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.