> ## Documentation Index
> Fetch the complete documentation index at: https://mintlify.com/kyutai-labs/unmute/llms.txt
> Use this file to discover all available pages before exploring further.

# Quick Start

> Get Unmute running locally in minutes with Docker Compose

## Prerequisites

Before you begin, ensure you have:

<AccordionGroup>
  <Accordion title="Hardware Requirements" icon="microchip">
    * **GPU**: CUDA-capable GPU with **16GB+ VRAM**
    * **Architecture**: x86\_64 only (no aarch64 support)
    * Single GPU is sufficient for basic setup
    * Multi-GPU setup recommended for production (see below)
  </Accordion>

  <Accordion title="Operating System" icon="desktop">
    * **Linux**: Any modern distribution
    * **Windows**: WSL 2 ([installation guide](https://ubuntu.com/desktop/wsl))
    * **macOS**: Not supported ([issue #74](https://github.com/kyutai-labs/unmute/issues/74))
    * **Windows Native**: Not supported ([issue #84](https://github.com/kyutai-labs/unmute/issues/84))
  </Accordion>

  <Accordion title="Software Requirements" icon="download">
    * **Docker Compose** ([install guide](https://docs.docker.com/compose/))
    * **NVIDIA Container Toolkit** ([install guide](https://docs.nvidia.com/datacenter/cloud-native/container-toolkit/latest/install-guide.html))
  </Accordion>
</AccordionGroup>

## Step 1: Verify NVIDIA Container Toolkit

Confirm your GPU is accessible to Docker:

```bash theme={null}
sudo docker run --rm --runtime=nvidia --gpus all ubuntu nvidia-smi
```

<Accordion title="Expected output">
  You should see your GPU information displayed, including:

  * GPU name and driver version
  * Memory usage and total VRAM
  * CUDA version

  If this fails, install the NVIDIA Container Toolkit before proceeding.
</Accordion>

## Step 2: Get Hugging Face Access

<Note>
  Unmute uses open-weight models from Hugging Face. You'll need a token to download them.
</Note>

<Steps>
  <Step title="Create Hugging Face Account">
    Sign up at [huggingface.co](https://huggingface.co) if you don't have an account
  </Step>

  <Step title="Accept Model License">
    By default, Unmute uses [Llama 3.2 1B Instruct](https://huggingface.co/meta-llama/Llama-3.2-1B-Instruct). Visit the model page and accept the license terms.

    <Tip>
      For better quality with more VRAM, use [Mistral Small 3.2 24B](https://huggingface.co/mistralai/Mistral-Small-3.2-24B-Instruct-2506) or [Gemma 3 12B](https://huggingface.co/google/gemma-3-12b-it)
    </Tip>
  </Step>

  <Step title="Generate Access Token">
    [Create a token](https://huggingface.co/docs/hub/en/security-tokens) with these settings:

    * **Type**: Fine-grained
    * **Permission**: Read access to contents of all public gated repos you can access

    <Warning>
      Never use tokens with write access when deploying publicly. If compromised, attackers could modify your Hugging Face content.
    </Warning>
  </Step>

  <Step title="Set Environment Variable">
    Add your token to your shell configuration:

    ```bash theme={null}
    echo 'export HUGGING_FACE_HUB_TOKEN=hf_your_token_here' >> ~/.bashrc
    source ~/.bashrc
    ```

    Verify it's set:

    ```bash theme={null}
    echo $HUGGING_FACE_HUB_TOKEN
    ```
  </Step>
</Steps>

## Step 3: Clone Repository

```bash theme={null}
git clone https://github.com/kyutai-labs/unmute.git
cd unmute
```

## Step 4: Configure Memory (Optional)

<Note>
  The default `docker-compose.yml` uses Llama 3.2 1B which requires 16GB VRAM. If you have memory issues, adjust these settings:
</Note>

<CodeGroup>
  ```yaml docker-compose.yml (Default - 16GB VRAM) theme={null}
  llm:
    image: vllm/vllm-openai:v0.11.0
    command:
      [
        "--model=meta-llama/Llama-3.2-1B-Instruct",
        "--max-model-len=1536",
        "--dtype=bfloat16",
        "--gpu-memory-utilization=0.4",
      ]
  ```

  ```yaml docker-compose.yml (Higher Quality - 24GB+ VRAM) theme={null}
  llm:
    image: vllm/vllm-openai:v0.11.0
    command:
      [
        "--model=mistralai/Mistral-Small-3.2-24B-Instruct-2506",
        "--max-model-len=4096",
        "--dtype=bfloat16",
        "--gpu-memory-utilization=0.8",
      ]
  ```

  ```yaml docker-compose.yml (Lower Memory - Adjust utilization) theme={null}
  llm:
    command:
      [
        "--model=meta-llama/Llama-3.2-1B-Instruct",
        "--max-model-len=1024",      # Reduce for shorter conversations
        "--gpu-memory-utilization=0.3",  # Lower GPU memory usage
      ]
  ```
</CodeGroup>

## Step 5: Launch Unmute

Start all services with a single command:

```bash theme={null}
docker compose up --build
```

<Steps>
  <Step title="First Run (10-15 minutes)">
    Docker will:

    * Build container images
    * Download models from Hugging Face (\~8GB)
    * Initialize services

    <Info>
      Models are cached in `./volumes/hf-cache/` so subsequent starts are much faster (30-60 seconds)
    </Info>
  </Step>

  <Step title="Wait for Services">
    Monitor the logs. Services are ready when you see:

    ```
    unmute-backend-1  | INFO:     Application startup complete.
    unmute-frontend-1 | ✓ Ready in 3.2s
    unmute-llm-1      | INFO:     Uvicorn running on http://0.0.0.0:8000
    unmute-stt-1      | Listening on 0.0.0.0:8080
    unmute-tts-1      | Listening on 0.0.0.0:8080
    ```
  </Step>

  <Step title="Access Unmute">
    Open your browser to:

    **[http://localhost:80](http://localhost:80)**

    <Tip>
      If port 80 is in use, edit `docker-compose.yml` and change `"80:80"` to `"3000:80"` under the `traefik` service, then access via `http://localhost:3000`
    </Tip>
  </Step>
</Steps>

## Step 6: Start Talking

<Steps>
  <Step title="Grant Microphone Access">
    Your browser will request microphone permission. Click "Allow".
  </Step>

  <Step title="Select a Character">
    Choose from the available voices and personalities:

    * **Watercooler**: Casual small talk
    * **Quiz show**: Interactive trivia
    * **Gertrude**: Life advice and sympathy
    * More voices available in `voices.yaml`
  </Step>

  <Step title="Click Connect">
    The system establishes WebSocket connections and initializes the conversation.
  </Step>

  <Step title="Speak Naturally">
    Start talking! The bot will:

    * Transcribe your speech in real-time
    * Generate contextual responses
    * Speak back to you with the selected voice
  </Step>
</Steps>

<Tip>
  **Keyboard Shortcuts**:

  * Press **S** to toggle subtitles for both user and bot
  * Press **D** for debug mode (requires enabling `ALLOW_DEV_MODE` in `useKeyboardShortcuts.ts`)
</Tip>

## Multi-GPU Configuration

<Note>
  Running STT, TTS, and LLM on separate GPUs reduces TTS latency from \~750ms to \~450ms.
</Note>

If you have 3+ GPUs, edit `docker-compose.yml` to assign dedicated GPUs:

```yaml docker-compose.yml theme={null}
stt:
  # ...existing config...
  deploy:
    resources:
      reservations:
        devices:
          - driver: nvidia
            count: 1
            capabilities: [gpu]

tts:
  # ...existing config...
  deploy:
    resources:
      reservations:
        devices:
          - driver: nvidia
            count: 1
            capabilities: [gpu]

llm:
  # ...existing config...
  deploy:
    resources:
      reservations:
        devices:
          - driver: nvidia
            count: 1
            capabilities: [gpu]
```

By default, all services share available GPUs. This configuration ensures each service gets its own GPU.

## Remote Access via SSH

<AccordionGroup>
  <Accordion title="Docker Compose (Port 80)" icon="docker">
    Forward port 80 from remote to local port 3333:

    ```bash theme={null}
    ssh -N -L 3333:localhost:80 unmute-box
    ```

    Then access via **[http://localhost:3333](http://localhost:3333)**

    <Info>
      Browsers require localhost or HTTPS for microphone access. Direct HTTP access to `http://unmute-box:80` won't work.
    </Info>
  </Accordion>

  <Accordion title="Why Port Forwarding is Required" icon="shield">
    Modern browsers block microphone access over HTTP unless:

    * The origin is `localhost` or `127.0.0.1`
    * The connection uses HTTPS

    Port forwarding makes the remote server appear as `localhost` to your browser.
  </Accordion>
</AccordionGroup>

## Common Issues

<AccordionGroup>
  <Accordion title="Out of Memory Errors" icon="memory">
    Reduce memory usage in `docker-compose.yml`:

    ```yaml theme={null}
    llm:
      command:
        [
          "--model=meta-llama/Llama-3.2-1B-Instruct",
          "--max-model-len=1024",           # Lower from 1536
          "--gpu-memory-utilization=0.3",   # Lower from 0.4
        ]
    ```

    Or increase batch sizes for TTS/STT to share memory better (higher latency).
  </Accordion>

  <Accordion title="Models Not Downloading" icon="cloud-arrow-down">
    Check your Hugging Face token:

    ```bash theme={null}
    echo $HUGGING_FACE_HUB_TOKEN
    ```

    Verify you accepted the model license on Hugging Face.
  </Accordion>

  <Accordion title="Port Already in Use" icon="plug">
    Change the port in `docker-compose.yml`:

    ```yaml theme={null}
    traefik:
      ports:
        - "8080:80"  # Use port 8080 instead
    ```
  </Accordion>

  <Accordion title="WebSocket Connection Failed" icon="wifi">
    Ensure all services are running:

    ```bash theme={null}
    docker compose ps
    ```

    Check backend logs:

    ```bash theme={null}
    docker compose logs backend
    ```
  </Accordion>
</AccordionGroup>

## Next Steps

<CardGroup cols={2}>
  <Card title="Customize Voices" icon="microphone" href="/customization/voices">
    Add custom voices and modify character personalities
  </Card>

  <Card title="Use External LLMs" icon="cloud" href="/customization/llm">
    Connect to OpenAI, Ollama, or other LLM providers
  </Card>

  <Card title="Production Deployment" icon="server" href="/deployment/swarm">
    Scale Unmute with Docker Swarm for production workloads
  </Card>

  <Card title="Development Guide" icon="code" href="/development/overview">
    Contribute to Unmute or build custom frontends
  </Card>
</CardGroup>

## Stop Unmute

To stop all services:

```bash theme={null}
docker compose down
```

To also remove downloaded models and caches:

```bash theme={null}
docker compose down -v
rm -rf volumes/
```

<Warning>
  Need help? Open an issue on [GitHub](https://github.com/kyutai-labs/unmute/issues) - the Kyutai team actively supports Docker Compose deployments.
</Warning>


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.