Skip to main content
While Unmute doesn’t currently have built-in tool calling support, you can integrate function calling capabilities by wrapping the LLM server with a tool-calling layer.

Overview

Tool calling allows the LLM to invoke external functions, such as:
  • Weather API lookups
  • Database queries
  • Web searches
  • Smart home controls
  • Calendar operations
  • Custom business logic
The key insight is that tool calling can be implemented transparently at the LLM layer, making it invisible to the Unmute backend.

Architecture

The tool-calling proxy:
  1. Receives text generation requests from Unmute backend
  2. Forwards to the LLM with tool definitions
  3. Detects when the LLM wants to call a tool
  4. Executes the tool and injects results
  5. Continues generation with tool results
  6. Returns final response to Unmute

Implementation Approach

The recommended approach is to create a FastAPI server that:
  1. Exposes an OpenAI-compatible API endpoint
  2. Wraps your LLM server (VLLM, Ollama, etc.)
  3. Intercepts tool calls and executes them
  4. Streams results back to Unmute

Why This Works

Unmute expects streaming text responses from an OpenAI-compatible endpoint. As long as your proxy server provides this interface, Unmute doesn’t need to know about the tool calling happening behind the scenes.

Step-by-Step Implementation

1

Create a Tool-Calling Proxy Server

Create a new FastAPI application that wraps your LLM:
2

Add Streaming Support

For real-time voice, streaming is essential:
3

Configure Unmute to Use Your Proxy

Update docker-compose.yml to point to your proxy server:
4

Test Tool Calling

Start a conversation and ask the assistant to use a tool:
The tool call happens transparently behind the scenes.

Example Tools

Weather Lookup

Database Query

LLM Server Compatibility

Tool calling support varies by LLM server:

VLLM

Supports OpenAI-compatible function calling format
Enable with --enable-auto-tool-choice and --tool-call-parser:

Ollama

Supports tools parameter in recent versions

OpenAI API

Native support for function calling
No proxy needed - OpenAI handles tool calls directly.

Advanced Patterns

Multi-Step Tool Chains

Allow the LLM to call multiple tools in sequence:

Conditional Tool Availability

Show different tools based on context:

Error Handling

Considerations

Latency Impact: Tool calls add latency to responses. For voice conversations, this can be noticeable. Consider:
  • Using fast APIs
  • Caching tool results
  • Setting reasonable timeouts
  • Limiting tool call depth
Security: Validate all tool inputs and outputs. Never execute arbitrary code or SQL without sanitization.
Streaming: Continue streaming text to Unmute while executing tools in the background to maintain responsiveness.

Community Contributions

Tool calling support would make a great contribution to the Unmute project! If you build a robust tool-calling proxy server, consider:
  1. Opening a pull request to add it to the main repository
  2. Documenting your approach for others
  3. Sharing example tool implementations
See the GitHub discussion for more details and community input.

Reference Implementation

For a complete working example, check out these resources:

Next Steps

External LLM

Configure Unmute to use different LLM providers

Custom Frontend

Build your own client using the WebSocket protocol