Skip to main content

Overview

Unmute includes comprehensive monitoring capabilities using Prometheus for metrics collection and Grafana for visualization. The monitoring stack tracks latency, throughput, error rates, and resource utilization across all services.

Architecture

Prometheus scrapes metrics from all services and stores time-series data. Grafana queries Prometheus to display real-time dashboards.

Metrics Collection

Available Metrics

Unmute exposes metrics via the Prometheus client library (defined in unmute/metrics.py):

Session Metrics

STT (Speech-to-Text) Metrics

TTS (Text-to-Speech) Metrics

LLM (Language Model) Metrics

Error Metrics

Histogram Buckets

Metrics use predefined buckets for accurate percentile calculations:

Prometheus Setup

Docker Compose Configuration

For production deployments, add Prometheus to your Docker Swarm stack:

Prometheus Configuration

Create services/prometheus/prometheus.yml:

Service Labels

Label services to expose metrics:

Grafana Setup

Docker Configuration

Data Source Configuration

Create services/grafana/provisioning/datasources/datasources.yaml:

Dashboard Provisioning

Create services/grafana/provisioning/dashboards/dashboards.yaml:

Key Dashboards

System Overview Dashboard

Active Sessions:
Total Sessions (Rate):
Error Rate:

Latency Dashboard

STT Latency (p95):
TTS Latency (p95):
LLM Latency (p95):
Average STT Latency:

Throughput Dashboard

STT Words Per Second:
TTS Words Per Second:
LLM Tokens Per Second:

Service Health Dashboard

STT Service Misses:
TTS Service Misses:
Active STT Sessions:
Active TTS Sessions:
Active LLM Sessions:

User Behavior Dashboard

Session Duration (Average):
Interruption Rate:
Average Request Length:
Average Reply Length:

Accessing Dashboards

Local Development

Access Grafana at http://localhost:3000 (default credentials: admin/admin).

Production Deployment

For unmute.sh deployment with Traefik:
Access at: https://grafana.unmute.sh

Alerting

Example Alert Rules

Create services/prometheus/alerts.yml:
Add to Prometheus configuration:

Health Checks

Backend Health Endpoint

Unmute exposes a health check endpoint:
Response:

Load Testing

Use the built-in load test client to validate monitoring:
Watch metrics in Grafana during the test to verify collection.

Production Monitoring URLs

From unmute.sh deployment: All monitoring services are protected by OAuth authentication (Google) via Traefik Forward Auth.

Best Practices

  1. Set appropriate scrape intervals: 5s for real-time, 30s for cost savings
  2. Use retention policies: Configure Prometheus to retain data for 30-90 days
  3. Monitor percentiles, not just averages: p95 and p99 reveal tail latencies
  4. Set up alerts: Proactive notification prevents outages
  5. Archive long-term data: Export to long-term storage (e.g., S3) for historical analysis

Troubleshooting

Metrics not appearing

Check:
  1. Service has prometheus-port label
  2. Prometheus can reach the service (check targets page)
  3. Metrics endpoint returns data: curl http://backend/metrics

High cardinality warnings

Cause: Too many unique label combinations Solution: Avoid using user IDs or session IDs as labels. Use counters instead.

Missing histograms

Check: Bucket configuration matches expected latency ranges. Add buckets if values exceed defined ranges.

Next Steps