ca3e66f17a7d7cb1e25e120219fb9fc5293af4fc
XTTS2 OpenAI-Compatible TTS Server
High-Performance Text-to-Speech Platform with FastAPI, Redis Queueing, GPU Workers & OpenAI Speech API Compatibility.
🔥 Übersicht
Dieses Projekt stellt eine vollständig modulare, skalierbare und produktionsreife lokale Text-to-Speech-Plattform bereit. Die API ist vollständig kompatibel zur OpenAI Speech API, unterstützt XTTS2 Zero-Shot Voice Cloning, mehrere Audioformate und verteilte GPU-Worker.
Zielsetzung:
- Hochqualitative TTS-Synthese für Spiele, Voice-Overs und AI-Produktion
- Zero-Shot Voice Cloning mit XTTS2
- OpenAI-kompatible Endpoints als Drop-in Replacement
- Skalierbare Worker-Architektur für hohe Lasten
🧱 Architektur
┌────────────────────┐ ┌───────────────────────────────┐
│ FastAPI Gateway │◀──────┤ Redis Queue + Cache │
│ - OpenAI API │ └───────────────────────────────┘
│ - Rate Limits │
│ - Port Auto-Select │ ┌───────────────────────────────┐
│ - Voice Registry │──────▶│ Worker (XTTS2) │
└────────────────────┘ │ - GPU/CPU Auto Detect │
│ - Zero-Shot Voice Cloning │
└───────────────────────────────┘
🚀 Features
Core Features
- OpenAI-kompatible Speech API
- Zero-Shot Voice Cloning mit XTTS2
- Unterstützung für
wav,mp3,ogg - Dynamische Portwahl (8000–8100) mit Fallback
- Registry für permanente Stimmen
- Queue-basierte Worker-Architektur (Redis)
- GPU Auto-Detection für Worker
Deployment Features
- Multi-Stage Docker Images (Gateway & Worker)
- Makefile PRO für Build, Deploy, Scaling & Testing
- Logs + Portfile + Healthchecks
- Lazy Load der XTTS2-Modelle
📦 Installation
Voraussetzungen
- Docker & Docker Compose
- NVIDIA Container Runtime (für GPU-Worker)
- Linux oder macOS (Windows WSL2 möglich)
Start in 3 Schritten
make build
make up
make status
Die dynamische Portwahl speichert den genutzten Port in:
gateway/port.txt
gateway/logs/gateway.log
📡 OpenAI-kompatible Endpunkte
POST /v1/audio/speech
Request Beispiel:
{
"model": "xtts-v2",
"input": "Hello, hero",
"voice": "auto",
"format": "wav"
}
Antwort:
- binares Audio
- Content-Type abhängig vom Format
POST /v1/voices/register
- Registriert permanente Stimmen
- Unterstützt Samples per URL oder Base64
GET /health
- Healthcheck für Monitoring & CI
🧪 Selftest
make selftest
Prüft:
- Redis erreichbar
- Gateway erreichbar
- Worker zieht Jobs
- Mini-Synthese erfolgreich
📁 Projektstruktur
project/
├── gateway/
│ ├── main.py
│ ├── api/
│ ├── core/
│ ├── voices/
│ ├── logs/
│ └── Dockerfile
│
├── worker/
│ ├── main.py
│ ├── engine/
│ ├── core/
│ └── Dockerfile
│
├── scripts/
│ └── find_port.py
│
├── Makefile
├── docker-compose.yml
└── README.md
🔐 Sicherheit & Best Practices
- Optional: API Keys für öffentliche Deployments
- Rate Limits im Gateway aktivierbar
- HTTPS über Reverse Proxy
- Keine sensiblen Voice-Daten einchecken
- Worker nur intern erreichbar halten
🛠 Roadmap
- Prometheus & Grafana Monitoring
- Business-Level Logging (JSON Logs)
- Support für weitere Modelle (F5, Kokoro, Piper)
- WebSocket Realtime TTS Output
🧑💻 Maintainer
Stephan W. – Architektur & Betrieb
Description
Languages
Python
78.4%
Makefile
18.6%
Dockerfile
3%