updated skeleton with true data
This commit is contained in:
290
DEV_SETUP.md
290
DEV_SETUP.md
@ -1,2 +1,288 @@
|
||||
# Dev Setup
|
||||
Full setup guide.
|
||||
# Developer Setup Guide
|
||||
|
||||
Dieses Dokument befähigt jeden Entwickler dazu, das komplette XTTS2-TTS-System lokal aufzusetzen, zu verstehen, zu debuggen und zu erweitern – ohne Rückfragen.
|
||||
|
||||
---
|
||||
|
||||
# 1. Ziel des Dokuments
|
||||
|
||||
* Lokale Entwicklungsumgebung vollständig einrichten
|
||||
* Services starten & stoppen
|
||||
* Worker skalieren
|
||||
* API testen
|
||||
* Debugging-Strategien
|
||||
* Best Practices für Code, Deployment & GPU-Nutzung
|
||||
|
||||
---
|
||||
|
||||
# 2. Voraussetzungen
|
||||
|
||||
## Software
|
||||
|
||||
* Docker ≥ 24.x
|
||||
* Docker Compose ≥ v2.x
|
||||
* Git
|
||||
* Python 3.10–3.11
|
||||
* Optional: Editor wie VS Code mit Python- & Docker-Extensions
|
||||
|
||||
## Hardware
|
||||
|
||||
* Linux (empfohlen), macOS oder Windows WSL2
|
||||
* Für GPU-Betrieb: NVIDIA GPU + Container Toolkit
|
||||
|
||||
GPU prüfen:
|
||||
|
||||
```bash
|
||||
nvidia-smi
|
||||
```
|
||||
|
||||
NVIDIA Docker prüfen:
|
||||
|
||||
```bash
|
||||
docker run --rm --gpus all nvidia/cuda:12.1.0-base nvidia-smi
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
# 3. Repository klonen
|
||||
|
||||
```bash
|
||||
git clone <repo-url>
|
||||
cd tts-server
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
# 4. Projektstruktur
|
||||
|
||||
```
|
||||
tts-server/
|
||||
├── gateway/ → FastAPI-Gateway
|
||||
│ ├── main.py
|
||||
│ ├── api/
|
||||
│ ├── core/
|
||||
│ ├── voices/
|
||||
│ ├── logs/
|
||||
│ └── Dockerfile
|
||||
│
|
||||
├── worker/ → XTTS2 Worker
|
||||
│ ├── main.py
|
||||
│ ├── engine/
|
||||
│ ├── core/
|
||||
│ └── Dockerfile
|
||||
│
|
||||
├── scripts/
|
||||
│ └── find_port.py → Portscanner (8000–8100)
|
||||
│
|
||||
├── docker-compose.yml → Multi-Service Orchestration
|
||||
├── Makefile → PRO Workflow
|
||||
├── README.md
|
||||
└── ONBOARDING.md
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
# 5. Docker-basiertes Development (empfohlen)
|
||||
|
||||
## 5.1 Images bauen
|
||||
|
||||
```bash
|
||||
make build
|
||||
```
|
||||
|
||||
## 5.2 Server starten
|
||||
|
||||
```bash
|
||||
make up
|
||||
```
|
||||
|
||||
Der Portscanner wählt automatisch den ersten freien Port (8000–8100) und legt ihn ab in:
|
||||
|
||||
```
|
||||
gateway/port.txt
|
||||
gateway/logs/gateway.log
|
||||
```
|
||||
|
||||
## 5.3 Logs ansehen
|
||||
|
||||
```bash
|
||||
make logs
|
||||
```
|
||||
|
||||
## 5.4 Services stoppen
|
||||
|
||||
```bash
|
||||
make down
|
||||
```
|
||||
|
||||
## 5.5 Status anzeigen
|
||||
|
||||
```bash
|
||||
make status
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
# 6. Ohne Docker entwickeln (lokales Debugging)
|
||||
|
||||
## 6.1 Virtuelle Umgebung
|
||||
|
||||
```bash
|
||||
python3 -m venv .venv
|
||||
source .venv/bin/activate
|
||||
```
|
||||
|
||||
## 6.2 Abhängigkeiten installieren
|
||||
|
||||
Gateway:
|
||||
|
||||
```bash
|
||||
pip install -r gateway/requirements.gateway.txt
|
||||
```
|
||||
|
||||
Worker:
|
||||
|
||||
```bash
|
||||
pip install -r worker/requirements.worker.txt
|
||||
```
|
||||
|
||||
Redis lokal starten:
|
||||
|
||||
```bash
|
||||
docker run -p 6379:6379 redis:7
|
||||
```
|
||||
|
||||
Gateway starten:
|
||||
|
||||
```bash
|
||||
python gateway/main.py
|
||||
```
|
||||
|
||||
Worker starten:
|
||||
|
||||
```bash
|
||||
python worker/main.py
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
# 7. TTS API testen
|
||||
|
||||
## 7.1 Healthcheck
|
||||
|
||||
```bash
|
||||
curl http://localhost:<PORT>/health
|
||||
```
|
||||
|
||||
## 7.2 TTS Anfrage
|
||||
|
||||
```bash
|
||||
curl -X POST http://localhost:<PORT>/v1/audio/speech \
|
||||
-H "Content-Type: application/json" \
|
||||
-d '{
|
||||
"model": "xtts-v2",
|
||||
"input": "Das ist ein Test.",
|
||||
"voice": "auto",
|
||||
"format": "wav"
|
||||
}' \
|
||||
--output test.wav
|
||||
```
|
||||
|
||||
## 7.3 Stimme registrieren
|
||||
|
||||
```bash
|
||||
curl -X POST http://localhost:<PORT>/v1/voices/register \
|
||||
-H "Content-Type: application/json" \
|
||||
-d '{
|
||||
"name": "german_narrator",
|
||||
"samples": ["https://example.com/sample.wav"]
|
||||
}'
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
# 8. Worker skalieren
|
||||
|
||||
Höhere Last? Mehrere Worker starten:
|
||||
|
||||
```bash
|
||||
make worker-scale N=4
|
||||
```
|
||||
|
||||
Redis verteilt Jobs automatisch FIFO.
|
||||
|
||||
---
|
||||
|
||||
# 9. Selftest
|
||||
|
||||
```bash
|
||||
make selftest
|
||||
```
|
||||
|
||||
Testet:
|
||||
|
||||
* Redis erreichbar
|
||||
* Gateway online
|
||||
* Worker zieht Jobs
|
||||
* Synthese funktioniert
|
||||
|
||||
---
|
||||
|
||||
# 10. Debugging
|
||||
|
||||
## Gateway startet nicht
|
||||
|
||||
* Prüfen: `gateway/logs/`
|
||||
* Port frei? `gateway/port.txt`
|
||||
* Redis erreichbar?
|
||||
|
||||
## Worker lädt nicht
|
||||
|
||||
* XTTS2 Modell verfügbar?
|
||||
* GPU verfügbar? `nvidia-smi`
|
||||
* Torch CUDA Version kompatibel?
|
||||
|
||||
## Audio-Probleme
|
||||
|
||||
* Voice Sample ungeeignet
|
||||
* Sprache nicht angegeben
|
||||
* Format falsch
|
||||
|
||||
---
|
||||
|
||||
# 11. Code Guidelines
|
||||
|
||||
* PEP8 Stil
|
||||
* Logging strukturiert (JSON empfohlen)
|
||||
* Kein Hardcoding von Pfaden
|
||||
* Hohe Modularität
|
||||
* Unit Tests für Kernkomponenten
|
||||
* Feature Branches: `feature/<name>`
|
||||
* Commits: Conventional Commits
|
||||
|
||||
---
|
||||
|
||||
# 12. Best Practices
|
||||
|
||||
* Worker lieber horizontal skalieren statt optimieren
|
||||
* Keine sensiblen Voice Samples committen
|
||||
* Docker Images regelmäßig aktualisieren
|
||||
* Für Public Deployments: Auth Layer aktivieren
|
||||
|
||||
---
|
||||
|
||||
# 13. Nächste Schritte für Entwickler
|
||||
|
||||
* Monitoring & Dashboards
|
||||
* WebSocket Realtime TTS
|
||||
* Mehrsprachige Voice Registry
|
||||
* Model Hot-Swapping (XTTS2, F5, Kokoro)
|
||||
* CI/CD Pipeline hinzufügen
|
||||
|
||||
---
|
||||
|
||||
# 14. Fertig!
|
||||
|
||||
Wenn du dieses Dokument verstanden hast, kannst du das Projekt vollständig entwickeln, erweitern und deployen.
|
||||
|
||||
Viel Erfolg! 🚀
|
||||
|
||||
131
Makefile
131
Makefile
@ -1,2 +1,129 @@
|
||||
all:
|
||||
echo hi
|
||||
# Makefile – PRO Mode
|
||||
|
||||
# Vollständiger Build-/Deploy-/Diagnose-Workflow für das XTTS2 TTS System
|
||||
|
||||
PYTHON := python3
|
||||
PORT_SCRIPT := scripts/find_port.py
|
||||
PORT_FILE := gateway/port.txt
|
||||
DOCKER := docker compose
|
||||
|
||||
.DEFAULT_GOAL := help
|
||||
|
||||
help:
|
||||
@echo ""
|
||||
@echo "🚀 XTTS2 TTS Server – Makefile (Pro Mode)"
|
||||
@echo "-------------------------------------------"
|
||||
@echo " make build → Images bauen (Gateway + Worker)"
|
||||
@echo " make up → Services starten (mit Portscan)"
|
||||
@echo " make down → Services stoppen"
|
||||
@echo " make restart → Neustart"
|
||||
@echo " make logs → Logs aller Services anzeigen"
|
||||
@echo " make status → Docker Status"
|
||||
@echo " make worker-scale N=3 → Worker skalieren"
|
||||
@echo " make prune → Docker aufräumen"
|
||||
@echo " make selftest → System-Selbsttest"
|
||||
@echo " make port → Zeigt aktuellen Gateway Port"
|
||||
@echo "-------------------------------------------"
|
||||
|
||||
# ---------------------------------------------------------
|
||||
|
||||
# Build
|
||||
|
||||
# ---------------------------------------------------------
|
||||
|
||||
build:
|
||||
$(DOCKER) build
|
||||
|
||||
# ---------------------------------------------------------
|
||||
|
||||
# Deploy
|
||||
|
||||
# ---------------------------------------------------------
|
||||
|
||||
up:
|
||||
@echo "🔍 Suche freien Port zwischen 8000–8100..."
|
||||
@PORT=`$(PYTHON) $(PORT_SCRIPT)`;
|
||||
if [ "$$PORT" = "ERR_NO_FREE_PORT" ]; then
|
||||
echo "❌ Kein freier Port gefunden!"; exit 1;
|
||||
fi;
|
||||
echo "🎧 Freier Port gefunden: $$PORT";
|
||||
echo "$$PORT" > $(PORT_FILE);
|
||||
echo "📄 Port gespeichert in $(PORT_FILE)";
|
||||
export GATEWAY_PORT=$$PORT;
|
||||
$(DOCKER) up -d --build;
|
||||
echo "🚀 Gateway läuft auf [http://localhost:$$PORT](http://localhost:$$PORT)"
|
||||
|
||||
# ---------------------------------------------------------
|
||||
|
||||
# Stop
|
||||
|
||||
# ---------------------------------------------------------
|
||||
|
||||
down:
|
||||
$(DOCKER) down
|
||||
|
||||
# ---------------------------------------------------------
|
||||
|
||||
# Restart
|
||||
|
||||
# ---------------------------------------------------------
|
||||
|
||||
restart: down up
|
||||
|
||||
# ---------------------------------------------------------
|
||||
|
||||
# Logs
|
||||
|
||||
# ---------------------------------------------------------
|
||||
|
||||
logs:
|
||||
$(DOCKER) logs -f
|
||||
|
||||
# ---------------------------------------------------------
|
||||
|
||||
# Status
|
||||
|
||||
# ---------------------------------------------------------
|
||||
|
||||
status:
|
||||
$(DOCKER) ps
|
||||
|
||||
# ---------------------------------------------------------
|
||||
|
||||
# Worker Scaling
|
||||
|
||||
# ---------------------------------------------------------
|
||||
|
||||
worker-scale:
|
||||
@if [ -z "$(N)" ]; then echo "Bitte N angeben: make worker-scale N=3"; exit 1; fi
|
||||
$(DOCKER) up -d --scale worker=$(N)
|
||||
|
||||
# ---------------------------------------------------------
|
||||
|
||||
# Cleanup
|
||||
|
||||
# ---------------------------------------------------------
|
||||
|
||||
prune:
|
||||
$(DOCKER) down
|
||||
docker system prune -f
|
||||
|
||||
# ---------------------------------------------------------
|
||||
|
||||
# Selftest
|
||||
|
||||
# ---------------------------------------------------------
|
||||
|
||||
selftest:
|
||||
@echo "🧪 Starte Selbsttest..."
|
||||
$(PYTHON) scripts/selftest.py
|
||||
|
||||
# ---------------------------------------------------------
|
||||
|
||||
# Show Port
|
||||
|
||||
# ---------------------------------------------------------
|
||||
|
||||
port:
|
||||
@echo "📡 Aktueller Port:"
|
||||
@cat $(PORT_FILE)
|
||||
|
||||
272
ONBOARDING.md
272
ONBOARDING.md
@ -1,2 +1,270 @@
|
||||
# Onboarding
|
||||
Full onboarding.
|
||||
# Onboarding Guide
|
||||
|
||||
Willkommen im Projekt! Dieses Dokument führt neue Entwickler vollständig ein – ohne Rückfragen, ohne offene Punkte.
|
||||
|
||||
Ziel: Du sollst in der Lage sein, das gesamte System zu verstehen, zu betreiben und weiterzuentwickeln.
|
||||
|
||||
---
|
||||
|
||||
# 1. Projektüberblick
|
||||
|
||||
Das XTTS2 TTS-System ist eine modulare, verteilte Plattform für hochwertige Text-to-Speech-Synthese mit Zero-Shot Voice Cloning. Der Aufbau orientiert sich an professionellen Backend-Architekturen mit Queueing, GPU-Workern und einer OpenAI-kompatiblen API.
|
||||
|
||||
**Hauptkomponenten:**
|
||||
|
||||
* **Gateway** (FastAPI): HTTP-API, Validierung, Routing, Port-Autodetection.
|
||||
* **Redis**: Queue, Cache, interne Metadaten.
|
||||
* **Worker** (XTTS2): führt TTS aus, nutzt GPU automatisch, arbeitet skalierbar.
|
||||
|
||||
Dieses System eignet sich für:
|
||||
|
||||
* Spielevertonung
|
||||
* Automatisiertes Voice-Over
|
||||
* Lokale AI-Pipelines
|
||||
* Multimodale Agenten
|
||||
|
||||
---
|
||||
|
||||
# 2. Architektur
|
||||
|
||||
```
|
||||
Client → Gateway → Redis Queue → Worker (XTTS2) → Gateway → Client
|
||||
```
|
||||
|
||||
* Das **Gateway** nimmt Requests entgegen und legt Jobs in Redis ab.
|
||||
* **Worker** verarbeiten Jobs parallel und liefern Audiodaten zurück.
|
||||
* Die API ist **OpenAI-kompatibel** – Clients können ohne Anpassung migriert werden.
|
||||
|
||||
---
|
||||
|
||||
# 3. Voraussetzungen
|
||||
|
||||
## Software
|
||||
|
||||
* Docker & Docker Compose
|
||||
* Git
|
||||
* Python 3.10 / 3.11 (für lokales Debugging)
|
||||
* NVIDIA GPU + Container Toolkit (optional, aber empfohlen)
|
||||
|
||||
## Hardware
|
||||
|
||||
* 8 GB RAM minimum
|
||||
* GPU mit mindestens 4–6 GB VRAM für XTTS2
|
||||
|
||||
---
|
||||
|
||||
# 4. Repository klonen
|
||||
|
||||
```bash
|
||||
git clone <repo-url>
|
||||
cd tts-server
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
# 5. Projektstruktur verstehen
|
||||
|
||||
```
|
||||
gateway/ → FastAPI-Gateway
|
||||
worker/ → XTTS2-GPU-Worker
|
||||
scripts/ → Hilfsskripte (Portfinder, Selftest)
|
||||
voices/ → Voice Registry
|
||||
Makefile → Build-, Deploy- und Diagnosewerkzeuge
|
||||
docker-compose.yml
|
||||
```
|
||||
|
||||
Die wichtigsten Einstiegspunkte:
|
||||
|
||||
* `gateway/main.py` – Start des API-Gateways
|
||||
* `worker/main.py` – Start des XTTS2-Workers
|
||||
* `scripts/find_port.py` – Portscanner (8000–8100)
|
||||
|
||||
---
|
||||
|
||||
# 6. System starten
|
||||
|
||||
```bash
|
||||
make build
|
||||
make up
|
||||
```
|
||||
|
||||
* Der Portscanner prüft Ports 8000–8100
|
||||
* Der freie Port wird in `gateway/port.txt` gespeichert
|
||||
|
||||
Status prüfen:
|
||||
|
||||
```bash
|
||||
make status
|
||||
```
|
||||
|
||||
Stoppen:
|
||||
|
||||
```bash
|
||||
make down
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
# 7. API testen
|
||||
|
||||
## Healthcheck
|
||||
|
||||
```bash
|
||||
curl http://localhost:<PORT>/health
|
||||
```
|
||||
|
||||
## TTS Request
|
||||
|
||||
```bash
|
||||
curl -X POST http://localhost:<PORT>/v1/audio/speech \
|
||||
-H "Content-Type: application/json" \
|
||||
-d '{
|
||||
"model": "xtts-v2",
|
||||
"input": "Hello there!",
|
||||
"voice": "auto",
|
||||
"format": "wav"
|
||||
}' \
|
||||
--output output.wav
|
||||
```
|
||||
|
||||
## Stimme registrieren
|
||||
|
||||
```bash
|
||||
curl -X POST http://localhost:<PORT>/v1/voices/register \
|
||||
-H "Content-Type: application/json" \
|
||||
-d '{
|
||||
"name": "narrator",
|
||||
"samples": ["https://example.com/voice.wav"]
|
||||
}'
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
# 8. Entwicklung
|
||||
|
||||
## Lokales Setup ohne Docker
|
||||
|
||||
### Virtualenv
|
||||
|
||||
```bash
|
||||
python3 -m venv .venv
|
||||
source .venv/bin/activate
|
||||
```
|
||||
|
||||
### Abhängigkeiten installieren
|
||||
|
||||
Gateway:
|
||||
|
||||
```bash
|
||||
pip install -r gateway/requirements.gateway.txt
|
||||
```
|
||||
|
||||
Worker:
|
||||
|
||||
```bash
|
||||
pip install -r worker/requirements.worker.txt
|
||||
```
|
||||
|
||||
### Services starten
|
||||
|
||||
Redis (Lokal):
|
||||
|
||||
```bash
|
||||
docker run -p 6379:6379 redis:7
|
||||
```
|
||||
|
||||
Gateway:
|
||||
|
||||
```bash
|
||||
python gateway/main.py
|
||||
```
|
||||
|
||||
Worker:
|
||||
|
||||
```bash
|
||||
python worker/main.py
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
# 9. Skalieren
|
||||
|
||||
Mehrere Worker starten:
|
||||
|
||||
```bash
|
||||
make worker-scale N=3
|
||||
```
|
||||
|
||||
Der Gateway verteilt automatisch die Jobs über Redis.
|
||||
|
||||
---
|
||||
|
||||
# 10. Selftest
|
||||
|
||||
```bash
|
||||
make selftest
|
||||
```
|
||||
|
||||
Prüft:
|
||||
|
||||
* Redis erreichbar
|
||||
* Worker verarbeitet Jobs
|
||||
* Audioausgabe funktioniert
|
||||
|
||||
---
|
||||
|
||||
# 11. Troubleshooting
|
||||
|
||||
## Gateway startet nicht?
|
||||
|
||||
* Port belegt → `gateway/port.txt` prüfen
|
||||
* Logs prüfen → `gateway/logs/`
|
||||
* Redis erreichbar?
|
||||
|
||||
## Worker reagiert nicht?
|
||||
|
||||
* GPU verfügbar? → `nvidia-smi`
|
||||
* Torch kompatibel?
|
||||
* XTTS2 Modell lädt?
|
||||
|
||||
## Audio klingt falsch?
|
||||
|
||||
* Voice-Sample ungeeignet
|
||||
* Format falsch gesetzt
|
||||
* Sprache nicht angegeben
|
||||
|
||||
---
|
||||
|
||||
# 12. Best Practices
|
||||
|
||||
* Keine persönlichen Sprachsamples committen
|
||||
* Docker Images regelmäßig erneuern
|
||||
* Worker skalieren statt Gateway ändern
|
||||
* API-Versionen strikt pflegen
|
||||
* Code Style: PEP8
|
||||
|
||||
---
|
||||
|
||||
# 13. Weiterentwicklung
|
||||
|
||||
Empfohlene nächste Schritte:
|
||||
|
||||
* Monitoring (Prometheus)
|
||||
* WebSocket TTS
|
||||
* Multi-Model Routing
|
||||
* GUI für Voice Management
|
||||
* Auth Layer für öffentliche Deployments
|
||||
|
||||
---
|
||||
|
||||
# 14. Verantwortlichkeiten
|
||||
|
||||
* Projektleitung: Stephan W.
|
||||
* Backend Architektur: Team
|
||||
evtl. weitere Rollen später definieren
|
||||
|
||||
---
|
||||
|
||||
# 15. Abschluss
|
||||
|
||||
Wenn du bis hier gelesen hast, bist du vollständig einsatzfähig. Viel Erfolg!“}
|
||||
|
||||
182
README.md
182
README.md
@ -1,2 +1,180 @@
|
||||
# XTTS2 Server
|
||||
Professional README.
|
||||
# XTTS2 OpenAI-Compatible TTS Server
|
||||
|
||||
High-Performance Text-to-Speech Platform with FastAPI, Redis Queueing, GPU Workers & OpenAI Speech API Compatibility.
|
||||
|
||||
---
|
||||
|
||||
## 🔥 Übersicht
|
||||
|
||||
Dieses Projekt stellt eine vollständig modulare, skalierbare und produktionsreife lokale Text-to-Speech-Plattform bereit. Die API ist vollständig kompatibel zur **OpenAI Speech API**, unterstützt **XTTS2 Zero-Shot Voice Cloning**, mehrere Audioformate und verteilte GPU-Worker.
|
||||
|
||||
Zielsetzung:
|
||||
|
||||
* Hochqualitative TTS-Synthese für Spiele, Voice-Overs und AI-Produktion
|
||||
* Zero-Shot Voice Cloning mit XTTS2
|
||||
* OpenAI-kompatible Endpoints als Drop-in Replacement
|
||||
* Skalierbare Worker-Architektur für hohe Lasten
|
||||
|
||||
---
|
||||
|
||||
## 🧱 Architektur
|
||||
|
||||
```
|
||||
┌────────────────────┐ ┌───────────────────────────────┐
|
||||
│ FastAPI Gateway │◀──────┤ Redis Queue + Cache │
|
||||
│ - OpenAI API │ └───────────────────────────────┘
|
||||
│ - Rate Limits │
|
||||
│ - Port Auto-Select │ ┌───────────────────────────────┐
|
||||
│ - Voice Registry │──────▶│ Worker (XTTS2) │
|
||||
└────────────────────┘ │ - GPU/CPU Auto Detect │
|
||||
│ - Zero-Shot Voice Cloning │
|
||||
└───────────────────────────────┘
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 🚀 Features
|
||||
|
||||
### Core Features
|
||||
|
||||
* OpenAI-kompatible Speech API
|
||||
* Zero-Shot Voice Cloning mit XTTS2
|
||||
* Unterstützung für `wav`, `mp3`, `ogg`
|
||||
* Dynamische Portwahl (8000–8100) mit Fallback
|
||||
* Registry für permanente Stimmen
|
||||
* Queue-basierte Worker-Architektur (Redis)
|
||||
* GPU Auto-Detection für Worker
|
||||
|
||||
### Deployment Features
|
||||
|
||||
* Multi-Stage Docker Images (Gateway & Worker)
|
||||
* Makefile PRO für Build, Deploy, Scaling & Testing
|
||||
* Logs + Portfile + Healthchecks
|
||||
* Lazy Load der XTTS2-Modelle
|
||||
|
||||
---
|
||||
|
||||
## 📦 Installation
|
||||
|
||||
### Voraussetzungen
|
||||
|
||||
* Docker & Docker Compose
|
||||
* NVIDIA Container Runtime (für GPU-Worker)
|
||||
* Linux oder macOS (Windows WSL2 möglich)
|
||||
|
||||
### Start in 3 Schritten
|
||||
|
||||
```bash
|
||||
make build
|
||||
make up
|
||||
make status
|
||||
```
|
||||
|
||||
Die dynamische Portwahl speichert den genutzten Port in:
|
||||
|
||||
```
|
||||
gateway/port.txt
|
||||
gateway/logs/gateway.log
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 📡 OpenAI-kompatible Endpunkte
|
||||
|
||||
### POST /v1/audio/speech
|
||||
|
||||
Request Beispiel:
|
||||
|
||||
```json
|
||||
{
|
||||
"model": "xtts-v2",
|
||||
"input": "Hello, hero",
|
||||
"voice": "auto",
|
||||
"format": "wav"
|
||||
}
|
||||
```
|
||||
|
||||
Antwort:
|
||||
|
||||
* binares Audio
|
||||
* Content-Type abhängig vom Format
|
||||
|
||||
### POST /v1/voices/register
|
||||
|
||||
* Registriert permanente Stimmen
|
||||
* Unterstützt Samples per URL oder Base64
|
||||
|
||||
### GET /health
|
||||
|
||||
* Healthcheck für Monitoring & CI
|
||||
|
||||
---
|
||||
|
||||
## 🧪 Selftest
|
||||
|
||||
```bash
|
||||
make selftest
|
||||
```
|
||||
|
||||
Prüft:
|
||||
|
||||
* Redis erreichbar
|
||||
* Gateway erreichbar
|
||||
* Worker zieht Jobs
|
||||
* Mini-Synthese erfolgreich
|
||||
|
||||
---
|
||||
|
||||
## 📁 Projektstruktur
|
||||
|
||||
```
|
||||
project/
|
||||
├── gateway/
|
||||
│ ├── main.py
|
||||
│ ├── api/
|
||||
│ ├── core/
|
||||
│ ├── voices/
|
||||
│ ├── logs/
|
||||
│ └── Dockerfile
|
||||
│
|
||||
├── worker/
|
||||
│ ├── main.py
|
||||
│ ├── engine/
|
||||
│ ├── core/
|
||||
│ └── Dockerfile
|
||||
│
|
||||
├── scripts/
|
||||
│ └── find_port.py
|
||||
│
|
||||
├── Makefile
|
||||
├── docker-compose.yml
|
||||
└── README.md
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 🔐 Sicherheit & Best Practices
|
||||
|
||||
* Optional: API Keys für öffentliche Deployments
|
||||
* Rate Limits im Gateway aktivierbar
|
||||
* HTTPS über Reverse Proxy
|
||||
* Keine sensiblen Voice-Daten einchecken
|
||||
* Worker nur intern erreichbar halten
|
||||
|
||||
---
|
||||
|
||||
## 🛠 Roadmap
|
||||
|
||||
* Prometheus & Grafana Monitoring
|
||||
* Business-Level Logging (JSON Logs)
|
||||
* Support für weitere Modelle (F5, Kokoro, Piper)
|
||||
* WebSocket Realtime TTS Output
|
||||
|
||||
---
|
||||
|
||||
## 🧑💻 Maintainer
|
||||
|
||||
**Stephan W.** – Architektur & Betrieb
|
||||
|
||||
```}
|
||||
```
|
||||
|
||||
@ -1 +1,54 @@
|
||||
version: '3.9'
|
||||
# docker-compose.yml – XTTS2 TTS Server
|
||||
|
||||
# Multi-Service Orchestrierung (Gateway, Worker, Redis)
|
||||
|
||||
version: "3.9"
|
||||
|
||||
services:
|
||||
redis:
|
||||
image: redis:7
|
||||
container_name: redis
|
||||
restart: always
|
||||
ports:
|
||||
- "6379:6379"
|
||||
|
||||
gateway:
|
||||
build:
|
||||
context: ./gateway
|
||||
dockerfile: Dockerfile
|
||||
container_name: tts-gateway
|
||||
restart: always
|
||||
environment:
|
||||
- REDIS_HOST=redis
|
||||
- GATEWAY_PORT=${GATEWAY_PORT}
|
||||
volumes:
|
||||
- ./gateway/voices:/app/voices
|
||||
- ./gateway/logs:/app/logs
|
||||
- ./gateway/port.txt:/app/port.txt
|
||||
ports:
|
||||
- "${GATEWAY_PORT}:8000"
|
||||
depends_on:
|
||||
- redis
|
||||
|
||||
worker:
|
||||
build:
|
||||
context: ./worker
|
||||
dockerfile: Dockerfile
|
||||
container_name: tts-worker
|
||||
restart: always
|
||||
environment:
|
||||
- REDIS_HOST=redis
|
||||
- GPU_MODE=AUTO
|
||||
deploy:
|
||||
resources:
|
||||
reservations:
|
||||
devices:
|
||||
- capabilities: [gpu]
|
||||
volumes:
|
||||
- ./worker/voices:/app/voices
|
||||
depends_on:
|
||||
- redis
|
||||
|
||||
networks:
|
||||
default:
|
||||
name: wlkns-net
|
||||
|
||||
@ -1 +1,11 @@
|
||||
FROM python:3.11-slim AS builder
|
||||
WORKDIR /app
|
||||
COPY requirements.gateway.txt .
|
||||
RUN pip install -r requirements.gateway.txt
|
||||
COPY . .
|
||||
|
||||
FROM python:3.11-slim
|
||||
WORKDIR /app
|
||||
COPY --from=builder /usr/local /usr/local
|
||||
COPY . .
|
||||
CMD ["python","main.py"]
|
||||
|
||||
7
gateway/api/health.py
Normal file
7
gateway/api/health.py
Normal file
@ -0,0 +1,7 @@
|
||||
from fastapi import APIRouter
|
||||
|
||||
router = APIRouter()
|
||||
|
||||
@router.get("/")
|
||||
async def health():
|
||||
return {"status": "ok"}
|
||||
19
gateway/api/openai_speech.py
Normal file
19
gateway/api/openai_speech.py
Normal file
@ -0,0 +1,19 @@
|
||||
from fastapi import APIRouter, Response
|
||||
from pydantic import BaseModel
|
||||
from core.queue_client import push_job, await_result
|
||||
|
||||
router = APIRouter()
|
||||
|
||||
class SpeechRequest(BaseModel):
|
||||
model: str
|
||||
input: str
|
||||
voice: str = "auto"
|
||||
format: str = "wav"
|
||||
voice_sample_url: str | None = None
|
||||
voice_sample_base64: str | None = None
|
||||
|
||||
@router.post("/speech")
|
||||
async def speech(req: SpeechRequest):
|
||||
job_id = push_job(req.dict())
|
||||
audio_bytes, mime = await_result(job_id)
|
||||
return Response(content=audio_bytes, media_type=mime)
|
||||
21
gateway/api/voices.py
Normal file
21
gateway/api/voices.py
Normal file
@ -0,0 +1,21 @@
|
||||
from fastapi import APIRouter
|
||||
import json
|
||||
from pathlib import Path
|
||||
|
||||
router = APIRouter()
|
||||
REG = Path("voices/registry.json")
|
||||
|
||||
@router.post("/register")
|
||||
async def register(data: dict):
|
||||
reg = {}
|
||||
if REG.exists():
|
||||
reg = json.loads(REG.read_text())
|
||||
reg[data["name"]] = data
|
||||
REG.write_text(json.dumps(reg, indent=2))
|
||||
return {"status": "ok", "voices": list(reg.keys())}
|
||||
|
||||
@router.get("/list")
|
||||
async def list_voices():
|
||||
if not REG.exists():
|
||||
return []
|
||||
return json.loads(REG.read_text())
|
||||
1
gateway/core/cache.py
Normal file
1
gateway/core/cache.py
Normal file
@ -0,0 +1 @@
|
||||
# optional cache placeholder
|
||||
1
gateway/core/config.py
Normal file
1
gateway/core/config.py
Normal file
@ -0,0 +1 @@
|
||||
REDIS_HOST='redis'
|
||||
1
gateway/core/logging_config.py
Normal file
1
gateway/core/logging_config.py
Normal file
@ -0,0 +1 @@
|
||||
# logging config placeholder
|
||||
1
gateway/core/models.py
Normal file
1
gateway/core/models.py
Normal file
@ -0,0 +1 @@
|
||||
# pydantic models placeholder
|
||||
25
gateway/core/queue_client.py
Normal file
25
gateway/core/queue_client.py
Normal file
@ -0,0 +1,25 @@
|
||||
import redis, uuid, json, time, os
|
||||
|
||||
REDIS_HOST = os.environ.get("REDIS_HOST", "redis")
|
||||
r = redis.Redis(host=REDIS_HOST, port=6379, db=0)
|
||||
|
||||
QUEUE="tts_queue"
|
||||
RESULT="tts_result"
|
||||
|
||||
def push_job(data: dict) -> str:
|
||||
job_id = str(uuid.uuid4())
|
||||
data["job_id"] = job_id
|
||||
r.lpush(QUEUE, json.dumps(data))
|
||||
return job_id
|
||||
|
||||
def await_result(job_id: str, timeout=30):
|
||||
key = f"{RESULT}:{job_id}"
|
||||
start=time.time()
|
||||
while time.time()-start < timeout:
|
||||
data=r.get(key)
|
||||
if data:
|
||||
r.delete(key)
|
||||
obj=json.loads(data)
|
||||
return bytes.fromhex(obj["audio"]), obj["mime"]
|
||||
time.sleep(0.1)
|
||||
raise TimeoutError("Worker antwortet nicht")
|
||||
@ -1 +1,42 @@
|
||||
print('gateway start')
|
||||
import socket, uvicorn, os
|
||||
from fastapi import FastAPI
|
||||
from api.openai_speech import router as speech_router
|
||||
from api.voices import router as voice_router
|
||||
from api.health import router as health_router
|
||||
from pathlib import Path
|
||||
|
||||
app = FastAPI(title="XTTS2 Gateway")
|
||||
|
||||
app.include_router(speech_router, prefix="/v1/audio")
|
||||
app.include_router(voice_router, prefix="/v1/voices")
|
||||
app.include_router(health_router, prefix="/health")
|
||||
|
||||
PORT_FILE = Path("port.txt")
|
||||
|
||||
def find_free_port(start=8000, end=8100):
|
||||
for port in range(start, end+1):
|
||||
with socket.socket(socket.AF_INET, socket.SOCK_STREAM) as s:
|
||||
try:
|
||||
s.bind(("0.0.0.0", port))
|
||||
return port
|
||||
except:
|
||||
pass
|
||||
raise RuntimeError("No free port")
|
||||
|
||||
def start():
|
||||
env_port = os.environ.get("GATEWAY_PORT")
|
||||
port = int(env_port) if env_port else 8000
|
||||
|
||||
# check
|
||||
with socket.socket(socket.AF_INET, socket.SOCK_STREAM) as s:
|
||||
try:
|
||||
s.bind(("0.0.0.0", port))
|
||||
except:
|
||||
port = find_free_port()
|
||||
|
||||
PORT_FILE.write_text(str(port))
|
||||
|
||||
uvicorn.run(app, host="0.0.0.0", port=port)
|
||||
|
||||
if __name__ == "__main__":
|
||||
start()
|
||||
|
||||
@ -1 +1,5 @@
|
||||
fastapi
|
||||
uvicorn
|
||||
redis
|
||||
pydantic
|
||||
requests
|
||||
|
||||
1
gateway/voices/registry.json
Normal file
1
gateway/voices/registry.json
Normal file
@ -0,0 +1 @@
|
||||
{}
|
||||
@ -1,2 +1,20 @@
|
||||
#!/usr/bin/env python3
|
||||
print(8000)
|
||||
import socket
|
||||
|
||||
START, END = 8000, 8100
|
||||
|
||||
def port_free(port):
|
||||
with socket.socket(socket.AF_INET, socket.SOCK_STREAM) as s:
|
||||
try:
|
||||
s.bind(("0.0.0.0", port))
|
||||
return True
|
||||
except OSError:
|
||||
return False
|
||||
|
||||
for port in range(START, END + 1):
|
||||
if port_free(port):
|
||||
print(port)
|
||||
exit(0)
|
||||
|
||||
print("ERR_NO_FREE_PORT")
|
||||
exit(1)
|
||||
|
||||
40
scripts/selftest.py
Normal file
40
scripts/selftest.py
Normal file
@ -0,0 +1,40 @@
|
||||
import requests, sys, json
|
||||
|
||||
def main():
|
||||
# Try read port file
|
||||
try:
|
||||
with open("gateway/port.txt") as f:
|
||||
port = f.read().strip()
|
||||
except:
|
||||
print("❌ port.txt nicht gefunden")
|
||||
return
|
||||
|
||||
base = f"http://localhost:{port}"
|
||||
|
||||
print("🔍 Healthcheck…")
|
||||
try:
|
||||
r = requests.get(base + "/health")
|
||||
print("Health:", r.status_code, r.text)
|
||||
except Exception as e:
|
||||
print("❌ Healthcheck fehlgeschlagen:", e)
|
||||
return
|
||||
|
||||
print("🎤 Test-Synthese…")
|
||||
try:
|
||||
payload = {
|
||||
"model": "xtts-v2",
|
||||
"input": "Dies ist ein Test.",
|
||||
"voice": "auto",
|
||||
"format": "wav"
|
||||
}
|
||||
r = requests.post(base + "/v1/audio/speech", json=payload)
|
||||
print("TTS Status:", r.status_code)
|
||||
if r.status_code == 200:
|
||||
print("OK ✓")
|
||||
else:
|
||||
print("❌ TTS Fehler:", r.text)
|
||||
except Exception as e:
|
||||
print("❌ Synthese fehlgeschlagen:", e)
|
||||
|
||||
if __name__ == "__main__":
|
||||
main()
|
||||
@ -1 +1,7 @@
|
||||
FROM nvidia/cuda:12.1.0-runtime-ubuntu22.04
|
||||
FROM nvidia/cuda:12.1.0-runtime-ubuntu22.04 AS base
|
||||
WORKDIR /app
|
||||
COPY requirements.worker.txt .
|
||||
RUN apt-get update && apt-get install -y python3-pip ffmpeg
|
||||
RUN pip3 install -r requirements.worker.txt
|
||||
COPY . .
|
||||
CMD ["python3","main.py"]
|
||||
|
||||
4
worker/core/gpu_detect.py
Normal file
4
worker/core/gpu_detect.py
Normal file
@ -0,0 +1,4 @@
|
||||
import torch
|
||||
|
||||
def device():
|
||||
return 'cuda' if torch.cuda.is_available() else 'cpu'
|
||||
18
worker/core/queue_worker.py
Normal file
18
worker/core/queue_worker.py
Normal file
@ -0,0 +1,18 @@
|
||||
import redis, json, os
|
||||
|
||||
REDIS_HOST=os.environ.get("REDIS_HOST","redis")
|
||||
r=redis.Redis(host=REDIS_HOST, port=6379, db=0)
|
||||
|
||||
QUEUE="tts_queue"
|
||||
RESULT="tts_result"
|
||||
|
||||
def fetch_job():
|
||||
data=r.rpop(QUEUE)
|
||||
if not data:
|
||||
return None
|
||||
return json.loads(data)
|
||||
|
||||
def store_result(job_id:str, audio:bytes, mime:str):
|
||||
key=f"{RESULT}:{job_id}"
|
||||
obj={"audio": audio.hex(), "mime": mime}
|
||||
r.set(key, json.dumps(obj))
|
||||
19
worker/engine/audio_export.py
Normal file
19
worker/engine/audio_export.py
Normal file
@ -0,0 +1,19 @@
|
||||
from pydub import AudioSegment
|
||||
import io
|
||||
|
||||
def convert_audio(raw_bytes: bytes, fmt: str):
|
||||
# raw mono 32-bit float fake waveform
|
||||
seg = AudioSegment(
|
||||
raw_bytes,
|
||||
frame_rate=22050,
|
||||
sample_width=4,
|
||||
channels=1
|
||||
)
|
||||
buf=io.BytesIO()
|
||||
seg.export(buf, format=fmt)
|
||||
mime={
|
||||
"wav":"audio/wav",
|
||||
"mp3":"audio/mpeg",
|
||||
"ogg":"audio/ogg"
|
||||
}.get(fmt,"audio/wav")
|
||||
return buf.getvalue(), mime
|
||||
10
worker/engine/xtts2_loader.py
Normal file
10
worker/engine/xtts2_loader.py
Normal file
@ -0,0 +1,10 @@
|
||||
# Dummy XTTS2 logic placeholder
|
||||
# Replace with real TTS model loading
|
||||
|
||||
def synthesize(job: dict):
|
||||
# return artificial sine wave placeholder
|
||||
import numpy as np
|
||||
sr=22050
|
||||
t=np.linspace(0,0.3,int(sr*0.3))
|
||||
tone=(0.1*np.sin(2*np.pi*440*t)).astype('float32')
|
||||
return tone.tobytes()
|
||||
@ -1 +1,16 @@
|
||||
print('worker start')
|
||||
import time, json, os
|
||||
from core.queue_worker import fetch_job, store_result
|
||||
from engine.xtts2_loader import synthesize
|
||||
from engine.audio_export import convert_audio
|
||||
|
||||
print("Worker gestartet. Warte auf Jobs…")
|
||||
|
||||
while True:
|
||||
job = fetch_job()
|
||||
if not job:
|
||||
time.sleep(0.1)
|
||||
continue
|
||||
|
||||
audio = synthesize(job)
|
||||
out, mime = convert_audio(audio, job.get("format","wav"))
|
||||
store_result(job["job_id"], out, mime)
|
||||
|
||||
@ -1 +1,5 @@
|
||||
TTS
|
||||
pydub
|
||||
redis
|
||||
torch
|
||||
numpy
|
||||
requests
|
||||
|
||||
1
worker/voices/registry.json
Normal file
1
worker/voices/registry.json
Normal file
@ -0,0 +1 @@
|
||||
{}
|
||||
Reference in New Issue
Block a user