feat: Add XTTS v2 support, refactor Docker/GPU infra, and improve Piper engine

- Add XTTS v2 configuration to .env.example
- Refactor Dockerfile to multi-stage build with CUDA 12.1 support
- Update Makefile with Kokoro and XTTS test environment targets
- Refactor Piper engine (app/engines/piper.py) to use python module execution
- Add comprehensive documentation for Kokoro and XTTS plans
- Add helper scripts and patches for build process
This commit is contained in:
2025-12-13 11:37:58 +01:00
parent d6d1fe9d23
commit 91887ae296
19 changed files with 1625 additions and 46 deletions

View File

@ -0,0 +1,492 @@
# Kokoro TTS Engine Implementation Plan
## Overview
Implement Kokoro-82M as a new TTS engine in AudioEngineHub. Kokoro is a lightweight, high-performance open-weight TTS model with 82 million parameters that delivers quality comparable to models 5-15× its size.
## Background Research
### Key Features
- **Size**: 82 million parameters (extremely lightweight)
- **Performance**: ~210× real-time on RTX 4090, ~90× real-time on RTX 3090 Ti
- **Quality**: Took first place in TTS Spaces Arena, outperforming XTTS v2 (467M) and MetaVoice (1.2B)
- **Audio**: 24kHz high-fidelity output
- **License**: Apache 2.0 (open-source, commercial use allowed)
- **Languages**: 8 languages (English US/UK, French, Spanish, Japanese, Chinese, Italian, Portuguese, Hindi, Korean)
- **Voices**: 54 voices available
### Sources
- [Kokoro-82M Hugging Face](https://huggingface.co/hexgrad/Kokoro-82M)
- [Kokoro TTS Official Site](https://kokorotts.net/)
- [VOICES.md - Complete Voice List](https://huggingface.co/hexgrad/Kokoro-82M/blob/main/VOICES.md)
- [GitHub Repository](https://github.com/hexgrad/kokoro)
- [Analytics Vidhya Article](https://www.analyticsvidhya.com/blog/2025/01/kokoro-82m/)
---
## Implementation Strategy
### Phase 1: Research & Setup
#### 1.1 Model Investigation
- [x] Research Kokoro TTS architecture and capabilities
- [x] Identify Python library: `kokoro>=0.9.2`
- [x] Document voice list (54 voices across 8 languages)
- [ ] Test Kokoro locally to understand API
#### 1.2 Dependency Analysis
**Required packages:**
```bash
kokoro>=0.9.2
soundfile
phonemizer
torch
transformers
scipy
munch
```
**System dependencies:**
```bash
espeak-ng # Required for phonemization
```
---
### Phase 2: Engine Implementation
#### 2.1 Create `app/engines/kokoro.py`
**Architecture:**
```python
class KokoroEngine(TTSEngineBase):
def __init__(self):
# Initialize Kokoro pipeline
# Handle GPU/CPU detection
# Load model from Hugging Face
async def synthesize(text, speaker, model, fmt):
# Generate audio using KPipeline
# Handle voice selection
# Convert to requested format (wav/ogg/mp3)
# Return temp file path
def list_models(self):
# Return available language models
# Options: 'a' (American English), 'b' (British English), etc.
def list_voices(self, model):
# Return 54 available voices
# Filter by language if model specified
def healthcheck(self):
# Check if kokoro library is available
# Verify model is loaded
# Return status
async def selftest(self):
# Run quick synthesis test
# Verify audio generation works
```
**Key Implementation Details:**
1. **Model Selection:**
- Kokoro uses `lang_code` parameter (e.g., 'a' = American English, 'b' = British English)
- Map this to "models" concept in our API
- Models: `kokoro-en-us`, `kokoro-en-gb`, `kokoro-fr`, `kokoro-es`, `kokoro-ja`, `kokoro-zh`, `kokoro-it`, `kokoro-pt`, `kokoro-hi`, `kokoro-ko`
2. **Voice Selection:**
- 54 voices available (see VOICES.md)
- Popular voices: `af_alloy`, `af_bella`, `af_sarah`, `af_sky`, `af_nova`, etc.
- Each voice has quality grade and language support
3. **Audio Generation:**
```python
from kokoro import KPipeline
pipeline = KPipeline(lang_code='a')
generator = pipeline(text, voice='af_heart')
for gs, ps, audio in generator:
# audio is numpy array at 24kHz
# Save to temp file
```
4. **Format Conversion:**
- Native output: 24kHz WAV
- Use ffmpeg (already available) for OGG/MP3 conversion
- Reuse `_run_ffmpeg_blocking()` pattern from Piper engine
5. **Error Handling:**
- Apply all bug fixes from Piper engine (timeouts, temp file cleanup, logging)
- Handle GPU out-of-memory gracefully (fallback to CPU)
- Validate voice exists before synthesis
---
### Phase 3: Docker Integration
#### 3.1 Update Dockerfile
**Add dependencies to `Dockerfile`:**
```dockerfile
# Install espeak-ng for Kokoro phonemization
RUN apt-get update && \
apt-get install -y --no-install-recommends espeak-ng && \
apt-get clean && \
rm -rf /var/lib/apt/lists/*
# Install Kokoro Python dependencies
RUN pip install --no-cache-dir \
kokoro>=0.9.2 \
soundfile \
phonemizer \
torch \
transformers \
scipy \
munch
```
#### 3.2 Model Download Strategy
**Options:**
**Option A: Download on first use (lazy loading)**
- Models auto-download from Hugging Face (~200MB per model)
- Advantage: No pre-download needed
- Disadvantage: First synthesis will be slow
**Option B: Pre-download in Docker build**
```dockerfile
# Pre-download Kokoro model during build
RUN python3 -c "from kokoro import KPipeline; KPipeline(lang_code='a')"
```
**Option C: Volume mount like Piper**
```yaml
# docker-compose.yml
volumes:
- ./models/kokoro:/home/appuser/.cache/huggingface
```
**Recommendation**: Option A for MVP, Option C for production
---
### Phase 4: Configuration
#### 4.1 Update `app/config.py`
```python
class Settings(BaseSettings):
# ... existing settings ...
# Kokoro Engine Configuration
KOKORO_DEVICE: str = "cuda" # or "cpu"
KOKORO_TIMEOUT_SECONDS: int = 30
KOKORO_DEFAULT_LANG: str = "a" # American English
```
#### 4.2 Update `.env.example`
```bash
# Kokoro Engine Configuration
KOKORO_DEVICE=cuda # cuda or cpu
KOKORO_TIMEOUT_SECONDS=30
KOKORO_DEFAULT_LANG=a
```
#### 4.3 Update `app/main.py`
```python
ALL_ENGINES = {
"piper": PiperEngine,
"styletts": StyleTTSEngine,
"chattts": ChatTTSEngine,
"f5-tts": F5TTSEngine,
"kokoro": KokoroEngine, # Add this
}
```
---
### Phase 5: Voice Metadata
#### 5.1 Create `app/engines/kokoro_voices.py`
Store voice metadata for better UX:
```python
KOKORO_VOICES = {
"af_alloy": {
"gender": "F",
"language": ["en-us"],
"quality": "high",
"description": "Clear, professional female voice"
},
"af_bella": {
"gender": "F",
"language": ["en-us", "en-gb"],
"quality": "high",
"description": "Warm, expressive female voice"
},
# ... all 54 voices
}
```
Download full voice list from: https://huggingface.co/hexgrad/Kokoro-82M/blob/main/VOICES.md
---
### Phase 6: Testing
#### 6.1 Unit Tests
Create `tests/test_kokoro_engine.py`:
```python
import pytest
from app.engines.kokoro import KokoroEngine
@pytest.mark.asyncio
async def test_kokoro_synthesis():
engine = KokoroEngine()
audio_path = await engine.synthesize(
"Hello, this is Kokoro TTS.",
speaker="af_bella",
model="kokoro-en-us",
fmt="wav"
)
assert os.path.exists(audio_path)
assert os.path.getsize(audio_path) > 0
def test_kokoro_list_voices():
engine = KokoroEngine()
voices = engine.list_voices("kokoro-en-us")
assert len(voices) > 0
assert "af_bella" in voices
def test_kokoro_healthcheck():
engine = KokoroEngine()
health = engine.healthcheck()
assert health["status"] in ["ok", "not_available"]
```
#### 6.2 Integration Tests
```bash
# Test via API
curl -X POST http://localhost:8000/tts \
-H "Content-Type: application/json" \
-d '{
"text": "Hello from Kokoro TTS!",
"engine": "kokoro",
"model": "kokoro-en-us",
"speaker": "af_bella",
"format": "ogg"
}'
```
#### 6.3 Performance Benchmarks
Test synthesis speed:
- Short text (10 words): Target <0.5s on RTX 3090
- Medium text (100 words): Target <2s on RTX 3090
- Long text (1000 words): Target <15s on RTX 3090
---
### Phase 7: Documentation
#### 7.1 Update README.md
Add Kokoro to supported engines:
```markdown
## Supported TTS Engines
- **Piper** - Fast, lightweight, ONNX-based TTS
- **Kokoro** - 82M parameter high-quality TTS (NEW!)
- **StyleTTS** - Expressive style-based TTS (planned)
- **ChatTTS** - Conversational TTS (planned)
- **F5-TTS** - Advanced flow-based TTS (planned)
```
#### 7.2 Update CLAUDE.md
Add Kokoro engine details:
```markdown
### Kokoro Engine (`app/engines/kokoro.py`)
- Uses `kokoro` Python library (KPipeline)
- 54 voices across 8 languages
- Outputs 24kHz audio natively
- Extremely fast (~90x real-time on consumer GPU)
- Model size: 82M parameters (~200MB download)
```
#### 7.3 Create Kokoro Usage Guide
Create `docs/kokoro-guide.md`:
- Voice selection guide
- Language support matrix
- Performance optimization tips
- GPU vs CPU mode comparison
- Troubleshooting common issues
---
## Implementation Checklist
### Phase 1: Research & Setup
- [x] Research Kokoro capabilities
- [x] Document API and dependencies
- [ ] Test Kokoro locally outside Docker
### Phase 2: Engine Implementation
- [ ] Create `app/engines/kokoro.py`
- [ ] Implement `synthesize()` method
- [ ] Implement `list_models()` method
- [ ] Implement `list_voices()` method
- [ ] Implement `healthcheck()` method
- [ ] Implement `selftest()` method
- [ ] Create `app/engines/kokoro_voices.py` metadata file
### Phase 3: Docker Integration
- [ ] Update Dockerfile with dependencies
- [ ] Add espeak-ng system package
- [ ] Add Kokoro Python packages
- [ ] Test Docker build
- [ ] Verify GPU access in container
### Phase 4: Configuration
- [ ] Update `app/config.py` with Kokoro settings
- [ ] Update `.env.example` with Kokoro variables
- [ ] Add KokoroEngine to `app/main.py` ALL_ENGINES
- [ ] Update active engines in `.env`
### Phase 5: Testing
- [ ] Create unit tests
- [ ] Test synthesis with various voices
- [ ] Test format conversion (wav/ogg/mp3)
- [ ] Test caching behavior
- [ ] Performance benchmarks
- [ ] Memory usage profiling
### Phase 6: Documentation
- [ ] Update README.md
- [ ] Update CLAUDE.md
- [ ] Create Kokoro usage guide
- [ ] Document voice selection
- [ ] Add troubleshooting section
### Phase 7: Deployment
- [ ] Test in development environment
- [ ] Rebuild Docker image
- [ ] Update docker-compose.yml if needed
- [ ] Test API endpoints
- [ ] Commit changes to git
- [ ] Tag release
---
## Key Decisions
### 1. Model Organization
**Decision**: Map Kokoro lang_codes to model names
- `kokoro-en-us` → lang_code='a'
- `kokoro-en-gb` → lang_code='b'
- etc.
**Rationale**: Maintains consistency with existing API structure
### 2. Voice Naming
**Decision**: Use Kokoro's native voice names (e.g., `af_bella`)
**Rationale**:
- Avoids confusion with remapping
- Documented in official VOICES.md
- Users can reference official docs
### 3. GPU Support
**Decision**: Support both GPU and CPU with fallback
**Rationale**:
- GPU provides 90x real-time performance
- CPU fallback ensures it works on all systems
- Configurable via KOKORO_DEVICE env var
### 4. Model Download Strategy
**Decision**: Lazy loading on first use
**Rationale**:
- Smaller Docker image
- Only download models that are actually used
- Can switch to pre-download later if needed
---
## Potential Issues & Solutions
### Issue 1: Large Model Download
**Problem**: Model is ~200MB per language
**Solution**:
- Lazy loading (download on first use)
- Cache in Docker volume
- Document expected download time
### Issue 2: GPU Memory Usage
**Problem**: May require 2-4GB VRAM
**Solution**:
- Implement memory monitoring
- Graceful fallback to CPU
- Document GPU requirements
### Issue 3: First Synthesis Slow
**Problem**: Model loading + compilation takes time
**Solution**:
- Warm up during healthcheck
- Keep model loaded in memory
- Document expected first-run delay
### Issue 4: Voice Compatibility
**Problem**: Not all voices work with all languages
**Solution**:
- Validate voice-language compatibility
- Return clear error messages
- Document voice language support
---
## Success Criteria
- [ ] Kokoro engine passes all unit tests
- [ ] API endpoints work correctly (`/tts`, `/models`, `/speakers`)
- [ ] Audio quality matches expected output
- [ ] Synthesis speed: >50x real-time on RTX 3090
- [ ] Memory usage: <4GB VRAM for single synthesis
- [ ] Cache versioning works correctly
- [ ] No temp file leaks
- [ ] No hung processes
- [ ] Documentation complete and accurate
- [ ] Voice wizard integration works
---
## Timeline Estimate
**Total**: 8-12 hours
- **Phase 1 (Research)**: ✅ Complete (2 hours)
- **Phase 2 (Implementation)**: 3-4 hours
- **Phase 3 (Docker)**: 1-2 hours
- **Phase 4 (Config)**: 0.5 hours
- **Phase 5 (Testing)**: 2-3 hours
- **Phase 6 (Documentation)**: 1-2 hours
- **Phase 7 (Deployment)**: 1 hour
---
## References
- [Kokoro-82M Hugging Face](https://huggingface.co/hexgrad/Kokoro-82M)
- [Kokoro GitHub Repository](https://github.com/hexgrad/kokoro)
- [VOICES.md - Complete Voice List](https://huggingface.co/hexgrad/Kokoro-82M/blob/main/VOICES.md)
- [Kokoro TTS Official Website](https://kokorotts.net/)
- [Kokoro-82M Analytics Vidhya Article](https://www.analyticsvidhya.com/blog/2025/01/kokoro-82m/)
- [Kokoro TTS Live Demo](https://huggingface.co/spaces/hexgrad/Kokoro-TTS)