Files
audio-engine-hub/KOKORO_IMPLEMENTATION_PLAN.md
stephan 91887ae296 feat: Add XTTS v2 support, refactor Docker/GPU infra, and improve Piper engine
- Add XTTS v2 configuration to .env.example
- Refactor Dockerfile to multi-stage build with CUDA 12.1 support
- Update Makefile with Kokoro and XTTS test environment targets
- Refactor Piper engine (app/engines/piper.py) to use python module execution
- Add comprehensive documentation for Kokoro and XTTS plans
- Add helper scripts and patches for build process
2025-12-13 11:37:58 +01:00

493 lines
13 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

# Kokoro TTS Engine Implementation Plan
## Overview
Implement Kokoro-82M as a new TTS engine in AudioEngineHub. Kokoro is a lightweight, high-performance open-weight TTS model with 82 million parameters that delivers quality comparable to models 5-15× its size.
## Background Research
### Key Features
- **Size**: 82 million parameters (extremely lightweight)
- **Performance**: ~210× real-time on RTX 4090, ~90× real-time on RTX 3090 Ti
- **Quality**: Took first place in TTS Spaces Arena, outperforming XTTS v2 (467M) and MetaVoice (1.2B)
- **Audio**: 24kHz high-fidelity output
- **License**: Apache 2.0 (open-source, commercial use allowed)
- **Languages**: 8 languages (English US/UK, French, Spanish, Japanese, Chinese, Italian, Portuguese, Hindi, Korean)
- **Voices**: 54 voices available
### Sources
- [Kokoro-82M Hugging Face](https://huggingface.co/hexgrad/Kokoro-82M)
- [Kokoro TTS Official Site](https://kokorotts.net/)
- [VOICES.md - Complete Voice List](https://huggingface.co/hexgrad/Kokoro-82M/blob/main/VOICES.md)
- [GitHub Repository](https://github.com/hexgrad/kokoro)
- [Analytics Vidhya Article](https://www.analyticsvidhya.com/blog/2025/01/kokoro-82m/)
---
## Implementation Strategy
### Phase 1: Research & Setup
#### 1.1 Model Investigation
- [x] Research Kokoro TTS architecture and capabilities
- [x] Identify Python library: `kokoro>=0.9.2`
- [x] Document voice list (54 voices across 8 languages)
- [ ] Test Kokoro locally to understand API
#### 1.2 Dependency Analysis
**Required packages:**
```bash
kokoro>=0.9.2
soundfile
phonemizer
torch
transformers
scipy
munch
```
**System dependencies:**
```bash
espeak-ng # Required for phonemization
```
---
### Phase 2: Engine Implementation
#### 2.1 Create `app/engines/kokoro.py`
**Architecture:**
```python
class KokoroEngine(TTSEngineBase):
def __init__(self):
# Initialize Kokoro pipeline
# Handle GPU/CPU detection
# Load model from Hugging Face
async def synthesize(text, speaker, model, fmt):
# Generate audio using KPipeline
# Handle voice selection
# Convert to requested format (wav/ogg/mp3)
# Return temp file path
def list_models(self):
# Return available language models
# Options: 'a' (American English), 'b' (British English), etc.
def list_voices(self, model):
# Return 54 available voices
# Filter by language if model specified
def healthcheck(self):
# Check if kokoro library is available
# Verify model is loaded
# Return status
async def selftest(self):
# Run quick synthesis test
# Verify audio generation works
```
**Key Implementation Details:**
1. **Model Selection:**
- Kokoro uses `lang_code` parameter (e.g., 'a' = American English, 'b' = British English)
- Map this to "models" concept in our API
- Models: `kokoro-en-us`, `kokoro-en-gb`, `kokoro-fr`, `kokoro-es`, `kokoro-ja`, `kokoro-zh`, `kokoro-it`, `kokoro-pt`, `kokoro-hi`, `kokoro-ko`
2. **Voice Selection:**
- 54 voices available (see VOICES.md)
- Popular voices: `af_alloy`, `af_bella`, `af_sarah`, `af_sky`, `af_nova`, etc.
- Each voice has quality grade and language support
3. **Audio Generation:**
```python
from kokoro import KPipeline
pipeline = KPipeline(lang_code='a')
generator = pipeline(text, voice='af_heart')
for gs, ps, audio in generator:
# audio is numpy array at 24kHz
# Save to temp file
```
4. **Format Conversion:**
- Native output: 24kHz WAV
- Use ffmpeg (already available) for OGG/MP3 conversion
- Reuse `_run_ffmpeg_blocking()` pattern from Piper engine
5. **Error Handling:**
- Apply all bug fixes from Piper engine (timeouts, temp file cleanup, logging)
- Handle GPU out-of-memory gracefully (fallback to CPU)
- Validate voice exists before synthesis
---
### Phase 3: Docker Integration
#### 3.1 Update Dockerfile
**Add dependencies to `Dockerfile`:**
```dockerfile
# Install espeak-ng for Kokoro phonemization
RUN apt-get update && \
apt-get install -y --no-install-recommends espeak-ng && \
apt-get clean && \
rm -rf /var/lib/apt/lists/*
# Install Kokoro Python dependencies
RUN pip install --no-cache-dir \
kokoro>=0.9.2 \
soundfile \
phonemizer \
torch \
transformers \
scipy \
munch
```
#### 3.2 Model Download Strategy
**Options:**
**Option A: Download on first use (lazy loading)**
- Models auto-download from Hugging Face (~200MB per model)
- Advantage: No pre-download needed
- Disadvantage: First synthesis will be slow
**Option B: Pre-download in Docker build**
```dockerfile
# Pre-download Kokoro model during build
RUN python3 -c "from kokoro import KPipeline; KPipeline(lang_code='a')"
```
**Option C: Volume mount like Piper**
```yaml
# docker-compose.yml
volumes:
- ./models/kokoro:/home/appuser/.cache/huggingface
```
**Recommendation**: Option A for MVP, Option C for production
---
### Phase 4: Configuration
#### 4.1 Update `app/config.py`
```python
class Settings(BaseSettings):
# ... existing settings ...
# Kokoro Engine Configuration
KOKORO_DEVICE: str = "cuda" # or "cpu"
KOKORO_TIMEOUT_SECONDS: int = 30
KOKORO_DEFAULT_LANG: str = "a" # American English
```
#### 4.2 Update `.env.example`
```bash
# Kokoro Engine Configuration
KOKORO_DEVICE=cuda # cuda or cpu
KOKORO_TIMEOUT_SECONDS=30
KOKORO_DEFAULT_LANG=a
```
#### 4.3 Update `app/main.py`
```python
ALL_ENGINES = {
"piper": PiperEngine,
"styletts": StyleTTSEngine,
"chattts": ChatTTSEngine,
"f5-tts": F5TTSEngine,
"kokoro": KokoroEngine, # Add this
}
```
---
### Phase 5: Voice Metadata
#### 5.1 Create `app/engines/kokoro_voices.py`
Store voice metadata for better UX:
```python
KOKORO_VOICES = {
"af_alloy": {
"gender": "F",
"language": ["en-us"],
"quality": "high",
"description": "Clear, professional female voice"
},
"af_bella": {
"gender": "F",
"language": ["en-us", "en-gb"],
"quality": "high",
"description": "Warm, expressive female voice"
},
# ... all 54 voices
}
```
Download full voice list from: https://huggingface.co/hexgrad/Kokoro-82M/blob/main/VOICES.md
---
### Phase 6: Testing
#### 6.1 Unit Tests
Create `tests/test_kokoro_engine.py`:
```python
import pytest
from app.engines.kokoro import KokoroEngine
@pytest.mark.asyncio
async def test_kokoro_synthesis():
engine = KokoroEngine()
audio_path = await engine.synthesize(
"Hello, this is Kokoro TTS.",
speaker="af_bella",
model="kokoro-en-us",
fmt="wav"
)
assert os.path.exists(audio_path)
assert os.path.getsize(audio_path) > 0
def test_kokoro_list_voices():
engine = KokoroEngine()
voices = engine.list_voices("kokoro-en-us")
assert len(voices) > 0
assert "af_bella" in voices
def test_kokoro_healthcheck():
engine = KokoroEngine()
health = engine.healthcheck()
assert health["status"] in ["ok", "not_available"]
```
#### 6.2 Integration Tests
```bash
# Test via API
curl -X POST http://localhost:8000/tts \
-H "Content-Type: application/json" \
-d '{
"text": "Hello from Kokoro TTS!",
"engine": "kokoro",
"model": "kokoro-en-us",
"speaker": "af_bella",
"format": "ogg"
}'
```
#### 6.3 Performance Benchmarks
Test synthesis speed:
- Short text (10 words): Target <0.5s on RTX 3090
- Medium text (100 words): Target <2s on RTX 3090
- Long text (1000 words): Target <15s on RTX 3090
---
### Phase 7: Documentation
#### 7.1 Update README.md
Add Kokoro to supported engines:
```markdown
## Supported TTS Engines
- **Piper** - Fast, lightweight, ONNX-based TTS
- **Kokoro** - 82M parameter high-quality TTS (NEW!)
- **StyleTTS** - Expressive style-based TTS (planned)
- **ChatTTS** - Conversational TTS (planned)
- **F5-TTS** - Advanced flow-based TTS (planned)
```
#### 7.2 Update CLAUDE.md
Add Kokoro engine details:
```markdown
### Kokoro Engine (`app/engines/kokoro.py`)
- Uses `kokoro` Python library (KPipeline)
- 54 voices across 8 languages
- Outputs 24kHz audio natively
- Extremely fast (~90x real-time on consumer GPU)
- Model size: 82M parameters (~200MB download)
```
#### 7.3 Create Kokoro Usage Guide
Create `docs/kokoro-guide.md`:
- Voice selection guide
- Language support matrix
- Performance optimization tips
- GPU vs CPU mode comparison
- Troubleshooting common issues
---
## Implementation Checklist
### Phase 1: Research & Setup
- [x] Research Kokoro capabilities
- [x] Document API and dependencies
- [ ] Test Kokoro locally outside Docker
### Phase 2: Engine Implementation
- [ ] Create `app/engines/kokoro.py`
- [ ] Implement `synthesize()` method
- [ ] Implement `list_models()` method
- [ ] Implement `list_voices()` method
- [ ] Implement `healthcheck()` method
- [ ] Implement `selftest()` method
- [ ] Create `app/engines/kokoro_voices.py` metadata file
### Phase 3: Docker Integration
- [ ] Update Dockerfile with dependencies
- [ ] Add espeak-ng system package
- [ ] Add Kokoro Python packages
- [ ] Test Docker build
- [ ] Verify GPU access in container
### Phase 4: Configuration
- [ ] Update `app/config.py` with Kokoro settings
- [ ] Update `.env.example` with Kokoro variables
- [ ] Add KokoroEngine to `app/main.py` ALL_ENGINES
- [ ] Update active engines in `.env`
### Phase 5: Testing
- [ ] Create unit tests
- [ ] Test synthesis with various voices
- [ ] Test format conversion (wav/ogg/mp3)
- [ ] Test caching behavior
- [ ] Performance benchmarks
- [ ] Memory usage profiling
### Phase 6: Documentation
- [ ] Update README.md
- [ ] Update CLAUDE.md
- [ ] Create Kokoro usage guide
- [ ] Document voice selection
- [ ] Add troubleshooting section
### Phase 7: Deployment
- [ ] Test in development environment
- [ ] Rebuild Docker image
- [ ] Update docker-compose.yml if needed
- [ ] Test API endpoints
- [ ] Commit changes to git
- [ ] Tag release
---
## Key Decisions
### 1. Model Organization
**Decision**: Map Kokoro lang_codes to model names
- `kokoro-en-us` → lang_code='a'
- `kokoro-en-gb` → lang_code='b'
- etc.
**Rationale**: Maintains consistency with existing API structure
### 2. Voice Naming
**Decision**: Use Kokoro's native voice names (e.g., `af_bella`)
**Rationale**:
- Avoids confusion with remapping
- Documented in official VOICES.md
- Users can reference official docs
### 3. GPU Support
**Decision**: Support both GPU and CPU with fallback
**Rationale**:
- GPU provides 90x real-time performance
- CPU fallback ensures it works on all systems
- Configurable via KOKORO_DEVICE env var
### 4. Model Download Strategy
**Decision**: Lazy loading on first use
**Rationale**:
- Smaller Docker image
- Only download models that are actually used
- Can switch to pre-download later if needed
---
## Potential Issues & Solutions
### Issue 1: Large Model Download
**Problem**: Model is ~200MB per language
**Solution**:
- Lazy loading (download on first use)
- Cache in Docker volume
- Document expected download time
### Issue 2: GPU Memory Usage
**Problem**: May require 2-4GB VRAM
**Solution**:
- Implement memory monitoring
- Graceful fallback to CPU
- Document GPU requirements
### Issue 3: First Synthesis Slow
**Problem**: Model loading + compilation takes time
**Solution**:
- Warm up during healthcheck
- Keep model loaded in memory
- Document expected first-run delay
### Issue 4: Voice Compatibility
**Problem**: Not all voices work with all languages
**Solution**:
- Validate voice-language compatibility
- Return clear error messages
- Document voice language support
---
## Success Criteria
- [ ] Kokoro engine passes all unit tests
- [ ] API endpoints work correctly (`/tts`, `/models`, `/speakers`)
- [ ] Audio quality matches expected output
- [ ] Synthesis speed: >50x real-time on RTX 3090
- [ ] Memory usage: <4GB VRAM for single synthesis
- [ ] Cache versioning works correctly
- [ ] No temp file leaks
- [ ] No hung processes
- [ ] Documentation complete and accurate
- [ ] Voice wizard integration works
---
## Timeline Estimate
**Total**: 8-12 hours
- **Phase 1 (Research)**: ✅ Complete (2 hours)
- **Phase 2 (Implementation)**: 3-4 hours
- **Phase 3 (Docker)**: 1-2 hours
- **Phase 4 (Config)**: 0.5 hours
- **Phase 5 (Testing)**: 2-3 hours
- **Phase 6 (Documentation)**: 1-2 hours
- **Phase 7 (Deployment)**: 1 hour
---
## References
- [Kokoro-82M Hugging Face](https://huggingface.co/hexgrad/Kokoro-82M)
- [Kokoro GitHub Repository](https://github.com/hexgrad/kokoro)
- [VOICES.md - Complete Voice List](https://huggingface.co/hexgrad/Kokoro-82M/blob/main/VOICES.md)
- [Kokoro TTS Official Website](https://kokorotts.net/)
- [Kokoro-82M Analytics Vidhya Article](https://www.analyticsvidhya.com/blog/2025/01/kokoro-82m/)
- [Kokoro TTS Live Demo](https://huggingface.co/spaces/hexgrad/Kokoro-TTS)