- Add XTTS v2 configuration to .env.example - Refactor Dockerfile to multi-stage build with CUDA 12.1 support - Update Makefile with Kokoro and XTTS test environment targets - Refactor Piper engine (app/engines/piper.py) to use python module execution - Add comprehensive documentation for Kokoro and XTTS plans - Add helper scripts and patches for build process
16 KiB
Kokoro TTS Implementation Status
Date Started: 2025-12-05 Last Updated: 2025-12-05 00:48 CET Status: 🟡 In Progress (Phase 2 & 3 partially complete)
Quick Summary
We are implementing Kokoro-82M TTS engine as a new engine in AudioEngineHub. Kokoro is a lightweight 82M parameter model that delivers quality comparable to models 5-15× its size, with ~90× real-time performance on consumer GPUs.
Progress: ~40% complete
- ✅ Phase 1: Research & Planning (100%)
- ✅ Phase 2: Engine Implementation (100%)
- 🟡 Phase 3: Docker Integration (50%)
- ⏳ Phase 4: Configuration (0%)
- ⏳ Phase 5: Testing (0%)
- ⏳ Phase 6: Documentation (0%)
What We've Completed
✅ Phase 1: Research & Planning
- Researched Kokoro TTS capabilities and architecture
- Identified Python library (
kokoro>=0.9.2) - Documented all 54 voices across 8 languages
- Created comprehensive implementation plan in
KOKORO_IMPLEMENTATION_PLAN.md
Key Findings:
- Kokoro uses 82M parameters (extremely lightweight)
- Performance: ~90× real-time on RTX 3090 Ti, ~210× on RTX 4090
- Output: 24kHz high-fidelity audio
- License: Apache 2.0 (open-source, commercial use allowed)
- 54 voices across 8 languages (EN-US, EN-GB, JA, ZH, ES, FR, HI, IT, PT)
✅ Phase 2: Engine Implementation (Complete)
-
Created
app/engines/kokoro.py- Full engine implementation with:- Complete
KokoroEngineclass implementingTTSEngineBase synthesize()method with all bug fixes from Piper (timeouts, cleanup, logging)- GPU/CPU support with automatic fallback
- Language code mapping (10 models: kokoro-en-us, kokoro-en-gb, etc.)
- Format conversion (WAV/OGG/MP3) via ffmpeg
- Comprehensive error handling and logging
- Model pipeline caching for performance
- Complete
-
Created
app/engines/kokoro_voices.py- Voice metadata with:- Complete list of all 54 Kokoro voices
- Voice metadata (gender, language, description)
- Language-based voice filtering
- Helper functions:
get_voices_for_model(),get_voice_info()
File Locations:
/home/stephan/Projekte/KI/AudioEngineHub/app/engines/kokoro.py(316 lines)/home/stephan/Projekte/KI/AudioEngineHub/app/engines/kokoro_voices.py(222 lines)
🟡 Phase 3: Docker Integration (50% Complete)
-
Updated
Dockerfile- Added espeak-ng system dependency- Line 17: Added
espeak-ngto apt-get install
- Line 17: Added
-
Update
requirements.txt- Need to add Kokoro Python dependencies- NEXT STEP: Add these lines to requirements.txt:
# Kokoro TTS Engine kokoro>=0.9.2 soundfile phonemizer scipy munch
- NEXT STEP: Add these lines to requirements.txt:
What's Left To Do
⏳ Phase 4: Configuration (Not Started)
1. Update app/config.py
Add Kokoro configuration settings:
class Settings(BaseSettings):
# ... existing settings ...
# Kokoro Engine Configuration
KOKORO_DEVICE: str = "cuda" # or "cpu"
KOKORO_TIMEOUT_SECONDS: int = 30
File: /home/stephan/Projekte/KI/AudioEngineHub/app/config.py
Location: Add after line 25 (after FFMPEG_TIMEOUT_SECONDS)
2. Update app/main.py
Register Kokoro engine in the engine registry:
from app.engines.kokoro import KokoroEngine # Add this import
ALL_ENGINES = {
"piper": PiperEngine,
"styletts": StyleTTSEngine,
"chattts": ChatTTSEngine,
"f5-tts": F5TTSEngine,
"kokoro": KokoroEngine, # Add this line
}
File: /home/stephan/Projekte/KI/AudioEngineHub/app/main.py
Location: Line 28 (import), Line 38 (registry)
3. Update .env.example
Add Kokoro configuration documentation:
# Kokoro Engine Configuration
KOKORO_DEVICE=cuda # cuda or cpu
KOKORO_TIMEOUT_SECONDS=30
File: /home/stephan/Projekte/KI/AudioEngineHub/.env.example
Location: Add after FFMPEG_TIMEOUT_SECONDS
4. Optional: Update .env
To enable Kokoro by default:
ACTIVE_ENGINES='["piper", "kokoro"]'
⏳ Phase 5: Testing (Not Started)
1. Test Import and Basic Functionality
# Test engine loads without errors
python3 -c "from app.engines.kokoro import KokoroEngine; print(KokoroEngine().healthcheck())"
Expected output:
{
"status": "not_available", # OK if kokoro not installed yet
"engine": "kokoro",
"library_available": false,
"error": "Kokoro library not installed..."
}
2. Install Kokoro Dependencies (Outside Docker First)
# Install espeak-ng
sudo apt-get install espeak-ng
# Install Python packages
pip install kokoro>=0.9.2 soundfile phonemizer scipy munch
3. Test Synthesis Locally
# Run engine selftest
cd /home/stephan/Projekte/KI/AudioEngineHub
python3 app/engines/kokoro.py
Expected output:
Healthcheck: {'status': 'ok', 'engine': 'kokoro', ...}
Models: ['kokoro-en-us', 'kokoro-en-gb', ...]
Voices: ['af_alloy', 'af_aoede', 'af_bella', ...] ...
Selftest: {'selftest': True, 'models': [...], 'voices_count': 54}
4. Test via API
# Start development server
make dev-up
# Test API endpoint
curl -X POST http://localhost:8000/tts \
-H "Content-Type: application/json" \
-d '{"text":"Hello from Kokoro!","engine":"kokoro","model":"kokoro-en-us","speaker":"af_bella","format":"ogg"}'
5. Test with Voice Wizard
# Launch voice wizard and test preview with Kokoro voices
cd /home/stephan/Games/University/tts
./scripts/voice-assignment-wizard.sh
⏳ Phase 6: Documentation (Not Started)
1. Update README.md
Add Kokoro to supported engines list:
## Supported TTS Engines
- **Piper** - Fast, lightweight, ONNX-based TTS with 100+ voices
- **Kokoro** - 82M parameter high-quality TTS with 54 voices across 8 languages 🆕
- **StyleTTS** - Expressive style-based TTS (coming soon)
- **ChatTTS** - Conversational TTS (coming soon)
- **F5-TTS** - Advanced flow-based TTS (coming soon)
2. Update CLAUDE.md
Add Kokoro engine section:
### Kokoro Engine (`app/engines/kokoro.py`)
- **Library**: `kokoro` Python package (KPipeline)
- **Models**: 10 language models (en-us, en-gb, ja, zh, es, fr, hi, it, pt, ko)
- **Voices**: 54 voices with gender and language metadata
- **Output**: 24kHz audio natively (WAV), converts to OGG/MP3 via ffmpeg
- **Performance**: ~90× real-time on RTX 3090 Ti
- **Model Size**: 82M parameters (~200MB download per language)
- **Voice Metadata**: `app/engines/kokoro_voices.py` contains all 54 voices with metadata
3. Create Usage Guide (Optional)
Create docs/kokoro-usage.md with:
- Voice selection guide
- Language support matrix
- Performance tips (GPU vs CPU)
- Troubleshooting common issues
Implementation Checklist
Phase 1: Research ✅
- Research Kokoro capabilities
- Document API and dependencies
- Create implementation plan
Phase 2: Engine Implementation ✅
- Create
app/engines/kokoro.py - Implement
synthesize()method - Implement
list_models()method - Implement
list_voices()method - Implement
healthcheck()method - Implement
selftest()method - Create
app/engines/kokoro_voices.py
Phase 3: Docker Integration 🟡
- Update Dockerfile (added espeak-ng)
- → NEXT: Update requirements.txt (add Kokoro packages)
- Test Docker build
- Verify GPU access in container (if available)
Phase 4: Configuration ⏳
- Update
app/config.pywith Kokoro settings - Update
.env.examplewith Kokoro variables - Add KokoroEngine import to
app/main.py - Add Kokoro to ALL_ENGINES in
app/main.py - Optional: Update
.envto enable Kokoro
Phase 5: Testing ⏳
- Test engine imports without errors
- Install Kokoro dependencies locally
- Run selftest (
python3 app/engines/kokoro.py) - Test synthesis via API
- Test with voice wizard
- Verify caching works correctly
- Check for temp file leaks
- Performance benchmarks
Phase 6: Documentation ⏳
- Update README.md
- Update CLAUDE.md
- Optional: Create Kokoro usage guide
- Document voice selection
- Add troubleshooting section
Phase 7: Deployment ⏳
- Rebuild Docker image
- Test in development environment
- Commit changes to git
- Push to remote
- Tag release (optional)
Files Created/Modified
✅ Created Files
KOKORO_IMPLEMENTATION_PLAN.md- Complete implementation plan (280 lines)app/engines/kokoro.py- Kokoro engine implementation (316 lines)app/engines/kokoro_voices.py- Voice metadata (222 lines)KOKORO_IMPLEMENTATION_STATUS.md- This file
✅ Modified Files
Dockerfile- Added espeak-ng dependency (line 17)
⏳ Files To Modify
requirements.txt- Add Kokoro Python packagesapp/config.py- Add Kokoro configurationapp/main.py- Register Kokoro engine.env.example- Document Kokoro configREADME.md- Add Kokoro to engines listCLAUDE.md- Add Kokoro engine details
Next Session: Action Plan
Step 1: Complete Docker Integration (5 minutes)
cd /home/stephan/Projekte/KI/AudioEngineHub
# Add to requirements.txt (after line 11):
cat >> requirements.txt << 'EOF'
# Kokoro TTS Engine
kokoro>=0.9.2
soundfile
phonemizer
scipy
munch
EOF
Step 2: Complete Configuration (10 minutes)
- Edit
app/config.py- add KOKORO_DEVICE and KOKORO_TIMEOUT_SECONDS - Edit
app/main.py- import KokoroEngine and add to ALL_ENGINES - Edit
.env.example- document new config options
Step 3: Test Locally (15 minutes)
- Install dependencies:
sudo apt-get install espeak-ng && pip install kokoro>=0.9.2 soundfile phonemizer scipy munch - Run selftest:
python3 app/engines/kokoro.py - Start server:
make dev-up(will rebuild Docker image) - Test API synthesis with Kokoro
- Test voice wizard preview
Step 4: Documentation (10 minutes)
- Update README.md with Kokoro
- Update CLAUDE.md with Kokoro engine details
- Commit all changes to git
Total Time Remaining: ~40 minutes
Key Technical Details
Language Code Mapping
KOKORO_LANG_CODES = {
"kokoro-en-us": "a", # American English
"kokoro-en-gb": "b", # British English
"kokoro-fr": "fr", # French
"kokoro-es": "es", # Spanish
"kokoro-ja": "ja", # Japanese
"kokoro-zh": "zh", # Chinese
"kokoro-it": "it", # Italian
"kokoro-pt": "pt", # Portuguese
"kokoro-hi": "hi", # Hindi
"kokoro-ko": "ko", # Korean
}
Voice Organization
- 54 total voices across 8 languages
- Naming convention:
{language}{gender}_{name}af_= American Femaleam_= American Malebf_= British Femalebm_= British Malejf_/jm_= Japanese F/Mzf_/zm_= Chinese F/M- etc.
- Most popular: af_bella, af_sarah, af_sky, am_adam, am_michael
API Usage Example
{
"text": "Hello from Kokoro TTS!",
"engine": "kokoro",
"model": "kokoro-en-us",
"speaker": "af_bella",
"format": "ogg"
}
Known Issues & Considerations
1. Model Download on First Use
- Issue: First synthesis will be slow (~30-60s) due to model download
- Size: ~200MB per language model
- Location: Models cached in
~/.cache/huggingface/ - Solution: Expected behavior, document in README
2. GPU vs CPU Performance
- GPU: ~90× real-time (RTX 3090 Ti)
- CPU: ~5-10× real-time (estimated)
- Fallback: Code supports both, configurable via KOKORO_DEVICE env var
3. Voice-Language Compatibility
- Issue: Not all voices work with all languages
- Solution:
kokoro_voices.pyfilters voices by language - API:
/speakers?engine=kokoro&model=kokoro-en-usreturns only compatible voices
4. Dependencies
- espeak-ng: Required system package for phonemization
- soundfile: Required for WAV file I/O
- phonemizer: Required for text-to-phoneme conversion
- scipy, munch: Required by Kokoro library
Testing Commands Reference
Local Testing (Outside Docker)
# Install system dependency
sudo apt-get install espeak-ng
# Install Python packages
pip install kokoro>=0.9.2 soundfile phonemizer scipy munch
# Test engine
cd /home/stephan/Projekte/KI/AudioEngineHub
python3 app/engines/kokoro.py
# Test voice metadata
python3 -c "from app.engines.kokoro_voices import ALL_VOICES, get_voices_for_model; print(f'Total voices: {len(ALL_VOICES)}'); print(f'EN-US voices: {get_voices_for_model(\"kokoro-en-us\")}')"
Docker Testing
# Rebuild and start
make down
make dev-up
# Check logs
make logs
# Health check
make health-check
# Test synthesis
curl -X POST http://localhost:8000/tts \
-H "Content-Type: application/json" \
-d '{"text":"Testing Kokoro TTS engine","engine":"kokoro","model":"kokoro-en-us","speaker":"af_bella","format":"ogg"}' | jq
Voice Wizard Testing
cd /home/stephan/Games/University/tts
./scripts/voice-assignment-wizard.sh
# Should show Kokoro engine in dropdown
# Should list 54 voices (or subset by language)
# Preview should work with all Kokoro voices
Success Criteria
Before marking this implementation as complete, verify:
- Kokoro engine passes healthcheck
- All 10 language models are listed in
/modelsendpoint - 54 voices are accessible via
/speakersendpoint - Voice filtering by model/language works correctly
- Synthesis produces valid audio files
- Synthesis speed: >50× real-time on GPU (if available)
- Format conversion works (WAV/OGG/MP3)
- Cache versioning works correctly
- No temp file leaks after synthesis
- No hung processes under load
- Voice wizard integration works
- Documentation is complete and accurate
Resources & References
Official Documentation
- Kokoro-82M Hugging Face
- VOICES.md - Complete Voice List
- Kokoro GitHub Repository
- Kokoro Official Website
Technical Articles
Installation Guides
Timeline & Estimates
| Phase | Status | Time Spent | Remaining | Total |
|---|---|---|---|---|
| Phase 1: Research | ✅ Complete | 2h | 0h | 2h |
| Phase 2: Implementation | ✅ Complete | 2h | 0h | 2h |
| Phase 3: Docker | 🟡 50% | 0.5h | 0.5h | 1h |
| Phase 4: Configuration | ⏳ Pending | 0h | 0.5h | 0.5h |
| Phase 5: Testing | ⏳ Pending | 0h | 1h | 1h |
| Phase 6: Documentation | ⏳ Pending | 0h | 0.5h | 0.5h |
| Total | 40% Complete | 4.5h | 2.5h | 7h |
Original Estimate: 8-12 hours Current Progress: ~5 hours spent, ~2-3 hours remaining On Track: Yes, ahead of schedule
Questions for User (Next Session)
-
GPU Availability: Do you have a CUDA-capable GPU? This affects:
- Performance expectations (90× vs 5× real-time)
- Default KOKORO_DEVICE setting (cuda vs cpu)
-
Language Priority: Which languages do you need most?
- Only EN-US model will be tested initially
- Other languages can be tested on demand
-
Docker vs Local: Prefer testing locally first or directly in Docker?
- Local testing is faster for iteration
- Docker testing validates full deployment
-
Voice Wizard: Should we integrate Kokoro metadata into the voice wizard UI?
- Could show language, gender, description for each voice
- Requires changes to audioengine_client.py
End of Status Report
Last updated: 2025-12-05 00:48 CET Next session: Continue with Step 1 (Complete Docker Integration)