# Kokoro TTS Engine Implementation Plan ## Overview Implement Kokoro-82M as a new TTS engine in AudioEngineHub. Kokoro is a lightweight, high-performance open-weight TTS model with 82 million parameters that delivers quality comparable to models 5-15× its size. ## Background Research ### Key Features - **Size**: 82 million parameters (extremely lightweight) - **Performance**: ~210× real-time on RTX 4090, ~90× real-time on RTX 3090 Ti - **Quality**: Took first place in TTS Spaces Arena, outperforming XTTS v2 (467M) and MetaVoice (1.2B) - **Audio**: 24kHz high-fidelity output - **License**: Apache 2.0 (open-source, commercial use allowed) - **Languages**: 8 languages (English US/UK, French, Spanish, Japanese, Chinese, Italian, Portuguese, Hindi, Korean) - **Voices**: 54 voices available ### Sources - [Kokoro-82M Hugging Face](https://huggingface.co/hexgrad/Kokoro-82M) - [Kokoro TTS Official Site](https://kokorotts.net/) - [VOICES.md - Complete Voice List](https://huggingface.co/hexgrad/Kokoro-82M/blob/main/VOICES.md) - [GitHub Repository](https://github.com/hexgrad/kokoro) - [Analytics Vidhya Article](https://www.analyticsvidhya.com/blog/2025/01/kokoro-82m/) --- ## Implementation Strategy ### Phase 1: Research & Setup #### 1.1 Model Investigation - [x] Research Kokoro TTS architecture and capabilities - [x] Identify Python library: `kokoro>=0.9.2` - [x] Document voice list (54 voices across 8 languages) - [ ] Test Kokoro locally to understand API #### 1.2 Dependency Analysis **Required packages:** ```bash kokoro>=0.9.2 soundfile phonemizer torch transformers scipy munch ``` **System dependencies:** ```bash espeak-ng # Required for phonemization ``` --- ### Phase 2: Engine Implementation #### 2.1 Create `app/engines/kokoro.py` **Architecture:** ```python class KokoroEngine(TTSEngineBase): def __init__(self): # Initialize Kokoro pipeline # Handle GPU/CPU detection # Load model from Hugging Face async def synthesize(text, speaker, model, fmt): # Generate audio using KPipeline # Handle voice selection # Convert to requested format (wav/ogg/mp3) # Return temp file path def list_models(self): # Return available language models # Options: 'a' (American English), 'b' (British English), etc. def list_voices(self, model): # Return 54 available voices # Filter by language if model specified def healthcheck(self): # Check if kokoro library is available # Verify model is loaded # Return status async def selftest(self): # Run quick synthesis test # Verify audio generation works ``` **Key Implementation Details:** 1. **Model Selection:** - Kokoro uses `lang_code` parameter (e.g., 'a' = American English, 'b' = British English) - Map this to "models" concept in our API - Models: `kokoro-en-us`, `kokoro-en-gb`, `kokoro-fr`, `kokoro-es`, `kokoro-ja`, `kokoro-zh`, `kokoro-it`, `kokoro-pt`, `kokoro-hi`, `kokoro-ko` 2. **Voice Selection:** - 54 voices available (see VOICES.md) - Popular voices: `af_alloy`, `af_bella`, `af_sarah`, `af_sky`, `af_nova`, etc. - Each voice has quality grade and language support 3. **Audio Generation:** ```python from kokoro import KPipeline pipeline = KPipeline(lang_code='a') generator = pipeline(text, voice='af_heart') for gs, ps, audio in generator: # audio is numpy array at 24kHz # Save to temp file ``` 4. **Format Conversion:** - Native output: 24kHz WAV - Use ffmpeg (already available) for OGG/MP3 conversion - Reuse `_run_ffmpeg_blocking()` pattern from Piper engine 5. **Error Handling:** - Apply all bug fixes from Piper engine (timeouts, temp file cleanup, logging) - Handle GPU out-of-memory gracefully (fallback to CPU) - Validate voice exists before synthesis --- ### Phase 3: Docker Integration #### 3.1 Update Dockerfile **Add dependencies to `Dockerfile`:** ```dockerfile # Install espeak-ng for Kokoro phonemization RUN apt-get update && \ apt-get install -y --no-install-recommends espeak-ng && \ apt-get clean && \ rm -rf /var/lib/apt/lists/* # Install Kokoro Python dependencies RUN pip install --no-cache-dir \ kokoro>=0.9.2 \ soundfile \ phonemizer \ torch \ transformers \ scipy \ munch ``` #### 3.2 Model Download Strategy **Options:** **Option A: Download on first use (lazy loading)** - Models auto-download from Hugging Face (~200MB per model) - Advantage: No pre-download needed - Disadvantage: First synthesis will be slow **Option B: Pre-download in Docker build** ```dockerfile # Pre-download Kokoro model during build RUN python3 -c "from kokoro import KPipeline; KPipeline(lang_code='a')" ``` **Option C: Volume mount like Piper** ```yaml # docker-compose.yml volumes: - ./models/kokoro:/home/appuser/.cache/huggingface ``` **Recommendation**: Option A for MVP, Option C for production --- ### Phase 4: Configuration #### 4.1 Update `app/config.py` ```python class Settings(BaseSettings): # ... existing settings ... # Kokoro Engine Configuration KOKORO_DEVICE: str = "cuda" # or "cpu" KOKORO_TIMEOUT_SECONDS: int = 30 KOKORO_DEFAULT_LANG: str = "a" # American English ``` #### 4.2 Update `.env.example` ```bash # Kokoro Engine Configuration KOKORO_DEVICE=cuda # cuda or cpu KOKORO_TIMEOUT_SECONDS=30 KOKORO_DEFAULT_LANG=a ``` #### 4.3 Update `app/main.py` ```python ALL_ENGINES = { "piper": PiperEngine, "styletts": StyleTTSEngine, "chattts": ChatTTSEngine, "f5-tts": F5TTSEngine, "kokoro": KokoroEngine, # Add this } ``` --- ### Phase 5: Voice Metadata #### 5.1 Create `app/engines/kokoro_voices.py` Store voice metadata for better UX: ```python KOKORO_VOICES = { "af_alloy": { "gender": "F", "language": ["en-us"], "quality": "high", "description": "Clear, professional female voice" }, "af_bella": { "gender": "F", "language": ["en-us", "en-gb"], "quality": "high", "description": "Warm, expressive female voice" }, # ... all 54 voices } ``` Download full voice list from: https://huggingface.co/hexgrad/Kokoro-82M/blob/main/VOICES.md --- ### Phase 6: Testing #### 6.1 Unit Tests Create `tests/test_kokoro_engine.py`: ```python import pytest from app.engines.kokoro import KokoroEngine @pytest.mark.asyncio async def test_kokoro_synthesis(): engine = KokoroEngine() audio_path = await engine.synthesize( "Hello, this is Kokoro TTS.", speaker="af_bella", model="kokoro-en-us", fmt="wav" ) assert os.path.exists(audio_path) assert os.path.getsize(audio_path) > 0 def test_kokoro_list_voices(): engine = KokoroEngine() voices = engine.list_voices("kokoro-en-us") assert len(voices) > 0 assert "af_bella" in voices def test_kokoro_healthcheck(): engine = KokoroEngine() health = engine.healthcheck() assert health["status"] in ["ok", "not_available"] ``` #### 6.2 Integration Tests ```bash # Test via API curl -X POST http://localhost:8000/tts \ -H "Content-Type: application/json" \ -d '{ "text": "Hello from Kokoro TTS!", "engine": "kokoro", "model": "kokoro-en-us", "speaker": "af_bella", "format": "ogg" }' ``` #### 6.3 Performance Benchmarks Test synthesis speed: - Short text (10 words): Target <0.5s on RTX 3090 - Medium text (100 words): Target <2s on RTX 3090 - Long text (1000 words): Target <15s on RTX 3090 --- ### Phase 7: Documentation #### 7.1 Update README.md Add Kokoro to supported engines: ```markdown ## Supported TTS Engines - **Piper** - Fast, lightweight, ONNX-based TTS - **Kokoro** - 82M parameter high-quality TTS (NEW!) - **StyleTTS** - Expressive style-based TTS (planned) - **ChatTTS** - Conversational TTS (planned) - **F5-TTS** - Advanced flow-based TTS (planned) ``` #### 7.2 Update CLAUDE.md Add Kokoro engine details: ```markdown ### Kokoro Engine (`app/engines/kokoro.py`) - Uses `kokoro` Python library (KPipeline) - 54 voices across 8 languages - Outputs 24kHz audio natively - Extremely fast (~90x real-time on consumer GPU) - Model size: 82M parameters (~200MB download) ``` #### 7.3 Create Kokoro Usage Guide Create `docs/kokoro-guide.md`: - Voice selection guide - Language support matrix - Performance optimization tips - GPU vs CPU mode comparison - Troubleshooting common issues --- ## Implementation Checklist ### Phase 1: Research & Setup - [x] Research Kokoro capabilities - [x] Document API and dependencies - [ ] Test Kokoro locally outside Docker ### Phase 2: Engine Implementation - [ ] Create `app/engines/kokoro.py` - [ ] Implement `synthesize()` method - [ ] Implement `list_models()` method - [ ] Implement `list_voices()` method - [ ] Implement `healthcheck()` method - [ ] Implement `selftest()` method - [ ] Create `app/engines/kokoro_voices.py` metadata file ### Phase 3: Docker Integration - [ ] Update Dockerfile with dependencies - [ ] Add espeak-ng system package - [ ] Add Kokoro Python packages - [ ] Test Docker build - [ ] Verify GPU access in container ### Phase 4: Configuration - [ ] Update `app/config.py` with Kokoro settings - [ ] Update `.env.example` with Kokoro variables - [ ] Add KokoroEngine to `app/main.py` ALL_ENGINES - [ ] Update active engines in `.env` ### Phase 5: Testing - [ ] Create unit tests - [ ] Test synthesis with various voices - [ ] Test format conversion (wav/ogg/mp3) - [ ] Test caching behavior - [ ] Performance benchmarks - [ ] Memory usage profiling ### Phase 6: Documentation - [ ] Update README.md - [ ] Update CLAUDE.md - [ ] Create Kokoro usage guide - [ ] Document voice selection - [ ] Add troubleshooting section ### Phase 7: Deployment - [ ] Test in development environment - [ ] Rebuild Docker image - [ ] Update docker-compose.yml if needed - [ ] Test API endpoints - [ ] Commit changes to git - [ ] Tag release --- ## Key Decisions ### 1. Model Organization **Decision**: Map Kokoro lang_codes to model names - `kokoro-en-us` → lang_code='a' - `kokoro-en-gb` → lang_code='b' - etc. **Rationale**: Maintains consistency with existing API structure ### 2. Voice Naming **Decision**: Use Kokoro's native voice names (e.g., `af_bella`) **Rationale**: - Avoids confusion with remapping - Documented in official VOICES.md - Users can reference official docs ### 3. GPU Support **Decision**: Support both GPU and CPU with fallback **Rationale**: - GPU provides 90x real-time performance - CPU fallback ensures it works on all systems - Configurable via KOKORO_DEVICE env var ### 4. Model Download Strategy **Decision**: Lazy loading on first use **Rationale**: - Smaller Docker image - Only download models that are actually used - Can switch to pre-download later if needed --- ## Potential Issues & Solutions ### Issue 1: Large Model Download **Problem**: Model is ~200MB per language **Solution**: - Lazy loading (download on first use) - Cache in Docker volume - Document expected download time ### Issue 2: GPU Memory Usage **Problem**: May require 2-4GB VRAM **Solution**: - Implement memory monitoring - Graceful fallback to CPU - Document GPU requirements ### Issue 3: First Synthesis Slow **Problem**: Model loading + compilation takes time **Solution**: - Warm up during healthcheck - Keep model loaded in memory - Document expected first-run delay ### Issue 4: Voice Compatibility **Problem**: Not all voices work with all languages **Solution**: - Validate voice-language compatibility - Return clear error messages - Document voice language support --- ## Success Criteria - [ ] Kokoro engine passes all unit tests - [ ] API endpoints work correctly (`/tts`, `/models`, `/speakers`) - [ ] Audio quality matches expected output - [ ] Synthesis speed: >50x real-time on RTX 3090 - [ ] Memory usage: <4GB VRAM for single synthesis - [ ] Cache versioning works correctly - [ ] No temp file leaks - [ ] No hung processes - [ ] Documentation complete and accurate - [ ] Voice wizard integration works --- ## Timeline Estimate **Total**: 8-12 hours - **Phase 1 (Research)**: ✅ Complete (2 hours) - **Phase 2 (Implementation)**: 3-4 hours - **Phase 3 (Docker)**: 1-2 hours - **Phase 4 (Config)**: 0.5 hours - **Phase 5 (Testing)**: 2-3 hours - **Phase 6 (Documentation)**: 1-2 hours - **Phase 7 (Deployment)**: 1 hour --- ## References - [Kokoro-82M Hugging Face](https://huggingface.co/hexgrad/Kokoro-82M) - [Kokoro GitHub Repository](https://github.com/hexgrad/kokoro) - [VOICES.md - Complete Voice List](https://huggingface.co/hexgrad/Kokoro-82M/blob/main/VOICES.md) - [Kokoro TTS Official Website](https://kokorotts.net/) - [Kokoro-82M Analytics Vidhya Article](https://www.analyticsvidhya.com/blog/2025/01/kokoro-82m/) - [Kokoro TTS Live Demo](https://huggingface.co/spaces/hexgrad/Kokoro-TTS)