Files
audio-engine-hub/KOKORO_IMPLEMENTATION_STATUS.md
stephan 91887ae296 feat: Add XTTS v2 support, refactor Docker/GPU infra, and improve Piper engine
- Add XTTS v2 configuration to .env.example
- Refactor Dockerfile to multi-stage build with CUDA 12.1 support
- Update Makefile with Kokoro and XTTS test environment targets
- Refactor Piper engine (app/engines/piper.py) to use python module execution
- Add comprehensive documentation for Kokoro and XTTS plans
- Add helper scripts and patches for build process
2025-12-13 11:37:58 +01:00

16 KiB
Raw Blame History

Kokoro TTS Implementation Status

Date Started: 2025-12-05 Last Updated: 2025-12-05 00:48 CET Status: 🟡 In Progress (Phase 2 & 3 partially complete)


Quick Summary

We are implementing Kokoro-82M TTS engine as a new engine in AudioEngineHub. Kokoro is a lightweight 82M parameter model that delivers quality comparable to models 5-15× its size, with ~90× real-time performance on consumer GPUs.

Progress: ~40% complete

  • ✅ Phase 1: Research & Planning (100%)
  • ✅ Phase 2: Engine Implementation (100%)
  • 🟡 Phase 3: Docker Integration (50%)
  • ⏳ Phase 4: Configuration (0%)
  • ⏳ Phase 5: Testing (0%)
  • ⏳ Phase 6: Documentation (0%)

What We've Completed

✅ Phase 1: Research & Planning

  • Researched Kokoro TTS capabilities and architecture
  • Identified Python library (kokoro>=0.9.2)
  • Documented all 54 voices across 8 languages
  • Created comprehensive implementation plan in KOKORO_IMPLEMENTATION_PLAN.md

Key Findings:

  • Kokoro uses 82M parameters (extremely lightweight)
  • Performance: ~90× real-time on RTX 3090 Ti, ~210× on RTX 4090
  • Output: 24kHz high-fidelity audio
  • License: Apache 2.0 (open-source, commercial use allowed)
  • 54 voices across 8 languages (EN-US, EN-GB, JA, ZH, ES, FR, HI, IT, PT)

✅ Phase 2: Engine Implementation (Complete)

  • Created app/engines/kokoro.py - Full engine implementation with:

    • Complete KokoroEngine class implementing TTSEngineBase
    • synthesize() method with all bug fixes from Piper (timeouts, cleanup, logging)
    • GPU/CPU support with automatic fallback
    • Language code mapping (10 models: kokoro-en-us, kokoro-en-gb, etc.)
    • Format conversion (WAV/OGG/MP3) via ffmpeg
    • Comprehensive error handling and logging
    • Model pipeline caching for performance
  • Created app/engines/kokoro_voices.py - Voice metadata with:

    • Complete list of all 54 Kokoro voices
    • Voice metadata (gender, language, description)
    • Language-based voice filtering
    • Helper functions: get_voices_for_model(), get_voice_info()

File Locations:

  • /home/stephan/Projekte/KI/AudioEngineHub/app/engines/kokoro.py (316 lines)
  • /home/stephan/Projekte/KI/AudioEngineHub/app/engines/kokoro_voices.py (222 lines)

🟡 Phase 3: Docker Integration (50% Complete)

  • Updated Dockerfile - Added espeak-ng system dependency

    • Line 17: Added espeak-ng to apt-get install
  • Update requirements.txt - Need to add Kokoro Python dependencies

    • NEXT STEP: Add these lines to requirements.txt:
      # Kokoro TTS Engine
      kokoro>=0.9.2
      soundfile
      phonemizer
      scipy
      munch
      

What's Left To Do

⏳ Phase 4: Configuration (Not Started)

1. Update app/config.py

Add Kokoro configuration settings:

class Settings(BaseSettings):
    # ... existing settings ...

    # Kokoro Engine Configuration
    KOKORO_DEVICE: str = "cuda"  # or "cpu"
    KOKORO_TIMEOUT_SECONDS: int = 30

File: /home/stephan/Projekte/KI/AudioEngineHub/app/config.py Location: Add after line 25 (after FFMPEG_TIMEOUT_SECONDS)

2. Update app/main.py

Register Kokoro engine in the engine registry:

from app.engines.kokoro import KokoroEngine  # Add this import

ALL_ENGINES = {
    "piper": PiperEngine,
    "styletts": StyleTTSEngine,
    "chattts": ChatTTSEngine,
    "f5-tts": F5TTSEngine,
    "kokoro": KokoroEngine,  # Add this line
}

File: /home/stephan/Projekte/KI/AudioEngineHub/app/main.py Location: Line 28 (import), Line 38 (registry)

3. Update .env.example

Add Kokoro configuration documentation:

# Kokoro Engine Configuration
KOKORO_DEVICE=cuda  # cuda or cpu
KOKORO_TIMEOUT_SECONDS=30

File: /home/stephan/Projekte/KI/AudioEngineHub/.env.example Location: Add after FFMPEG_TIMEOUT_SECONDS

4. Optional: Update .env

To enable Kokoro by default:

ACTIVE_ENGINES='["piper", "kokoro"]'

⏳ Phase 5: Testing (Not Started)

1. Test Import and Basic Functionality

# Test engine loads without errors
python3 -c "from app.engines.kokoro import KokoroEngine; print(KokoroEngine().healthcheck())"

Expected output:

{
  "status": "not_available",  # OK if kokoro not installed yet
  "engine": "kokoro",
  "library_available": false,
  "error": "Kokoro library not installed..."
}

2. Install Kokoro Dependencies (Outside Docker First)

# Install espeak-ng
sudo apt-get install espeak-ng

# Install Python packages
pip install kokoro>=0.9.2 soundfile phonemizer scipy munch

3. Test Synthesis Locally

# Run engine selftest
cd /home/stephan/Projekte/KI/AudioEngineHub
python3 app/engines/kokoro.py

Expected output:

Healthcheck: {'status': 'ok', 'engine': 'kokoro', ...}
Models: ['kokoro-en-us', 'kokoro-en-gb', ...]
Voices: ['af_alloy', 'af_aoede', 'af_bella', ...] ...
Selftest: {'selftest': True, 'models': [...], 'voices_count': 54}

4. Test via API

# Start development server
make dev-up

# Test API endpoint
curl -X POST http://localhost:8000/tts \
  -H "Content-Type: application/json" \
  -d '{"text":"Hello from Kokoro!","engine":"kokoro","model":"kokoro-en-us","speaker":"af_bella","format":"ogg"}'

5. Test with Voice Wizard

# Launch voice wizard and test preview with Kokoro voices
cd /home/stephan/Games/University/tts
./scripts/voice-assignment-wizard.sh

⏳ Phase 6: Documentation (Not Started)

1. Update README.md

Add Kokoro to supported engines list:

## Supported TTS Engines

- **Piper** - Fast, lightweight, ONNX-based TTS with 100+ voices
- **Kokoro** - 82M parameter high-quality TTS with 54 voices across 8 languages 🆕
- **StyleTTS** - Expressive style-based TTS (coming soon)
- **ChatTTS** - Conversational TTS (coming soon)
- **F5-TTS** - Advanced flow-based TTS (coming soon)

2. Update CLAUDE.md

Add Kokoro engine section:

### Kokoro Engine (`app/engines/kokoro.py`)
- **Library**: `kokoro` Python package (KPipeline)
- **Models**: 10 language models (en-us, en-gb, ja, zh, es, fr, hi, it, pt, ko)
- **Voices**: 54 voices with gender and language metadata
- **Output**: 24kHz audio natively (WAV), converts to OGG/MP3 via ffmpeg
- **Performance**: ~90× real-time on RTX 3090 Ti
- **Model Size**: 82M parameters (~200MB download per language)
- **Voice Metadata**: `app/engines/kokoro_voices.py` contains all 54 voices with metadata

3. Create Usage Guide (Optional)

Create docs/kokoro-usage.md with:

  • Voice selection guide
  • Language support matrix
  • Performance tips (GPU vs CPU)
  • Troubleshooting common issues

Implementation Checklist

Phase 1: Research ✅

  • Research Kokoro capabilities
  • Document API and dependencies
  • Create implementation plan

Phase 2: Engine Implementation ✅

  • Create app/engines/kokoro.py
  • Implement synthesize() method
  • Implement list_models() method
  • Implement list_voices() method
  • Implement healthcheck() method
  • Implement selftest() method
  • Create app/engines/kokoro_voices.py

Phase 3: Docker Integration 🟡

  • Update Dockerfile (added espeak-ng)
  • → NEXT: Update requirements.txt (add Kokoro packages)
  • Test Docker build
  • Verify GPU access in container (if available)

Phase 4: Configuration ⏳

  • Update app/config.py with Kokoro settings
  • Update .env.example with Kokoro variables
  • Add KokoroEngine import to app/main.py
  • Add Kokoro to ALL_ENGINES in app/main.py
  • Optional: Update .env to enable Kokoro

Phase 5: Testing ⏳

  • Test engine imports without errors
  • Install Kokoro dependencies locally
  • Run selftest (python3 app/engines/kokoro.py)
  • Test synthesis via API
  • Test with voice wizard
  • Verify caching works correctly
  • Check for temp file leaks
  • Performance benchmarks

Phase 6: Documentation ⏳

  • Update README.md
  • Update CLAUDE.md
  • Optional: Create Kokoro usage guide
  • Document voice selection
  • Add troubleshooting section

Phase 7: Deployment ⏳

  • Rebuild Docker image
  • Test in development environment
  • Commit changes to git
  • Push to remote
  • Tag release (optional)

Files Created/Modified

✅ Created Files

  1. KOKORO_IMPLEMENTATION_PLAN.md - Complete implementation plan (280 lines)
  2. app/engines/kokoro.py - Kokoro engine implementation (316 lines)
  3. app/engines/kokoro_voices.py - Voice metadata (222 lines)
  4. KOKORO_IMPLEMENTATION_STATUS.md - This file

✅ Modified Files

  1. Dockerfile - Added espeak-ng dependency (line 17)

⏳ Files To Modify

  1. requirements.txt - Add Kokoro Python packages
  2. app/config.py - Add Kokoro configuration
  3. app/main.py - Register Kokoro engine
  4. .env.example - Document Kokoro config
  5. README.md - Add Kokoro to engines list
  6. CLAUDE.md - Add Kokoro engine details

Next Session: Action Plan

Step 1: Complete Docker Integration (5 minutes)

cd /home/stephan/Projekte/KI/AudioEngineHub

# Add to requirements.txt (after line 11):
cat >> requirements.txt << 'EOF'

# Kokoro TTS Engine
kokoro>=0.9.2
soundfile
phonemizer
scipy
munch
EOF

Step 2: Complete Configuration (10 minutes)

  1. Edit app/config.py - add KOKORO_DEVICE and KOKORO_TIMEOUT_SECONDS
  2. Edit app/main.py - import KokoroEngine and add to ALL_ENGINES
  3. Edit .env.example - document new config options

Step 3: Test Locally (15 minutes)

  1. Install dependencies: sudo apt-get install espeak-ng && pip install kokoro>=0.9.2 soundfile phonemizer scipy munch
  2. Run selftest: python3 app/engines/kokoro.py
  3. Start server: make dev-up (will rebuild Docker image)
  4. Test API synthesis with Kokoro
  5. Test voice wizard preview

Step 4: Documentation (10 minutes)

  1. Update README.md with Kokoro
  2. Update CLAUDE.md with Kokoro engine details
  3. Commit all changes to git

Total Time Remaining: ~40 minutes


Key Technical Details

Language Code Mapping

KOKORO_LANG_CODES = {
    "kokoro-en-us": "a",   # American English
    "kokoro-en-gb": "b",   # British English
    "kokoro-fr": "fr",     # French
    "kokoro-es": "es",     # Spanish
    "kokoro-ja": "ja",     # Japanese
    "kokoro-zh": "zh",     # Chinese
    "kokoro-it": "it",     # Italian
    "kokoro-pt": "pt",     # Portuguese
    "kokoro-hi": "hi",     # Hindi
    "kokoro-ko": "ko",     # Korean
}

Voice Organization

  • 54 total voices across 8 languages
  • Naming convention: {language}{gender}_{name}
    • af_ = American Female
    • am_ = American Male
    • bf_ = British Female
    • bm_ = British Male
    • jf_/jm_ = Japanese F/M
    • zf_/zm_ = Chinese F/M
    • etc.
  • Most popular: af_bella, af_sarah, af_sky, am_adam, am_michael

API Usage Example

{
  "text": "Hello from Kokoro TTS!",
  "engine": "kokoro",
  "model": "kokoro-en-us",
  "speaker": "af_bella",
  "format": "ogg"
}

Known Issues & Considerations

1. Model Download on First Use

  • Issue: First synthesis will be slow (~30-60s) due to model download
  • Size: ~200MB per language model
  • Location: Models cached in ~/.cache/huggingface/
  • Solution: Expected behavior, document in README

2. GPU vs CPU Performance

  • GPU: ~90× real-time (RTX 3090 Ti)
  • CPU: ~5-10× real-time (estimated)
  • Fallback: Code supports both, configurable via KOKORO_DEVICE env var

3. Voice-Language Compatibility

  • Issue: Not all voices work with all languages
  • Solution: kokoro_voices.py filters voices by language
  • API: /speakers?engine=kokoro&model=kokoro-en-us returns only compatible voices

4. Dependencies

  • espeak-ng: Required system package for phonemization
  • soundfile: Required for WAV file I/O
  • phonemizer: Required for text-to-phoneme conversion
  • scipy, munch: Required by Kokoro library

Testing Commands Reference

Local Testing (Outside Docker)

# Install system dependency
sudo apt-get install espeak-ng

# Install Python packages
pip install kokoro>=0.9.2 soundfile phonemizer scipy munch

# Test engine
cd /home/stephan/Projekte/KI/AudioEngineHub
python3 app/engines/kokoro.py

# Test voice metadata
python3 -c "from app.engines.kokoro_voices import ALL_VOICES, get_voices_for_model; print(f'Total voices: {len(ALL_VOICES)}'); print(f'EN-US voices: {get_voices_for_model(\"kokoro-en-us\")}')"

Docker Testing

# Rebuild and start
make down
make dev-up

# Check logs
make logs

# Health check
make health-check

# Test synthesis
curl -X POST http://localhost:8000/tts \
  -H "Content-Type: application/json" \
  -d '{"text":"Testing Kokoro TTS engine","engine":"kokoro","model":"kokoro-en-us","speaker":"af_bella","format":"ogg"}' | jq

Voice Wizard Testing

cd /home/stephan/Games/University/tts
./scripts/voice-assignment-wizard.sh

# Should show Kokoro engine in dropdown
# Should list 54 voices (or subset by language)
# Preview should work with all Kokoro voices

Success Criteria

Before marking this implementation as complete, verify:

  • Kokoro engine passes healthcheck
  • All 10 language models are listed in /models endpoint
  • 54 voices are accessible via /speakers endpoint
  • Voice filtering by model/language works correctly
  • Synthesis produces valid audio files
  • Synthesis speed: >50× real-time on GPU (if available)
  • Format conversion works (WAV/OGG/MP3)
  • Cache versioning works correctly
  • No temp file leaks after synthesis
  • No hung processes under load
  • Voice wizard integration works
  • Documentation is complete and accurate

Resources & References

Official Documentation

Technical Articles

Installation Guides


Timeline & Estimates

Phase Status Time Spent Remaining Total
Phase 1: Research ✅ Complete 2h 0h 2h
Phase 2: Implementation ✅ Complete 2h 0h 2h
Phase 3: Docker 🟡 50% 0.5h 0.5h 1h
Phase 4: Configuration ⏳ Pending 0h 0.5h 0.5h
Phase 5: Testing ⏳ Pending 0h 1h 1h
Phase 6: Documentation ⏳ Pending 0h 0.5h 0.5h
Total 40% Complete 4.5h 2.5h 7h

Original Estimate: 8-12 hours Current Progress: ~5 hours spent, ~2-3 hours remaining On Track: Yes, ahead of schedule


Questions for User (Next Session)

  1. GPU Availability: Do you have a CUDA-capable GPU? This affects:

    • Performance expectations (90× vs 5× real-time)
    • Default KOKORO_DEVICE setting (cuda vs cpu)
  2. Language Priority: Which languages do you need most?

    • Only EN-US model will be tested initially
    • Other languages can be tested on demand
  3. Docker vs Local: Prefer testing locally first or directly in Docker?

    • Local testing is faster for iteration
    • Docker testing validates full deployment
  4. Voice Wizard: Should we integrate Kokoro metadata into the voice wizard UI?

    • Could show language, gender, description for each voice
    • Requires changes to audioengine_client.py

End of Status Report

Last updated: 2025-12-05 00:48 CET Next session: Continue with Step 1 (Complete Docker Integration)