- Add XTTS v2 configuration to .env.example - Refactor Dockerfile to multi-stage build with CUDA 12.1 support - Update Makefile with Kokoro and XTTS test environment targets - Refactor Piper engine (app/engines/piper.py) to use python module execution - Add comprehensive documentation for Kokoro and XTTS plans - Add helper scripts and patches for build process
11 KiB
CLAUDE.md
This file provides guidance to Claude Code (claude.ai/code) when working with code in this repository.
Project Overview
AudioEngineHub is a local-first, modular, multi-engine Text-to-Speech (TTS) server built with FastAPI. It provides a unified API to interact with various TTS engines (Piper, Kokoro, StyleTTS, ChatTTS, F5-TTS) with caching, audio format conversion, and Docker deployment support.
Architecture
Core Components
FastAPI Application Factory (app/main.py)
- Uses
create_app()factory pattern to dynamically load engines based on configuration - Maintains
app.ENGINE_REGISTRYdict mapping engine names to instantiated engine objects - Engines are loaded from
settings.ACTIVE_ENGINESat startup - All available engines defined in
ALL_ENGINESdict (lines 35-41) - Startup event checks for models directory availability with retry logic (lines 192-213)
Engine Architecture (app/engines/)
- All engines inherit from
TTSEngineBaseabstract class - Required methods:
synthesize(),list_models(),list_voices(),healthcheck(),selftest() - All
synthesize()methods are async and return file paths to generated audio - Engines are stateless and instantiated once per app lifecycle
Configuration (app/config.py)
- Uses pydantic-settings for type-safe configuration
ACTIVE_ENGINESis a Set[str] loaded from.envfile as JSON array- Default model path:
/models/piper/{model_name}/{model_name}.onnx - Audio cache directory:
/home/appuser/app/asset/audio
Utility Modules
app/utils/text.py: Text chunking for long inputsapp/utils/audio.py: Audio concatenation via ffmpegapp/utils/cache.py: Cache key generation based on request parameters
Volume Mounts
The docker-compose.yml mounts three critical paths:
./app:/home/appuser/app- Hot-reload for development./app/models:/models- TTS model files (not included in image)./asset:/home/appuser/app/asset- Persistent audio cache
CRITICAL: The model mount is ./app/models:/models (maps to /models inside container). The engine code uses absolute path /models/piper/ to locate models. Do NOT use ./models as the source - models must be in the app/models/ directory on the host.
Model Structure
Piper models follow this directory structure:
app/models/piper/{model_name}/
├── {model_name}.onnx # Model file
└── {model_name}.onnx.json # Config with speaker_id_map
The speaker_id_map in the JSON config maps speaker names to numeric IDs. See app/engines/piper.py:30-47 for speaker resolution logic.
Development Commands
Local Development
# Build and run from local source (auto-finds free port)
make dev-up
# Run with pre-built registry image
make up
# Stop containers
make down
# View logs
make logs
# Access container shell
make shell
Testing
# Run pytest with coverage (requires .venv activated or auto-activates)
make test
# Health check on running container
make health-check
Container Registry Workflow
# Login (one-time setup)
docker login git.wlkns.org
# Build, tag, and push to registry
make push
# Pull latest from registry
make pull
Key Implementation Details
TTS Request Flow
- Request arrives at
/ttsendpoint (line 69) - Engine lookup and healthcheck (lines 72-78)
- Model and speaker validation (lines 81-87)
- Cache key generation and lookup (lines 90-108)
- If not cached: synthesize (optionally with chunking) (lines 111-122)
- Cache result and return URL or base64 (lines 125-143)
Audio Format Handling
IMPORTANT: Piper engine currently returns WAV only. The ffmpeg conversion code is commented out at lines 99-118 in app/engines/piper.py. This was done to debug an ffmpeg issue. When re-enabling format conversion:
- Uncomment lines 103-118
- Remove line 101 (direct WAV return)
- Ensure ffmpeg is available in container
Critical: Piper Engine Input Method (app/engines/piper.py)
The Piper engine uses --input-file with a temporary file instead of stdin. Previous attempts to use --stdin with asyncio.subprocess.communicate(input=...) caused issues. Current implementation (lines 71-96):
- Creates temporary text file with input text
- Passes file path via
--input-fileflag to piper executable - Cleans up temp file in finally block
Known Issue: There were reports of the Piper engine occasionally synthesizing incorrect text or causing server crashes. If investigating synthesis issues:
- Check the temporary file creation/cleanup in
piper.py:71-96 - Verify the text is being written correctly to the temp file
- Test manually inside container:
echo "text" | piper --model /models/piper/{model}/{model}.onnx --output_file test.wav
Async Patterns
- All synthesis operations are async
- Blocking I/O (file operations, ffmpeg) wrapped in
asyncio.to_thread() - Concurrent chunk synthesis via
asyncio.gather()(line 115)
Engine Registration
To add a new engine:
- Create class in
app/engines/inheriting fromTTSEngineBase - Add to
ALL_ENGINESdict inapp/main.py - Add engine name to
ACTIVE_ENGINESin.env
Testing Strategy
Tests use pytest with async support and fixtures in tests/conftest.py:
app_client: TestClient for the FastAPI apppiper_engine: Direct access to PiperEngine instance from registry
Run tests with make test which sets PYTHONPATH and activates venv automatically.
Configuration Notes
.env Format:
ACTIVE_ENGINES='["piper", "styletts"]' # JSON array as string
HOST=0.0.0.0
PORT=8000
Registry Configuration (in Makefile):
REGISTRY=git.wlkns.orgUSERNAME=stephanIMAGE_NAME=audio-engine-hubTAG=latest
Common Patterns
Adding a New Endpoint
Follow the pattern in app/main.py:
- Add route to the app instance inside
create_app() - Use
app.ENGINE_REGISTRYto access active engines - Handle engine not found with 404
- Validate models/speakers before processing
Working with Engines
# Access from registry
engine = app.ENGINE_REGISTRY.get("piper")
# Always check health before use
health = engine.healthcheck()
if health.get("status") != "ok":
# Handle unhealthy engine
# Get available resources
models = engine.list_models()
voices = engine.list_voices(model_name)
# Synthesize
audio_path = await engine.synthesize(text, speaker=voice, model=model, fmt="wav")
Docker Multi-Stage Build
The Dockerfile uses a two-stage build:
- Builder stage: Creates venv and installs Python dependencies
- Runner stage: Slim image with ffmpeg, copies venv and app code, runs as unprivileged
appuser
Production uses gunicorn with uvicorn workers (2 workers default).
Critical Issues & Lessons Learned
App Instance Must Be at Module Level
The FastAPI app instance MUST be created at module level in app/main.py (line 227: app = create_app()). Do not only create it inside if __name__ == "__main__" block, or uvicorn/gunicorn will fail with "Attribute 'app' not found" error.
Startup Race Condition with Volume Mounts
The startup event handler (app/main.py:192-213) waits for /models/piper directory to exist and be non-empty. This prevents race conditions where the app starts before Docker has finished mounting volumes. Max retries: 10, retry delay: 2 seconds.
Model Path Must Be Absolute
Engine implementations use absolute path /models/piper/ not relative paths. This ensures consistency regardless of working directory and matches the Docker volume mount structure.
Engine Status
Currently Working:
piper: Functional, uses real Piper TTS executable with ONNX modelskokoro: Fully functional, uses Kokoro-82M TTS library with 54 voices across 8 languages
Dummy Implementations (for testing only):
styletts: Returns hardcoded model/voice lists, generates dummy audiochattts: Dummy implementationf5-tts: Implemented but may be inactive by default
Kokoro Engine (app/engines/kokoro.py)
Overview:
- Uses
kokoroPython library (KPipeline) for high-performance TTS synthesis - 82M parameter model delivering ~90× real-time performance on RTX 3090 Ti
- 54 voices across 8 languages: EN-US, EN-GB, FR, ES, JA, ZH, IT, PT, HI, KO
- Outputs 24kHz audio natively (WAV), converts to OGG/MP3 via ffmpeg
- Model size: ~200MB per language, auto-downloaded from Hugging Face on first use
- Voice metadata:
app/engines/kokoro_voices.pycontains all 54 voices with metadata
Architecture:
class KokoroEngine(TTSEngineBase):
def __init__(self):
# GPU/CPU detection via settings.KOKORO_DEVICE
# Pipeline instances cached per language code
async def synthesize(text, speaker, model, fmt):
# Uses KPipeline for synthesis (24kHz native output)
# Format conversion via ffmpeg (_run_ffmpeg_blocking)
# Returns temp file path to generated audio
def list_models(self):
# Returns 10 language models:
# kokoro-en-us, kokoro-en-gb, kokoro-fr, kokoro-es,
# kokoro-ja, kokoro-zh, kokoro-it, kokoro-pt, kokoro-hi, kokoro-ko
def list_voices(self, model=None):
# Returns all 54 voices or filtered by language
# Uses get_voices_for_model() from kokoro_voices.py
Language Code Mapping: The engine maps model names to Kokoro's internal language codes:
kokoro-en-us→'a'(American English)kokoro-en-gb→'b'(British English)kokoro-fr→'fr'(French)kokoro-es→'es'(Spanish)kokoro-ja→'ja'(Japanese)kokoro-zh→'zh'(Chinese)kokoro-it→'it'(Italian)kokoro-pt→'pt'(Portuguese)kokoro-hi→'hi'(Hindi)kokoro-ko→'ko'(Korean)
Voice Organization (app/engines/kokoro_voices.py):
- All 54 voices documented with metadata (gender, language, description)
- Naming convention:
{language}{gender}_{name}(e.g.,af_bella,am_adam) - Popular voices:
af_bella,af_sarah,af_sky,am_adam,am_michael - Helper functions:
get_voices_for_model(),get_voice_info()
Model Download & Caching:
- Models auto-download from Hugging Face on first synthesis
- Cached in
~/.cache/huggingface/(inside container) - First synthesis may take 30-60s due to download + compilation
- Subsequent syntheses are fast (~90× real-time on GPU)
Configuration (app/config.py):
KOKORO_DEVICE: "cuda" or "cpu" (default: "cuda")KOKORO_TIMEOUT_SECONDS: Synthesis timeout (default: 30)
Error Handling:
- Applies all bug fixes from Piper engine (timeouts, temp file cleanup, logging)
- Graceful GPU fallback if CUDA unavailable
- Voice validation before synthesis
- Comprehensive error logging with context
Debugging Tips
Port Conflicts: The Makefile automatically finds free ports starting from 8000. Run with sudo make dev-up for most reliable port detection.
Container Crashes: If the container crashes during synthesis:
- Check logs:
make logs - Disable uvicorn reload in docker-compose.yml command to see full tracebacks
- Test piper executable directly inside container:
make shell
Model Loading Issues: If list_models() returns empty:
- Verify volume mount in docker-compose.yml points to
./app/models:/models - Check models exist on host:
ls app/models/piper/ - Check inside container:
docker exec -it audio-engine-hub_app ls -la /models/piper/
Connection Reset Errors in Client: Add small delay (time.sleep(1)) before making requests if experiencing ConnectionResetError. This indicates server needs time to fully initialize.