This commit fixes all critical bugs identified in the Piper engine that were causing server crashes, hung processes, resource leaks, and audio synthesis failures. ## Bug Fixes **Bug #1: Invalid --stdin CLI flag** - Removed non-existent --stdin flag from piper command - Piper reads stdin by default, flag was causing undefined behavior **Bug #2: Temp file race conditions** - Replaced NamedTemporaryFile context manager with tempfile.mkstemp() - Prevents file handle locking issues on some systems **Bug #3: Missing process timeouts** - Added asyncio.wait_for() with 30s timeout for piper synthesis - Added 60s timeout for ffmpeg conversion - Prevents hung processes from accumulating and exhausting memory **Bug #4: Orphaned temp files** - Implemented comprehensive temp file tracking list - Added cleanup in finally block to ensure all temp files are removed - Prevents /tmp/ from filling with orphaned audio files **Bug #5: FFmpeg error suppression** - Removed quiet=True from ffmpeg calls - Added capture_stdout and capture_stderr for full error context - Improved error messages with actual ffmpeg output **Bug #6: No config file validation** - Added upfront validation for .onnx and .onnx.json files - Validates speaker exists in config before synthesis - Provides clear error messages with available speakers list **Bug #7: Poor error context** - Added comprehensive logging with DEBUG/INFO/ERROR/WARNING levels - Error messages now include full command, model, speaker, text preview - Added success logging with file sizes and synthesis details ## New Features **Cache Versioning System** - Added CACHE_VERSION constant to automatically invalidate cache on bug fixes - Old cached files (with bugs) are automatically bypassed - Includes cleanup_old_cache_files() utility for removing stale cache - Version history documented in code comments **Configurable Timeouts** - PIPER_TIMEOUT_SECONDS: defaults to 30s (configurable via .env) - FFMPEG_TIMEOUT_SECONDS: defaults to 60s (configurable via .env) **Path Fixes** - Fixed audio file path validation in main.py (Bug causing 404s) - Updated AUDIO_CACHE_DIR to use relative paths for better portability **Documentation** - Added CLAUDE.md for future AI assistant context - Updated .env.example with new timeout configuration options ## Testing - ✅ Health check passed - ✅ Synthesis working correctly - ✅ No temp file leaks - ✅ No hung processes - ✅ Cache versioning prevents old buggy files from being served - ✅ Voice preview in wizard working correctly 🤖 Generated with [Claude Code](https://claude.com/claude-code) Co-Authored-By: Claude <noreply@anthropic.com>
8.6 KiB
CLAUDE.md
This file provides guidance to Claude Code (claude.ai/code) when working with code in this repository.
Project Overview
AudioEngineHub is a local-first, modular, multi-engine Text-to-Speech (TTS) server built with FastAPI. It provides a unified API to interact with various TTS engines (Piper, StyleTTS, ChatTTS, F5-TTS) with caching, audio format conversion, and Docker deployment support.
Architecture
Core Components
FastAPI Application Factory (app/main.py)
- Uses
create_app()factory pattern to dynamically load engines based on configuration - Maintains
app.ENGINE_REGISTRYdict mapping engine names to instantiated engine objects - Engines are loaded from
settings.ACTIVE_ENGINESat startup - All available engines defined in
ALL_ENGINESdict (lines 34-39) - Startup event checks for models directory availability with retry logic (lines 192-213)
Engine Architecture (app/engines/)
- All engines inherit from
TTSEngineBaseabstract class - Required methods:
synthesize(),list_models(),list_voices(),healthcheck(),selftest() - All
synthesize()methods are async and return file paths to generated audio - Engines are stateless and instantiated once per app lifecycle
Configuration (app/config.py)
- Uses pydantic-settings for type-safe configuration
ACTIVE_ENGINESis a Set[str] loaded from.envfile as JSON array- Default model path:
/models/piper/{model_name}/{model_name}.onnx - Audio cache directory:
/home/appuser/app/asset/audio
Utility Modules
app/utils/text.py: Text chunking for long inputsapp/utils/audio.py: Audio concatenation via ffmpegapp/utils/cache.py: Cache key generation based on request parameters
Volume Mounts
The docker-compose.yml mounts three critical paths:
./app:/home/appuser/app- Hot-reload for development./app/models:/models- TTS model files (not included in image)./asset:/home/appuser/app/asset- Persistent audio cache
CRITICAL: The model mount is ./app/models:/models (maps to /models inside container). The engine code uses absolute path /models/piper/ to locate models. Do NOT use ./models as the source - models must be in the app/models/ directory on the host.
Model Structure
Piper models follow this directory structure:
app/models/piper/{model_name}/
├── {model_name}.onnx # Model file
└── {model_name}.onnx.json # Config with speaker_id_map
The speaker_id_map in the JSON config maps speaker names to numeric IDs. See app/engines/piper.py:30-47 for speaker resolution logic.
Development Commands
Local Development
# Build and run from local source (auto-finds free port)
make dev-up
# Run with pre-built registry image
make up
# Stop containers
make down
# View logs
make logs
# Access container shell
make shell
Testing
# Run pytest with coverage (requires .venv activated or auto-activates)
make test
# Health check on running container
make health-check
Container Registry Workflow
# Login (one-time setup)
docker login git.wlkns.org
# Build, tag, and push to registry
make push
# Pull latest from registry
make pull
Key Implementation Details
TTS Request Flow
- Request arrives at
/ttsendpoint (line 69) - Engine lookup and healthcheck (lines 72-78)
- Model and speaker validation (lines 81-87)
- Cache key generation and lookup (lines 90-108)
- If not cached: synthesize (optionally with chunking) (lines 111-122)
- Cache result and return URL or base64 (lines 125-143)
Audio Format Handling
IMPORTANT: Piper engine currently returns WAV only. The ffmpeg conversion code is commented out at lines 99-118 in app/engines/piper.py. This was done to debug an ffmpeg issue. When re-enabling format conversion:
- Uncomment lines 103-118
- Remove line 101 (direct WAV return)
- Ensure ffmpeg is available in container
Critical: Piper Engine Input Method (app/engines/piper.py)
The Piper engine uses --input-file with a temporary file instead of stdin. Previous attempts to use --stdin with asyncio.subprocess.communicate(input=...) caused issues. Current implementation (lines 71-96):
- Creates temporary text file with input text
- Passes file path via
--input-fileflag to piper executable - Cleans up temp file in finally block
Known Issue: There were reports of the Piper engine occasionally synthesizing incorrect text or causing server crashes. If investigating synthesis issues:
- Check the temporary file creation/cleanup in
piper.py:71-96 - Verify the text is being written correctly to the temp file
- Test manually inside container:
echo "text" | piper --model /models/piper/{model}/{model}.onnx --output_file test.wav
Async Patterns
- All synthesis operations are async
- Blocking I/O (file operations, ffmpeg) wrapped in
asyncio.to_thread() - Concurrent chunk synthesis via
asyncio.gather()(line 115)
Engine Registration
To add a new engine:
- Create class in
app/engines/inheriting fromTTSEngineBase - Add to
ALL_ENGINESdict inapp/main.py - Add engine name to
ACTIVE_ENGINESin.env
Testing Strategy
Tests use pytest with async support and fixtures in tests/conftest.py:
app_client: TestClient for the FastAPI apppiper_engine: Direct access to PiperEngine instance from registry
Run tests with make test which sets PYTHONPATH and activates venv automatically.
Configuration Notes
.env Format:
ACTIVE_ENGINES='["piper", "styletts"]' # JSON array as string
HOST=0.0.0.0
PORT=8000
Registry Configuration (in Makefile):
REGISTRY=git.wlkns.orgUSERNAME=stephanIMAGE_NAME=audio-engine-hubTAG=latest
Common Patterns
Adding a New Endpoint
Follow the pattern in app/main.py:
- Add route to the app instance inside
create_app() - Use
app.ENGINE_REGISTRYto access active engines - Handle engine not found with 404
- Validate models/speakers before processing
Working with Engines
# Access from registry
engine = app.ENGINE_REGISTRY.get("piper")
# Always check health before use
health = engine.healthcheck()
if health.get("status") != "ok":
# Handle unhealthy engine
# Get available resources
models = engine.list_models()
voices = engine.list_voices(model_name)
# Synthesize
audio_path = await engine.synthesize(text, speaker=voice, model=model, fmt="wav")
Docker Multi-Stage Build
The Dockerfile uses a two-stage build:
- Builder stage: Creates venv and installs Python dependencies
- Runner stage: Slim image with ffmpeg, copies venv and app code, runs as unprivileged
appuser
Production uses gunicorn with uvicorn workers (2 workers default).
Critical Issues & Lessons Learned
App Instance Must Be at Module Level
The FastAPI app instance MUST be created at module level in app/main.py (line 227: app = create_app()). Do not only create it inside if __name__ == "__main__" block, or uvicorn/gunicorn will fail with "Attribute 'app' not found" error.
Startup Race Condition with Volume Mounts
The startup event handler (app/main.py:192-213) waits for /models/piper directory to exist and be non-empty. This prevents race conditions where the app starts before Docker has finished mounting volumes. Max retries: 10, retry delay: 2 seconds.
Model Path Must Be Absolute
Engine implementations use absolute path /models/piper/ not relative paths. This ensures consistency regardless of working directory and matches the Docker volume mount structure.
Engine Status
Currently Working:
piper: Functional, uses real Piper TTS executable with ONNX models
Dummy Implementations (for testing only):
styletts: Returns hardcoded model/voice lists, generates dummy audiochattts: Dummy implementationf5-tts: Implemented but may be inactive by default
Debugging Tips
Port Conflicts: The Makefile automatically finds free ports starting from 8000. Run with sudo make dev-up for most reliable port detection.
Container Crashes: If the container crashes during synthesis:
- Check logs:
make logs - Disable uvicorn reload in docker-compose.yml command to see full tracebacks
- Test piper executable directly inside container:
make shell
Model Loading Issues: If list_models() returns empty:
- Verify volume mount in docker-compose.yml points to
./app/models:/models - Check models exist on host:
ls app/models/piper/ - Check inside container:
docker exec -it audio-engine-hub_app ls -la /models/piper/
Connection Reset Errors in Client: Add small delay (time.sleep(1)) before making requests if experiencing ConnectionResetError. This indicates server needs time to fully initialize.