stephan c06fd677dc fix: Resolve 7 critical bugs in Piper TTS engine and add cache versioning
This commit fixes all critical bugs identified in the Piper engine that were
causing server crashes, hung processes, resource leaks, and audio synthesis
failures.

## Bug Fixes

**Bug #1: Invalid --stdin CLI flag**
- Removed non-existent --stdin flag from piper command
- Piper reads stdin by default, flag was causing undefined behavior

**Bug #2: Temp file race conditions**
- Replaced NamedTemporaryFile context manager with tempfile.mkstemp()
- Prevents file handle locking issues on some systems

**Bug #3: Missing process timeouts**
- Added asyncio.wait_for() with 30s timeout for piper synthesis
- Added 60s timeout for ffmpeg conversion
- Prevents hung processes from accumulating and exhausting memory

**Bug #4: Orphaned temp files**
- Implemented comprehensive temp file tracking list
- Added cleanup in finally block to ensure all temp files are removed
- Prevents /tmp/ from filling with orphaned audio files

**Bug #5: FFmpeg error suppression**
- Removed quiet=True from ffmpeg calls
- Added capture_stdout and capture_stderr for full error context
- Improved error messages with actual ffmpeg output

**Bug #6: No config file validation**
- Added upfront validation for .onnx and .onnx.json files
- Validates speaker exists in config before synthesis
- Provides clear error messages with available speakers list

**Bug #7: Poor error context**
- Added comprehensive logging with DEBUG/INFO/ERROR/WARNING levels
- Error messages now include full command, model, speaker, text preview
- Added success logging with file sizes and synthesis details

## New Features

**Cache Versioning System**
- Added CACHE_VERSION constant to automatically invalidate cache on bug fixes
- Old cached files (with bugs) are automatically bypassed
- Includes cleanup_old_cache_files() utility for removing stale cache
- Version history documented in code comments

**Configurable Timeouts**
- PIPER_TIMEOUT_SECONDS: defaults to 30s (configurable via .env)
- FFMPEG_TIMEOUT_SECONDS: defaults to 60s (configurable via .env)

**Path Fixes**
- Fixed audio file path validation in main.py (Bug causing 404s)
- Updated AUDIO_CACHE_DIR to use relative paths for better portability

**Documentation**
- Added CLAUDE.md for future AI assistant context
- Updated .env.example with new timeout configuration options

## Testing

- ✅ Health check passed
- ✅ Synthesis working correctly
- ✅ No temp file leaks
- ✅ No hung processes
- ✅ Cache versioning prevents old buggy files from being served
- ✅ Voice preview in wizard working correctly

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-Authored-By: Claude <noreply@anthropic.com>
2025-12-05 00:38:00 +01:00
2025-12-04 11:58:36 +01:00
2025-12-04 11:58:36 +01:00
2025-12-04 11:58:36 +01:00
2025-12-04 11:58:36 +01:00
2025-12-04 11:58:36 +01:00

AudioEngineHub

AudioEngineHub is a local-first, modular, multi-engine Text-to-Speech (TTS) server designed for homelabs and automation. It provides a single, unified API to interact with various TTS engines like Piper and StyleTTS.

Features

  • Multi-Engine Support: Easily switch between different TTS engines.
  • Configurable Engines: Activate or deactivate engines on the fly via a simple configuration file.
  • Caching: Caches generated audio to save resources and provide faster responses for repeated requests.
  • Dockerized: Runs in a containerized environment for easy setup and dependency management.
  • Automatic Port Finding: Automatically finds and uses a free port when building locally.
  • Container Registry Support: Pre-configured to push to and pull from a container registry.

Getting Started

This guide covers local development. For information on using the container registry, see the "Container Registry" section below.

Prerequisites

Downloading Models (Crucial Step!)

The Docker image for AudioEngineHub does not include the large TTS model files to keep the image small and portable. You need to manually download the models for the engines you wish to use and place them in the correct local directory. The docker-compose.yml then makes these models available to the container via a volume mount.

Piper Models

Instructions:

  1. Go to the link above and navigate to a voice you want to use (e.g., en/en_GB/vctk/medium/).
  2. For each voice, you need to download two files:
    • The .onnx model file (e.g., en_GB-vctk-medium.onnx)
    • The corresponding .onnx.json configuration file (e.g., en_GB-vctk-medium.onnx.json)
  3. Create a directory for the voice inside your local app/models/piper/ directory. The directory name must match the model's base name (e.g., en_GB-vctk-medium).
  4. Place both downloaded files into that new directory.

Example: Setting up en_GB-vctk-medium:

Your local directory structure should look like this:

AudioEngineHub/
├── app/
│   ├── models/
│   │   ├── piper/
│   │   │   ├── en_GB-vctk-medium/  <-- This directory's name MUST match the model name
│   │   │   │   ├── en_GB-vctk-medium.onnx
│   │   │   │   └── en_GB-vctk-medium.onnx.json
│   │   └── styletts/ # Placeholder, no external models currently needed
│   └── ...
├── ...

StyleTTS Models

The styletts engine is currently a placeholder (dummy implementation) and does not require external model downloads at this time. Its list_models() method provides hardcoded model names.

Local Development Setup

  1. Clone the repository:

    git clone <repository_url>
    cd AudioEngineHub
    
  2. Configure the environment: Create a .env file by copying the example file:

    cp .env.example .env
    

    Open the .env file and configure the ACTIVE_ENGINES list to include the engines you want to use. Make sure the model directories exist for activated engines (e.g., if you enable piper, ensure its models are downloaded). For example:

    ACTIVE_ENGINES='["piper", "styletts"]'
    
  3. Build and start the container: Use the make dev-up command to build the Docker image from your local source and start the service.

    make dev-up
    

    This command will automatically find a free port, build the image, and run the application.

    Note: For the most reliable port detection, it is recommended to run the command with sudo:

    sudo make dev-up
    

Container Registry

The project is configured to work with the container registry at git.wlkns.org.

Pushing an Image

  1. Log in to the Registry: You only need to do this once per machine.

    docker login git.wlkns.org
    
  2. Push the Image: This command will build your image, tag it correctly, and push it to the registry.

    make push
    

Pulling and Running an Image

  1. Pull the Image: To download the latest image from the registry:

    make pull
    
  2. Run the Image: This command will start the application using the pre-built image from the registry (pulling it if necessary).

    make up
    

Usage

Endpoints

  • POST /tts: The main endpoint to synthesize text to speech.
  • GET /health: Check the health of the API and the status of the loaded engines.
  • GET /engines: List the currently active engines.
  • GET /models: List the available models for each active engine. _ GET /speakers: List the available speakers for a given engine and model.

Makefile Commands

The project includes a Makefile with several commands to simplify development and management:

  • make dev-up: Build the image from local source and start the application. Recommended for development.
  • make up: Start the application using the image from the container registry (pulls if not present).
  • make down: Stop the application container(s).
  • make logs: View the application logs.
  • make health-check: Run a sanity check to ensure the deployed container is healthy and all engines are "ok".
  • make pull: Pull the latest image from the container registry.
  • make push: Build, tag, and push the image to the container registry.
  • make test: Run the pytest test suite.
  • make help: Display a list of all available commands.

Configuration

The application is configured through the .env file in the root of the project.

  • ACTIVE_ENGINES: A comma-separated list of strings specifying which TTS engines to activate. Available engines are defined in app/main.py.
  • HOST: The host address for the server (defaults to 0.0.0.0).
  • PORT: The internal port for the server (defaults to 8000).
  • IMAGE_NAME: The name of the Docker image to build (defaults to audioenginehub).
Description
Dies ist ein Python basierter TTS Server der mehrere engines zur verfühung stellt.
Readme 192 KiB
Languages
Python 87.8%
Makefile 8.6%
Dockerfile 1.8%
Shell 1.8%