Files
audio-engine-hub/session_resumee.md
stephan a466d8e00f fix: Resolve model loading and deployment issues
This commit fixes several critical issues that prevented the service from deploying correctly and loading the TTS models.

- **Fix Model Loading:** Corrected the volume mount path in  to point to the correct  source directory. The  engine code was also updated to use absolute paths () inside the container, making the model loading mechanism robust.
- **Update Session Resume:** The  has been updated to reflect the final debugging steps and the successful resolution of all issues.
2025-12-04 22:24:58 +01:00

8.2 KiB

Session Resumee - AudioEngineHub Project Refactoring

Date: Donnerstag, 4. Dezember 2025

Objective: Refactor the AudioEngineHub project to follow best practices for containerization, error handling, and project structure, based on a provided guide.md document.


Initial Project State & Overview

The project was a modular FastAPI-based TTS server with a TTSEngineBase interface. Key initial findings:

  • piper was functional.
  • styletts and chattts were dummy implementations.
  • f5_tts was implemented but inactive.
  • Configuration was scattered and hardcoded.
  • Error handling was basic.
  • Tests (pytest and unittest) existed but were limited and inconsistent.
  • No containerization strategy was in place, leading to potential dependency hell.

Refactoring Phase 1: Robustness & Configuration

  1. Centralized Configuration:

    • Replaced hardcoded values with pydantic-settings for .env file management.
    • config.py created (later moved to app/config.py).
    • IMAGE_NAME in Makefile was also user-configurable.
    • requirements.txt was updated with pydantic-settings.
    • docker-compose.yml was updated to use .env.
  2. Configurable Engines Feature:

    • Implemented dynamic ENGINE_REGISTRY loading based on ACTIVE_ENGINES setting in .env.
    • Allows easy activation/deactivation of TTS engines.
  3. Robust Error Handling:

    • Implemented comprehensive input validation in the /tts endpoint (checking engine, model, speaker existence).
    • Added dependency checks (e.g., ffmpeg, piper executables) to engines, reporting HTTP 503 for unavailable engines.
    • Secured tempfile.mktemp usage by replacing it with tempfile.NamedTemporaryFile.
    • Wrapped synthesis logic in try-except blocks to catch and propagate engine-specific errors as HTTP 500.
  4. Asynchronous Operations:

    • Changed TTSEngineBase.synthesize and selftest to async.
    • Refactored all concrete engine implementations (piper, f5_tts, styletts, chattts) to use async def methods.
    • Updated app/main.py's /tts endpoint to be async and use await for engine calls and asyncio.gather for concurrent chunk synthesis.
    • Wrapped blocking I/O (file ops, ffmpeg) and CPU-bound tasks in asyncio.to_thread.
    • Updated tests/test_f5_tts.py to correctly await async calls.

Refactoring Phase 2: Containerization & Workflow (Based on guide.md)

  1. Integrated Makefile:

    • Created a Makefile with targets for build, run, stop, logs, shell, up, down, test, tag, push, clean, help.
    • Included robust shell functions for port checking (check_traefik_ports_free, check_app_port_free).
    • Set SHELL := /bin/bash in Makefile to ensure correct shell interpretation.
    • Ensured PYTHONPATH=$(PWD) is set for make test.
  2. Adopted Multi-Stage Dockerfile:

    • Implemented a multi-stage Dockerfile (builder/runner stages).
    • builder stage creates a Python virtual environment and installs requirements.txt (including gunicorn).
    • runner stage uses python:3.11-slim, installs runtime system dependencies (ffmpeg), creates an unprivileged appuser, and sets the production CMD to gunicorn with uvicorn workers.
  3. Refactored Project Structure (app/ package):

    • Created an app/ directory.
    • Moved main.py, config.py, engines/, models/, utils/ into app/.
    • Created app/__init__.py.
    • Updated all Python import paths (from app.config import settings, from app.engines.piper import PiperEngine, etc.).
    • Updated internal references in engine files (e.g., model_dir, voices_dir).
  4. Refined docker-compose.yml with Traefik:

    • Integrated traefik service for dynamic reverse proxying during local development.
    • Modified app service with Traefik labels and connected both services to a web network.
    • Adjusted Docker volumes mounts to match the new app/ structure (e.g., ./models:/home/appuser/app/models).
    • Updated app service command for Uvicorn hot-reloading in dev.
  5. Centralized Testing Workflow:

    • Removed test_run.sh.
    • Integrated make test for running pytest --cov=. app/ tests/.
  6. Dedicated Documentation:

    • Created docs/ directory.
    • Moved guide.md to docs/guide.md.

Verification & Troubleshooting

  • Tests: All unit/integration tests (make test) are passing.
  • Local Run (Virtual Env): Initial local runs (python app/main.py) failed due to ModuleNotFoundError (fixed by python -m app.main) and PermissionError (fixed by needing to mock or redirect settings.AUDIO_CACHE_DIR for local direct execution, but not strictly needed for successful app execution through uvicorn). The current approach is to verify in Docker.
  • Docker Compose:
    • Initial make up failures were due to an outdated docker-compose client (1.29.2) and later, Makefile syntax issues (fixed by setting SHELL := /bin/bash and fixing macros).
    • docker-compose was eventually updated to the docker compose CLI plugin (v5.0.0).
    • make up command finally succeeded in bringing up containers.
    • The curl http://localhost/health command returned 404 page not found. This indicates a potential routing issue with Traefik or the application not being responsive on the expected path within the container. (This is the last unresolved issue).

Final Debugging & Resolution (Post-Refactoring)

After the major refactoring, the service was unable to start and was returning 404 errors. The final debugging session addressed these issues.

  1. FastAPI App Startup Fix:

    • Problem: The Uvicorn server was failing with Attribute "app" not found in module "app.main".
    • Solution: The app instance was only created inside the if __name__ == "__main__" block. Moved app = create_app() to the module's global scope in app/main.py so Uvicorn could find it.
  2. Traefik & Docker Compose Issues:

    • Problem: The initial docker-compose.yml included a Traefik service for reverse proxying, which was causing multiple issues (invalid container names due to missing .env variables, Docker client API version errors). The user also clarified that they use an existing reverse proxy and did not want a new Traefik container deployed.
    • Solution: Removed the traefik service entirely from docker-compose.yml. The app service's port was exposed directly to the host.
  3. Model Loading Failure:

    • Problem: The piper engine was not loading any models, returning an empty list and causing /tts requests to fail with a 422 Unprocessable Entity error.
    • Diagnosis: The root cause was an incorrect volume mount. The docker-compose.yml was attempting to mount a non-existent, empty ./models directory from the host. The actual models were located in ./app/models.
    • Solution:
      1. Corrected the docker-compose.yml volume mount to point to the correct source directory: - ./app/models:/models.
      2. Updated the app/engines/piper.py code to use the absolute path /models/piper/ to look for models inside the container, making the configuration more robust and independent of the working directory.
  4. Developer Experience (DX) Improvements:

    • Automatic Port Finding: The Makefile was enhanced. The make up and make run commands now automatically find a free port starting from 8000, preventing port conflicts.
    • Health Check Command: A make health-check target was added. It runs a sanity check against the deployed container's /health endpoint to verify that the service is up and all engines are reporting an "ok" status.
    • Documentation:
      • The README.md was completely rewritten to provide a clear and up-to-date guide for getting started, usage, and available make commands.
      • A note was added to docs/guide.md to clarify that it describes an older, more advanced setup and to point readers to the new README.md.

Final Status: The service is now fully deployable via make up and passes make health-check. All identified startup issues and bugs have been resolved.