Files
audio-engine-hub/README.md
stephan 06a484d0aa docs: Add model downloading instructions to README
This commit adds a crucial "Downloading Models" section to the README.md.

- Explains that models are not included in the Docker image and must be downloaded manually.
- Provides detailed instructions for downloading Piper models from Hugging Face, including an example for the  voice.
- Clarifies the expected directory structure for placing the downloaded models.
- Mentions that StyleTTS is currently a placeholder and does not require model downloads.
2025-12-04 23:04:08 +01:00

154 lines
6.1 KiB
Markdown

# AudioEngineHub
AudioEngineHub is a local-first, modular, multi-engine Text-to-Speech (TTS) server designed for homelabs and automation. It provides a single, unified API to interact with various TTS engines like Piper and StyleTTS.
## Features
- **Multi-Engine Support:** Easily switch between different TTS engines.
- **Configurable Engines:** Activate or deactivate engines on the fly via a simple configuration file.
- **Caching:** Caches generated audio to save resources and provide faster responses for repeated requests.
- **Dockerized:** Runs in a containerized environment for easy setup and dependency management.
- **Automatic Port Finding:** Automatically finds and uses a free port when building locally.
- **Container Registry Support:** Pre-configured to push to and pull from a container registry.
## Getting Started
This guide covers local development. For information on using the container registry, see the "Container Registry" section below.
### Prerequisites
- [Docker](https://docs.docker.com/get-docker/)
- [Docker Compose](https://docs.docker.com/compose/install/)
### Downloading Models (Crucial Step!)
The Docker image for AudioEngineHub does *not* include the large TTS model files to keep the image small and portable. You need to **manually download** the models for the engines you wish to use and place them in the correct local directory. The `docker-compose.yml` then makes these models available to the container via a volume mount.
#### Piper Models
* **Source:** [https://huggingface.co/rhasspy/piper-voices/tree/main](https://huggingface.co/rhasspy/piper-voices/tree/main)
**Instructions:**
1. Go to the link above and navigate to a voice you want to use (e.g., `en/en_GB/vctk/medium/`).
2. For each voice, you need to download two files:
* The `.onnx` model file (e.g., `en_GB-vctk-medium.onnx`)
* The corresponding `.onnx.json` configuration file (e.g., `en_GB-vctk-medium.onnx.json`)
3. Create a directory for the voice inside your local `app/models/piper/` directory. The directory name must match the model's base name (e.g., `en_GB-vctk-medium`).
4. Place both downloaded files into that new directory.
**Example: Setting up `en_GB-vctk-medium`:**
Your local directory structure should look like this:
```
AudioEngineHub/
├── app/
│ ├── models/
│ │ ├── piper/
│ │ │ ├── en_GB-vctk-medium/ <-- This directory's name MUST match the model name
│ │ │ │ ├── en_GB-vctk-medium.onnx
│ │ │ │ └── en_GB-vctk-medium.onnx.json
│ │ └── styletts/ # Placeholder, no external models currently needed
│ └── ...
├── ...
```
#### StyleTTS Models
The `styletts` engine is currently a placeholder (dummy implementation) and does not require external model downloads at this time. Its `list_models()` method provides hardcoded model names.
### Local Development Setup
1. **Clone the repository:**
```bash
git clone <repository_url>
cd AudioEngineHub
```
2. **Configure the environment:**
Create a `.env` file by copying the example file:
```bash
cp .env.example .env
```
Open the `.env` file and configure the `ACTIVE_ENGINES` list to include the engines you want to use. Make sure the model directories exist for activated engines (e.g., if you enable `piper`, ensure its models are downloaded). For example:
```
ACTIVE_ENGINES='["piper", "styletts"]'
```
3. **Build and start the container:**
Use the `make dev-up` command to build the Docker image from your local source and start the service.
```bash
make dev-up
```
This command will automatically find a free port, build the image, and run the application.
> **Note:** For the most reliable port detection, it is recommended to run the command with `sudo`:
> ```bash
> sudo make dev-up
> ```
## Container Registry
The project is configured to work with the container registry at `git.wlkns.org`.
### Pushing an Image
1. **Log in to the Registry:**
You only need to do this once per machine.
```bash
docker login git.wlkns.org
```
2. **Push the Image:**
This command will build your image, tag it correctly, and push it to the registry.
```bash
make push
```
### Pulling and Running an Image
1. **Pull the Image:**
To download the latest image from the registry:
```bash
make pull
```
2. **Run the Image:**
This command will start the application using the pre-built image from the registry (pulling it if necessary).
```bash
make up
```
## Usage
### Endpoints
- `POST /tts`: The main endpoint to synthesize text to speech.
- `GET /health`: Check the health of the API and the status of the loaded engines.
- `GET /engines`: List the currently active engines.
- `GET /models`: List the available models for each active engine.
_ `GET /speakers`: List the available speakers for a given engine and model.
### Makefile Commands
The project includes a `Makefile` with several commands to simplify development and management:
- `make dev-up`: Build the image from local source and start the application. Recommended for development.
- `make up`: Start the application using the image from the container registry (pulls if not present).
- `make down`: Stop the application container(s).
- `make logs`: View the application logs.
- `make health-check`: Run a sanity check to ensure the deployed container is healthy and all engines are "ok".
- `make pull`: Pull the latest image from the container registry.
- `make push`: Build, tag, and push the image to the container registry.
- `make test`: Run the `pytest` test suite.
- `make help`: Display a list of all available commands.
## Configuration
The application is configured through the `.env` file in the root of the project.
- `ACTIVE_ENGINES`: A comma-separated list of strings specifying which TTS engines to activate. Available engines are defined in `app/main.py`.
- `HOST`: The host address for the server (defaults to `0.0.0.0`).
- `PORT`: The internal port for the server (defaults to `8000`).
- `IMAGE_NAME`: The name of the Docker image to build (defaults to `audioenginehub`).