This commit adds a crucial "Downloading Models" section to the README.md. - Explains that models are not included in the Docker image and must be downloaded manually. - Provides detailed instructions for downloading Piper models from Hugging Face, including an example for the voice. - Clarifies the expected directory structure for placing the downloaded models. - Mentions that StyleTTS is currently a placeholder and does not require model downloads.
154 lines
6.1 KiB
Markdown
154 lines
6.1 KiB
Markdown
# AudioEngineHub
|
|
|
|
AudioEngineHub is a local-first, modular, multi-engine Text-to-Speech (TTS) server designed for homelabs and automation. It provides a single, unified API to interact with various TTS engines like Piper and StyleTTS.
|
|
|
|
## Features
|
|
|
|
- **Multi-Engine Support:** Easily switch between different TTS engines.
|
|
- **Configurable Engines:** Activate or deactivate engines on the fly via a simple configuration file.
|
|
- **Caching:** Caches generated audio to save resources and provide faster responses for repeated requests.
|
|
- **Dockerized:** Runs in a containerized environment for easy setup and dependency management.
|
|
- **Automatic Port Finding:** Automatically finds and uses a free port when building locally.
|
|
- **Container Registry Support:** Pre-configured to push to and pull from a container registry.
|
|
|
|
## Getting Started
|
|
|
|
This guide covers local development. For information on using the container registry, see the "Container Registry" section below.
|
|
|
|
### Prerequisites
|
|
|
|
- [Docker](https://docs.docker.com/get-docker/)
|
|
- [Docker Compose](https://docs.docker.com/compose/install/)
|
|
|
|
### Downloading Models (Crucial Step!)
|
|
|
|
The Docker image for AudioEngineHub does *not* include the large TTS model files to keep the image small and portable. You need to **manually download** the models for the engines you wish to use and place them in the correct local directory. The `docker-compose.yml` then makes these models available to the container via a volume mount.
|
|
|
|
#### Piper Models
|
|
|
|
* **Source:** [https://huggingface.co/rhasspy/piper-voices/tree/main](https://huggingface.co/rhasspy/piper-voices/tree/main)
|
|
|
|
**Instructions:**
|
|
|
|
1. Go to the link above and navigate to a voice you want to use (e.g., `en/en_GB/vctk/medium/`).
|
|
2. For each voice, you need to download two files:
|
|
* The `.onnx` model file (e.g., `en_GB-vctk-medium.onnx`)
|
|
* The corresponding `.onnx.json` configuration file (e.g., `en_GB-vctk-medium.onnx.json`)
|
|
3. Create a directory for the voice inside your local `app/models/piper/` directory. The directory name must match the model's base name (e.g., `en_GB-vctk-medium`).
|
|
4. Place both downloaded files into that new directory.
|
|
|
|
**Example: Setting up `en_GB-vctk-medium`:**
|
|
|
|
Your local directory structure should look like this:
|
|
```
|
|
AudioEngineHub/
|
|
├── app/
|
|
│ ├── models/
|
|
│ │ ├── piper/
|
|
│ │ │ ├── en_GB-vctk-medium/ <-- This directory's name MUST match the model name
|
|
│ │ │ │ ├── en_GB-vctk-medium.onnx
|
|
│ │ │ │ └── en_GB-vctk-medium.onnx.json
|
|
│ │ └── styletts/ # Placeholder, no external models currently needed
|
|
│ └── ...
|
|
├── ...
|
|
```
|
|
|
|
#### StyleTTS Models
|
|
|
|
The `styletts` engine is currently a placeholder (dummy implementation) and does not require external model downloads at this time. Its `list_models()` method provides hardcoded model names.
|
|
|
|
### Local Development Setup
|
|
|
|
1. **Clone the repository:**
|
|
```bash
|
|
git clone <repository_url>
|
|
cd AudioEngineHub
|
|
```
|
|
|
|
2. **Configure the environment:**
|
|
Create a `.env` file by copying the example file:
|
|
```bash
|
|
cp .env.example .env
|
|
```
|
|
Open the `.env` file and configure the `ACTIVE_ENGINES` list to include the engines you want to use. Make sure the model directories exist for activated engines (e.g., if you enable `piper`, ensure its models are downloaded). For example:
|
|
```
|
|
ACTIVE_ENGINES='["piper", "styletts"]'
|
|
```
|
|
|
|
3. **Build and start the container:**
|
|
Use the `make dev-up` command to build the Docker image from your local source and start the service.
|
|
```bash
|
|
make dev-up
|
|
```
|
|
This command will automatically find a free port, build the image, and run the application.
|
|
|
|
> **Note:** For the most reliable port detection, it is recommended to run the command with `sudo`:
|
|
> ```bash
|
|
> sudo make dev-up
|
|
> ```
|
|
|
|
## Container Registry
|
|
|
|
The project is configured to work with the container registry at `git.wlkns.org`.
|
|
|
|
### Pushing an Image
|
|
|
|
1. **Log in to the Registry:**
|
|
You only need to do this once per machine.
|
|
```bash
|
|
docker login git.wlkns.org
|
|
```
|
|
|
|
2. **Push the Image:**
|
|
This command will build your image, tag it correctly, and push it to the registry.
|
|
```bash
|
|
make push
|
|
```
|
|
|
|
### Pulling and Running an Image
|
|
|
|
1. **Pull the Image:**
|
|
To download the latest image from the registry:
|
|
```bash
|
|
make pull
|
|
```
|
|
|
|
2. **Run the Image:**
|
|
This command will start the application using the pre-built image from the registry (pulling it if necessary).
|
|
```bash
|
|
make up
|
|
```
|
|
|
|
## Usage
|
|
|
|
### Endpoints
|
|
|
|
- `POST /tts`: The main endpoint to synthesize text to speech.
|
|
- `GET /health`: Check the health of the API and the status of the loaded engines.
|
|
- `GET /engines`: List the currently active engines.
|
|
- `GET /models`: List the available models for each active engine.
|
|
_ `GET /speakers`: List the available speakers for a given engine and model.
|
|
|
|
### Makefile Commands
|
|
|
|
The project includes a `Makefile` with several commands to simplify development and management:
|
|
|
|
- `make dev-up`: Build the image from local source and start the application. Recommended for development.
|
|
- `make up`: Start the application using the image from the container registry (pulls if not present).
|
|
- `make down`: Stop the application container(s).
|
|
- `make logs`: View the application logs.
|
|
- `make health-check`: Run a sanity check to ensure the deployed container is healthy and all engines are "ok".
|
|
- `make pull`: Pull the latest image from the container registry.
|
|
- `make push`: Build, tag, and push the image to the container registry.
|
|
- `make test`: Run the `pytest` test suite.
|
|
- `make help`: Display a list of all available commands.
|
|
|
|
## Configuration
|
|
|
|
The application is configured through the `.env` file in the root of the project.
|
|
|
|
- `ACTIVE_ENGINES`: A comma-separated list of strings specifying which TTS engines to activate. Available engines are defined in `app/main.py`.
|
|
- `HOST`: The host address for the server (defaults to `0.0.0.0`).
|
|
- `PORT`: The internal port for the server (defaults to `8000`).
|
|
- `IMAGE_NAME`: The name of the Docker image to build (defaults to `audioenginehub`).
|