docs: runbook deploiement kokoro-tts, MAJ ports-registry et _INDEX (ticket infra-2026-09-029)

This commit is contained in:
Gemini
2026-09-09 01:25:23 +01:00
parent 0fcbf112fc
commit cdee34c3d2
3 changed files with 145 additions and 0 deletions
+143
View File
@@ -0,0 +1,143 @@
# Deploiement serveur TTS local Kokoro-82M sur NAS
**Instance auteur** : gemini-ops
**Date** : 2026-09-09
**Tags** : kokoro, tts, speech, nas, docker, audio, hermes, dsh
**Statut** : valide
---
## Probleme
Fournir aux agents locaux (Hermes NAS, Hermes VPS, DSH) un service Text-to-Speech (TTS) haute qualite, local, souverain et gratuit pour la synthese vocale, notamment en francais, sans dependre d'API externes payantes ni surcharger les ressources CPU/RAM du NAS.
---
## Contexte et contraintes
- **NAS Synology DS920+** : DSM 7.2 (Linux 4.4 x86_64), execution sur CPU uniquement (pas de GPU dedie).
- **Empreinte memoire** : Apres liberation de la RAM via la bascule nyora-doc/convert vers VPS (Ticket infra-2026-09-028), le NAS dispose de plus de 9 Go de RAM libre. Kokoro necessite ~1.5 Go de RAM.
- **Port choisi** : Port hôte `3088` mappe vers `8880` dans le conteneur.
- **Image Docker** : `hwdsl2/kokoro-server:latest` (OpenAI-compatible `/v1/audio/speech`).
- **Reseau Docker** : Attache au reseau Docker `n8n` pour acces direct des conteneurs locaux (Hermes TT, Hermes Nyora, Hermes Perso) sans passer par le bridge NAT externe.
- **Routage Bifrost proscrit** : Bifrost (v1.6.11) ne gere pas le format `/v1/audio/speech` (reserve aux modeles LLM text/chat/embedding). Les clients doivent contacter Kokoro directement.
- **Acces distant VPS** : Via Tailscale (IP NAS `100.86.197.88:3088`).
---
## Ce qui NE fonctionne PAS
| Tentative | Erreur obtenue | Raison de l'echec |
|-----------|----------------|-------------------|
| Passer par Bifrost (`http://bifrost:8080/v1/audio/speech`) | HTTP 404 / route non geree | Bifrost v1.6.11 ne supporte pas l'API speech OpenAI, uniquement chat/completions/embeddings. |
| Demarrage direct sans persistance de cache | Re-telechargement de 320 Mo a chaque demarrage | Sans montage sur `/var/lib/kokoro`, le modele PyTorch est telecharge dans la couche overlay ephemere. |
| Timeout premier demarrage | `Error: Kokoro TTS server did not become ready within 300 seconds` | Sur CPU NAS x86_64, le tout premier demarrage (chargement des libs lourdes `transformers`/`torch` + download HuggingFace) peut atteindre la limite des 300s de `run.sh`. Une fois le cache present dans le volume `/var/lib/kokoro`, le demarrage se fait en ~90s. |
---
## Solution validee
### 1. Structure des dossiers et compose
Dossier : `/volume1/docker/kokoro-tts`
Volume de persistance : `/volume1/docker/kokoro-tts/data`
Fichier `/volume1/docker/kokoro-tts/docker-compose.yml` :
```yaml
services:
kokoro-tts:
image: hwdsl2/kokoro-server:latest
container_name: kokoro-tts
restart: unless-stopped
ports:
- "3088:8880"
environment:
- KOKORO_VOICE=ff_siwis
- KOKORO_LANG_CODE=f
- KOKORO_SPEED=1.0
- KOKORO_PORT=8880
- KOKORO_API_KEY=kokoro-tts-89b3f421e90d7c2a
- KOKORO_LOG_LEVEL=INFO
volumes:
- /volume1/docker/kokoro-tts/data:/var/lib/kokoro
networks:
- n8n
healthcheck:
test: ["CMD-SHELL", "curl -sf http://127.0.0.1:8880/health >/dev/null 2>&1 || exit 1"]
interval: 30s
timeout: 10s
retries: 3
start_period: 120s
networks:
n8n:
external: true
```
### 2. URLs d'acces pour les consommateurs
- **Agents internes NAS (reseau `n8n`)** :
`http://kokoro-tts:8880/v1/audio/speech`
Healthcheck : `http://kokoro-tts:8880/health`
- **Clients LAN / localhost NAS** :
`http://192.168.100.33:3088/v1/audio/speech` (ou `http://127.0.0.1:3088`)
- **Agents VPS (Hermes Nabil, DSH) via Tailscale** :
`http://100.86.197.88:3088/v1/audio/speech`
Healthcheck : `http://100.86.197.88:3088/health`
**Authentification requise** :
Header : `Authorization: Bearer kokoro-tts-89b3f421e90d7c2a`
Content-Type : `application/json`
---
## Verification
### 1. Healthcheck
```bash
curl -s http://127.0.0.1:3088/health
# Resultat : {"status":"ok","engine":"kokoro"}
```
### 2. Controle de l'authentification
```bash
# Sans token -> HTTP 401
curl -s -i http://127.0.0.1:3088/v1/voices | head -n 1
# HTTP/1.1 401 Unauthorized
# Avec token -> HTTP 200
curl -s -i http://127.0.0.1:3088/v1/voices -H "Authorization: Bearer kokoro-tts-89b3f421e90d7c2a" | head -n 1
# HTTP/1.1 200 OK
```
### 3. Generation reelle depuis le NAS
```bash
curl -s -w '%{http_code}\n' -X POST http://127.0.0.1:3088/v1/audio/speech \
-H "Authorization: Bearer kokoro-tts-89b3f421e90d7c2a" \
-H "Content-Type: application/json" \
-d '{"model": "tts-1", "input": "Bonjour, test de synthese vocale sur le NAS.", "voice": "ff_siwis"}' \
-o /tmp/test_nas.mp3
# Resultat : HTTP 200, fichier MP3 de 23 Ko genere
```
### 4. Generation reelle depuis le VPS (via Tailscale)
```bash
ssh vps-gemini "curl -s -w '%{http_code}\n' -X POST http://100.86.197.88:3088/v1/audio/speech \
-H 'Authorization: Bearer kokoro-tts-89b3f421e90d7c2a' \
-H 'Content-Type: application/json' \
-d '{\"model\": \"tts-1\", \"input\": \"Test de synthese vocale depuis le VPS vers le NAS avec Kokoro.\", \"voice\": \"ff_siwis\"}' \
-o /tmp/test_kokoro_vps.mp3 && ls -lh /tmp/test_kokoro_vps.mp3 && file /tmp/test_kokoro_vps.mp3"
# Resultat : HTTP 200, fichier MP3 22 Ko (ID3 version 2.4.0, MPEG ADTS, layer III, 32 kbps, 24 kHz)
```
---
## References
- Ticket Baserow : infra-2026-09-029 (Id 30)
- Image GitHub : https://github.com/hwdsl2/docker-kokoro
- Modele HuggingFace : https://huggingface.co/hexgrad/Kokoro-82M