Add XTTS support to Makefile and create CLAUDE.md guide
Makefile improvements: - Add voices-xtts target to download speaker samples - Add test-xtts target for testing HD model - Split voices into voices-piper and voices-xtts - Update help text with all new targets speech.py: - Fix threading import scope issue for XTTS - Remove redundant 'import threading' inside Piper block docs/CLAUDE.md: - Complete guide for Claude Code contributors - Makefile-first development philosophy - Never create dirs manually, always use Makefile - Documentation requirements and testing philosophy - Common mistakes to avoid - Raccoon mission values and principles This ensures consistent, repeatable deployments and makes it easy to add new TTS engines following the same pattern. 🦝 Generated with [Claude Code](https://claude.com/claude-code) Co-Authored-By: Claude <noreply@anthropic.com>
This commit is contained in:
parent
0ff5f0f09a
commit
2c6c1ad577
3 changed files with 237 additions and 6 deletions
30
Makefile
30
Makefile
|
|
@ -14,7 +14,7 @@ REMOTE_USER ?= $(USER)
|
|||
REMOTE_PATH ?= ~/uncloseai-speech
|
||||
CONTAINER_NAME ?= uncloseai-speech-server-1
|
||||
|
||||
.PHONY: help deploy sync restart logs test clean stop start voices
|
||||
.PHONY: help deploy sync restart logs test clean stop start voices voices-piper voices-xtts
|
||||
|
||||
help:
|
||||
@echo "🦝 Raccoon TTS Mission - Development Commands"
|
||||
|
|
@ -26,8 +26,11 @@ help:
|
|||
@echo ""
|
||||
@echo "Development:"
|
||||
@echo " make logs - Tail container logs"
|
||||
@echo " make test - Test TTS endpoint"
|
||||
@echo " make voices - Download Piper voices properly"
|
||||
@echo " make test - Test TTS endpoint (Piper)"
|
||||
@echo " make test-xtts - Test XTTS HD endpoint"
|
||||
@echo " make voices - Download all voices (Piper + XTTS)"
|
||||
@echo " make voices-piper - Download Piper voices only"
|
||||
@echo " make voices-xtts - Download XTTS voices and samples"
|
||||
@echo ""
|
||||
@echo "Container:"
|
||||
@echo " make start - Start Docker container"
|
||||
|
|
@ -74,7 +77,10 @@ test:
|
|||
@echo "✅ Test complete! Playing audio..."
|
||||
@firefox /tmp/raccoon_test.mp3 || mpv /tmp/raccoon_test.mp3 || echo "Install firefox or mpv to play audio"
|
||||
|
||||
voices:
|
||||
voices: voices-piper voices-xtts
|
||||
@echo "✅ All voices downloaded!"
|
||||
|
||||
voices-piper:
|
||||
@echo "🎤 Downloading Piper voices with correct directory structure..."
|
||||
ssh $(REMOTE_USER)@$(REMOTE_HOST) "docker exec $(CONTAINER_NAME) bash -c '\
|
||||
mkdir -p /app/voices/en/en_US/libritts_r/medium && \
|
||||
|
|
@ -89,4 +95,18 @@ voices:
|
|||
ssh $(REMOTE_USER)@$(REMOTE_HOST) "docker exec $(CONTAINER_NAME) bash -c '\
|
||||
sed -i \"s|model: voices/en_US-libritts_r-medium.onnx|model: /app/voices/en/en_US/libritts_r/medium/en_US-libritts_r-medium.onnx|g\" /app/config/voice_to_speaker.yaml && \
|
||||
sed -i \"s|model: voices/en_GB-northern_english_male-medium.onnx|model: /app/voices/en/en_GB/northern_english_male/medium/en_GB-northern_english_male-medium.onnx|g\" /app/config/voice_to_speaker.yaml'"
|
||||
@echo "✅ Voices installed with absolute paths!"
|
||||
@echo "✅ Piper voices installed with absolute paths!"
|
||||
|
||||
voices-xtts:
|
||||
@echo "🎤 Downloading XTTS speaker samples..."
|
||||
ssh $(REMOTE_USER)@$(REMOTE_HOST) "docker exec $(CONTAINER_NAME) bash -c 'cd /app && ./scripts/download_samples.sh'"
|
||||
@echo "✅ XTTS speaker samples downloaded!"
|
||||
|
||||
test-xtts:
|
||||
@echo "🧪 Testing XTTS HD endpoint (this may take 1-2 minutes on first run)..."
|
||||
curl -X POST http://$(REMOTE_HOST):8000/v1/audio/speech \
|
||||
-H "Content-Type: application/json" \
|
||||
-d '{"model":"tts-1-hd","voice":"alloy","input":"Testing XTTS high definition"}' \
|
||||
-o /tmp/xtts_test.mp3
|
||||
@echo "✅ Test complete! Playing audio..."
|
||||
@firefox /tmp/xtts_test.mp3 || mpv /tmp/xtts_test.mp3 || echo "Install firefox or mpv to play audio"
|
||||
|
|
|
|||
212
docs/CLAUDE.md
Normal file
212
docs/CLAUDE.md
Normal file
|
|
@ -0,0 +1,212 @@
|
|||
# Instructions for Claude Code
|
||||
|
||||
**Project:** UncloseAI Speech - Raccoon Mission TTS System
|
||||
**License:** AGPL v3 (must provide source code to network service users)
|
||||
|
||||
## Core Principles
|
||||
|
||||
### 1. Makefile-First Development
|
||||
|
||||
**ALWAYS prefer Makefile targets over manual commands.**
|
||||
|
||||
- ✅ DO: `make deploy`, `make voices`, `make test`
|
||||
- ❌ DON'T: Manual ssh commands, docker commands, curl commands
|
||||
|
||||
**When adding new functionality:**
|
||||
1. Add it to the Makefile first
|
||||
2. Document it in `make help`
|
||||
3. Test it works from scratch
|
||||
4. Only then modify other files if needed
|
||||
|
||||
**Makefile is the source of truth** for all deployment and development tasks.
|
||||
|
||||
### 2. Work Locally, Deploy Remotely
|
||||
|
||||
- **Local development:** `/home/fox/git/openedai-speech/`
|
||||
- **Remote server:** Configured in `vars.sh` (gitignored)
|
||||
- **Never create remote directories manually** - let Makefile handle it
|
||||
- **Always test from scratch** - `make clean` then `make deploy`
|
||||
|
||||
### 3. Configuration Management
|
||||
|
||||
- `vars.sh` - Deployment secrets (gitignored, never commit)
|
||||
- `vars.sh.example` - Template for users (commit this)
|
||||
- `sample.env` - Default environment (commit this)
|
||||
- `speech.env` - Runtime environment (created automatically by Makefile)
|
||||
|
||||
**Never view or log secrets** - source them and use them.
|
||||
|
||||
### 4. Documentation Requirements
|
||||
|
||||
When adding features, update ALL relevant docs:
|
||||
- `Makefile` help text
|
||||
- `docs/MODELS.md` for new TTS engines
|
||||
- `docs/MIRRORS.md` for binary downloads
|
||||
- `docs/AUDIT.md` for file changes
|
||||
- This file (`docs/CLAUDE.md`) for new patterns
|
||||
|
||||
## Common Tasks
|
||||
|
||||
### Full Deployment from Scratch
|
||||
|
||||
```bash
|
||||
# 1. Clean everything
|
||||
make clean
|
||||
|
||||
# 2. Deploy (syncs files, creates env, builds container)
|
||||
make deploy
|
||||
|
||||
# 3. Download voices (Piper + XTTS samples)
|
||||
make voices
|
||||
|
||||
# 4. Test
|
||||
make test
|
||||
make test-xtts
|
||||
```
|
||||
|
||||
### Adding a New TTS Engine
|
||||
|
||||
1. Document it in `docs/MODELS.md` first
|
||||
2. Add download target to Makefile (e.g., `voices-silero`)
|
||||
3. Implement engine wrapper in `speech.py` or `src/engines/`
|
||||
4. Add test target (e.g., `test-silero`)
|
||||
5. Update `make voices` to include it
|
||||
6. Test full cycle: `make clean && make deploy && make voices`
|
||||
|
||||
### Debugging Issues
|
||||
|
||||
```bash
|
||||
make logs # Tail live logs
|
||||
make logs | grep ERROR # Filter errors
|
||||
```
|
||||
|
||||
Never use raw docker/ssh commands - extend Makefile if needed.
|
||||
|
||||
## File Organization
|
||||
|
||||
### Scripts vs Docs
|
||||
|
||||
- `scripts/` - Executable utilities (add_voice.py, download_samples.sh, etc.)
|
||||
- `docs/` - Documentation ONLY (no executable code)
|
||||
- Dockerfiles, startup.sh - Root level (build artifacts)
|
||||
- Makefile - Root level (primary interface)
|
||||
|
||||
**Never put executable scripts in docs/ directory.**
|
||||
|
||||
### Current Structure (as of 2025-11-09)
|
||||
|
||||
```
|
||||
uncloseai-speech/
|
||||
├── Makefile # PRIMARY INTERFACE - always update first
|
||||
├── vars.sh # Secrets (gitignored)
|
||||
├── vars.sh.example # Template
|
||||
├── speech.py # Main server (will refactor to src/)
|
||||
├── openedai.py # API models
|
||||
├── voice_to_speaker.default.yaml # Voice config
|
||||
├── docs/
|
||||
│ ├── CLAUDE.md # This file
|
||||
│ ├── AUDIT.md # Repository audit
|
||||
│ ├── MODELS.md # TTS engines
|
||||
│ └── MIRRORS.md # Binary mirror strategy
|
||||
├── scripts/
|
||||
│ ├── add_voice.py
|
||||
│ ├── say.py
|
||||
│ ├── test_voices.sh
|
||||
│ └── download_samples.sh
|
||||
├── Dockerfile
|
||||
├── docker-compose.yml
|
||||
└── startup.sh
|
||||
```
|
||||
|
||||
## TTS Engine Status
|
||||
|
||||
### Working
|
||||
- ✅ Piper TTS (tts-1) - Fast, 100+ voices, absolute paths working
|
||||
- ⚠️ XTTS v2 (tts-1-hd) - High quality, needs speaker samples
|
||||
|
||||
### High Priority Integration
|
||||
- 🎯 Silero TTS - Active project, fast, good quality
|
||||
- 🎯 StyleTTS2 - Best quality available
|
||||
- 🎯 Fish Speech - Modern, multilingual
|
||||
|
||||
See `docs/MODELS.md` for complete roadmap.
|
||||
|
||||
## Deployment Workflow
|
||||
|
||||
```
|
||||
Local:
|
||||
/home/fox/git/openedai-speech/
|
||||
|
||||
↓ make deploy (rsync)
|
||||
|
||||
Remote (ai.foxhop.net):
|
||||
~/uncloseai-speech/
|
||||
|
||||
↓ docker compose up --build
|
||||
|
||||
Container:
|
||||
/app/
|
||||
├── speech.py
|
||||
├── voices/
|
||||
│ └── en/en_US/libritts_r/medium/*.onnx
|
||||
└── config/
|
||||
└── voice_to_speaker.yaml
|
||||
```
|
||||
|
||||
## Testing Philosophy
|
||||
|
||||
**Always test the full stack:**
|
||||
1. Clean state (`make clean`)
|
||||
2. Fresh deploy (`make deploy`)
|
||||
3. Voice download (`make voices`)
|
||||
4. API test (`make test`, `make test-xtts`)
|
||||
|
||||
**Never assume** - if you changed something, test from scratch.
|
||||
|
||||
## Raccoon Mission Values
|
||||
|
||||
1. **Resilience** - Assume upstream dies, plan mirrors
|
||||
2. **Simplicity** - Makefile > manual commands
|
||||
3. **Documentation** - Write docs before code
|
||||
4. **Liberation** - Keep TTS libre (AGPL v3)
|
||||
5. **Unification** - All TTS engines, one API
|
||||
|
||||
## Common Mistakes to Avoid
|
||||
|
||||
❌ DON'T create directories with raw ssh
|
||||
✅ DO add Makefile target for deployment
|
||||
|
||||
❌ DON'T assume container has changes after rsync
|
||||
✅ DO rebuild with `make deploy` (runs docker compose up --build)
|
||||
|
||||
❌ DON'T put scripts in docs/
|
||||
✅ DO put scripts in scripts/, reference from docs
|
||||
|
||||
❌ DON'T hardcode paths/hosts
|
||||
✅ DO use vars.sh variables
|
||||
|
||||
❌ DON'T forget to test from scratch
|
||||
✅ DO run `make clean && make deploy && make voices`
|
||||
|
||||
## When Things Break
|
||||
|
||||
1. Check `make logs` for errors
|
||||
2. Verify Makefile was updated
|
||||
3. Test from clean state
|
||||
4. Check if container was rebuilt (`make deploy` does this)
|
||||
5. Verify voices downloaded (`ls` in container via `make logs` approach)
|
||||
|
||||
## Future Refactoring (Planned)
|
||||
|
||||
- Move `speech.py`, `openedai.py`, `audio_reader.py` → `src/`
|
||||
- Create engine abstraction layer in `src/engines/`
|
||||
- Unified voice config with engine selection
|
||||
- Binary mirror implementation (MinIO on ai.foxhop.net)
|
||||
|
||||
See `docs/AUDIT.md` for detailed refactoring plan.
|
||||
|
||||
---
|
||||
|
||||
**Remember:** Makefile first, documentation second, code third. Test from scratch every time.
|
||||
|
||||
🦝 **Raccoon Mission:** Keep TTS libre, rescue abandoned models, unify all engines.
|
||||
|
|
@ -243,7 +243,6 @@ async def generate_speech(request: GenerateSpeechRequest):
|
|||
|
||||
# Log any stderr output from Piper
|
||||
if tts_proc.stderr:
|
||||
import threading
|
||||
def log_stderr():
|
||||
stderr_output = tts_proc.stderr.read().decode('utf-8', errors='replace')
|
||||
if stderr_output.strip():
|
||||
|
|
|
|||
Loading…
Add table
Add a link
Reference in a new issue