uncloseai-speech/docs/AUDIT.md
Russell Ballestrini 7559e56d0c Standardize project naming to uncloseai-speech across all files
- Add CHANGELOG.md with full version history (moved from README)
- Update all documentation to use lowercase 'uncloseai-speech' project name
- Update organization references to lowercase 'uncloseai' (not 'UncloseAI')
- Add Brand Identity section to docs/CLAUDE.md with naming guidelines
- Update speech.py argparse description to match branding
- Update README.md headers and sections with consistent naming
- Update all model documentation with consistent branding

Files updated:
- CHANGELOG.md (new file)
- README.md (changelog reference, server options, multilingual section)
- speech.py (--workers argument, branding in argparse)
- Makefile (header comment)
- docs/CLAUDE.md (Brand Identity section)
- docs/MODELS.md
- docs/MIRRORS.md
- docs/AUDIT.md
- docs/models/coqui-tts.md
- docs/research/tts-models-overview.md

Branding standard:
- Project: uncloseai-speech (lowercase, hyphenated)
- Organization: uncloseai (lowercase, one word)

🦝 Generated with Claude Code
2025-11-10 05:23:34 -05:00

343 lines
11 KiB
Markdown

# uncloseai-speech Repository Audit
**Date:** 2025-11-09
**Mission:** Raccoon TTS - Build a unified, resilient TTS system from abandoned open source projects
## Current Repository Structure
### Core Service Files (KEEP - Essential)
- `speech.py` (18K) - Main FastAPI server, handles OpenAI-compatible TTS API
- `openedai.py` (5.8K) - OpenAI API models and data structures
- `audio_reader.py` (3.7K) - Audio processing utilities
- `voice_to_speaker.default.yaml` - Default voice configuration mapping
- `pre_process_map.default.yaml` - Text preprocessing rules
### Docker & Deployment (KEEP - Essential)
- `Dockerfile` - Main docker build (Piper + XTTS with NVIDIA GPU)
- `Dockerfile.min` - Minimal build (Piper only, no GPU)
- `docker-compose.yml` - Standard deployment
- `docker-compose.min.yml` - Minimal deployment
- `docker-compose.rocm.yml` - AMD ROCm GPU support
- `Makefile` - Raccoon deployment system (recently added)
- `vars.sh.example` - Deployment configuration template
### Requirements Files (KEEP - Essential)
- `requirements.txt` - Main Python dependencies
- `requirements-min.txt` - Minimal dependencies (Piper only)
- `requirements-rocm.txt` - AMD ROCm dependencies
### Startup Scripts (KEEP - Essential)
- `startup.sh` - Main startup with voice downloads
- `startup.min.sh` - Minimal startup
- `download_voices_tts-1.sh` - Download Piper voices
- `download_voices_tts-1-hd.sh` - Download XTTS HD voices
### CLI Utilities (KEEP - Useful)
- `say.py` (3.5K) - Command-line TTS client using OpenAI API
- `add_voice.py` (2.8K) - Tool to add custom voices to config
- `test_voices.sh` (2.1K) - Test script for all voice models
### Windows Batch Files (EVALUATE - Questionable)
- `startup.bat` (185B) - Windows startup
- `download_samples.bat` (230B) - Download sample WAVs
- `download_voices_tts-1.bat` (323B) - Download voices on Windows
- `download_voices_tts-1-hd.bat` (199B) - Download HD voices on Windows
**Assessment:** These are only useful if we support Windows deployment. Most raccoon missions run on Linux servers. Could move to `docs/windows/` or remove entirely.
### Sample Scripts (EVALUATE - Non-Essential)
- `download_samples.sh` (205B) - Downloads OpenAI sample WAV files
- `boot_docker_compose_ubuntu.sh` (623B) - One-time NVIDIA Docker setup for Ubuntu
**Assessment:** `download_samples.sh` is only for comparing against OpenAI's original voices - move to `docs/tools/`. The boot script is one-time setup - move to `docs/setup/`.
### Documentation (KEEP - Update)
- `README.md` (16.6K) - Main documentation
- `LICENSE` - **AGPL v3** (GNU Affero General Public License v3)
- **Important:** Network services must provide source code to users
- All modifications must be shared under AGPL v3
- Perfect for raccoon mission - keeps TTS libre forever!
- `sample.env` - Environment variable examples
### CI/CD (KEEP)
- `.github/workflows/build-docker.yml` - GitHub Actions for building Docker images
### Configuration (KEEP)
- `.gitignore` - Git ignore patterns
- `config/config_files_will_go_here.txt` - Placeholder
## Files to Remove or Relocate
### REMOVE (Windows-specific, low value):
- `startup.bat`
- `download_samples.bat`
- `download_voices_tts-1.bat`
- `download_voices_tts-1-hd.bat`
### MOVE to `docs/tools/`:
- `download_samples.sh` (useful for voice comparison)
- `test_voices.sh` (useful for testing)
### MOVE to `docs/setup/`:
- `boot_docker_compose_ubuntu.sh` (one-time setup instructions)
## Abandoned TTS Models to Integrate
Based on the raccoon mission to rescue abandoned TTS projects:
### 1. **Piper TTS** (Currently Integrated)
- **Status:** ✅ Working with absolute paths
- **Original:** rhasspy/piper (abandoned)
- **Fork:** OHF-Voice/piper1-gpl v1.3.0 (no PyPI package)
- **Current:** Using PyPI piper-tts v1.2.0
- **Models:** ~100+ voices from HuggingFace rhasspy/piper-voices
- **Action:** Consider creating our own PyPI package from OHF-Voice fork
### 2. **Coqui XTTS** (Currently Integrated)
- **Status:** ✅ Integrated as tts-1-hd
- **Original:** coqui-ai/TTS (company shut down, repo archived)
- **Package:** coqui-tts (still on PyPI)
- **Models:** XTTS v2 from HuggingFace
- **Action:** Mirror XTTS v2 models, explore community forks
### 3. **Mozilla TTS** (Predecessor to Coqui)
- **Status:** ❌ Not integrated, abandoned
- **Original:** mozilla/TTS (archived, became Coqui)
- **Action:** Evaluate if any unique models worth preserving
### 4. **Silero TTS**
- **Status:** ❌ Not integrated
- **Repo:** snakers4/silero-models (still active!)
- **Models:** Fast, multilingual, small footprint
- **Languages:** Russian, English, German, Spanish, French
- **Action:** HIGH PRIORITY - Add Silero support
### 5. **Bark (Suno AI)**
- **Status:** ❌ Not integrated
- **Repo:** suno-ai/bark (archived, company pivoted)
- **Features:** Music, sound effects, non-speech audio
- **Action:** Evaluate for unique use cases
### 6. **Tortoise TTS**
- **Status:** ❌ Not integrated
- **Repo:** neonbjb/tortoise-tts (low activity)
- **Features:** High quality, very slow generation
- **Action:** Low priority - too slow for production
### 7. **StyleTTS2**
- **Status:** ❌ Not integrated
- **Repo:** yl4579/StyleTTS2 (active research project)
- **Features:** State-of-art quality, prosody
- **Action:** HIGH PRIORITY - Best quality available
### 8. **MetaVoice**
- **Status:** ❌ Not integrated
- **Repo:** metavoiceio/metavoice-src (partially abandoned)
- **Features:** Long-form TTS, emotional control
- **Action:** Evaluate community forks
### 9. **Fish Speech**
- **Status:** ❌ Not integrated
- **Repo:** fishaudio/fish-speech (active)
- **Features:** Fast, multilingual
- **Action:** MEDIUM PRIORITY - Active project
### 10. **Kokoro TTS**
- **Status:** ❌ Not integrated
- **Repo:** hexgrad/kokoro (new, active)
- **Features:** Fast, small, quality
- **Action:** MEDIUM PRIORITY - Promising new project
## Binary Mirror Strategy
### Problem
- HuggingFace models can disappear
- Upstream repositories get archived/deleted
- Need resilience for production systems
### Solution: Multi-Tier Mirror System
#### Tier 1: Primary Sources (Current)
- HuggingFace Hub
- PyPI packages
- GitHub releases
#### Tier 2: UncloseAI Mirror (Proposed)
- Self-hosted model repository
- Options:
- MinIO/S3-compatible storage
- LFS-backed Git repository
- Dedicated file server with rsync
#### Tier 3: Fallback (Proposed)
- Archive.org uploads for critical models
- IPFS pinning for decentralization
- Torrent magnets for popular models
### Priority Models to Mirror
1. **Piper voices** - ~2GB from rhasspy/piper-voices
2. **XTTS v2** - ~1.8GB model + config
3. **Silero models** - ~100MB per language
4. **StyleTTS2 weights** - TBD size
### Implementation Plan
- [ ] Set up MinIO on ai.foxhop.net
- [ ] Create download script with fallback URLs
- [ ] Document mirror usage in docs
- [ ] Automated sync from upstream sources
## Recommended Refactoring
### 1. Directory Structure
```
uncloseai-speech/
├── docs/
│ ├── AUDIT.md (this file)
│ ├── MODELS.md (TTS engine documentation)
│ ├── MIRRORS.md (Binary mirror setup)
│ ├── API.md (API documentation)
│ └── setup/
│ └── nvidia-docker-ubuntu.md
├── scripts/
│ ├── add_voice.py (add custom voices to config)
│ ├── say.py (CLI TTS client)
│ ├── test_voices.sh (test all voices)
│ ├── download_samples.sh (download OpenAI sample WAVs)
│ ├── download_voices_tts-1.sh (Piper voice downloader)
│ └── download_voices_tts-1-hd.sh (XTTS voice downloader)
├── src/ (future refactor)
│ ├── engines/
│ │ ├── piper.py
│ │ ├── xtts.py
│ │ ├── silero.py (NEW)
│ │ └── base.py (NEW)
│ ├── api/
│ │ ├── openedai.py
│ │ └── server.py
│ └── utils/
│ └── audio_reader.py
├── config/
│ ├── voice_to_speaker.default.yaml
│ └── pre_process_map.default.yaml
├── Dockerfile
├── Dockerfile.min
├── docker-compose.yml
├── docker-compose.min.yml
├── docker-compose.rocm.yml
├── startup.sh
├── startup.min.sh
├── Makefile
├── speech.py (current monolith - will refactor to src/)
├── openedai.py (will move to src/api/)
├── audio_reader.py (will move to src/utils/)
├── requirements.txt
├── requirements-min.txt
├── requirements-rocm.txt
└── README.md
```
### 2. Engine Abstraction
Create a base engine class that all TTS engines implement:
```python
# src/engines/base.py
from abc import ABC, abstractmethod
from typing import BinaryIO, Optional
class TTSEngine(ABC):
@abstractmethod
def synthesize(self, text: str, voice: str,
speed: float = 1.0,
language: Optional[str] = None) -> BinaryIO:
"""Synthesize speech from text"""
pass
@abstractmethod
def list_voices(self) -> list[str]:
"""Return available voices for this engine"""
pass
@abstractmethod
def download_models(self, mirror_urls: list[str] = None):
"""Download required models with fallback mirrors"""
pass
```
### 3. Configuration Schema
Extend voice config to support multiple engines:
```yaml
# Future voice_to_speaker.yaml format
tts-1:
alloy:
engine: piper # or xtts, silero, styletts2
model: /app/voices/en/en_US/libritts_r/medium/en_US-libritts_r-medium.onnx
speaker: 79
nova:
engine: silero
model: v3_1_ru
speaker: aidar
tts-1-hd:
alloy:
engine: xtts
model: xtts
speaker: /app/voices/alloy.wav
custom-voice:
engine: styletts2
model: /app/models/styletts2-base
speaker: /app/voices/custom.wav
language: en
```
## Next Steps
### Immediate (This Session)
1. ✅ Create docs directory
2. ✅ Write audit document
3. [ ] Remove Windows .bat files
4. [ ] Relocate tools to docs/tools/
5. [ ] Create MODELS.md documenting each TTS engine
### Short Term (Next Few Sessions)
1. [ ] Implement engine abstraction layer
2. [ ] Add Silero TTS support
3. [ ] Set up model mirror on ai.foxhop.net
4. [ ] Test all existing functionality still works
### Medium Term
1. [ ] Add StyleTTS2 support
2. [ ] Add Fish Speech support
3. [ ] Implement fallback mirror system
4. [ ] Create automated model sync
### Long Term
1. [ ] Add Bark for sound effects
2. [ ] Implement voice cloning
3. [ ] Multi-language auto-detection
4. [ ] Performance benchmarking dashboard
## Raccoon Mission Status
🦝 **Mission:** Rescue abandoned TTS models and build a resilient unified system
**Current Status:**
- ✅ Piper TTS working with absolute paths
- ✅ XTTS integrated
- ✅ Deployment system (Makefile + vars.sh)
- ✅ Renamed to uncloseai-speech
- 📝 Repository audit complete
- 🔄 Documentation in progress
**Key Risks:**
- Upstream model URLs can break (need mirrors)
- PyPI packages could disappear (need vendoring strategy)
- HuggingFace rate limits (need caching)
**Mitigation:**
- Build multi-tier mirror system
- Document all model sources
- Create offline installation option