uncloseai-speech/docs/MIRRORS.md
Russell Ballestrini 7559e56d0c Standardize project naming to uncloseai-speech across all files
- Add CHANGELOG.md with full version history (moved from README)
- Update all documentation to use lowercase 'uncloseai-speech' project name
- Update organization references to lowercase 'uncloseai' (not 'UncloseAI')
- Add Brand Identity section to docs/CLAUDE.md with naming guidelines
- Update speech.py argparse description to match branding
- Update README.md headers and sections with consistent naming
- Update all model documentation with consistent branding

Files updated:
- CHANGELOG.md (new file)
- README.md (changelog reference, server options, multilingual section)
- speech.py (--workers argument, branding in argparse)
- Makefile (header comment)
- docs/CLAUDE.md (Brand Identity section)
- docs/MODELS.md
- docs/MIRRORS.md
- docs/AUDIT.md
- docs/models/coqui-tts.md
- docs/research/tts-models-overview.md

Branding standard:
- Project: uncloseai-speech (lowercase, hyphenated)
- Organization: uncloseai (lowercase, one word)

🦝 Generated with Claude Code
2025-11-10 05:23:34 -05:00

10 KiB

Binary Mirror Strategy

Purpose: Ensure uncloseai-speech keeps working even if upstream model sources disappear

The Problem

Upstream Fragility

  • HuggingFace repos can be deleted
  • GitHub releases disappear when repos are archived
  • PyPI packages can be yanked
  • Companies shut down and take their models offline
  • Rate limiting breaks automated deployments

Real Examples

  • Coqui AI: Company shut down 2024, repo archived
  • Rhasspy Piper: Original project abandoned
  • Suno Bark: Company pivoted, model archived
  • ⚠️ XTTS-v2: Depends on archived Coqui repo

Multi-Tier Mirror Architecture

Tier 1: Upstream Sources (Primary)

Always try upstream first - they're fastest and most up-to-date.

Sources:

  • HuggingFace Hub (huggingface.co)
  • PyPI (pypi.org)
  • GitHub Releases
  • Official project websites

Advantages:

  • Latest versions
  • Fast CDN delivery
  • Community validation

Disadvantages:

  • Can disappear
  • Rate limits
  • Requires internet

Tier 2: UncloseAI Mirror (Secondary)

Self-hosted mirror under our control.

Location: ai.foxhop.net Storage: MinIO S3-compatible object storage Capacity: 100GB allocated for models

Setup:

# Install MinIO on ai.foxhop.net
docker run -d \
  -p 9000:9000 \
  -p 9001:9001 \
  --name minio \
  -v /data/minio:/data \
  -e "MINIO_ROOT_USER=admin" \
  -e "MINIO_ROOT_PASSWORD=<secure_password>" \
  minio/minio server /data --console-address ":9001"

# Create bucket for models
mc alias set unclose http://ai.foxhop.net:9000 admin <password>
mc mb unclose/tts-models
mc policy set download unclose/tts-models

Directory Structure:

tts-models/
├── piper/
│   ├── v1.0.0/
│   │   ├── en/
│   │   │   ├── en_US/
│   │   │   │   └── libritts_r/
│   │   │   │       └── medium/
│   │   │   │           ├── en_US-libritts_r-medium.onnx
│   │   │   │           └── en_US-libritts_r-medium.onnx.json
│   │   └── voices.json (metadata)
├── xtts/
│   └── v2.0.3/
│       ├── model.pth
│       ├── config.json
│       ├── vocab.json
│       └── README.md
├── silero/
│   └── v4/
│       ├── en_v4.pt
│       ├── ru_v4.pt
│       └── models.json
├── styletts2/
│   └── libritts/
│       ├── checkpoint.pt
│       └── config.yml
└── metadata.json (master index)

Sync Script:

#!/bin/bash
# scripts/sync_models_to_mirror.sh
# Sync upstream models to UncloseAI mirror

set -euo pipefail

MIRROR_URL="http://ai.foxhop.net:9000/tts-models"
TEMP_DIR="/tmp/model_sync"

# Sync Piper voices
sync_piper() {
    echo "Syncing Piper models..."
    for voice in en_US-libritts_r-medium en_GB-northern_english_male-medium; do
        wget -P "$TEMP_DIR/piper/" \
            "https://huggingface.co/rhasspy/piper-voices/resolve/v1.0.0/en/en_US/libritts_r/medium/${voice}.onnx" \
            "https://huggingface.co/rhasspy/piper-voices/resolve/v1.0.0/en/en_US/libritts_r/medium/${voice}.onnx.json"
    done
    mc cp --recursive "$TEMP_DIR/piper/" unclose/tts-models/piper/v1.0.0/
}

# Sync XTTS
sync_xtts() {
    echo "Syncing XTTS v2..."
    # Use huggingface-cli or git lfs
    git clone https://huggingface.co/coqui/XTTS-v2 "$TEMP_DIR/xtts"
    mc cp --recursive "$TEMP_DIR/xtts/" unclose/tts-models/xtts/v2.0.3/
}

# Sync Silero
sync_silero() {
    echo "Syncing Silero models..."
    wget -P "$TEMP_DIR/silero/" \
        "https://models.silero.ai/models/tts/en/v4_en.pt" \
        "https://models.silero.ai/models/tts/ru/v4_ru.pt"
    mc cp --recursive "$TEMP_DIR/silero/" unclose/tts-models/silero/v4/
}

sync_piper
sync_xtts
sync_silero

echo "✅ Mirror sync complete"

Advantages:

  • Under our control
  • No rate limits
  • Fast local access
  • Can modify models

Disadvantages:

  • Maintenance overhead
  • Storage costs
  • Single point of failure (us)

Tier 3: Archive.org (Tertiary)

Public archive for critical models.

Purpose: Long-term preservation, public good

What to Archive:

  • Piper voice pack (full 2GB)
  • XTTS-v2 weights
  • Key Silero models
  • StyleTTS2 checkpoints

Upload Process:

# Install internet archive CLI
pip install internetarchive

# Configure
ia configure

# Upload critical model
ia upload uncloseai-piper-voices-v1.0.0 \
    piper_voices.tar.gz \
    --metadata="title:Piper TTS Voices v1.0.0" \
    --metadata="description:Complete Piper TTS voice collection from rhasspy/piper-voices" \
    --metadata="subject:text-to-speech;tts;piper;neural-tts" \
    --metadata="creator:uncloseai Raccoon Mission" \
    --metadata="date:2025-11-09"

Advantages:

  • Permanent storage
  • Public access
  • Free
  • Trusted platform

Disadvantages:

  • Slow downloads
  • No control over availability
  • Upload limits

Tier 4: IPFS (Experimental)

Decentralized storage for the future.

Purpose: Censorship-resistant, distributed

Implementation:

# Pin critical models to IPFS
ipfs add -r piper_voices/
# Output: QmXXXXXXXXXXXXXXXX

# Pin via Pinata or other service
curl -X POST "https://api.pinata.cloud/pinning/pinByHash" \
  -H "pinata_api_key: YOUR_KEY" \
  -d '{"hashToPin":"QmXXXXXXXXXXXXXXXX"}'

Advantages:

  • Decentralized
  • Censorship resistant
  • Content-addressed

Disadvantages:

  • Slow
  • Requires pinning service
  • Less reliable

Download Strategy with Fallbacks

Smart Downloader

# src/utils/model_downloader.py
from typing import List, Optional
import requests
import logging

logger = logging.getLogger(__name__)

class ModelDownloader:
    """Download models with automatic fallback to mirrors"""

    def __init__(self):
        self.mirrors = [
            "https://huggingface.co/rhasspy/piper-voices/resolve/v1.0.0/",
            "http://ai.foxhop.net:9000/tts-models/piper/v1.0.0/",
            "https://archive.org/download/uncloseai-piper-voices-v1.0.0/",
        ]

    def download(self, model_path: str, output_path: str) -> bool:
        """Try each mirror until successful"""
        for mirror_url in self.mirrors:
            full_url = f"{mirror_url}{model_path}"
            logger.info(f"Trying {full_url}...")

            try:
                response = requests.get(full_url, stream=True, timeout=30)
                if response.status_code == 200:
                    with open(output_path, 'wb') as f:
                        for chunk in response.iter_content(chunk_size=8192):
                            f.write(chunk)
                    logger.info(f"✅ Downloaded from {mirror_url}")
                    return True
            except Exception as e:
                logger.warning(f"❌ Failed {mirror_url}: {e}")
                continue

        logger.error(f"All mirrors failed for {model_path}")
        return False

Configuration

# config/mirrors.yaml
mirrors:
  piper:
    - https://huggingface.co/rhasspy/piper-voices/resolve/v1.0.0/
    - http://ai.foxhop.net:9000/tts-models/piper/v1.0.0/
    - https://archive.org/download/uncloseai-piper-voices-v1.0.0/
    - ipfs://QmXXXXXXXXXXXXXXXX/

  xtts:
    - https://huggingface.co/coqui/XTTS-v2/resolve/main/
    - http://ai.foxhop.net:9000/tts-models/xtts/v2.0.3/
    - https://archive.org/download/uncloseai-xtts-v2/

  silero:
    - https://models.silero.ai/models/tts/
    - http://ai.foxhop.net:9000/tts-models/silero/v4/
    - https://github.com/snakers4/silero-models/releases/download/

retry:
  max_attempts: 3
  timeout_seconds: 30
  backoff_multiplier: 2

Implementation Checklist

Phase 1: Setup Mirror Infrastructure

  • Deploy MinIO on ai.foxhop.net
  • Create tts-models bucket
  • Set up public read access
  • Configure DNS/CDN (optional)

Phase 2: Initial Sync

  • Download all Piper voices (2GB)
  • Download XTTS-v2 (1.8GB)
  • Upload to MinIO mirror
  • Test download from mirror

Phase 3: Implement Fallback Logic

  • Create ModelDownloader class
  • Add mirror config to config/mirrors.yaml
  • Update download scripts to use fallbacks
  • Add mirror health checks

Phase 4: Archive Critical Models

  • Upload Piper voices to Archive.org
  • Upload XTTS-v2 to Archive.org
  • Document archive locations
  • Test restoration from archive

Phase 5: Automation

  • Create sync script (scripts/sync_models.sh)
  • Set up cron job for weekly sync
  • Monitor mirror disk usage
  • Alert on upstream changes

Phase 6: Future Engines

  • Add Silero to mirror
  • Add StyleTTS2 to mirror
  • Add Fish Speech to mirror

Monitoring and Maintenance

Health Checks

#!/bin/bash
# scripts/check_mirrors.sh
# Verify all mirrors are accessible

MODELS=(
    "piper/v1.0.0/en/en_US/libritts_r/medium/en_US-libritts_r-medium.onnx"
    "xtts/v2.0.3/model.pth"
)

for model in "${MODELS[@]}"; do
    echo "Checking $model..."

    # Check HuggingFace
    curl -I "https://huggingface.co/rhasspy/piper-voices/resolve/v1.0.0/$model" | head -n 1

    # Check our mirror
    curl -I "http://ai.foxhop.net:9000/tts-models/$model" | head -n 1

    echo "---"
done

Storage Usage

# Monitor MinIO usage
mc du unclose/tts-models

# Expected:
# piper/: 2GB
# xtts/: 1.8GB
# silero/: 500MB
# Total: ~5GB

Cost Estimation

Storage (100GB allocated)

  • MinIO on existing server: $0 (using spare disk)
  • Bandwidth: $0 (self-hosted, unlimited)

Archive.org

  • Storage: $0 (free)
  • Bandwidth: $0 (free)

IPFS Pinning (Optional)

  • Pinata: $20/month for 100GB
  • Or self-host: $0

Total Cost: $0-20/month


Recovery Scenarios

Scenario 1: HuggingFace is down

  1. Downloader tries HF, gets timeout
  2. Falls back to ai.foxhop.net mirror
  3. Download succeeds in 30 seconds

Scenario 2: Our mirror is down

  1. Downloader tries ai.foxhop.net, fails
  2. Falls back to Archive.org
  3. Download succeeds in 2 minutes (slower)

Scenario 3: Total internet failure

  1. Models already cached in /app/voices/
  2. Service continues with cached models
  3. No downloads needed for operation

Scenario 4: Apocalypse (all servers gone)

  1. Restore from Archive.org archive
  2. Restore from IPFS if configured
  3. Restore from torrents if distributed
  4. Rebuild from source if absolutely necessary

Last Updated: 2025-11-09 Raccoon Status: 🦝 Building resilient caches like storing nuts for winter