uncloseai-speech/TESTING.md

7.1 KiB

Testing Guide: Silero TTS Integration

Status: Code validated, syntax verified Docker: Not available in dev environment - deployment testing required

Pre-Deployment Validation

Code Validation

✅ Python syntax validated (speech.py)
✅ Requirements.txt format verified
✅ 14 packages defined in requirements.txt
✅ F-string syntax error fixed

Changes Summary

  • Integrated: Silero TTS (tts-1-silero model)
  • Prepared: Chatterbox and Kokoro dependencies
  • Added: 6 new Makefile targets for model downloads and testing

Deployment Testing Instructions

Since Docker is not available in the development environment, follow these steps on your deployment server:

1. Pull Latest Changes

cd ~/uncloseai-speech  # or your deployment path
git pull origin claude/implement-models-docs-011CUxXuNMytPjEr5vsvcboo

2. Rebuild Container (Using Makefile)

# Option A: Full rebuild with restart
make restart

# Option B: Manual rebuild
docker compose up -d --build

3. Monitor Build Logs

# Watch the build process
docker compose logs -f

# Or use Makefile
make logs

Expected Output:

✓ Installing fastapi, uvicorn, loguru
✓ Installing piper-tts>=1.2.0
✓ Installing coqui-tts[languages]
✓ Installing transformers>=4.35.0
✓ Installing huggingface-hub[cli]
✓ Installing torch, torchaudio
✓ Cloning chatterbox from GitHub (may take 2-5 mins)
✓ Server starting on 0.0.0.0:8000

4. Verify Models Available

curl http://localhost:8000/v1/models

Expected Response:

{
  "data": [
    {"id": "tts-1", "object": "model"},
    {"id": "tts-1-hd", "object": "model"},
    {"id": "tts-1-silero", "object": "model"}
  ]
}

5. Test Silero TTS (Fast CPU-friendly synthesis)

Test 1: Basic Synthesis

curl -X POST http://localhost:8000/v1/audio/speech \
  -H "Content-Type: application/json" \
  -d '{"model":"tts-1-silero","voice":"alloy","input":"Testing Silero fast synthesis"}' \
  -o test_silero.mp3

# Play the audio
mpv test_silero.mp3

Test 2: Different Voices

# Test all 6 voices (alloy, echo, fable, onyx, nova, shimmer)
for voice in alloy echo fable onyx nova shimmer; do
  echo "Testing voice: $voice"
  curl -X POST http://localhost:8000/v1/audio/speech \
    -H "Content-Type: application/json" \
    -d "{\"model\":\"tts-1-silero\",\"voice\":\"$voice\",\"input\":\"This is the $voice voice\"}" \
    -o "test_silero_${voice}.mp3"
done

Test 3: Speed Control

curl -X POST http://localhost:8000/v1/audio/speech \
  -H "Content-Type: application/json" \
  -d '{"model":"tts-1-silero","voice":"alloy","input":"Testing speed control","speed":1.5}' \
  -o test_silero_fast.mp3

Test 4: Download Silero Models (Optional Pre-caching)

# Pre-download models for 5 languages
make voices-silero

# Or manually inside container
docker exec uncloseai-speech-server-1 python3 -c "
import torch
for lang in ['en', 'ru', 'de', 'es', 'fr']:
    model, *_ = torch.hub.load('snakers4/silero-models', model='silero_tts', language=lang)
    print(f'Downloaded Silero {lang}')
"

6. Compare Model Performance

Test all three engines:

# Piper (tts-1) - Very fast, CPU
time curl -X POST http://localhost:8000/v1/audio/speech \
  -H "Content-Type: application/json" \
  -d '{"model":"tts-1","voice":"alloy","input":"Performance test"}' \
  -o test_piper.mp3

# Silero (tts-1-silero) - Fast, CPU
time curl -X POST http://localhost:8000/v1/audio/speech \
  -H "Content-Type: application/json" \
  -d '{"model":"tts-1-silero","voice":"alloy","input":"Performance test"}' \
  -o test_silero.mp3

# XTTS (tts-1-hd) - Slower, GPU recommended
time curl -X POST http://localhost:8000/v1/audio/speech \
  -H "Content-Type: application/json" \
  -d '{"model":"tts-1-hd","voice":"alloy","input":"Performance test"}' \
  -o test_xtts.mp3

Expected Performance:

  • Piper: 0.5-1 second (RTF ~0.05x)
  • Silero: 1-2 seconds (RTF ~0.1x)
  • XTTS: 3-5 seconds (RTF ~0.3x)

7. Test Error Handling

Test invalid model:

curl -X POST http://localhost:8000/v1/audio/speech \
  -H "Content-Type: application/json" \
  -d '{"model":"invalid","voice":"alloy","input":"Test"}' \
  -v

Expected: HTTP 400 with error message

Test invalid voice:

curl -X POST http://localhost:8000/v1/audio/speech \
  -H "Content-Type: application/json" \
  -d '{"model":"tts-1-silero","voice":"invalid","input":"Test"}' \
  -v

Expected: HTTP 400 or 503 with voice error

Troubleshooting

Build Fails on Chatterbox

If cloning chatterbox fails (GitHub rate limit or network):

# Comment out Chatterbox temporarily
sed -i 's/^git+https:\/\/github.com\/resemble-ai\/chatterbox.git/# &/' requirements.txt
docker compose up -d --build

Chatterbox is not yet integrated into speech.py, so it's safe to skip for now.

Silero Model Download Slow

First synthesis will download Silero models (~50-100MB). Subsequent calls will be fast.

# Pre-cache during deployment
make voices-silero

Out of Memory

Silero runs on CPU and uses minimal memory (~500MB). If issues occur:

# Check container memory
docker stats uncloseai-speech-server-1

# Restart if needed
make restart

Check Logs

# Full logs
docker compose logs

# Follow logs in real-time
make logs

# Filter for errors
docker compose logs | grep -i error

Success Criteria

Container builds without errors All 3 models listed in /v1/models Silero synthesis works (tts-1-silero) Response time < 2 seconds for Silero Audio quality is clear and natural All 6 voices work (alloy, echo, fable, onyx, nova, shimmer) No memory leaks after 10+ requests

Next Steps After Successful Deployment

  1. Integrate Chatterbox (emotion control)

    • Implement chatterbox_wrapper in speech.py
    • Add model handler for tts-1-chatter
    • Test emotion parameters
  2. Integrate Kokoro (fast decoder)

    • Implement kokoro_wrapper in speech.py
    • Add model handler for tts-1-kokoro
    • Test performance vs Silero
  3. Create Detailed Silero Documentation

    • Write docs/models/silero-tts.md
    • Document all 117 English speakers
    • Add multilingual examples
  4. Performance Benchmarking

    • Test all models under load
    • Measure memory usage over time
    • Compare audio quality subjectively

Files Modified in This Integration

requirements.txt              - Added Silero comments, Chatterbox, Kokoro, huggingface-hub
speech.py                     - Added silero_wrapper, tts-1-silero handler, model registration
voice_to_speaker.default.yaml - Added tts-1-silero voice mappings
Makefile                      - Added 6 new targets (voices-silero, test-silero, etc.)
docs/MODELS.md                - Updated with Silero integration status

Rollback Instructions

If deployment fails:

# Stop current container
docker compose down

# Checkout previous working commit
git checkout 0073e87^  # Parent of Silero integration

# Rebuild
docker compose up -d --build

Last Updated: 2025-11-09 Branch: claude/implement-models-docs-011CUxXuNMytPjEr5vsvcboo Status: Ready for deployment testing