Replace the fragile client-side RMS / Web Audio pause detection (which stalled on sentence 0 whenever the audio context was suspended or the MSE duration was unknown) with exact per-sentence timing from the speech service. For tts-1-f5, fetch via the new timestamps mode and light each sentence by comparing audio.currentTime to the returned start/end ms; falls back to proportional mapping if the rendered and spoken sentence counts differ. Cache now stores sentences for correct replay glow. Non-F5 models keep streaming playback with no glow. |
||
|---|---|---|
| .. | ||
| auth.html | ||
| base.html | ||
| browse.html | ||
| chat.html | ||
| index.html | ||
| profile.html | ||
| search.html | ||