zebra-report/docs/tickets/0001-fedora-chrome-cannot-hear-speakers.md
Russell Ballestrini 4cedd85dc6
zebra-spaces: re-apply sink on every attach + per-tick ctx.sink telemetry (kill blanka-chrome sink drift)
After caa0548 (fixes A + B for ticket 0001), blanka-chrome was still
silent. Fresh telemetry pinned a third failure mode:

  21:31:31  audio via AudioContext 47e5 sink=c38572ec…   ← good
  …leave + rejoin…
  21:31:33  audio via AudioContext 47e5 sink=default     ← drifted

No `audioCtx sink → c38572ec…` log line between the two attaches.
applySinkToAudioCtx was only invoked inside the `if (!audioCtx)` branch
in attachAudioStreamViaWorklet — a persistent audioCtx whose .sinkId
getter returned '' (system default) after a leave/rejoin never got
its sink re-applied, so every subsequent attach emitted to the system
default speaker instead of the device the user picked in the dropdown.

Fix: when audioCtx already exists AND speakerDeviceId is set AND
audioCtx.sinkId !== speakerDeviceId, call applySinkToAudioCtx() to
restore routing. Idempotent (same-sink call is a noop).

Telemetry: every 5s tick now includes `ctx.sink=…` and `want=…` when
the active sink doesn't match the picked one, so drift is visible
without grepping for the rare attach event.

Test: attach re-applies sink when audioCtx persists across leave/rejoin
and the picked sink drifted. Sandbox now exposes setSpeakerDeviceId so
the test can drive the picked-device path through the real
attachAudioStreamViaWorklet branch.

Ticket 0001 updated with the second-pass telemetry and fix (C).
2026-06-07 17:42:12 -04:00

220 lines
9.8 KiB
Markdown

# 0001: fedora chrome cannot hear speakers
**Status:** in-progress
**Priority:** P0 (silent regression — every invariant green except audible output)
**Surface:** fedora chrome — observed under both listener and speaker roles
**First seen:** 2026-06-04 (or earlier — "for days")
**Reporter:** fox (blanka-chrome identity)
## Symptom
On fedora chrome, no audio from any speaker in the room is audible. Affects
every role chrome takes: as listener, no mic audio plays. As speaker, no
peer mic audio plays through the mesh+SFU path. Same chrome can publish
mic + camera + screen fine — only the receive-side audio is silent.
Other browsers on the same machine (firefox) hear the room correctly once
the same code is loaded — so the codebase is not uniformly broken.
## Telemetry / Repro
Signal-server log (`/var/log/zebra-spaces-signal.log`,
`CLIENT_LOG actor_handle="blanka-chrome"`) on 2026-06-06 ~23:24 UTC:
```
· role=speaker sub=connected/connected pub=connected/connected
mesh=2 sListen=0 streamMode=0 muted=1 xcr=off
aud.recv pkt=308 lost=0 bytes=197428 jitter=0.0080 level=0.000 jbuf=? lp=0.2s
aud.recv pkt=22075 lost=3 bytes=44147 jitter=0.3210 level=0.000 jbuf=? lp=0.9s
aud.recv pkt=44242 lost=4 bytes=751664 jitter=0.0260 level=0.000 jbuf=? lp=0.3s
```
- RTP arriving at ~50 pps on the active speaker (real Opus voice traffic)
- `level=0.000` on EVERY receiver
- `mesh=2` (both mesh peers connected)
- `sub=connected` and `pub=connected` (both SFU PCs healthy)
Earlier in the session:
```
audio attach uuid=2d4a target=0.5s (current pool=3)
audio via AudioContext 2d4a target=0.5s ctxState=running
jitter-buffer installed uuid=2d4a target=0.5s
jitter-buffer started uuid=2d4a target=0.5s
mesh stream swapped into worklet for 2d4a — buffer cushion now applies to mesh too
```
Every attach invariant green. Worklet `started` event fired. ctxState=running.
Yet the user hears nothing.
User report: chrome tab indicator shows NO audio playing.
## Hypothesis
**Primary (deployed 2026-06-06, awaiting retest):** the speaker-output
dropdown in the page calls `setSinkId` on `<audio>` elements only. The
worklet path (every listener AND every speaker monitor since c3ff58c routed
mesh through the worklet) runs through `audioCtx.destination`, which always
emits to the **system default** sink. If the user picks a non-default sink
in the dropdown, the worklet keeps emitting to the wrong device — silence
on the device they're monitoring.
Fox 2026-06-06 confirmed picking a non-default sink on fedora chrome. Firefox
defaults to system default which happens to match the right speaker for him,
which is why firefox music started working without needing this fix.
**Secondary (not yet ruled out):**
- AudioContext.setSinkId() may not be honored on this chrome version /
flag set → silent fallthrough to default
- PulseAudio per-app default routing for chrome may override the picked
sink at the OS layer
- The worklet's `started=true` may flip back to `false` on a network blip
and never re-arm if the queue can't refill to target (rearm guard fires)
## Fix
**Landed:** `467146a` 2026-06-06 — `applySinkToAudioCtx()` calls
`audioCtx.setSinkId(speakerDeviceId)` on creation and in the dropdown's
fan-out. Diagnostic log line now includes `sink=...` so the next
telemetry session shows which sink the worklet is routed to.
**Deployed:** `0f36e02` to www.unturf.com — blanka-chrome needs page reload
(it's a code-shell change, not SDP-level).
**Pinned by tests:**
- `audioCtx.setSinkId is called on creation so the worklet routes to the
user-picked sink (not just system default)`
- `audioCtx without setSinkId support (older browser) does not throw —
silent fallback to default`
## 2026-06-07 telemetry — sink hypothesis ruled out
Fresh `/var/log/zebra-spaces-signal.log` capture (00:20:01 — fxhp host
publishing real music, blanka-chrome speaker):
```
audio via AudioContext 085b target=0.5s ctxState=running sink=default
mesh stream swapped into worklet for 085b — buffer cushion now applies to mesh too
...
aud.recv pkt=14622 lost=1 bytes=9117575 jitter=0.0100 level=0.000 jbuf=? lp=0.1s
```
Comparison — fxhp-phone in the same room, same window, same publisher:
```
aud.recv pkt=11161 lost=1 bytes=7154201 jitter=0.0090 level=0.001 jbuf=0.49s lp=0.0s
```
- sink IS routed (deployed fix `467146a` confirmed live — later capture
shows `sink=c38572ec…` = the user-picked deviceId, NOT default).
- 50 pps RTP, 9.1 MB of stereo Opus arriving on blanka-chrome's SFU sub PC.
- `jbuf=?` = `jitterBufferEmittedCount === 0` = receiver **decodes
nothing**. There is no consumer wired to the SFU audio receiver.
- fxhp-phone on the same SFU stream has `jbuf=0.49s` + `level=0.001`
(native decoder running) — so the SFU is fine, the problem is per-
client wiring.
So: **sink routing was not the bug, or was only part of it.** With the
sink fix applied, fedora chrome still hears silence.
## Real root cause
Two coupled defects, both visible in `handleRemoteSfuTrack` (line 4700)
and the mesh `pc.ontrack` (line 6693):
**(A) `tracks=0` race at mic ontrack.** Chrome can fire `ontrack` where
`ev.streams[0]` exists but has 0 tracks at handler-time (the track is
added a microtask later, or MSID-supplant merges new + dead tracks).
Screen / camera / game already work around this by constructing a fresh
`new MediaStream([ev.track])` and passing that. **Mic does not.** It
caches `ev.streams[0]` directly into `sfuStreamsByPubHex` and hands the
same reference to `attachSfuTrack`. When that fires before the track
shows up on the stream, `attachAudioStreamViaWorklet` rejects with
`liveAudio.length === 0` → falls through to the silent `<audio>` path
→ logs `sfu attach b1a7 fresh=1 tracks=0` + `meter for b1a7: MediaStream
has no audio track` (both verbatim in today's capture).
**(B) Mesh swap orphans the SFU receiver.** Mesh `pc.ontrack` calls
`setWorkletStream(uuid, meshStream)` unconditionally — it does NOT
verify that mesh is actually emitting audio. The worklet's
`MediaStreamSource` is now bound to the mesh track. The SFU receiver
has no consumer, so its decoder doesn't run (matches `jbuf=?`,
`level=0`). If the mesh track is muted or DTX-silent, the listener
hears nothing — and there is no automatic SFU restore until mesh PC
actually transitions to `failed` (line 6739). Mesh that stays
`connected` but silent traps the listener forever.
Hard refresh + leave/enter fixes both: fresh PC negotiation → fresh
ontrack with the live track in the stream → attach succeeds on SFU
before mesh comes up → mesh swap eventually replaces source but by
then SFU receiver is fine and… actually no, this still ends up on
mesh. The reason hard-refresh works is that mesh negotiation is fresh
and the new mesh track arrives in an `unmuted` state because audio is
already flowing. The mesh swap then lands on a real audio source.
## Fix (in progress)
**(A)** In `handleRemoteSfuTrack` mic branch, mirror the screen/camera
pattern: `const s = new MediaStream([ev.track]); sfuStreamsByPubHex.set(pubHex, s); attachSfuTrack(uuid, s)`.
Guarantees the stream handed to `attachAudioStreamViaWorklet` contains
exactly the live track, every time.
**(B)** In mesh `pc.ontrack`, gate the `setWorkletStream` swap on
`!ev.track.muted` (or `addEventListener('unmute', swap)` if it starts
muted). If the mesh track never unmutes, keep SFU as the worklet
source. Add a watchdog: if a swapped-in mesh track returns to muted
for >5s, swap the worklet back to the cached SFU stream — same
pattern the `connectionState === 'failed'` path already uses at line
6755.
## 2026-06-07 second telemetry pass — sink DRIFT identified
After deploying (A) + (B), blanka-chrome was still silent. Fresh
telemetry pinned a different per-attach failure mode:
```
21:31:30 audio attach uuid=47e5 target=0.5s (current pool=1)
21:31:31 audio via AudioContext 47e5 target=0.5s ctxState=running sink=c38572ec…
21:31:31 jitter-buffer started uuid=47e5 target=0.5s
…leave/rejoin cycle…
21:31:33 audio attach uuid=47e5 target=0.5s (current pool=1)
21:31:33 audio via AudioContext 47e5 target=0.5s ctxState=running sink=default ← drifted
```
No `audioCtx sink → c38572ec…` log line ever appeared. So
`applySinkToAudioCtx()` either never ran or hit the
`typeof audioCtx.setSinkId !== 'function'` early-return at the moment
of execution. The earlier sink fix only invoked
`applySinkToAudioCtx()` inside the `if (!audioCtx)` branch — a
persistent audioCtx whose `.sinkId` getter returned `''` (default) was
never re-corrected.
## Fix (C) — re-apply sink on every attach when it drifted
`attachAudioStreamViaWorklet` now: if `audioCtx` already exists AND
`speakerDeviceId` is set AND `audioCtx.sinkId !== speakerDeviceId`,
call `applySinkToAudioCtx()`. Idempotent — same-sink call is a no-op.
Also added per-tick `ctx.sink=` to the telemetry stream so sink drift
is visible on every 5s line instead of needing to find the rare attach
log. When the active sink doesn't match the picked one, the line
includes `want=…` so the mismatch jumps out.
## Status notes
- **2026-06-07 (in-progress, fix C deployed):** sink-drift on
audioCtx-persistent path identified + patched. Need a third telemetry
pass to confirm the per-tick `ctx.sink=` line now stays on the
picked device across leave/rejoin.
- **2026-06-07 (fixes A + B deployed):** sink hypothesis from 467146a
ruled out by fresh telemetry; (A) tracks=0 race + (B) mesh-swap
orphan landed in caa0548. blanka-chrome was still silent, leading
to fix (C) above.
- **2026-06-06:** the Jun 5 cascade + flushSfuStreams prefix-match fix
(`0ce1339`) helped firefox but not chrome. Telemetry above ruled out
the FSM / attach defects.
- **2026-06-04 → 2026-06-05:** earlier cascade commits (2e74b92,
f479878, 144dd15, 61b21a6) tightened the audio attach path but did
not address this defect (they targeted double-audio on
fxhp-android-firefox).