startStream() created the <audio> element with muted=true and trusted
applyAudioMute() to unmute it after the user gesture. But that hook
became a listener-only no-op during the recent listener refactor, so
host/speaker self-monitor and host-driven DJ on remote speakers played
silently. Unmute on every startStream() call — the toggle click is
itself a gesture so we no longer need the muted-autoplay workaround.
JS error fox spotted via the new unhandledrejection forwarder:
audioPath is not defined
audioPath was a per-uuid map ('dj' | 'rtc') I added when listeners
auto-enrolled into HTTP DJ mode. The auto-enrol path was reverted
when fox pointed out phones can't autoplay HTTP audio, but four
audioPath.set/.clear references survived in startStream, stopStream,
and the btn-leave click handler. The leave handler hit
audioPath.clear() before the bye send and threw ReferenceError,
abandoning the rest of the cleanup (which is why the leave button
felt half-broken: bye did fire from somewhere earlier in the
handler, but post-leave UI never reset).
Removed every dead audioPath reference. WebRTC unmute is now handled
inline in stopStream where needed. applyAudioMute is the last
defined-but-unused holdover; keeping it since callers of attachSfu-
Track / startStream still reference the no-op via stale comments,
and it's cheap.
Browser HTML caching has been silently serving listeners old builds
even after deploy — log shows zero new-build markers from fox's
phone or laptop chrome session despite md5 matching on the
served file. Heuristic caching with no Cache-Control header gives
browsers freedom to keep the HTML for hours.
Adding the standard meta cache-bust triplet so the HTML is treated
as no-cache by every browser regardless of Caddy headers. JS / CSS
is inline so this covers the whole page in one go.
Fox: leave button still not working. Adding a logLine at the top of
the click handler so we can see whether the click ever reaches it,
and a window.onerror + unhandledrejection forwarder so any silent
JS exception that broke event-handler attachment surfaces in
CLIENT_LOG instead of dying in the browser console.
Next test: tap leave on the phone, then check the proxy log for
either 'leave-button: click fired' or a JS error line — that tells
us if the click is being lost (overlay / CSS) or if the handler
itself is throwing.
Two regressions from the audio-pool work:
1. Auto-rejoin was firing joinSpace() with no user activation, so
the audio pool couldn't be primed and listener phones landed
silent. The follow-up attempt (global click listener on document
capture-phase) intercepted the leave button and other UI clicks.
Now: auto-rejoin pre-fills the rendezvous code and surfaces
'click enter to resume — <code>' as a status — user clicks enter
to actually join. One extra tap is the price of working audio.
2. mute + leave buttons were always visible (just disabled) before
the user joined a space. Fox: they shouldn't be there when no
space is active. Both start hidden in the HTML; appear when the
welcome lands; disappear on leave / WS-tear.
CLIENT_LOG confirmed the listener-phone autoplay rejection:
rtc autoplay 72b026b4: The play method is not allowed by the user
agent or the platform in the current context, possibly because
the user denied permission.
attachSfuTrack runs several async hops after the entry-button click
(WS join → welcome → onRoleEntered → sfuSubscribe → SFU PC
negotiate → ontrack). Firefox Android's user-activation window has
aged out by then. Speakers dodge this because getUserMedia for the
mic counts as confirmed audio engagement; listeners have no
equivalent.
Fix: primeAudioOnGesture now also pre-creates a pool of 16 <audio>
elements and calls play() on each inside the entry-button gesture.
Firefox Android grants autoplay PER ELEMENT and the engagement bit
persists on the element after the click activation expires.
attachSfuTrack now leases from the pool (leaseAudioElement) instead
of creating fresh — the leased element is already blessed for
unmuted autoplay, so srcObject + play() succeed without a fresh
gesture. Pool falls back to fresh-element creation on exhaustion.
Fox: 'i thought you removed the streaming mode and the play button to
autoplay webrtc.' Right — the play button was meant to come out too.
WebRTC playback for listeners is now identical to the speaker inbound
path: attachSfuTrack creates an autoplay=true <audio>, sets srcObject,
calls play() once, logs a rejection if it happens. No second-gesture
UI, no listenerOutputMuted toggle, no swap-fresh trick.
The mute button stays disabled for listeners (no mic to mute) — the
same disabled state speakers see before they have a mic. Speaker
mic-mute path is unchanged.
Fox: 'no we regressed we have never gotten the stream mode to work
for phones only time phone has worked is falling back to webrtc which
is what we are using for speakers not DJ mode... :( listeners.'
The auto-DJ + swap-fresh-<audio> + muted-on-create combo broke the
one path that actually worked: WebRTC playback for listener phones.
Reverting to the original attachSfuTrack flow (autoplay=true,
playsInline, .play() with .catch logging) so WebRTC audio comes up
the same way it did for speakers all along.
DJ mode is intentionally NOT auto-enrolled for listeners — Firefox
Android refuses autoplay on every fresh <audio src=URL> and the
entry-button gesture has aged out by the time the SFU lazy-inits
its Ogg writer. The DJ stream toggle stays on speaker rows for the
host's room-wide control.
btn-mute as listener now just calls activateListenerAudio() which
re-fires .play() on every existing remoteAudio element inside the
click gesture. If attachSfuTrack's initial play() got swallowed,
this gesture lands the audio.
CLIENT_LOG made the failure mode unambiguous: after the play button
tap, the WebRTC audio elements report paused=false rs=4 muted=false
srcObject=true — the browser believes it is playing — but Firefox
Android emits zero output. Same shape as the camera autoplay-block
we fixed earlier with swapFreshVideoElement: an element created
muted=true and later unmuted is internally pinned to silent.
Fix mirrors that pattern. Inside the listener's btn-mute click
(real user gesture), every <audio> in remoteAudio / streamAudio
is replaced with a freshly-built unmuted element, the same
srcObject (for WebRTC) or src URL (for DJ stream) is reattached,
and play() fires inside the gesture context. The fresh element
has never been muted so the internal routing is clean.
Last click showed touched=4 and no play() rejections, yet phone is
silent — so the issue isn't the click handler returning early. It's
that the audio elements themselves are in some not-actually-playing
state we're not catching.
This commit logs per-element: paused, muted, readyState, srcObject
presence (rtc) / src + networkState (dj), and the path tag. Also
ALWAYS attempts play() in the click handler regardless of mute
decision so the user gesture isn't wasted.
Two regressions where the phone went silent across a leave + re-enter
cycle, with no way to recover short of a hard refresh:
1. Leave handler tore down WebRTC + members + tiles but left:
- streamMode (Set of streaming pubHexes)
- streamAudio (uuid -> <audio>)
- audioPath (uuid -> 'dj'|'rtc')
- autoEnrolRetryTimer running
On the next entry as listener, autoEnableDjModeForListener would
skip every pubHex because streamMode.has(pubHex) was still true
from the prior session → startStream never called → no HTTP /stream
request → phone falls back to WebRTC and stays there even when the
host's mic publish is live.
Fix: stopAutoEnrolRetryLoop + autoDisableDjModeForAll + remove every
stale streamAudio element from the DOM + clear all three maps as
part of the leave click handler.
2. primeAudioOnGesture had an audioPrimed=true short-circuit that
skipped the silent-WAV prime on re-entry. Mobile browsers can
suspend the audio session on tab background / leave, so a one-shot
prime doesn't carry across sessions. Drop the gate — every entry
click re-primes. The primer element self-removes 500ms after
play() resolves so we don't accumulate hidden <audio> elements.
Tests green (83 fsm + 16 zebra-spaces).
Server log was silent on what the phone was doing after my last
change. Adding logLine() in three places so the next attempt is
diagnoseable from CLIENT_LOG without asking fox to relay phone
screen content:
- autoEnableDjModeForListener: how many members it saw, how many
pubHexes it actually added to streamMode.
- applyAudioMute: how many <audio> elements got their muted state
touched, per-element play() rejections.
- btn-mute click as listener: current listenerOutputMuted +
sizes of remoteAudio / streamAudio / audioPath maps so we know
if the click is firing and what state it sees.
If the host's mic publish lands AFTER the listener's initial
autoEnableDjModeForListener pass (which runs 600ms after the listener
sees the peer-joined event), the /stream request 404s with
"no such publisher in room" and onFail removes the pubHex from
streamMode. No event re-triggers the auto-enrol after that — peer-
joined doesn't refire when an already-present member starts publishing.
Add a 4s interval retry loop while in listener mode. autoEnableDj-
ModeForListener is idempotent (skips already-enrolled pubHexes), so
the loop is cheap and only re-attempts the missing ones. Stops
automatically on role transition out of listener.
This was the root cause behind every "no music on phone" report so
far: the SFU logs show /stream 404s when the phone tried, then never
again after the host's mic publish completed. The retry loop closes
the race.
Tests green (83 fsm + 16 zebra-spaces).
Fox 2026-06-04, after pulling CLIENT_LOG from the proxy:
stream for 25cc0aaf7eb0 autoplay blocked: ... — staying on live WebRTC
autoplay blocked: ... — tap the tile to play
Firefox Android refuses .play() on every <audio> element created
after the entry-button gesture has aged out. Both the HTTP DJ stream
AND the WebRTC fallback were silently dead — the phone heard nothing.
Two-part fix:
1. Every remoteAudio + streamAudio <audio> element is now created
with muted=true. Muted autoplay has no gesture requirement on
any browser — decoding starts the moment the src is set, the
buffer fills, and audibility is gated entirely by a later
user gesture.
2. btn-mute is enabled for listeners (was disabled because there's
no mic to mute) and re-purposed as the audio-unlock toggle. The
click is the gesture. applyAudioMute() reads a per-uuid audio-
Path map ('dj' | 'rtc') and unmutes only the canonical path for
each peer so the WebRTC duplicate stays silenced while the DJ
stream plays.
Also picked up while I was in here:
- startStream onFail clears the pubHex from streamMode so the
next auto-enrol pass retries.
- stopStream sets audioPath='rtc' instead of poking remoteAudio
directly.
The tap-anywhere-to-resume queue was a regression — users saw a "tap
anywhere to start" message and tapping often did nothing, with no
feedback about why. Replace with the obvious-in-hindsight approach:
seize the user gesture from the entry button click itself.
primeAudioOnGesture() plays a 1-frame silent WAV through a hidden
<audio> element inside the entry click handler. Mobile Firefox /
Safari treat that as gesture-driven audio playback and grant the
page audio permission for the session. It also resumes any
suspended AudioContext (meter / chime path) the same way.
By the time auto-enrolment into DJ mode runs (async hops after
entry), the page already has audio permission — dynamic <audio>
elements created later play() without further user interaction.
No "tap anywhere", no queue, no flip-flop on stalled.
Bound to btn-enter click + rdv-code Enter key. Tap-to-resume queue
and pendingAutoplay/installTapResume helpers removed.
Tests green (83 fsm + 16 zebra-spaces).
When stream autoplay was blocked and the listener tapped to resume,
the retry called a.play() on a <audio> that had already been left in
error state by the browser-aborted prior load. play() on a stuck
element silently no-ops — listener saw the 'tap anywhere' message but
got no audio and no feedback.
Three fixes:
- call a.load() inside the gesture handler before a.play() to
restart the fetch from a clean state
- add 'click' to the listened events alongside pointerdown/
touchstart — Firefox Android only grants gesture activation on
click in some configurations
- log retry result: 'stream resumed after tap' on success,
'stream tap-retry failed: <msg>' on rejection, so we can see
what's happening instead of silent failure
Tests green (83 fsm + 16 zebra-spaces).
Mobile Firefox/Safari refuse audio.play() when the gesture activation
window has expired between the entry-button tap and auto-enrolment
into DJ mode. Previously this surfaced as "autoplay blocked" in the
log and the listener was silently stuck on WebRTC.
Now: when play() rejects, the <audio> element goes into a
pendingAutoplay set and one global pointerdown/touchstart listener
on document retries every queued element. Once they all play the
listener self-uninstalls. Idempotent install — adding more elements
during the pending window just enqueues them.
Listener sees: 'stream for <pubhex> — tap anywhere to start (msg)'
in the log; one tap later, every queued stream resumes and the
log fills with 'stream on for ...' lines.
WebRTC stays unmuted during the wait so the listener still hears
the live (possibly choppy) audio — they're not left in silence
between the autoplay block and their tap.
Fox saw repeated "stream stalled" and "autoplay blocked: aborted at
user's request" log lines on both auto-enrolled mobile listeners and
host self-monitor clicks. Three coupled causes:
1. applySinkTo() was called sync-fire-and-forget right before setting
.src and calling .play(). setSinkId() can re-init the media
pipeline; when it landed during the in-flight load it aborted the
request — Firefox surfaces that as
"The fetching process for the media resource was aborted by the
user agent at the user's request." on the play() promise. We
wrongly logged that as autoplay-blocked. Fix: await applySinkTo()
BEFORE assigning .src.
2. The 'stalled' event handler called unmuteWebRtcOnFail every time
it fired. 'stalled' fires constantly during normal mobile-cellular
buffering and isn't a terminal failure. Each fire ping-ponged the
audio path between WebRTC (unmuted) and the still-loading HTTP
stream. Drop the handler — only 'error' and a rejected play()
promise indicate real failure.
3. A second toggle-on for the same pubHex (auto-enroll race when
peer-joined fires during initial enrol) reassigned .src on the
same <audio> element, aborting the prior load with the same abort
error. Make startStream idempotent: if the element already has
our wantUrl and no .error, return immediately.
Plus two cheap mobile-friendly attrs on the dynamic <audio>:
- preload="auto" so buffering starts before play() (gives the
user-gesture window time to last past the initial fill)
- playsInline so iOS/Safari doesn't escalate to a fullscreen player
startStream is now async — callers (toggleStreamFor,
autoEnableDjModeForListener) fire-and-forget the returned promise,
which is fine since all error paths are already caught inside.
Tests green (83 fsm + 16 zebra-spaces).
Page-side companion to proxy.unturf.com 90a5e96.
- modMute(uuid): signs + sends {type:"mute",target,epoch,sig}.
- Mute button on speaker / cohost rows (host always; cohost can
mute speakers only). Lives in .acts-primary beside the role
transitions; kick/ban stay on the .acts-removal row below.
- On peer-force-muted from server:
- If I'm the target: disable my mic tracks, set muted=true,
persist to sessionStorage, broadcast new mic-state, show
"you were muted by X. Click unmute to talk again." Banner —
the unmute button is the user's own action, not locked.
- For others: mark the target visually muted right away so
the roster reflects state without waiting on a mic-state
broadcast.
The .stream-toggle CSS class and `strm` grid column were already in
place but the JS still appended the button into .acts-primary —
which lives in the bottom-row mod-actions stack. That meant every
speaker row (including self-monitor on non-mod rows) grew a third
sub-row just to host the single ◉ glyph, adding scroll height.
Render it as a sibling of micEl in the row so it lands in the `strm`
grid column on the TOP row, alongside the mic icon. Single character
text (○ / ◉) keeps the column tight, and the .stream-toggle styles
already shipped pick it up.
For rows that don't qualify (listeners, non-host-viewers on other
speakers), the streamEl is null → not appended → strm column
collapses to its 1.4rem track width. No row-height penalty.
The mod-actions stack stays as-is for actual mod actions (mic-invite,
role transitions, kick, ban). Self rows and non-mod self-monitor rows
no longer carry an empty mod-actions row at all.
Fox: 'the buttons should be horizontall not vertical on invite and
kick and ban etc.' The previous grid template
`repeat(auto-fit, max-content)` was treated by browsers as a single
column when no fixed sizing function was provided, so every button
ended up on its own row.
Swap to the older inline-block + margin pattern: .acts-primary and
.acts-removal are block-level rows, their <button> children flow
horizontally and naturally wrap to a second line only when the
panel is too narrow to hold them. No flex.
Fox 2026-06-03: 'the listener phone doesn't hear anything since your last
push.' startStream() was muting the WebRTC <audio> for the speaker BEFORE
calling .play() on the fresh HTTP <audio>. On mobile the fresh <audio>'s
autoplay was often blocked (the entry-button gesture had aged out for
freshly-created elements), so the failure mode was: WebRTC silenced +
HTTP not playing = total silence.
Now WebRTC stays unmuted until the HTTP <audio> fires 'playing'. Any
autoplay/error/stalled path unmutes WebRTC and surfaces the reason in
the log, so we fall back to live WebRTC instead of going dead.
play() promise rejections are also logged now (were silenced) so next
session we can tell exactly which speakers' streams failed to start.
Per fox: 'lock them into DJ mode unconditionally.' Stream button is
no longer rendered on listener rows for any other peer. Visibility
rule reverts to {self|host} — same as the post-kick/ban gating —
and listeners get auto-enrolled into the HTTP Ogg/Opus path without
any escape hatch. Host can still flip individual speakers room-wide,
and every speaker still has their self-monitor button.
Fox: 'I want the sound to be epic and to fucking choppy ever.'
Two changes that compose:
1. Default-on DJ mode for listeners. On role landing (or transition
INTO listener), auto-flip every speaker's playback path off the
WebRTC subscribe and onto the existing HTTP Ogg/Opus tap from
the SFU. The browser's <audio> element keeps a deep media buffer
(~30s in Chrome) that absorbs glitches WebRTC can't. Speakers
stay on WebRTC for low-latency conversational audio. Listener
can still opt-out per-speaker via the stream button — gating
expanded from {self|host} to {self|host|listener}.
Auto-enrol triggers at:
- onRoleEntered (initial join as listener)
- peer-joined (new speaker arrives while we're listening)
- role-change of OTHER peer (they became speaker)
- onRoleChanged self speaker → listener
Auto-disable: onRoleChanged self listener → speaker
(cannot afford 2s+ buffer when you have to talk back).
2. RECV_PLAYOUT_DELAY_SEC 2.0 → 4.0 on both mesh and SFU receive
paths. Conversational latency is now ~4s for speakers; this is
the cheap-but-real win against chop on the WebRTC path for
everyone who still uses it. Music/DJ rooms generally don't care
about 4s — the room's already committed to a stream-mode 2s+.
Four corrections in one pass — all from fox's live-room session
2026-06-03 after the kick/ban + heartbeat ship:
1. Stream button visibility was too generous. A speaker viewing
the host's row got a stream button — they shouldn't. New rule:
SELF row always (self-monitor), OTHER rows only when myRole
=== 'host'. Cohost gets self-only too; lift the gate to
isMod(myRole) if room-wide cohost stream control is wanted.
2. .acts-removal's grid-column:1/-1 collapsed the parent auto-fit
grid down to a single column on a row with both primary +
removal actions, so EVERY button stacked vertically. Refactor:
.mod-actions is now a row-stack of sibling groups
(.acts-primary + .acts-removal), each its own horizontal auto-
fit grid. Primary still wraps inside itself when the panel is
narrow; kick + ban are forced to a fresh line by the parent
grid-auto-rows.
3. playoutDelayHint bumped 0.7 → 2.0 (RECV_PLAYOUT_DELAY_SEC) on
both mesh + SFU receive paths. Conversational latency goes up
but the stream/HTTP-pull DJ mode already commits to ~2s, so
matching the WebRTC path makes the room consistent. Persistent
chop survived every SDP-side dial-back fox tried; this is the
last knob left at the receiver.
4. pagehide-bye now skips bfcache (event.persisted=true). Without
this guard, a phone going to lock screen / app-switch /
minimise fired bye → server evicted SFU PCs → page resumed but
audio stayed silent. Aliveness of a bfcached page is already
handled by the server-side aliveTTL (heartbeat stops during
bfcache → server arms hiccup grace at 45s).
Destructive controls shouldn't share a row with role-transition buttons —
the muscle-memory misclick where 'ban' sat next to '→ speaker' is exactly
the kind of mod ergonomics fox flagged. Kick + ban now render inside
a nested .acts-removal grid that spans the parent auto-fit row, so they
always wrap to a fresh line below stream / → cohost / → speaker /
→ listener regardless of available width.
The room-member mod-actions row was hardcoded
grid-auto-flow: column; grid-auto-columns: max-content;
which forces every button onto one non-wrapping line. After adding
the per-speaker "stream" toggle the strip got long enough
(stream + → cohost + → listener + kick + ban) to overflow the side
panel on common widths — kicked the page into horizontal scroll.
Swap to
grid-template-columns: repeat(auto-fit, max-content);
so buttons flow onto a second row when the panel is narrow. Still
grid (no flex per project rule). Doesn't change appearance on wide
panels — only kicks in when the button strip wouldn't fit on one line.
Drop the m.uuid !== myUUID guard on the per-speaker stream toggle so
the DJ can preview what listeners are actually hearing of their own
mic. Self-stream URL hits the SFU pulling our own pubkey — the ~2s
delayed playback is the cue we want to confirm the broadcast works.
Title attr warns about feedback risk on open speakers (closed
headphones like WH-1000XM5 are fine — earcups don't bleed into the
headset mic enough to feedback).
Mod buttons (kick/ban/promote/demote) still self-exclude — different
guard, kept intact.
For each speaker (other than self) the member row now carries an
"○ stream" / "◉ stream" button. Toggling it:
- Mutes the WebRTC remote audio for that uuid (so they don't
double-play through both paths)
- Attaches a parallel <audio src="SFU/stream?room=R&pub=PUBHEX">
that pulls the SFU's new Ogg/Opus broadcast tap as plain HTTP
- Routes through setSinkId so the speaker-output picker still
applies
Trade-off: WebRTC path is ~700ms latency but glitch-prone under
network loss / mixed-browser NACK gaps. The HTTP stream path is
~1.5–3s latency but bulletproof — the browser's <audio> jitter
buffer absorbs everything WebRTC can't. The sub.fm-style DJ
listening UX.
State is keyed by pubHex (not uuid) so it survives session-uuid
churn on rejoin. Cleanup hooks into tearPeer so peer-left tears
down the HTTP pull and the SFU stops fanning Ogg pages to a dead
client.
Requires the matching SFU change in proxy.unturf.com main
(GET /zebra-spaces-sfu/stream endpoint).
Choppy persisted after every other dial-back (NACK off → 510k→320k →
maxptime off). The only remaining new Chrome-publisher param was the
"fullband pin" (maxplaybackrate=48000, sprop-maxcapturerate=48000,
cbr=0). Drop those too and step bitrate back to the known-good 256k.
Chrome publisher's SDP now identical to pre-feature shape:
stereo=1; sprop-stereo=1; maxaveragebitrate=256000;
useinbandfec=1; usedtx=0
Kept:
- Firefox SDP regex fix — Firefox publishers now actually get the
music-mode params applied (was silently no-op before).
- UA-gated NACK — Chrome publishers advertise it (responder support);
Firefox publishers omit (no responder). "audio NACK as sender" log
confirms which side we're on per session.
- Speaker output picker (setSinkId UI).
Comment warns against re-adding fullband pins without re-verifying.
Three coupled changes after still-choppy reports on UA-gated NACK build:
1. 510k → 320k Opus everywhere in music mode. 320k is the transparent
listening-test threshold (Audible masters at this). 510k is spec
ceiling but real-world uploads can't sustain it alongside concurrent
screen 6M + camera 4M without dropping audio packets. Choppy at 510k
persisted after NACK rollback, so bitrate is the next dial.
2. Drop a=maxptime:120 advertisement. Tried to give Opus larger encode
windows for cleaner music at same bitrate. Chrome publishers may have
actually honored it and packed at higher ptime, which the receiver
side handled poorly (jitter buffer + playoutDelayHint=0.7 weren't
tuned for >20ms packetization). Default 20ms stays.
3. Fix preferStereoOpus regex no-op on Firefox publishers. Previous
matcher anchored on minptime=10 which Chrome emits but Firefox does
not — so for years, Firefox-published mic SDPs went out unmunged:
no stereo=1, no maxaveragebitrate cap, no fullband pin, no NACK gate.
Now: scan rtpmap for all opus/48000/2 PTs, update existing fmtp lines
in place, or insert a fresh fmtp if none exists (Firefox's case).
This is why fox saw "audio NACK as sender" diagnostic never fire on
Firefox publish — the function early-returned before reaching it.
Kept: fullband pins (no bandwidth cost), UA-gated NACK code (honest
per-browser advertisement), speaker output picker, 700ms playoutDelayHint.
Audio NACK is an asymmetric feature — the SENDER has to respond to
retransmit requests. Chrome implements both directions; Firefox
implements neither for audio. Blanket-advertising NACK in offers from
Firefox publishers caused Chrome receivers to wait for retransmits
that never arrived and skip audibly.
Solution: each browser advertises only what it can back up. Probe via
RTCRtpSender.getCapabilities('audio') and look for nack in Opus's
rtcpFeedback list. Chrome → true → advertise. Firefox → false → omit.
Cached after first call (capabilities are static per UA).
Chrome → Chrome: NACK advertised, both sides honor it ✓
Chrome → Firefox: NACK advertised, Firefox ignores (no NACK requests) ✓
Firefox → Chrome: NACK omitted, Chrome never NACK-waits → no skip ✓
Firefox → Firefox: NACK omitted, both sides ignore ✓
No LCD across mixed rooms — each side gets the best contract its own
browser can honor. Logs "audio NACK as sender: on|off" once per session
for diagnostic.
Audio NACK was advertised in preferStereoOpus(). Chrome receivers will
request retransmits when they see rtcp-fb:nack on the Opus PT, but
Firefox senders never reply (Mozilla never shipped the responder side).
Chrome's jitter buffer waits for recovery that never arrives, then
skips — audibly choppy in mixed Firefox→Chrome rooms.
useinbandfec + 700ms playoutDelayHint already cover the loss case
without protocol churn. Comment block warns future code not to re-add
audio NACK without verifying both sides actually implement it for the
negotiated PT.
Keeps: 510k Opus ceiling, fullband pin, cbr=0, maxptime=120, speaker
picker — none of those are the choppy source.
Codec headroom (Tier 1):
- Opus 256 kbps → 510 kbps (spec ceiling) at every music-mode site:
mic publish (SFU + mesh), screen audio, game-share audio. Both
SDP maxaveragebitrate and RTP-level encodings[0].maxBitrate raised.
- Pin Opus fullband: maxplaybackrate=48000, sprop-maxcapturerate=48000
so BWE pressure can't opportunistically narrow to 16/24 kHz.
- cbr=0 explicit (VBR — Opus only spends what it needs).
Loss resilience (Tier 2):
- Audio NACK feedback (a=rtcp-fb:<opus_pt> nack) injected after Opus
rtpmap. Reactive packet recovery, near-zero overhead, ignored by
browsers that don't honor it.
- a=maxptime:120 in music mode advertises tolerance for larger frames
from peers (more encode context = cleaner music at same bitrate).
Playback chain (Tier 3):
- New speaker output selector (setSinkId) so listeners can route peer
audio to studio monitors / external DAC. Hidden on Safari (no
setSinkId on HTMLMediaElement). Persists to localStorage; applied
to every remote <audio> on creation and on user switch.
Live sessions need leave+enter to pick up the new SDP (per CLAUDE.md —
existing RTCPeerConnections are locked to whatever was negotiated at
creation time).
make stamp updates web/zebra-spaces.html footer date + hashes.
Server-side counterpart to proxy.unturf.com edb1009. Page sends
{type:"alive"} every 15s via setInterval on the signal WS; backgrounded
tabs throttle this timer hard (1Hz Android, ~1/minute iOS) so a
truly frozen tab reliably misses 3 heartbeats and the server's
aliveTTL=45s fires.
Pairs with the just-shipped pagehide-bye for the two distinct
failure modes:
- clean tab-close → pagehide fires bye → instant strong-leave
- dirty close / OS-backgrounded → JS heartbeat stops → server
aliveTTL fires after ~45s → standard hiccup grace → peer-left.
Two moderation actions, two buttons:
- kick : evict the session, allow rejoin
- ban : evict + block the pubkey (old 'boot' semantics)
Renders 'X was kicked by Y' vs 'X was banned by Y' off the new
peer-booted.action field. Self-notice text follows the same split.
pagehide / beforeunload now sends a synchronous 'bye' so a closed
tab counts as a strong-leave (peer-left + SFU evict immediately)
instead of waiting 8s for the hiccup grace. Fox 2026-06-03 —
"will closed tab on phone but the audio kept playing until i
kicked him out". With the new bye on pagehide his tab-close will
trigger the same fast-path the leave button does.
beforeunload kept as a fallback for older desktop browsers that
fire it before pagehide; pagehide is the cross-mobile primary.
Live-room defect (2026-06-03): Will refreshed his browser; the
supplant fired a SECOND ontrack for kind=camera pub=Will at +16s.
The page swapped srcObject on the existing <video>, called play(),
got 'fetching process for the media resource was aborted by the
user agent at the user's request' — browser refused to start a new
playback session on the same element after its autoplay grant had
already been consumed. Tile sat black on the moderator's screen.
Fix: swapFreshVideoElement() — on supplant, replace the <video>
with a freshly-built one carrying the same attrs. A brand-new
<video> is eligible for muted-autoplay even when the prior one had
its play() rejected, so the supplant lands cleanly without needing
a tap. Applied to both the thumbnail and the spotlight tile.
Mesh test harness extracts the helper so renderVideoTile still
runs end-to-end under the sandbox.
- moderation: 'boot' button onclick now .catch'es and logs server
errors. Server now returns 'boot target not found (stale uuid?)'
instead of silently no-op'ing when the page's member roster lagged
the room — the moderator was clicking 'boot' and seeing nothing
happen because their page held a stale uuid.
- diagnostic: watchFirstFrame logs 'still black after 2500ms — no
keyframe?' on any fresh SFU video track (screen/camera/game) whose
decoder never unmutes. Pairs with the SFU's extended kfBurst so we
can tell next session which path actually broke when a camera tile
renders black.
multi-peer-mesh test harness extracts the new fn alongside
watchVideoTrackForRemoval so handleRemoteSfuTrack still runs end-to-end
in the headless sandbox.
Stamp date refresh on the other pages (no behavior change).
Telemetry showed Willdabeast's mesh PC (peer edc257b4...) failing every
~20s in a perfect loop for 4+ minutes — the old 1.5s flat retry was
firing connectToPeer over and over against a NAT we couldn't traverse.
Each cycle ate CPU, network, signaling churn, and contributed to the
robot-voice + lost-mic noise we've been chasing.
New behavior:
- 2s -> 4s -> 8s -> 16s backoff (capped) between attempts
- Hard cap at PEER_MESH_MAX_RETRIES = 4 attempts. After that,
peerMeshGiveUp.add(uuid): connectToPeer becomes a no-op for that
uuid and audio rides on SFU permanently.
- Successful 'connected' transition resets the retry counter so a
much-later transient blip gets a fresh budget.
- peer-left / peer-booted clear all retry state via
clearMeshRetryState() so a rejoin from the same uuid starts over.
Log lines surface the decision so the next pathological mesh peer
shows up in CLIENT_LOG as 'mesh failed Nx — giving up, audio stays
on SFU' rather than a wall of identical 'failed — reconnecting' lines.
Side benefit: less mesh churn means the SFU mic fallback path (the
2f1e481 + 441c086 fix) gets to settle and stay settled, so receivers
don't keep flipping their audio elements between mesh and SFU
streams under a failing peer.
Were in the right column inside .controls, which on mobile cramped the
latency rows and squeezed the share URL. Now they sit as siblings of
sec-log below the .page grid — full viewport width everywhere. Order
top-to-bottom under .page:
latency
link
log
Same DOM whether the controls panel is shown or hidden, so the
mobile + hide-panel layout no longer needs the cameras column to
absorb them awkwardly.
The role-change notices (you-are-now-a-speaker / stepped-down / moved-
to-listener / promoted-to-cohost / now-the-host) sat above the room
list — top-of-page real estate for what's actually a personal status
update. Move sec-notice to sit BELOW sec-listener-actions so the
message reads in the same gaze as the controls it affects.
Was: leave button cleared MUTE_STATE_KEY, so a rejoin via bless-reclaim
came back unmuted regardless of prior choice. Fox: 'when I leave and
rejoin the mic state is unmuted.'
Now: keep MUTE_STATE_KEY in sessionStorage on leave (same tab, same
identity → same preference). Next promotion or bless-reclaim restores
via applyMuteState, matching the hard-refresh path's behavior. The
local in-page `muted` variable still syncs to whatever's in storage
so the dormant-tab state is consistent.
Two more Web Audio tones, completing the doorman set:
- playToneLeave: E5 -> C5 descending major third (mirror of join's
C5 -> E5 ascending). Fires on peer-left (any role leaves the
room post hiccup-grace) and peer-booted (mod kicks someone),
skipping self on the boot case.
- playToneStepDown: single G4 low sine blip — softer than the
arrival/departure pair. Fires on role-change for OTHER peers
when canSpeak(prev) && !canSpeak(next): they're still in the
room, just lost mic privileges.
Full notification ladder:
C5 -> E5 guest arrived
A5 (triangle) hand raised
G4 speaker stepped down (still here)
E5 -> C5 guest left / was kicked
Was: tones skipped when document.hidden. That's wrong for the doorman
cohost use case fox flagged — the whole point is to know a guest just
arrived without watching the tab.
Remove the document.hidden gate. Also defensively resume() the
AudioContext if it slipped to 'suspended' (some mobile browsers
suspend on tab-hide) — keeps the timeline alive so the scheduled tone
actually fires.
Reports of robot-voice artifacts from a peer on a different Wi-Fi —
classic symptom of jitter-buffer under-runs (Opus PLC kicks in,
generates synthetic samples that sound vocoder-y). 400ms was tight
enough that bursty inter-arrival on cross-AP / cellular hand-off
networks would exhaust the cushion.
700ms costs barely-audible end-to-end latency (well under the 1s
people perceive as 'walkie talkie'), well above typical Wi-Fi peak
jitter, and matches both the SFU subscribe path and mesh peer path
so the room sounds consistent regardless of how the audio is being
delivered.
Bitrate / FEC / DTX settings unchanged — adding bandwidth would
make a lossy link worse, not better.
Two Web Audio synthesised tones, no asset shipping required:
- playToneJoin: C5 -> E5 ascending major-third chime (~340ms total),
fires on every peer-joined. Server broadcasts peer-joined to
everyone except the joiner via broadcastExcept, so existing members
always hear it for new arrivals only.
- playToneRaise: single A5 triangle blip (~260ms), fires on
hand-raised UNLESS it's the raiser's own hand (server broadcasts
hand-raised to all so we filter on m.uuid !== myUUID).
Tones share a lazily-instantiated AudioContext (the user already
gestured to enter the room, autoplay policy is satisfied). Skipped
when document.hidden so background tabs stay quiet. Short
exponential attack + decay envelope avoids click artifacts.
Two new logLine calls in the SFU mic ontrack handler:
- 'sfu mic skipped for <uuid> — mesh peer connected' when the mesh is
serving (expected silence)
- 'sfu mic taking over for <uuid> — mesh state=<failed|disconnected|
closed|...>' when SFU stream attaches over a dying mesh PC
Makes the 'why don't I hear them' question one grep away in CLIENT_LOG
next time the cascade happens.
CLIENT_LOG telemetry showed the phone re-promoted to speaker, published
to SFU successfully, but neither the host nor the other speaker heard
them. Cause: the page's ontrack handler skipped attaching SFU mic
whenever peers.has(uuid) — even if that peer's mesh PC was in
'failed' or 'disconnected' state from an earlier role-change cycle.
Fix two paths:
1. ontrack-side: only skip SFU mic when peers.get(uuid).connectionState
is actually 'connected'. A stale entry or a failing PC no longer
blocks the SFU fallback; receiver hears the publisher via SFU until
mesh actually delivers.
2. mesh-fails-side: when an existing mesh PC transitions to 'failed',
the audio element was bound to the dying mesh stream. Reach into
sfuStreamsByPubHex and re-attach the cached SFU stream so the
listener hears continuous audio while the mesh reconnect runs in
the background, instead of a silent gap.
Mesh stays the preferred path when it's actually working — only
takes over the audio binding via its own ontrack when 'connected'.
Bug surfaced in CLIENT_LOG telemetry from a real session: when the
host's screen publish PC hit ICE failure, the watchPublishPC auto-
rebuilder called sfuPublishScreen() which immediately failed with
'getDisplayMedia requires transient activation from a user gesture.'
The auto-call path has no click, so getDisplayMedia can never succeed
from there.
Match the game-share pattern: on 'failed' tear down the dead PC + log
a clear 'tap share screen to re-share' message + clear state via
sfuUnpublishScreen(). User clicks the share button → picker opens (the
click IS the gesture) → fresh publish PC.
Mic + camera keep their auto-rebuild because getUserMedia honors the
persisted permission grant, no gesture needed.
Boot is the noisier sibling of demote — booted users can't publish
either, so their video tiles must come down immediately. The SFU-side
eviction (server /internal/block + OnConnectionStateChange) propagates
'ended' eventually but lags enough that a kicked speaker's camera tile
stayed visible after the boot. Authoritative removal by pubkey, same
pattern as the role-change-demote fix from 1e94fa6.