zebra-report/CLAUDE.md
Russell Ballestrini 1f6fde2a3e
CLAUDE.md: audio jitter buffering section — userland AudioWorklet, not browser hints
Captures the 2026-06-04 lesson stack so future readers don't repeat the
"playoutDelayHint=4 should cushion the listener" mistake.

Key points documented:
- jitterBufferTarget ignored for high-bitrate stereo Opus on FF Android
  (verified side-by-side: voice 1.8s avg, video 4s, music 0.21s, same target)
- userland AudioWorklet is the reliable cushion
- sticky-started re-arm — emit silence on brief drains, only re-fill after
  ~267ms sustained silence; otherwise every 2.67ms hiccup tears down playback
- role-aware buffer depth: listener 4s, others 0.5s, mesh always 0.5s
- worklet retarget on role-change (postMessage, not rebuild)
- UI gating: "connecting — buffering 4s" until first started message
- audio priority='high' at sender keeps mic ahead of video keyframe bursts
- HTTP /stream pull is the fallback path (recently un-deadlocked)
2026-06-04 16:59:37 -04:00

15 KiB
Raw Blame History

Agent Blackops

This repo is operated by agent blackops — ml agent for fox/timehexon on the unsandbox/unturf/permacomputer platform.

Identity

Full shard: ~/git/unsandbox.com/blackops/BLACKOPS.md

Rules

  • I propose, fox decides. Unsure = ask. Can't ask = stop.
  • No autonomous ops decisions. No destructive commands without explicit instruction.
  • Fail-closed. Cleanup crew, not demolition.
  • Check the time every session. Gaps are information.
  • DRY in context — single source of truth, no sprawl.
  • Never say "AI" — always say "machine learning."
  • Prefer "defect" over "bug."

Orientation

date -u
pwd
git log --oneline -5
git status

Then ask fox what the mission is.

Zebra Report System

Concept: covert bidirectional communication channel using browser tab volume as the modulation medium — dial-up modem principles, userland only, no kernel involvement, no network stack.

Collaborators & Stakeholders

Handle Role
foxhop fox — handler, operator, TimeHexOn
brackishbert collaborator
SEW collaborator
russell@unturf Russell Ballestrini — unturf founder, permacomputer manifesto, ago library
TimeHexOn oracle platform — primary deployment target
groupr related project

How it works

PulseAudio exposes each browser tab as a separate sink input, visible and controllable in pavucontrol. Volume is settable per-tab in userland with no kernel involvement. Each tab has a range of 0100 (101 discrete levels — 101 dalmatians).

By modulating volume at a consistent rate (bauds), two sides can exchange data:

  • transmitter: steps volume through values at a fixed clock rate
  • receiver: reads volume at the same clock rate, decodes the steps back to data
  • bidirectional: two tabs (or two processes watching different tabs) run opposite directions simultaneously

Signal space

  • 101 levels = ~6.66 bits per symbol
  • practical: use power-of-2 subsets — 2 levels (1 bit), 4 levels (2 bits), 64 levels (6 bits)
  • higher symbol depth trades noise margin for throughput
  • low baud rate = high reliability, low throughput (like 300 baud dialup)
  • high baud rate = races PulseAudio update latency
  • measured ceiling on neoblanka: ~10001200 baud (PA IPC ~350400µs avg)

Binaries

Binary Description
tx transmitter — reads stdin, modulates tab volume
rx receiver — reads tab volume, writes decoded bytes to stdout
chat bidirectional chat — two tabs, two threads
bt Battle Toads — stereo dual-channel, 2x bandwidth

Project Battle Toads

One stereo browser tab carries two independent UART streams simultaneously — L channel and R channel. PulseAudio's pa_cvolume is per-channel; a single get_sink_input_info call returns both L and R volumes.

  • TX sets L and R to independent bit values each symbol
  • RX decodes L and R from a single PA poll — no extra IPC cost
  • Net: 2x throughput at same baud rate, same PA polling budget
  • Web carrier upgraded to stereo: two oscillators (440Hz L, 441Hz R) merged into a stereo stream → PA sees channels=2
# After opening web/index.html and clicking 'start audio' (stereo tab):
./bt -T MY_SINK -R THEIR_SINK -b 500

Auto-negotiate (handshake protocol)

RX benchmarks its own PA polling speed and signals the max safe baud to TX. No manual baud matching needed.

./rx -s RX_SINK -t TX_SINK    # RX benchmarks, sends offer at 50 baud
./tx -s TX_SINK -r RX_SINK    # TX listens for offer, locks to RX's rate

Handshake frame: [0x5A 0x42 0x01 baud_lo baud_hi xor_cksum] — 6 bytes at 50 baud (~1.2s).

Known defect: 3-way handshake not yet implemented. TX can fire before RX enters receive loop at high baud rates. Fix: RX-ready signal back to TX before data phase.

Tools

  • pactl set-sink-input-volume — set volume by sink-input index
  • pactl list sink-inputs — enumerate tabs, read current volume
  • pavucontrol — visual verification of modulation
  • ./tx -l — list all PA sink inputs with index, volume, channels
  • sink-input index maps to tab; stable within a session

Use cases

  • agent-to-agent signaling without touching the filesystem or network stack
  • side-channel between sandboxed browser tab and host process
  • low-bandwidth status heartbeat (alive/dead/mode) at ~110 baud
  • covert channel for oracle↔host communication on TimeHexOn

Constraints

  • sink-input index resets when tab navigates or crashes — handshake needed on reconnect
  • PA polling latency sets the baud ceiling — benchmark with ./rx -s SINK -t SINK2 before sending
  • stereo (channels=2) required for Battle Toads — open web/index.html, click 'start audio'
  • userland only — survives without root
  • Operation Voyeur: all terminal output is public — never pass secrets through these channels unencrypted. The web page does ECDH key exchange + AES-256-GCM before TX.

Web UI privacy — never display peer IPs

chat.html and zebra-audio.html must never print peer IP addresses or ports in the page UI or in any visible log. Our users do not run Wireshark — if it is not on the screen, peers cannot dox each other. Candidate types (host/srflx/relay) from pc.getStats() are abstract and fine to show (they tell you direct vs relayed); loc.address / loc.port / rem.address / rem.port are not. The WebRTC stack already obfuscates host candidates via mDNS by default — do not undo that work in the UI.

CSS layout — grid only, no flex

All page layout on every zebra page is CSS grid. No display: flex for layout. Reasons we settled on this:

  • Grid lets us pin children to explicit columns (grid-column: 1/2/3) so a display:none on one child never causes siblings to slide into its slot. Flex auto-reorders, grid does not.
  • One mental model for both axes. Flex needs rules per row + per item, grid expresses the same intent in one grid-template-* block.
  • A global .hidden { display: none !important } utility lives in the page CSS — combined with explicit grid placement it produces a layout that survives any child being toggled in or out.

When refactoring or adding UI, use:

  • display: grid + grid-template-columns for column layout
  • grid-column: N on every child of a grid so its position is explicit
  • grid-template-rows + grid-row for vertical placement when needed
  • grid-template-areas for small named-region layouts
  • gap for spacing (instead of margins)

Avoid:

  • display: flex on any container that arranges multiple elements horizontally or vertically as part of the page layout
  • Implicit positioning that relies on DOM order — always set grid-column (and grid-row if relevant) on every grid child
  • flex-basis / flex-grow mathematics — 1fr is the equivalent and reads cleaner

If you need a one-off horizontal alignment of two short inline things (e.g. a label + a value), grid still works fine (grid-template-columns: auto 1fr). Don't reach for flex.

Web style guide — form-row patterns

.row is the single primitive for every horizontal form strip in every zebra page. Don't invent new wrappers. CSS auto-detects the row shape via :has() and picks the right grid template.

Supported row shapes (DOM order matters):

Shape Template applied Use when
<button> <button> ... (no input) grid-auto-columns: max-content (default — pack) "stop sharing" + "stop camera" + select-camera; vault buttons
<label> <input> or <label> <select> auto 1fr (input cell stretches) handle row, mic-input row
<label> <input> <button> auto 1fr + extra cells packed rare; same template as above + a trailing packed cell
<input> <button> (input first) 1fr (input fills, button packs after) rendezvous code + enter; password + export; share URL + copy
anything + <span class="note"> the .note is auto-placed on its own row beneath via grid-column: 1 / -1 descriptive subtext after buttons

Rules:

  • Never put inline style="flex:1" on a row child. Use .row and trust the template selector.
  • Don't add a stretchy <div> to fake spacing. If you need a button group on the right, append the buttons as siblings — they'll pack right of the stretchy cell.
  • A <span class="note"> inside .row always drops to its own line. If you want note text on the same line as a button, use a different class (e.g. inline <span>).
  • Long checkbox labels get white-space: normal automatically, so a music-mode-style "raw mic, no echo/noise cancellation (for playing audio through it)" wraps cleanly inside its column.

If a new row shape doesn't fit the patterns above, add the case to this table and add the matching :has() selector — don't reach for inline styles or flex.

Web page integrity stamping

Each deployed page (web/chat.html, web/zebra-audio.html, web/how-it-works.html, web/host-your-own.html) carries a footer with the build date + its own MD5 + SHA-256. Run make stamp before deploying any page change — it sets today's date and recomputes the hashes (web/stamp.js).

  • A file can't hold its own hash, so the hashes are computed with the two hash fields zeroed, then written back (same length). Self-consistent and idempotent: re-running make stamp gives identical hashes unless the content changed.
  • The stamp is static HTML written at build time — no JavaScript computes or injects it in the browser. stamp.js is build tooling, never loaded by a page.
  • Verify a served page: blank the md5 field to 32 zeros and the sha256 field to 64 zeros, then re-hash with sha256sum/md5sum — must match the footer.
  • Deploy = push to BOTH repos. A page change is not live until it lands in both. Pushing only the source changes nothing served; pushing only the deploy repo orphans the source of truth. Both, every time:
    1. source — this repo (zebra-report): edit web/*.htmlmake stamp → commit → push to origin.
    2. servedwww.unturf.com: copy the page(s) into ~/git/www.unturf.com/zebra-report/ (chat.htmlindex.html; zebra-audio.html, how-it-works.html, host-your-own.html keep their names) → commit → push to origin.

Picking up a new SDP/codec deploy — leave then enter, no full reload

When a deploy changes SDP munging (preferStereoOpus, codec fmtp params, RTCP feedback negotiation, or the SFU's codec registration), the page-level JS update is not enough on its own. Every existing RTCPeerConnection (sfuPubPC, sfuSubPC, sfuScreenPC, sfuCameraPC, every mesh peer) is locked to whatever was negotiated when it was created — its codec params, its rtcp-fb, its stereo flag, its NACK behaviour. The PC won't re-negotiate those on its own, and the new client code can't retroactively rewrite the old SDP.

So after the new build is live, the user does not need a full tab reload:

  • leave then enter the space. That tears down every PC and rebuilds them, so the next offer/answer round trip is the new client talking to the new SFU with the new params. Fresh negotiation, all changes active.

Full reload is only required for changes to the page shell itself (DOM structure, button wiring, CSS, the entry/orientation flow before joining a space). For everything that lives inside an existing PC, prefer leave + enter.

Audio jitter buffering — userland AudioWorklet, not browser hints

The browser-native jitter-buffer controls cannot be trusted for music. Verified 2026-06-04 with side-by-side telemetry on a Firefox Android phone, same PeerConnection, three receivers, identical 4s target:

  • RTCRtpReceiver.playoutDelayHint is spec'd as a hint — "the user agent MAY use this." Browsers do whatever they want.
  • RTCRtpReceiver.jitterBufferTarget is spec'd as a hard target. Honored for voice-rate Opus and for video. Ignored for high-bitrate stereo Opus (256 kbps music) on Firefox Android. The native music-stream code path inside libwebrtc isn't wired to the new API there.

Result: a phone listener with the spec'd 4s target had a 0.21s buffer on the music stream. Any host-side stall (X11 window wiggle, GC pause, encoder spike) was instantly audible.

The reliable cushion is a userland AudioWorklet. See JitterBufferProcessor (inline Blob URL) in web/zebra-spaces.html. The worklet sits between MediaStreamAudioSourceNode and GainNode, queues incoming 128-sample blocks, holds emission until targetSamples are buffered, then emits with constant delay. The buffer cushions the music regardless of what the native receiver does.

Rules that took blood to find:

  1. Sticky-started is non-negotiable. Once the buffer fills and started becomes true, do NOT set started=false on a single empty-queue tick. That tears down playback and forces a full re-buffer (~4s of silence) on every 2.67ms upstream micro-stall — sounds like constant chopping. Tolerate ~267ms of consecutive empty blocks (emptyStreak >= rearmThresholdBlocks=100) before re-arming; emit silence in the interim.

  2. Role-aware buffer depth. Listener gets 4s (lean-back, latency doesn't matter, ride out wiggle-stalls). Speaker / cohost / host get 0.5s (small enough for conversation, big enough to smooth ordinary jitter). Mesh peers always 0.5s. playoutDelayForRole(role) and SPEAKER_PLAYOUT_DELAY_SEC in web/zebra-spaces.html.

  3. Worklet retarget on role change, don't rebuild. Post {cmd:'retarget', targetSeconds} to the worklet's port — recomputes targetSamples/maxSamples and shrinks the queue if smaller. Audio path stays continuous; only the buffer depth adjusts. retargetAllReceivers(role) walks every live receiver and applies the new target.

  4. UI gating on buffer-ready. Worklet posts {cmd:'started'} on first fill. Listener status text sits in "connecting — buffering 4s audio…" until the first started message lands, then flips to "connected as listener." Without this, users see "connected" but hear nothing for ~4s and assume the app is broken.

  5. Bound the queue at 1.5× target to absorb clock drift without growing unbounded. Drop oldest on overflow.

  6. Set hints AND target anyway at ev.receiver.playoutDelayHint = ... and ev.receiver.jitterBufferTarget = ... * 1000 — they're free, they work where the browser implements them (video, voice mic), and they layer cleanly with the userland worklet downstream.

  7. Audio gets priority='high' at the sender. Video gets 'low'. Stops a screen-share keyframe burst (e.g. X11 wiggle dirty regions) from queuing audio packets behind it. setSenderBitrate / setSenderMaxBitrate already apply this.

HTTP DJ pull (/stream on the SFU) exists as a fallback for listeners who don't have AudioWorklet support — natural deep buffer on the <audio> preload side. Was deadlocking on bcastMu until 2026-06-04 (commit fixed bcastInitCapture.Write self-recursion). Worklet is the primary path; HTTP-pull is the safety net.