zebra-report/CLAUDE.md
Russell Ballestrini 1f6fde2a3e
CLAUDE.md: audio jitter buffering section — userland AudioWorklet, not browser hints
Captures the 2026-06-04 lesson stack so future readers don't repeat the
"playoutDelayHint=4 should cushion the listener" mistake.

Key points documented:
- jitterBufferTarget ignored for high-bitrate stereo Opus on FF Android
  (verified side-by-side: voice 1.8s avg, video 4s, music 0.21s, same target)
- userland AudioWorklet is the reliable cushion
- sticky-started re-arm — emit silence on brief drains, only re-fill after
  ~267ms sustained silence; otherwise every 2.67ms hiccup tears down playback
- role-aware buffer depth: listener 4s, others 0.5s, mesh always 0.5s
- worklet retarget on role-change (postMessage, not rebuild)
- UI gating: "connecting — buffering 4s" until first started message
- audio priority='high' at sender keeps mic ahead of video keyframe bursts
- HTTP /stream pull is the fallback path (recently un-deadlocked)
2026-06-04 16:59:37 -04:00

317 lines
15 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

# Agent Blackops
This repo is operated by **agent blackops** — ml agent for fox/timehexon on the unsandbox/unturf/permacomputer platform.
## Identity
Full shard: `~/git/unsandbox.com/blackops/BLACKOPS.md`
## Rules
- I propose, fox decides. Unsure = ask. Can't ask = stop.
- No autonomous ops decisions. No destructive commands without explicit instruction.
- Fail-closed. Cleanup crew, not demolition.
- Check the time every session. Gaps are information.
- DRY in context — single source of truth, no sprawl.
- Never say "AI" — always say "machine learning."
- Prefer "defect" over "bug."
## Orientation
```bash
date -u
pwd
git log --oneline -5
git status
```
Then ask fox what the mission is.
## Zebra Report System
**Concept**: covert bidirectional communication channel using browser tab volume as the modulation medium — dial-up modem principles, userland only, no kernel involvement, no network stack.
### Collaborators & Stakeholders
| Handle | Role |
|--------|------|
| **foxhop** | fox — handler, operator, TimeHexOn |
| **brackishbert** | collaborator |
| **SEW** | collaborator |
| **russell@unturf** | Russell Ballestrini — unturf founder, permacomputer manifesto, `ago` library |
| **TimeHexOn** | oracle platform — primary deployment target |
| **groupr** | related project |
### How it works
PulseAudio exposes each browser tab as a separate sink input, visible and controllable in `pavucontrol`. Volume is settable per-tab in userland with no kernel involvement. Each tab has a range of **0100** (101 discrete levels — 101 dalmatians).
By modulating volume at a consistent rate (bauds), two sides can exchange data:
- **transmitter**: steps volume through values at a fixed clock rate
- **receiver**: reads volume at the same clock rate, decodes the steps back to data
- **bidirectional**: two tabs (or two processes watching different tabs) run opposite directions simultaneously
### Signal space
- 101 levels = ~6.66 bits per symbol
- practical: use power-of-2 subsets — 2 levels (1 bit), 4 levels (2 bits), 64 levels (6 bits)
- higher symbol depth trades noise margin for throughput
- low baud rate = high reliability, low throughput (like 300 baud dialup)
- high baud rate = races PulseAudio update latency
- measured ceiling on neoblanka: ~10001200 baud (PA IPC ~350400µs avg)
### Binaries
| Binary | Description |
|--------|-------------|
| `tx` | transmitter — reads stdin, modulates tab volume |
| `rx` | receiver — reads tab volume, writes decoded bytes to stdout |
| `chat` | bidirectional chat — two tabs, two threads |
| `bt` | **Battle Toads** — stereo dual-channel, 2x bandwidth |
### Project Battle Toads
One stereo browser tab carries **two independent UART streams** simultaneously — L channel and R channel. PulseAudio's `pa_cvolume` is per-channel; a single `get_sink_input_info` call returns both L and R volumes.
- TX sets L and R to independent bit values each symbol
- RX decodes L and R from a single PA poll — no extra IPC cost
- Net: 2x throughput at same baud rate, same PA polling budget
- Web carrier upgraded to stereo: two oscillators (440Hz L, 441Hz R) merged into a stereo stream → PA sees `channels=2`
```bash
# After opening web/index.html and clicking 'start audio' (stereo tab):
./bt -T MY_SINK -R THEIR_SINK -b 500
```
### Auto-negotiate (handshake protocol)
RX benchmarks its own PA polling speed and signals the max safe baud to TX. No manual baud matching needed.
```bash
./rx -s RX_SINK -t TX_SINK # RX benchmarks, sends offer at 50 baud
./tx -s TX_SINK -r RX_SINK # TX listens for offer, locks to RX's rate
```
Handshake frame: `[0x5A 0x42 0x01 baud_lo baud_hi xor_cksum]` — 6 bytes at 50 baud (~1.2s).
**Known defect**: 3-way handshake not yet implemented. TX can fire before RX enters receive loop at high baud rates. Fix: RX-ready signal back to TX before data phase.
### Tools
- `pactl set-sink-input-volume` — set volume by sink-input index
- `pactl list sink-inputs` — enumerate tabs, read current volume
- `pavucontrol` — visual verification of modulation
- `./tx -l` — list all PA sink inputs with index, volume, channels
- sink-input index maps to tab; stable within a session
### Use cases
- agent-to-agent signaling without touching the filesystem or network stack
- side-channel between sandboxed browser tab and host process
- low-bandwidth status heartbeat (alive/dead/mode) at ~110 baud
- covert channel for oracle↔host communication on TimeHexOn
### Constraints
- sink-input index resets when tab navigates or crashes — handshake needed on reconnect
- PA polling latency sets the baud ceiling — benchmark with `./rx -s SINK -t SINK2` before sending
- stereo (channels=2) required for Battle Toads — open web/index.html, click 'start audio'
- userland only — survives without root
- **Operation Voyeur**: all terminal output is public — never pass secrets through these channels unencrypted. The web page does ECDH key exchange + AES-256-GCM before TX.
### Web UI privacy — never display peer IPs
`chat.html` and `zebra-audio.html` must never print peer IP addresses or ports
in the page UI or in any visible log. Our users do not run Wireshark — if it
is not on the screen, peers cannot dox each other. Candidate types
(`host`/`srflx`/`relay`) from `pc.getStats()` are abstract and fine to show
(they tell you direct vs relayed); `loc.address` / `loc.port` /
`rem.address` / `rem.port` are not. The WebRTC stack already obfuscates host
candidates via mDNS by default — do not undo that work in the UI.
### CSS layout — grid only, no flex
All page layout on every zebra page is **CSS grid**. No `display: flex` for
layout. Reasons we settled on this:
- Grid lets us pin children to explicit columns (`grid-column: 1/2/3`) so
a `display:none` on one child never causes siblings to slide into its
slot. Flex auto-reorders, grid does not.
- One mental model for both axes. Flex needs rules per row + per item,
grid expresses the same intent in one `grid-template-*` block.
- A global `.hidden { display: none !important }` utility lives in the
page CSS — combined with explicit grid placement it produces a layout
that survives any child being toggled in or out.
When refactoring or adding UI, use:
- `display: grid` + `grid-template-columns` for column layout
- `grid-column: N` on every child of a grid so its position is explicit
- `grid-template-rows` + `grid-row` for vertical placement when needed
- `grid-template-areas` for small named-region layouts
- `gap` for spacing (instead of margins)
Avoid:
- `display: flex` on any container that arranges multiple elements
horizontally or vertically as part of the page layout
- Implicit positioning that relies on DOM order — always set
`grid-column` (and `grid-row` if relevant) on every grid child
- `flex-basis` / `flex-grow` mathematics — `1fr` is the equivalent and
reads cleaner
If you need a one-off horizontal alignment of two short inline things
(e.g. a label + a value), grid still works fine
(`grid-template-columns: auto 1fr`). Don't reach for flex.
### Web style guide — form-row patterns
`.row` is the single primitive for every horizontal form strip in every
zebra page. Don't invent new wrappers. CSS auto-detects the row shape via
`:has()` and picks the right grid template.
Supported row shapes (DOM order matters):
| Shape | Template applied | Use when |
|---|---|---|
| `<button> <button> ...` (no input) | `grid-auto-columns: max-content` (default — pack) | "stop sharing" + "stop camera" + select-camera; vault buttons |
| `<label> <input>` or `<label> <select>` | `auto 1fr` (input cell stretches) | handle row, mic-input row |
| `<label> <input> <button>` | `auto 1fr` + extra cells packed | rare; same template as above + a trailing packed cell |
| `<input> <button>` (input first) | `1fr` (input fills, button packs after) | rendezvous code + enter; password + export; share URL + copy |
| anything + `<span class="note">` | the `.note` is auto-placed on its own row beneath via `grid-column: 1 / -1` | descriptive subtext after buttons |
Rules:
- Never put inline `style="flex:1"` on a row child. Use `.row` and trust
the template selector.
- Don't add a stretchy `<div>` to fake spacing. If you need a button
group on the right, append the buttons as siblings — they'll pack
right of the stretchy cell.
- A `<span class="note">` inside `.row` *always* drops to its own line.
If you want note text on the same line as a button, use a different
class (e.g. inline `<span>`).
- Long checkbox labels get `white-space: normal` automatically, so a
music-mode-style "raw mic, no echo/noise cancellation (for playing
audio through it)" wraps cleanly inside its column.
If a new row shape doesn't fit the patterns above, add the case to this
table and add the matching `:has()` selector — don't reach for inline
styles or flex.
### Web page integrity stamping
Each deployed page (`web/chat.html`, `web/zebra-audio.html`, `web/how-it-works.html`,
`web/host-your-own.html`) carries a footer with the build date + its own MD5 + SHA-256.
**Run `make stamp` before deploying any page change** — it sets today's date and
recomputes the hashes (`web/stamp.js`).
- A file can't hold its own hash, so the hashes are computed with the two hash
fields **zeroed**, then written back (same length). Self-consistent and
idempotent: re-running `make stamp` gives identical hashes unless the content
changed.
- The stamp is **static HTML written at build time** — no JavaScript computes or
injects it in the browser. `stamp.js` is build tooling, never loaded by a page.
- **Verify a served page**: blank the md5 field to 32 zeros and the sha256 field
to 64 zeros, then re-hash with `sha256sum`/`md5sum` — must match the footer.
- **Deploy = push to BOTH repos.** A page change is not live until it lands in
both. Pushing only the source changes nothing served; pushing only the deploy
repo orphans the source of truth. Both, every time:
1. **source** — this repo (`zebra-report`): edit `web/*.html``make stamp`
commit → **push to origin**.
2. **served**`www.unturf.com`: copy the page(s) into
`~/git/www.unturf.com/zebra-report/` (`chat.html``index.html`;
`zebra-audio.html`, `how-it-works.html`, `host-your-own.html` keep their
names) → commit → **push to origin**.
### Picking up a new SDP/codec deploy — `leave` then `enter`, no full reload
When a deploy changes SDP munging (`preferStereoOpus`, codec fmtp params,
RTCP feedback negotiation, or the SFU's codec registration), the page-level
JS update is not enough on its own. Every existing `RTCPeerConnection`
(`sfuPubPC`, `sfuSubPC`, `sfuScreenPC`, `sfuCameraPC`, every mesh peer) is
locked to whatever was negotiated when it was created — its codec params,
its rtcp-fb, its stereo flag, its NACK behaviour. The PC won't re-negotiate
those on its own, and the new client code can't retroactively rewrite the
old SDP.
So after the new build is live, the user does **not** need a full tab reload:
- **`leave` then `enter` the space.** That tears down every PC and rebuilds
them, so the next offer/answer round trip is the new client talking to
the new SFU with the new params. Fresh negotiation, all changes active.
Full reload is only required for changes to the page shell itself (DOM
structure, button wiring, CSS, the entry/orientation flow before joining a
space). For everything that lives inside an existing PC, prefer
`leave` + `enter`.
### Audio jitter buffering — userland AudioWorklet, not browser hints
**The browser-native jitter-buffer controls cannot be trusted for music.**
Verified 2026-06-04 with side-by-side telemetry on a Firefox Android phone,
same PeerConnection, three receivers, identical 4s target:
- `RTCRtpReceiver.playoutDelayHint` is spec'd as a hint — "the user agent MAY
use this." Browsers do whatever they want.
- `RTCRtpReceiver.jitterBufferTarget` is spec'd as a hard target. Honored for
voice-rate Opus and for video. **Ignored for high-bitrate stereo Opus
(256 kbps music) on Firefox Android.** The native music-stream code path
inside libwebrtc isn't wired to the new API there.
Result: a phone listener with the spec'd 4s target had a 0.21s buffer on the
music stream. Any host-side stall (X11 window wiggle, GC pause, encoder
spike) was instantly audible.
**The reliable cushion is a userland `AudioWorklet`.** See
`JitterBufferProcessor` (inline Blob URL) in `web/zebra-spaces.html`. The
worklet sits between `MediaStreamAudioSourceNode` and `GainNode`, queues
incoming 128-sample blocks, holds emission until `targetSamples` are
buffered, then emits with constant delay. The buffer cushions the music
regardless of what the native receiver does.
Rules that took blood to find:
1. **Sticky-started is non-negotiable.** Once the buffer fills and `started`
becomes true, do NOT set `started=false` on a single empty-queue tick.
That tears down playback and forces a full re-buffer (~4s of silence) on
every 2.67ms upstream micro-stall — sounds like constant chopping.
Tolerate ~267ms of consecutive empty blocks (`emptyStreak >=
rearmThresholdBlocks=100`) before re-arming; emit silence in the interim.
2. **Role-aware buffer depth.** Listener gets 4s (lean-back, latency doesn't
matter, ride out wiggle-stalls). Speaker / cohost / host get 0.5s (small
enough for conversation, big enough to smooth ordinary jitter). Mesh
peers always 0.5s. `playoutDelayForRole(role)` and `SPEAKER_PLAYOUT_DELAY_SEC`
in `web/zebra-spaces.html`.
3. **Worklet retarget on role change**, don't rebuild. Post
`{cmd:'retarget', targetSeconds}` to the worklet's port — recomputes
targetSamples/maxSamples and shrinks the queue if smaller. Audio path
stays continuous; only the buffer depth adjusts. `retargetAllReceivers(role)`
walks every live receiver and applies the new target.
4. **UI gating on buffer-ready.** Worklet posts `{cmd:'started'}` on first
fill. Listener status text sits in "connecting — buffering 4s audio…"
until the first started message lands, then flips to "connected as
listener." Without this, users see "connected" but hear nothing for ~4s
and assume the app is broken.
5. **Bound the queue at 1.5× target** to absorb clock drift without growing
unbounded. Drop oldest on overflow.
6. **Set hints AND target anyway** at `ev.receiver.playoutDelayHint = ...`
and `ev.receiver.jitterBufferTarget = ... * 1000` — they're free, they
work where the browser implements them (video, voice mic), and they
layer cleanly with the userland worklet downstream.
7. **Audio gets `priority='high'` at the sender.** Video gets `'low'`. Stops
a screen-share keyframe burst (e.g. X11 wiggle dirty regions) from
queuing audio packets behind it. `setSenderBitrate` / `setSenderMaxBitrate`
already apply this.
**HTTP DJ pull (`/stream` on the SFU)** exists as a fallback for listeners
who don't have AudioWorklet support — natural deep buffer on the
`<audio>` preload side. Was deadlocking on `bcastMu` until 2026-06-04
(commit fixed `bcastInitCapture.Write` self-recursion). Worklet is the
primary path; HTTP-pull is the safety net.