diff --git a/docs/TICKETS.md b/docs/TICKETS.md index 846602a..a0884c1 100644 --- a/docs/TICKETS.md +++ b/docs/TICKETS.md @@ -66,6 +66,7 @@ Newest first. Update on every open/close. | ID | Title | Status | Opened | Directive | |----------|------------------------------------------------|-----------------------|------------|-----------| +| #000043 | Empirical tightening of T3 bound constants (#000036 Phase 2) | open · awaiting go/no-go (parks until v7 deployment surfaces measurement data) | 2026-05-10 | — | | #000042 | Term-aliases table (vocabulary-mismatch bridge) | closed · 13 rows live across geometry + classical-physics + arithmetic domains by 2026-05-10 | 2026-05-09 | — | | #000041 | Citation-aliases table (PD substitutes for proprietary cites) | closed · 74 rows live as of 2026-05-10 (count grew 40 → 54 → 74; Goldstein/Newton, Mendelson/Enderton/Jech/Landau/Gödel→{Russell IMP, Russell PoM, De Morgan, Boole, Cantor, Peano, Dedekind, SF-LF}, Stanley/Brualdi/Knuth → Bogart+Levin+Keller-Trotter, Dummit-Foote/Barendregt/Böhm-Jacopini → Judson/PLFA/SF, Kolmogorov → Grinstead-Snell+Laplace) | 2026-05-09 | — | | #000040 | Phase 5 resolver fix — phrase + content-token cascade (Hilbert terminology mismatch surfaced) | closed · cascade landed 2026-05-09; lift blocked by 1902-vs-modern vocab; follow-up #000042 | 2026-05-09 | — | @@ -111,4 +112,4 @@ Newest first. Update on every open/close. ## Next ID -`000043` +`000044` diff --git a/docs/tickets/ticket-000043-t3-bound-constant-tightening.md b/docs/tickets/ticket-000043-t3-bound-constant-tightening.md new file mode 100644 index 0000000..2fab7c4 --- /dev/null +++ b/docs/tickets/ticket-000043-t3-bound-constant-tightening.md @@ -0,0 +1,265 @@ +# Ticket #000043 — Empirical tightening of T3 bound constants + +**Status:** open · awaiting go/no-go +**Opened:** 2026-05-10 +**Scope:** Replace the conservative-by-construction +``C_B1 = C_B2 = C_B3 = 1`` constants in +``docs/soft-hash-channel-t3-bound.md`` with deployment-measured +values where possible. Each tightening path is a separate +empirical experiment; this ticket organizes them as a +single deferred work item that opens once at least one of the +prerequisite measurements becomes available. +**Audience:** maintainers of the v7 plastic-training spec; #000036 +follow-up; downstream operators sizing nonce-window length under +real workload data. +**Hard constraint:** the framework in #000036 §6 (closed-form +bound) does NOT change. Tightening replaces specific constants; +the call sites of ``t3_bound_calculator.py`` are unchanged. + +--- + +## 1. Problem statement + +#000036 §10 explicitly lists three open questions for empirical +tightening of the T3 per-window covert-channel bound. Each one +is a self-contained measurement experiment but they share the +same downstream consumer (the calculator's ``--c-b1`` / +``--c-b2`` / ``--c-b3`` overrides). Capturing them as a single +ticket prevents the work from getting lost across three +unrelated tickets and clarifies the gating dependencies. + +The framework already lands as #000036 Phase 1; this ticket is +the **Phase 2 work item** that #000036 §7 status calls out as +"parks for v7 deployment data". This ticket exists so that +parking has a concrete name future blackops shifts can reference. + +--- + +## 2. The three tightening paths + +### 2.1 C_B1 — gradient-bias channel efficiency + +**Status quo**: ``C_B1 = 1`` (data-processing inequality +ceiling). + +**Tightening hypothesis**: real loss surfaces have reduced +adversary-controllable directions due to Hessian rank +deficiency. The effective bandwidth is bounded by the +*number of low-eigenvalue Hessian directions* the adversary +can steer cheaply, not by the full parameter dimension. + +**Measurement path**: feed from #000034 Phase 1b. Once a real +v7 checkpoint exists and the Hessian-alignment probe runs +against it, the alignment-score distribution gives a direct +estimate of `C_B1`: + +- ``STRUCTURAL_ALIGNMENT`` verdict ⇒ `C_B1` close to 1 + (φ_linear is exploitable; T3 bound near worst case). +- ``NO_ALIGNMENT`` verdict ⇒ `C_B1 ≈ rank(H_low) / dim_d`, + typically much less than 1. +- ``ANTI_ALIGNED`` verdict ⇒ `C_B1` floors lower still. + +**Prerequisite**: #000034 Phase 1b lands (real v7 checkpoint ++ alignment probe run). No measurement is meaningful until +v7 plastic-training has a deployment target producing +representative training data. + +**Acceptance**: a numeric `C_B1` value derived from a Phase 1b +alignment-probe run, with confidence interval, plumbed back +into the `t3_bound_calculator.py` defaults (or surfaced as a +documented per-deployment override). + +### 2.2 C_B2 — LR-trajectory diversity + +**Status quo**: ``C_B2 = 1`` (full categorical-channel +capacity). + +**Tightening hypothesis**: many of the `R` LR levels in a +typical schedule produce parameter trajectories that converge +to similar `Θ`. The loss surface's basin geometry damps +small-LR perturbations. The effective channel symbol count +is the *number of LR-distinguishable trajectories*, not the +full grid size `R`. + +**Measurement path**: per-deployment 1-Wasserstein distance +between final-checkpoint distributions across the LR grid. +Concretely: + +1. Train `R` parallel mini-replicas (small budget — same + architecture, fewer steps) under each LR level. +2. Compute pairwise 1-Wasserstein distances between the + resulting `Θ` distributions. +3. Cluster — the effective LR-symbol count is the number of + clusters separated by ≥ ε (deployment-specific tolerance). +4. ``C_B2 = log₂(effective_symbols) / log₂(R)``. + +**Prerequisite**: #000034 Phase 1b (same v7 checkpoint +infrastructure) plus the parallel-replica training framework. +Out of scope today; opens when v7 deployment ramp-up provides +the compute budget for parallel-replica training. + +**Acceptance**: a numeric `C_B2` value derived from +1-Wasserstein clustering on a representative deployment, with +the cluster count + per-cluster representative-LR documented. + +### 2.3 C_B3 — SGD shuffle regime + +**Status quo**: ``C_B3 = 1`` (adversarial-order assumption). +Already substantially tighter than the naive ``log₂(N_b!)`` +bound thanks to the Bottou-Bousquet refinement in #000036 +§5, but conservative for typical SGD which is random-shuffle. + +**Tightening hypothesis**: random-shuffle SGD has +``C_B3 = O(1/√N_b)``, several orders of magnitude tighter +than the adversarial-order assumption. Cyclic-shuffle SGD +has `C_B3 = O(1)` per epoch. Adversarial-order SGD is the +worst case. + +**Measurement path**: classify the deployment's SGD shuffle +regime: + +- If random-shuffle (`torch.utils.data.DataLoader(shuffle=True)` + or equivalent): ``C_B3 ≈ 1/√N_b``. +- If cyclic / fixed-order: ``C_B3 ≈ 1`` (adversarial-order + effectively realized passively). +- If adversarial-order is the threat model (T3 controls + ordering): ``C_B3 = 1`` (current default; can't tighten). + +**Prerequisite**: deployment configuration audit. Cheaper than +Phase 1b — answerable from the training script + DataLoader +settings, no checkpoint required. + +**Acceptance**: documented `C_B3` value per deployment shuffle +regime, with the audit method captured so future operators can +re-run the classification on changed schedules. + +--- + +## 3. Implementation sketch + +### 3.1 Calculator override surface + +Already landed in #000036 Phase 1. ``t3_bound_calculator.py`` +accepts ``--c-b1``, ``--c-b2``, ``--c-b3`` flags as floats in +``[0, 1]``. Phase 2 doesn't change the calculator's interface; +it just produces the numeric values to plug in. + +### 3.2 Per-deployment configuration capture + +When a v7 deployment lands, capture the three constants in a +per-deployment config: + +```yaml +# example: my-deployment-t3-config.yaml +phi_alignment_verdict: NO_ALIGNMENT # from #000034 Phase 1b +c_b1: 0.05 # from § 2.1 measurement +c_b2: 0.42 # from § 2.2 measurement +c_b3: 0.025 # 1/√1024 random-shuffle +notes: + - "phi_alignment ran 2026-XX-XX on commit XXXXXXXX" + - "lr_diversity measured 8 LR levels, 4 clusters → 2 / 3 = 0.42" + - "DataLoader uses shuffle=True with batch_size=1024 → 1/sqrt(1024)" +``` + +Operator passes the file via ``--c-b1``/``--c-b2``/``--c-b3`` flags +or via a new ``--from-config FILE`` flag (deferred — not on the +critical path until at least one deployment's measurements land). + +### 3.3 Update path back into #000036 + +Once a measurement set is captured for any deployment, the +``docs/soft-hash-channel-t3-bound.md`` § 7 numeric examples gain +a fourth subsection (§ 7.4) showing the deployment-specific +bound with its tightened constants. This becomes the +calibration evidence § 10's open questions resolve against. + +--- + +## 4. Out of scope + +- **Implementing v7 plastic training** — this ticket consumes + v7 deployment data; doesn't produce it. +- **Adversarial-training experiments** to validate the + bound itself — covered separately by a future + empirical-validation ticket per #000036 § 10 item 5. +- **Generalizing to B4/B5 control surfaces** (adaptive + optimizer state manipulation, etc.) — also covered by + #000036 § 10, separately ticket-able when those surfaces + surface. +- **Implementing M2 (per-checkpoint nonce) in the v7 codebase** + — separate ticket; this ticket assumes M2 is in place. + +--- + +## 5. Risks + +- **Measurement results don't tighten the bound usefully.** If + a deployment turns out to have ``C_B1 ≈ 1`` (genuine + STRUCTURAL_ALIGNMENT), the tightening fails for that + deployment and the conservative bound stands. Mitigation: + the failure case is itself useful information (#000018 + M1-priority signal); not a wasted experiment. +- **Per-deployment measurements don't generalize.** Each + deployment may need its own constant set. Mitigation: this + is a feature not a bug — operators get deployment-specific + guidance, not a one-size-fits-all bound. +- **Measurement noise.** Lanczos eigenvalue estimation, + 1-Wasserstein clustering, and shuffle-regime classification + all have noise. Mitigation: report constants with confidence + intervals; operators choose the upper-bound CI for safety + margin. + +--- + +## 6. Acceptance criteria + +1. At least one of § 2.1 / § 2.2 / § 2.3 measurements + completes against a real v7 deployment. +2. The resulting numeric constant(s) are captured in a + per-deployment config or a documented update to + ``t3_bound_calculator.py`` defaults. +3. ``docs/soft-hash-channel-t3-bound.md`` § 7 gains a § + 7.4 numeric example showing the tightened bound for the + measured deployment. +4. ``docs/soft-hash-channel-t3-bound.md`` § 10 item(s) + corresponding to the measured constant(s) flip from "open + question" to "resolved 2026-XX-XX via #000043 measurement + on ". + +Partial closure (any one of § 2.1 / § 2.2 / § 2.3 individually) +is acceptable. The ticket can stay in-progress with one or two +measurements landed and others still parked. + +--- + +## 7. Status + +**Open · awaiting go/no-go.** Doc-only spec for the empirical +tightening work. Lands when: + +- v7 plastic-training has at least one active deployment, AND +- Either #000034 Phase 1b lands (unlocking § 2.1) OR a + cheap-path operator audit lands (unlocking § 2.3). + +The full ticket parks until both conditions fire. § 2.3 alone +can land without v7 deployment if an operator audits an +existing arborist deployment's DataLoader configuration. + +Closure criterion: at least one of (§ 2.1 / § 2.2 / § 2.3) +measurements lands a numeric constant + the corresponding +``soft-hash-channel-t3-bound.md`` § 10 open question flips +to resolved. + +--- + +## 8. References + +- ``docs/soft-hash-channel-t3-bound.md`` § 6 (the closed-form + bound being tightened) + § 10 (open questions enumerated). +- ``bench/scripts/t3_bound_calculator.py`` (override surface + already in place). +- Ticket #000036 (parent ticket; this one's a Phase 2 sub). +- Ticket #000034 (Phase 1b unlocks § 2.1 of this ticket). +- Ticket #000018 (parent analysis; § 9.3 closes via #000036 + framework, would close further with #000043's tightened + constants).