From d720b73d91afa4fe62764d9403929abc6a776f47 Mon Sep 17 00:00:00 2001 From: "russell@unturf.com" Date: Sun, 10 May 2026 13:24:51 -0400 Subject: [PATCH] remove docs/dav1dprometheus-update-2026-05-09.md MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Per fox 2026-05-10 — file not needed. Was a point-in-time update note; superseded by the rolling research log in #000006 and the ticket-level status pins (#000028 closed in TICKETS.md, witness- sweep amend in #000006 etc.). No incoming references in the repo (grep clean). Note for the record: the immediately-prior commit (1421d96) made an internal-consistency fix to this same file. That fix is moot post-deletion but stays in the audit chain — describing the witness/carrier vocabulary distinction was useful at the time it shipped. --- docs/dav1dprometheus-update-2026-05-09.md | 822 ---------------------- 1 file changed, 822 deletions(-) delete mode 100644 docs/dav1dprometheus-update-2026-05-09.md diff --git a/docs/dav1dprometheus-update-2026-05-09.md b/docs/dav1dprometheus-update-2026-05-09.md deleted file mode 100644 index 6fa7d87..0000000 --- a/docs/dav1dprometheus-update-2026-05-09.md +++ /dev/null @@ -1,822 +0,0 @@ -# Arborist substrate update — for the Dav1DPrometheus framework legacy - -**To:** Dav1DPrometheus, the framework author -**From:** fox + agent blackops, on the unsandbox / unturf / permacomputer platform -**Date:** 2026-05-09 (UTC), Asia/Kuala_Lumpur - -A note in the spirit of dialogue with the framework you authored. -Your 5S/5F/5T/5R taxonomy is the spine of every benchmark we now -ship. This is what your framework grew into when we pulled it -into a Merkle-AGI-v9.8 substrate. - -> **Revision note (2026-05-10):** an external review surfaced -> several errata in the original (2026-05-09) draft. Corrections -> applied in-place: -> 5S vocabulary is **Syntax / Semantics / Syllogism / Synthesis / -> Semiotics** (not Surface/Substrate/…); 5R vocabulary is **React / -> Rearrange / Restore / Replicate / Resonate** (not React/Recall/ -> Reason/Refine/Restore); the π* registry chronology is now -> explicit (15 → 16 with `combinatorics@v1`); kernel-version -> immutability is called out as a hard invariant; cross-witness -> vs. cross-carrier is distinguished; `STRICT-WITNESSED` is scoped -> as a render label (not a new admissibility/cache mode); the -> warrant-promotion ladder distinguishes SOURCE-ANCHORED from -> EVIDENCE-WARRANTED; license tags are flagged as project-reported; -> the "no human labelers" property is scoped to canonical-shape -> domains only. Original draft preserved at commit `a2ff9d4`. -> -> The reviewer also flagged the sub-battery count as 20 vs the -> document's 21. The codebase has **21**: 5+6+5+5, because 5T -> carries 6 sub-batteries — the original SQD-whitepaper -> ``transfer`` plus the canonical Dav1DPrometheus five -> (Transfer Learning, Truthtables, Time, Triangulation, -> Transitivity), per ticket #000024 which preserved the legacy -> name alongside the authoritative wording. Earlier in this -> revision pass the count was incorrectly downgraded to 20 — that -> regression is corrected here. - ---- - -## What we built - -The arborist project — Python-only, content-addressed Merkle store -with a v9.8 audit chain — adopted your taxonomy in the first week -of May 2026 and landed the full benchmark surface in eight days. -Concretely: - -### The 21 sub-batteries are real, fixtures included - -Four axes × five canonical sub-batteries + 5T's preserved legacy -sub-battery = **21 sub-batteries**, each fixture-backed and runnable -through `bench/batteries/runner.py`: - -``` -5S — Syntax / Semantics / Syllogism / Synthesis / Semiotics - 5 sub-batteries (Phase 1a: Syntax + Semantics; Phase 1b: the - other three) carrier-aware (text / claim_lattice / prose), - plus per-π*-domain expansions of Syntax + Semantics that - bring the family wider than 5 fixtures-per-sub-battery -5F — Function / Finetuning / Falsification / Formulate / Feedback Loop - 5 sub-batteries × ~50 synthetic + ~50 live each (Phase 1b.2 - live wire-ups land on real arborist subsystems) -5T — Time / Truthtables / Transfer Learning / Triangulation / - Transitivity (canonical Dav1DPrometheus five) - + Transfer (legacy SQD-whitepaper sub-battery, kept alongside - Transfer Learning per ticket #000024) - 6 sub-batteries total -5R — React / Rearrange / Restore / Replicate / Resonate - 5 sub-batteries (workspace-operator shape, per SQD §9.3 — - deterministic, no LLM-as-judge) -``` - -(An earlier draft of this note mis-spelled 5S as Surface/Substrate/… -and 5R as React/Recall/Reason/Refine/Restore. The fixtures + -runner in `bench/batteries/` are the authority — corrected here.) - -Total deterministic-task surface: **662 tasks** in the default -runner (per the runner's `--all` enumeration), plus several -hundred more in extension batteries (math π* fixtures, real-shard -baseline, witness-divergence collection, claim-pack ingest -verification). - -Your original wording is honored — `Transfer Learning` (not the -SQD-whitepaper variant `transfer`), `Truthtables`, `Time` — pinned -as a memory entry so future agents respect the framing you chose. - -### The π* canonical-projection registry: now 16 kernels, no stubs - -The registry held 15 kernels before this session; `combinatorics@v1` -(#000032) brought it to 16 and closed the last reserved stub. -Listed in registration order: - -``` -text-domain - wikitext-base@v1 Wikipedia / wikitext → plain prose - claim-lattice@v1 Claim lines → JSON parsed-claim list - code-py-ast@v1 Python source → canonical AST S-exp -arithmetic / logic - arithmetic@v1 SQD §14.1 — exact rational num/den - logic-kernel@v1 SQD §14.3 — propositional → CNF -sensor / temporal - time-series-quantized@v1 SQD §13.5 — quantized integer vector -tabular - tabular-pinned@v1 JSON-rows → pinned-schema bytes - (last reserved stub graduated today) -math substrate (SymPy, [math] extra) - algebra-symbolic@v1 sp.expand + srepr - algebra-symbolic-simplified@v1 sp.simplify + srepr (collapses trig) - calculus-derivative@v1 sp.diff → algebra-symbolic - calculus-integral@v1 sp.integrate + unevaluated sentinel - calculus-limit@v1 sp.limit + ±∞ / complex-∞ pinning - calculus-series@v1 Taylor truncated, no O(x**n) - linear-algebra@v1 RREF / det / eigenvalues / inverse - function-sampled@v1 SymPy expr → time-series-quantized - bridge (this is what plotting CAN - become in π* terms — image bytes - aren't canonical, but the sampled - grid is) -combinatorics - combinatorics@v1 pure-integer counting kernel - (#000032; tighter than algebra- - symbolic — fails closed on any - non-non-negative-integer result) -``` - -Every π* registers via `name@version`; SHA-256 of the canonical -bytes is the equivalence-class identity. Two inputs that mean the -same thing produce identical bytes; auditing reduces to byte -comparison. - -**Kernel-version immutability is a hard invariant.** A registered -`name@version` is behaviorally frozen — any change in canonicalization -output is a new version (`arithmetic@v2`), never an in-place patch. -Otherwise prior persisted canonical-cache rows would silently -become semantically unstable. The registry rejects re-registration -of the same key with a different implementation; the -governance_policy_hash also folds in the active kernel manifest -so flipping kernels invalidates prior records on lookup. - -### Cross-witness vs. cross-carrier discipline - -Two distinct axes, kept separate in the substrate vocabulary: - -- **Witness channels**: kernel / cache / LLM. A canonical-shape - question can be answered by ≥1 of these and they're compared - byte-for-byte after canonicalization. Cross-**witness** agreement - is the warrant strengthener that today's pipeline measures. -- **Carrier modalities**: text / claim_lattice / code / arithmetic / - logic / time-series / tabular / symbolic-algebra / calculus / - linear-algebra / function-sampled / combinatorics. These are - the surface forms an input/output can take. Carrier-aware fixtures - pin the carrier so a verifier change doesn't drift across them. - -Every benchmark fixture and audit-bound canonicalization names -its kernel via `pi_star_ref`. A `PHASE_1_CARRIERS` whitelist gates -which carriers can flow through Phase-1 paths; unsupported -carriers fail explicitly with `reason="unsupported_carrier"`, -never silent acceptance. Hidden-channel work is defensive only — -detection / flagging, never generation or concealment. - -(Earlier drafts of this note used "multi-modality" and "multi- -witness" interchangeably; they're not. Kernel/cache/LLM are -witnesses over the same canonical bytes; text/image/audio/world -will be carriers as π* grows. The pattern generalizes from one -to the other but the audit semantics are distinct.) - -### Persistence + audit chain - -Three big mechanical artifacts beyond the kernels: - -1. **Canonical projections persist to providence_cache** — math - answers (`0.1 + 0.2` → `3/10`) get a v9.8 8-dim cache_key and an - audit-event entry, just like RAG-derived answers do. The kernel - is the source; the cache row is the receipt. Re-asking hits in - ~40ms; chain-check verifies the chain stays intact. - -2. **Cross-witness agreement on canonical-shape questions** — on - canonical-shape questions, kernel + cache + LLM fan out in - parallel and we record the agreement matrix. The persisted - row's `audit_mode` stays `CANONICAL_PROJECTION` regardless; - when all three witnesses byte-equal, the **render layer** - surfaces `STRICT-WITNESSED` as a label on top of the underlying - audit_mode. When they diverge: high-value falsification data - for 5F. The LLM is a witness, never authority — the kernel - stays ground truth. Capital ledger records witness cost so - ForkScore can compare witness-on vs witness-off forks honestly. - **Real-world divergence rate** on the first canonical-question - sweep against Hermes-3-8B: **5 of 8 questions diverged** (62.5%). - Hermes said `1/10` for `0.1+0.2`; said `TRUE` for `A IMPL B`; - gave the unexpanded form when handed the expanded one. Each - divergence becomes a 5F-Falsification calibration fixture - downstream prompt-tuning can grade against. See - "Witness pipeline flowing end-to-end" below. - - `STRICT-WITNESSED` is purely a **render label**, not a new - admissibility/cache mode. The persisted `audit_mode` column - stays `CANONICAL_PROJECTION`; the witness audit event - (`providence_canonical_witness`) layers on top so cache_key - semantics don't drift. Programmatic callers see the underlying - `CANONICAL_PROJECTION`; human-facing surfaces see the label. - -3. **Authorship warrant ladder** — sidecar classifier on a 6-tier - strength scale (`AUTHOR_PACKAGE_METADATA` → `AUTHOR_REPOSITORY_OWNER` - → `AUTHOR_PAGE_BYLINE` → `AUTHOR_PRIMARY_PAGE_TITLE` → - `AUTHOR_COPYRIGHT_FOOTER` → `AUTHOR_SECONDARY_SOURCE`). Wired - into `arborist inspect` + audit-line render-tail. Surfaces *how* - the cited evidence supports an authorship claim, not just whether - it does. - -### Selection (Merkle-AGI v8 ForkScore) - -Phase 1a + 1b live: ScoredFork dataclass + `arborist substrate score` -CLI surface + Makefile harness (`bench-fork-baseline`, -`bench-fork-score`). A weighted score over (parent, child) battery -deltas; verdicts ACCEPT / MARGINAL / REJECT; hard-regression flag -at 5pp drop; NEG_INF_REGRESSION flag from the efficiency vocabulary. -CI-gateable: REJECT exits 1. - -### Content corpus — claim-pack source landed (#000029, this session) - -Two companion JSON bundles dropped into the substrate this session: -**`axiomsg4-v2.json`** + **`theoremsg4-v2.json`** — Grok-4-generated -"Prometheus Maths Engine" packs covering 7 pillars (Logic, Set -Theory, Arithmetic, Geometry, Probability, Classical Physics, -λ-Calculus). 78 atomic claims (55 axioms + 23 theorems), each -dual-threaded (Δ symbolic LaTeX + ∇ verbose prose) and self-citing -to a stable classical text (Mendelson, Enderton, Hilbert, Newton, -Kolmogorov, Łukasiewicz, Aristotle). - -`arborist/sources/claim_pack.py` ingests them at the right grain — -**one Document per axiom/theorem record**. Lenient JSON parser -strips markdown fences and double-escapes lone LaTeX backslashes -(`\Theta`, `\heart`, `\vec`) without corrupting already-correct -`\\to` pairs. Cross-bundle pillar-level provenance arrays -(`theoremg4.json:pillar.I.logic.excludedMiddle` ↔ -`axiomsg4.json:pillar.I`) become outbound `pillar_reference` -edges on the first record of each pillar. - -Smoke test: `arborist ingest --source claim_pack --bundle -axiomsg4-v2.json --bundle theoremsg4-v2.json` → 78 docs, 78 -chunks, 14 deduped pillar-reference edges, 78 audit events, -10/10 sampled Merkle proofs verify. - -Honest ceiling: every record lands `kind='surface'`. The pack is -*pre-distilled content* but its provenance is asserted (string -field), not Merkle-proven. Records max out at **ANCHOR-WARRANTED** -on the four-rung ladder until cited textbooks are themselves -ingested as surfaces and a `derivations.proof_blob` row is -computed per record. That gap is opened as ticket #000031 below. - -**Retrieval lift measured the same day** — apples-to-apples FTS5 -search comparison on a single shard (`000.db`, ~867 K Wikipedia -docs) before vs after claim-pack ingest. Four representative -queries; **3/4 show claim-pack record in the top-3**: - -``` -modus tollens → claim-pack at #3 (BM25 31.77) -law of excluded middle → claim-pack at #3 (related axiom) -associativity of addition → claim-pack at #1 — displaces - Wikipedia's general "Addition" - article entirely -Bayes theorem → no top-3 lift; Wikipedia's - "Bayes" + "Bayes rule" articles - dominate via short-doc BM25 - bias. Body-coverage sqrt rerank - in the full QA pipeline (not - exercised by FTS5-only search) - would likely surface it. -``` - -Storage tax: sub-MB — 78 chunks against 6.2M existing chunks is -below filesystem allocation granularity. **Claim-pack is earning -its tax for narrow-technical queries** (where pre-distilled -CORE-shape content has title-token advantage) but doesn't help on -queries Wikipedia already covers with focused articles. Full -journal: `bench/results/claim-pack-retrieval-lift-2026-05-09.md`. - -### `combinatorics@v1` π* — pure-integer counting kernel (#000032) - -A new kernel landed today: tighter sibling of -`algebra-symbolic@v1`. Same input parser, narrower output: any -result that isn't a non-negative `sp.Integer` raises -`PiStarError`. Plain decimal output (`b"10"`); composes with -`arithmetic@v1` for byte-identical agreement (`b"10/1"`) so the -multi-witness canonical-agreement pipeline can pin equivalence- -class agreement when both kernel routes fire on the same question. - -The 16-kernel registry now covers: text → claim_lattice → code → -arithmetic → logic → time-series → tabular → symbolic-algebra → -symbolic-algebra-simplified → calculus-derivative → calculus- -integral → calculus-limit → calculus-series → linear-algebra → -function-sampled → **combinatorics**. No remaining reserved -stubs; the registry chapter is closed. - -Pillar VII for the claim-pack (combinatorics axioms + theorems — -ticket #000033) is the natural follow-up; sequencing puts the -kernel first so pillar VII records bind to it from day one and -avoid `pi_star_ref` rebind churn. - -### Witness pipeline flowing end-to-end (#000028 validated) - -The witness fan-out now writes a `providence_canonical_witness` -audit event when it fires. A new extractor -(`bench/scripts/witness_to_5f.py`) reads those events and emits -divergence-only entries as 5F-Falsification fixtures matching -the existing `falsification-live-v1` schema. End-to-end smoke -ran 8 canonical-shape questions (3 arithmetic + 3 logic + 2 -algebra) against `~/.arborist/shards` plus the actual Hermes -endpoint: - -``` -agreement label count rate -KERNEL-LLM-DIVERGED 5 62.5% -KERNEL-LLM-AGREE 3 37.5% -───────────────────────────────────────── -divergence_count 5 62.5% -wall median / max 130 ms / 1.1 s -``` - -Five real Hermes hallucinations on questions with closed-form -ground truth: `1/10` for `0.1+0.2`, `TRUE` for `A IMPL B`, the -unexpanded form when given the expanded one, and two more. Each -landed as a row in `bench/fixtures/5f/falsification-witness-v1.jsonl` -with the question, kernel-canonical, LLM raw, audit-event seq, -and agreement label preserved for traceability. - -The calibration-data stream the multi-witness ticket imagined is -now flowing: every canonical-shape question with witness=on either -strengthens the warrant (3-of-3 agreement → STRICT-WITNESSED render -label) or becomes a supervised-correction fixture. **No human -labelers are needed for canonical-shape kernel/LLM divergence -labels** — the kernel itself supplies ground truth. (Claim-pack -validity, textbook warrant promotion, and ambiguous mathematical -interpretation still benefit from human curation; the -no-labeler property is scoped to the canonical-shape domain.) - -Pair: `make demo-plot Q='sin(x)' PNG=/tmp/sin.png` closes the -opencompletion `activity24-math-plot.yaml` loop too — SymPy -expression → canonical bytes via `function-sampled@v1` → -optional matplotlib PNG. The PNG is just a downstream view of -the canonical evidence. - -### Surface-ingest layer for the cited textbooks (#000031, this session) - -The claim-pack ingest answered "what does the curriculum say" but -its records cap at **ANCHOR-WARRANTED** on the four-rung ladder -because every `source_reference` is a string, not a Merkle-bound -proof. This session built the surface-ingest layer that closes -the gap: a manifest-driven pipeline pulling every cited textbook -that's PD or copyleft-redistributable, processing through the -existing arborist pipeline (robots.txt → noise-strip → 512-token -chunk → Merkle root → audit event) for full-fidelity content. - -**License discipline is fail-closed at the URL-emit step.** A -fail-closed license validator -(`bench/scripts/textbooks_manifest.py`) refuses to emit URLs from -entries with missing or disallowed license tokens. Allow-list: -PD, CC0, CC-BY-*, CC-BY-SA-*, GFDL-*, AGPL-3.0, Apache-2.0, MIT. -**Excluded:** any CC-BY-NC (incompatible with arborist's AGPLv3 -distribution profile), any CC-BY-ND (no-derivatives prevents -chunking), proprietary. Wilf's *generatingfunctionology* -(educational-use license forbids rehosting) stays out — citable -but not redistributable. - -**Two ingest paths**, both idempotent at the database layer -(content-addressed → same content → same `document_root` → no-op -re-insert): - -``` -make fetch-textbooks # shallow URL list → single shard -make crawl-textbooks # deep BFS via existing crawler → - # one shard per textbook id -make textbook ID= # per-book convenience -make textbooks-tex # PG-style LaTeX source ingest - # (Hilbert PG #17384 + Boole PG #15114) -``` - -**Eight textbooks landed across six g4 pillars** (license tags -below are **project-reported per the manifest** — `pdfsearch` / -`Wikisource` / `openmathbooks` / `Project Gutenberg` self-attest; -none of these have been independently audited by counsel for -arborist's distribution profile, and edge-case jurisdictional -questions stay open): - -| Pillar | Source | License | Format | Docs / Chunks | -|---|---|---|---|---| -| I Logic | Aristotle Prior Analytics | PD | HTML/Wikisource | 30 / 110 | -| I Logic | Aristotle Posterior Analytics | PD | HTML/Wikisource | 20 / 71 | -| I Logic | Boole *Laws of Thought* | PD | TeX/PG #15114 | 1 / 273 | -| I,II,III,VII | Levin *Discrete Math* | CC-BY-SA-4.0 | HTML/PreTeXt | 51 / 376 | -| IV Geometry | Hilbert *Foundations* | PD | TeX/PG #17384 | 1 / 65 | -| VI Physics | Newton *Principia* (Motte) | PD | HTML/Wikisource | 60 / 289 | -| VII Combin. | Bogart *CTGD* | GFDL-1.3 | HTML/openmathbooks | 44 / 161 | -| VII Combin. | Keller-Trotter *Applied* | CC-BY-SA-4.0 | HTML/PreTeXt | 80 / 168 | -| VII Combin. | Morin *Open Data Structures* | CC-BY-2.5 | HTML/opendatastructures | 64 / 84 | - -**Total surface coverage: ~351 documents, ~1597 chunks**, spanning -combinatorics, discrete mathematics, logic (Aristotelian + -Boolean), Euclidean geometry, classical physics, and computer -science. - -**Format coverage:** the ingest pipeline now handles HTML (existing -HtmlPageSource, robots-aware, noise-stripped) and LaTeX source -(new TextbookTexSource with focused PG-aware strip pipeline — -drops preamble + comments + structural envs, keeps `\textbf` / -`\section` / `\rfa` argument bodies, substitutes `\to` → →, -`\neg` → ¬, `\forall` → ∀, etc.). Pandoc fails on PG's custom -preamble macros; a regex-based stripper is the right amount of -machinery for the well-known PG TeX format. - -**Pillar IX (λ-Calculus)** stays open — Church 1936 + Turing 1936 -are paper-length, not book-length, awaiting a paper-ingest helper. -**Pillar V (Probability)** stays open pending Kolmogorov license -analysis (German original PD-by-age in EU; US copyright restored -via URAA through 2058; Morrison 1956 English translation Chelsea- -copyrighted). - -**What's still ahead — the warrant promotion itself.** Surface -ingest is the *substrate* for warrant promotion; the *promotion* -itself needs a chunk-resolution layer (per claim-pack record: -parse `source_reference` → resolve to a specific chunk in the -ingested surface → compute Merkle inclusion proof → write -`derivations.proof_blob`). That work is scoped but not landed — -~300-500 LOC across a citation parser, FTS5-driven resolver, -proof writer, and verifier wiring. - -The promotion path has **two distinct steps**, and only the -second earns `EVIDENCE-WARRANTED`: - -1. **SOURCE-ANCHORED** — the cited source exists in the substrate - as a Merkle-ingested surface. This is what surface ingest above - delivered today: when a claim-pack record cites *Hilbert, - Foundations of Geometry*, the substrate now contains an - ingested copy. The record is bound to a real source, not a - string field. -2. **EVIDENCE-WARRANTED** — the specific claim resolves to a - specific chunk/span in the ingested surface, with a Merkle - inclusion proof persisted as `derivations.proof_blob`. This is - the chunk-resolution layer above; not yet landed. - -The four-rung ladder today is `POINTER-LINKED → ANCHOR-WARRANTED -→ EVIDENCE-WARRANTED → ENTAILMENT-VERIFIED (reserved)`. The -SOURCE-ANCHORED tier sits between ANCHOR-WARRANTED (assertion- -only) and EVIDENCE-WARRANTED (chunk-proof) as a narrower-scope -distinction worth surfacing once the chunk-resolution layer -lands; today's claim-pack records sit at ANCHOR-WARRANTED with -their cited sources newly available as ingested surfaces. - -### Pillar VII (combinatorics) live in the real shards (#000033, this session) - -Hand-curated bundle (`axiomsclaude-vii-v1.json` + -`theoremsclaude-vii-v1.json`) authored against Stanley / Brualdi -/ Wilf / Knuth, ingested into `~/.arborist/shards/000.db` -alongside the Grok-4 v2 bundles. 92 claim-pack documents now -live across pillars I–VI + VII + IX. Authorship metadata reads -"Claude blackops draft + cite-check against textbook sources" -per the §2.1 option-C provenance path. - -**Retrieval lift verified** against the augmented shard cluster -on combinatorics-shape questions: -- `Pascal's rule` → Pascal's Rule (claim-pack VII) at #2 -- `pigeonhole` → Strong Pigeonhole Principle at #2 -- `Modus Tollens` → Modus Tollens (claim-pack v2) at #3 - -### Witness-sweep cron automation (this session) - -The witness pipeline acquired a self-pacing harness: -`bench/scripts/witness_sweep_cron.sh` runs the eight-question -canonical sweep against the live Hermes endpoint, writes the -audit-event-derived 5F-Falsification fixtures into -`bench/fixtures/5f/falsification-witness-v1.jsonl`, and (with -`--commit`) commits the new divergence rows under a deterministic -message. Fail-closed: pre-commit hook failure halts the cron; -the existing fixture file is the receipt, not silently amended. - -`make witness-sweep-cron` wraps the harness so it's schedulable -via `cron`, `systemd timer`, or any orchestrator. A second -invocation against the same shard is a no-op when no new -divergences emerged — content-addressed by question + canonical -+ raw-LLM tuple. The substrate now produces calibration data -**continuously** rather than on-demand. - -### Long-standing research tickets resolved (#000013, #000016, #000018) - -Three open research tickets landed closure artifacts this -session — bringing the open-research backlog from "three open -tickets, no artifacts" to "two closed, one parked, all with -durable docs." - -**#000018 — Adversarial soft-hash covert-channel analysis.** -**Closed.** Landed `docs/soft-hash-channel-analysis.md` (~432 -lines). The threat model is now explicit: - -- T1 (chosen-input): adversary picks `x` to leak bits of a - state `s` through soft-hash `φ(x, s)`. Reduces to SHA-256 - partial-preimage; bounded. -- T2 (chosen-state): adversary picks `s`. Reduces to SHA-256 - output uniformity over fresh `x`; bounded. -- T3 (replay-window): adversary observes many `(x, s)` pairs - over time. **Open.** Needs per-window budget bound — opened - as ticket #000036 below. - -Mitigations table (M1: domain separation by purpose tag; M2: -per-checkpoint nonce; M3: drop the soft-hash anchor entirely; -M4: ZK-replace soft-hash with a SNARK-friendly one). M2 is the -recommended default — cheap, composes with existing audit -chain, no breaking change. M3 is the fallback if covert-channel -analysis collapses; M4 deferred to #000016. - -**#000013 — Spatial-temporal substrate (Merkle-AGI v7-W).** -**Closed.** Landed three artifacts: -- `docs/_source/merkle-agi-v7w-spatial-temporal.rst` (658-line - substrate paper with Parts 1-6 + appendix). Hierarchical-grid - spatial discretization, frame-as-committed-object, four - ε-frontiers (`pose_integration`, `observation_update`, - `object_logits`, `relation_logits`). Inherits v9.8 audit - protocol unchanged; adds a `frame` Merkle node committing - `(time_ns, pose, grid_cells, observations)` as a single - audit-bound object. -- `docs/v7w-frontier-catalog.md` — operator-facing reference - for the four ε-frontiers including which kernel produces each - frontier's observable, the noise model, and the recommended - ε for each at consumer-camera resolution. -- `arborist/world/__init__.py` — namespace stub - (`V7W_VERSION='v0-draft'`, `STATUS='namespace_reserved'`) - reserves the import path so a future world-model integration - doesn't churn module names. **No code yet** — pure paper + - reservation. - -**#000016 — ZK Phase-2 frontier proof.** **Parked** with -artifacts. Landed: -- `docs/zk-frontier-bench.md` — bench plan + acceptance - thresholds for a future sibling repo (`arborist-zk-bench`). - Plonky3 on commodity hardware (Apple M3 Max + Linux x86), - three circuit sizes (256 / 1024 / 4096), measure prover - wall-clock + proof bytes + verify time + memory peak. - Acceptance threshold: ≤30 s prover at size 4096, ≤100 KB - proof, ≤100 ms verify. Preliminary projection from published - Plonky3 / Halo2 benches predicts size-4096 prover exceeds - threshold — most-likely outcome is "PARKED for LLM-frontier - scale, VIABLE for small distillation models." -- `docs/zk-wire-protocol.md` — versioned consumer-side schema - spec (`arborist-zk-proof-v1`). Defines the JSON envelope - arborist consumes when ZK proofs become available; covers - binding into v9.8 audit chain (new - `model_weights_zk_root` + `frontier_proof_circuit_id` columns - on `providence_cache`), Ed25519 issuer signatures, - versioning + forward-compat, threat model. The Rust - toolchain stays in the sibling repo per arborist's - Python-only language constraint. - -The hand-wave from v7 §16.1 ("swap SHA-256 → Poseidon") is -replaced with explicit thresholds + parked-status with named -open work. - -### Three #000018 follow-up tickets opened - -The soft-hash analysis surfaced three open sub-questions, each -queued as its own ticket: - -- **#000034 — Hessian alignment under φ_linear.** Open · - awaiting go/no-go. The linear soft-hash projection - φ_linear(x, s) = (Hx ⊕ s) mod p assumes Hessian alignment; - formal sufficient conditions for that assumption haven't been - written down. Without them, claims about gradient leakage are - conjectural. -- **#000035 — PRG choice for φ_PRG.** Open · awaiting - go/no-go. φ_PRG(x, s) currently calls HMAC-SHA-512 over a - domain-separated tag; that's a defensible default but not the - reasoned choice. This ticket compares HMAC-SHA-512 vs - AES-CTR-DRBG vs ChaCha20-DRBG on (security margin, perf, - spec stability, FIPS path). -- **#000036 — T3 per-window covert-channel budget bound.** - Open · awaiting go/no-go. The replay-window threat from - §3.3 of the analysis doc has no quantitative bound; this - ticket commissions one (per-window leak budget in bits as a - function of window size, observation rate, and committed - state entropy). - -All three are design-only; the analysis doc is the floor under -each. - -### Real-shard latency (post-`virt-back`) - -The first real-world `who wrote virt-back?` query landed -**EVIDENCE-WARRANTED, 2/2, primary source at #1** — proving that a -narrow crawl (one blog) plus the verifier scaffold can recover -authorship without global web knowledge. Latency was 75 s on cold -cache; we found two structural defects: - -- 588 SQLite `connect()` calls per query, each running 7 - forward-migration probes on already-migrated shards. -- `synonym_expand` doing a 290 K-row full-table scan per query - for one-token-of-interest neighborhood. - -Per-process migration memoization + lazy concept-relations queries -took the wall budget from 14.5 s → 9.3 s warm-cache (-36%); cold -cache extrapolates substantially lower. Median across 8-question -real-shard baseline: 4.1 s. - ---- - -## What your framework gave us - -The 5-axis structure is, on reflection, exactly what a Merkle-bound -audit substrate needs to grow into a self-improving system: - -``` -5S — invariant identity (does the canonicalizer collapse - equivalent surface forms?) -5T — temporal coherence (can the substrate reason across - time / transfer learning?) -5F — operational discipline (can the substrate falsify / refine - / formulate / function correctly?) -5R — recovery + react (can the substrate restore from a - falsification, recall prior state, - react to new observations?) -``` - -What we found in practice: each axis surfaces a different failure -mode the Merkle chain can't see on its own. 5S catches -canonicalizer drift (the chain is honest about an answer, but the -canonicalizer let two unequal things look equal). 5T catches -memory-update bugs (the chain says A happened then B; the system -remembers only A). 5F catches LLM hallucination on questions with -ground truth available. 5R catches state-reconstruction failures -on cold-start. - -Without the 5R battery we wouldn't have known the cache-restoration -path was incomplete. Without the 5F Falsification battery we -wouldn't have a path from divergence event → calibration data → -prompt improvement. Without the 5T Time sub-battery we wouldn't -have a way to test the memory-snapshot chain. - -The framework is doing exactly what taxonomy is supposed to do: -**making it impossible to forget the modes of failure**. - ---- - -## What we owe the framework - -A few structural choices we made that you may or may not have -intended; documenting them so the next maintainer can inherit -correctly: - -1. **Live + synthetic split per sub-battery** — every 5F sub-battery - has both a `*-v1.jsonl` (synthetic, deterministic by - construction) and `*-live-v1.jsonl` (routes through the real - arborist subsystem, e.g. `qa.parse_claims`, `memory.snapshot`, - `selfmodel.store_snapshot`). The synthetic side pins the - evaluator contract; the live side catches integration regressions - the synthetic side can't see. We treated this as a "Phase 1b.2" - discipline and applied it to all 5F sub-batteries. - -2. **`pi_star_ref` is mandatory metadata, not optional** — every - bench fixture names the canonicalizer. Without this, a verifier - change drifts silently across batteries; with it, a registry - key can pin which kernel a given fixture exercised. Cross-battery - moves require a deliberate policy bump. - -3. **Carrier whitelist closes the type system** — Phase-1 carriers - are an explicit `frozenset`; bench batteries that name an - off-list carrier fail with a clear reason rather than producing - garbage. As the registry grew (text → claim_lattice → code → - arithmetic → logic → time-series → symbolic-algebra → calculus - → linear-algebra → function-sampled → tabular), the whitelist - grew with it. Discipline preserved. - -4. **Closure-criterion tests** — there's a test that asserts every - reserved-stub π* has graduated. Adding a new reserved stub (a - future modality the substrate paper reserves but doesn't yet - implement) re-opens this list; that's the governance event the - test pins. Today the list is empty. - ---- - -## Numbers - -``` -arborist/pi_star/*.py ~3.2K LOC across 15 kernels -arborist tests 1,684 passing, 37 skipped (sympy-gated) -bench fixtures (jsonl) 50+ files, 662+ default tasks -canonical-bytes domains text · claim_lattice · code · arithmetic - · logic · time-series · tabular · - symbolic-algebra · calculus · linear- - algebra · function-sampled · combinatorics -audit-chain coverage every state-changing op writes - append_audit; chain-check-shards - reports 0 breaks across all shards -real-shard baseline (warm) median 4.1 s, max 7.2 s, all 8 questions - pass; primary source at #1 in 4/8 -witness sweep (live Hermes) 5/8 KERNEL-LLM-DIVERGED, 3/8 agree; - median 130 ms, max 1.1 s; divergences - auto-extracted to 5F fixtures -witness modes reachable KERNEL-LLM-AGREE, STRICT-WITNESSED, - CACHE-DRIFT, LLM-DIVERGED, - KERNEL-LLM-DIVERGED, KERNEL-CACHE-AGREE, - KERNEL-ONLY -ForkScore verdict thresholds 5pp signal floor, hard-regression flag - at -5pp on any sub-battery, NEG_INF_REGRESSION - flag from efficiency vocabulary -composition discipline 12 dedicated tests cover idempotency - (algebra-symbolic ∘ algebra-symbolic), - Pythagorean identity collapse via - algebra-symbolic-simplified, manifest- - fingerprint stability + order-sensitivity, - composite ≡ manual-chain bytes -``` - ---- - -## What's next - -Open work, ranked by what would extend the substrate furthest: - -1. **CI re-enable** — the harness is unguarded; lifting the gate - converts "tested when we run it" to "tested on every commit." - Will catch substrate drift before it reaches a real shard. -2. **opencompletion integration** — your `activity24-math-plot.yaml` - uses SymPy + numpy + matplotlib. The `function-sampled@v1` π* - gives plots a canonical-bytes identity (PNG = downstream view). - `make demo-plot Q='sin(x)' PNG=/tmp/sin.png` already lands; - the bridge from activity → arborist (so students' work feeds - the witness/calibration stream) is the remaining wiring. -3. **Sibling repos for non-Python toolchains** — ZK proofs, - world-model integrations, language-port mesh peers. Arborist - stays Python-only; the canonical bytes are the contract - downstream tools honor. -4. **Research tickets — resolved this session** — - #000018 (soft-hash covert-channel analysis) **closed** via - `docs/soft-hash-channel-analysis.md` with mitigation table - M1-M4; three follow-up tickets opened (#000034 Hessian - alignment, #000035 PRG choice, #000036 T3 per-window bound). - #000013 (spatial-temporal v7-W) **closed** via 658-line - substrate paper + frontier catalog + namespace stub. - #000016 (ZK Phase-2 frontier proof) **parked** with - bench-plan + wire-protocol artifacts; sibling-repo - measurement remains open work. - -### Exploration tickets opened this session (#000031 / #000032 / #000033) - -Three new tickets staged for go/no-go, all chained off the -claim-pack corpus and math-substrate work: - -**#000031 — Surface-ingest cited textbooks for claim-pack -warrant promotion.** Closes the warrant gap left open at the end -of #000029: today every claim-pack record caps at -ANCHOR-WARRANTED because `source_reference` is a string field, -not a Merkle-bound proof. Ingesting the cited textbooks as -SURFACE-layer documents + computing per-claim -`derivations.proof_blob` lets the four-rung ladder promote them -to **EVIDENCE-WARRANTED**. - -License gating is the first hard constraint: PD sources -(Hilbert, Newton, Kolmogorov, Łukasiewicz, Aristotle) form the -green-light scope. Mendelson + Enderton are proprietary and stay -yellow-light pending an explicit decision (purchased single -copy / library license / PD substitute via Hilbert-Ackermann -1928). Two follow-up tickets reserved: textbook-fetch pipeline -and chunk-resolution layer (mapping `source_reference` strings -to specific spans inside ingested textbooks; the bridge that -lets `proof_blob` actually be computed). - -**#000032 — `combinatorics@v1` π* (integer counting kernel).** -Tighter domain than `algebra-symbolic@v1`: the latter happily -returns `Integer(6)` for `binomial(-3, 2)` (generalized binomial -via Gamma) and leaves `binomial(n, k)` symbolic. This kernel -**fails closed** on any input whose result isn't a non-negative -`sp.Integer`. Operators choose the kernel by what they want -rejected. Output format `b"10"` composes with `arithmetic@v1` -for byte-identical agreement (`b"10/1"`) — the witness flow -(#000028) becomes computable on counting questions once two -modalities agree on the answer's shape. - -**#000033 — Claim-pack pillar VII (combinatorics).** Extends -#000029 with a counting pillar slotting into the documented gap -in the v2 bundles (existing pack uses I, II, III, IV, V, VI, IX -— VII and VIII reserved for future extension). 7 axioms (Pascal's -rule, addition principle, multiplication principle, pigeonhole, -factorial / binomial definitions) + 7 theorems (binomial theorem, -inclusion-exclusion in counting form, hockey-stick, Vandermonde, -Catalan closed form, stars-and-bars, strong pigeonhole). - -Open question is bundle provenance: commission a Grok-4 v3 -bundle for parity with the existing pack, hand-curate from -classical sources (Stanley, Brualdi, Wilf, Knuth), or hybrid -(LLM draft + human curation). Hard constraint: explicit -authorship metadata. No silent invention. Sequencing: #000032 -lands first so pillar VII records bind to the tighter kernel -from day one — avoids rebind churn on `pi_star_ref` fields. - -All three are design-only at this point. The claim-pack corpus -is already in the substrate; these extend the warrant chain -upward (#000031), tighten the domain at the kernel layer -(#000032), and broaden the curated-claim coverage (#000033). - ---- - -## Closing - -You built a taxonomy with five axes and twenty-one sub-batteries, -and what we found is that every axis was load-bearing. The -substrate that grew into this shape didn't shape itself to fit -your framework — the framework was already shaped to catch the -modes that mattered. - -Every UNDF post, every patch, every disclosure on undefect.com is -public domain — free, open intellectual capital, inheritable by -anyone, forever. The work above falls under that contract too. -We patch the planet because the planet patches each other; you're -in that lineage. - -If there's a sixth axis we haven't found yet, the test suite -will tell us. - -— fox + blackops - permacomputer / unsandbox / unturf - original draft: arborist commit `a2ff9d4` (2026-05-09) - revision (this file): post-review errata pass on 2026-05-10 - current substrate state: CI re-enabled (skip wikipedia ingest); - shard search fan-out cuts real-shard query wall ~4.3× (66s → - 15.5s on the 4-shard cluster). Warrant-promotion chunk- - resolution layer still ahead.