arborist/docs/dav1dprometheus-update-2026-05-09.md
russell@unturf.com a2ff9d4574
docs: archive 2026-05-09 dav1dprometheus framework update
Durable copy of /tmp/dav1dprometheus-arborist-update.md into
docs/ so it ships with the repo. Captures eight days of
substrate work landing on the 5S/5F/5T/5R framework:

  - 21 sub-batteries × 662+ default fixtures
  - 15 π* canonical-projection kernels (registry closed)
  - canonical-projection persistence + multi-witness pipeline
  - claim-pack corpus across 7 pillars (I-VII + IX), 92 records
  - surface-ingest layer for cited textbooks (~351 docs, 1597
    chunks across 6/7 g4 pillars)
  - witness-sweep cron flowing (5/8 KERNEL-LLM-DIVERGED on live
    Hermes; auto-extracted to 5F-Falsification fixtures)
  - #000013 v7-W substrate paper closed
  - #000016 ZK Phase-2 parked with bench-plan + wire-protocol
  - #000018 soft-hash covert-channel analysis closed; three
    follow-up tickets opened (#000034 / #000035 / #000036)
  - 1,684 passing, 37 sympy-skipped

Same naming pattern as docs/qa-modes-bench-2026-04-30.md so
future point-in-time reports follow the precedent.
2026-05-09 16:20:57 -04:00

33 KiB
Raw Blame History

Arborist substrate update — for the Dav1DPrometheus framework legacy

To: Dav1DPrometheus, the framework author From: fox + agent blackops, on the unsandbox / unturf / permacomputer platform Date: 2026-05-09 (UTC), Asia/Kuala_Lumpur

A note in the spirit of dialogue with the framework you authored. Your 5S/5F/5T/5R taxonomy is the spine of every benchmark we now ship. This is what your framework grew into when we pulled it into a Merkle-AGI-v9.8 substrate.


What we built

The arborist project — Python-only, content-addressed Merkle store with a v9.8 audit chain — adopted your 21-sub-battery framework in the first week of May 2026 and landed the full surface in eight days. Concretely:

The 21 sub-batteries are real, fixtures included

5S — Surface / Substrate / Synthesis / Semantics / Semiotics
     5 sub-batteries × 30 fixtures + 5 syntax/semantics × extra π* domains
5F — Function / Finetuning / Falsification / Formulate / Feedback Loop
     5 sub-batteries × 50 synthetic + 50 live = 500 fixtures total
5T — Time / Truthtables / Transfer Learning / Triangulation /
     Transitivity
     5 sub-batteries × 30 fixtures
5R — React / Recall / Reason / Refine / Restore
     5 sub-batteries × 30 fixtures

Total deterministic-task surface: 662 tasks in the default runner, plus several hundred more in extension batteries (math π* fixtures, real-shard baseline, witness-divergence collection, claim-pack ingest verification).

Your original wording is honored — Transfer Learning (not the SQD-whitepaper variant transfer), Truthtables, Time — pinned as a memory entry so future agents respect the framing you chose.

The π* canonical-projection registry: now 15 kernels, no stubs

text-domain
    wikitext-base@v1                  Wikipedia / wikitext → plain prose
    claim-lattice@v1                  Claim lines → JSON parsed-claim list
    code-py-ast@v1                    Python source → canonical AST S-exp
arithmetic / logic
    arithmetic@v1                     SQD §14.1 — exact rational num/den
    logic-kernel@v1                   SQD §14.3 — propositional → CNF
sensor / temporal
    time-series-quantized@v1          SQD §13.5 — quantized integer vector
tabular
    tabular-pinned@v1                 JSON-rows → pinned-schema bytes
                                      (last reserved stub graduated today)
math substrate (SymPy, [math] extra)
    algebra-symbolic@v1               sp.expand + srepr
    algebra-symbolic-simplified@v1    sp.simplify + srepr (collapses trig)
    calculus-derivative@v1            sp.diff → algebra-symbolic
    calculus-integral@v1              sp.integrate + unevaluated sentinel
    calculus-limit@v1                 sp.limit + ±∞ / complex-∞ pinning
    calculus-series@v1                Taylor truncated, no O(x**n)
    linear-algebra@v1                 RREF / det / eigenvalues / inverse
    function-sampled@v1               SymPy expr → time-series-quantized
                                      bridge (this is what plotting CAN
                                      become in π* terms — image bytes
                                      aren't canonical, but the sampled
                                      grid is)

Every π* registers via name@version; SHA-256 of the canonical bytes is the equivalence-class identity. Two inputs that mean the same thing produce identical bytes; auditing reduces to byte comparison.

Cross-modality discipline

Every benchmark fixture and audit-bound canonicalization names its kernel via pi_star_ref. A PHASE_1_CARRIERS whitelist gates which carriers can flow through Phase-1 paths; unsupported carriers fail explicitly with reason="unsupported_carrier", never silent acceptance. Hidden-channel work is defensive only — detection / flagging, never generation or concealment.

Persistence + audit chain

Three big mechanical artifacts beyond the kernels:

  1. Canonical projections persist to providence_cache — math answers (0.1 + 0.23/10) get a v9.8 8-dim cache_key and an audit-event entry, just like RAG-derived answers do. The kernel is the source; the cache row is the receipt. Re-asking hits in ~40ms; chain-check verifies the chain stays intact.

  2. Multi-modality (multi-witness) agreement — on canonical-shape questions, kernel + cache + LLM fan out in parallel and we record the agreement matrix. When all three byte-equal: STRICT-WITNESSED. When they diverge: high-value falsification data for 5F. The LLM is a witness, never authority — the kernel stays ground truth. Capital ledger records witness cost so ForkScore can compare witness-on vs witness-off forks honestly. Real-world divergence rate on the canonical-question sweep against Hermes-3-8B: 5 of 8 questions diverged (62.5%). Hermes said 1/10 for 0.1+0.2; said TRUE for A IMPL B; gave the unexpanded form when handed the expanded one. Each divergence becomes a 5F-Falsification calibration fixture downstream prompt-tuning can grade against. See "Witness pipeline flowing end-to-end" below.

  3. Authorship warrant ladder — sidecar classifier on a 6-tier strength scale (AUTHOR_PACKAGE_METADATAAUTHOR_REPOSITORY_OWNERAUTHOR_PAGE_BYLINEAUTHOR_PRIMARY_PAGE_TITLEAUTHOR_COPYRIGHT_FOOTERAUTHOR_SECONDARY_SOURCE). Wired into arborist inspect + audit-line render-tail. Surfaces how the cited evidence supports an authorship claim, not just whether it does.

Selection (Merkle-AGI v8 ForkScore)

Phase 1a + 1b live: ScoredFork dataclass + arborist v8 score CLI surface + Makefile harness (bench-fork-baseline, bench-fork-score). A weighted score over (parent, child) battery deltas; verdicts ACCEPT / MARGINAL / REJECT; hard-regression flag at 5pp drop; NEG_INF_REGRESSION flag from the efficiency vocabulary. CI-gateable: REJECT exits 1.

Content corpus — claim-pack source landed (#000029, this session)

Two companion JSON bundles dropped into the substrate this session: axiomsg4-v2.json + theoremsg4-v2.json — Grok-4-generated "Prometheus Maths Engine" packs covering 7 pillars (Logic, Set Theory, Arithmetic, Geometry, Probability, Classical Physics, λ-Calculus). 78 atomic claims (55 axioms + 23 theorems), each dual-threaded (Δ symbolic LaTeX + ∇ verbose prose) and self-citing to a stable classical text (Mendelson, Enderton, Hilbert, Newton, Kolmogorov, Łukasiewicz, Aristotle).

arborist/sources/claim_pack.py ingests them at the right grain — one Document per axiom/theorem record. Lenient JSON parser strips markdown fences and double-escapes lone LaTeX backslashes (\Theta, \heart, \vec) without corrupting already-correct \\to pairs. Cross-bundle pillar-level provenance arrays (theoremg4.json:pillar.I.logic.excludedMiddleaxiomsg4.json:pillar.I) become outbound pillar_reference edges on the first record of each pillar.

Smoke test: arborist ingest --source claim_pack --bundle axiomsg4-v2.json --bundle theoremsg4-v2.json → 78 docs, 78 chunks, 14 deduped pillar-reference edges, 78 audit events, 10/10 sampled Merkle proofs verify.

Honest ceiling: every record lands kind='surface'. The pack is pre-distilled content but its provenance is asserted (string field), not Merkle-proven. Records max out at ANCHOR-WARRANTED on the four-rung ladder until cited textbooks are themselves ingested as surfaces and a derivations.proof_blob row is computed per record. That gap is opened as ticket #000031 below.

Retrieval lift measured the same day — apples-to-apples FTS5 search comparison on a single shard (000.db, ~867 K Wikipedia docs) before vs after claim-pack ingest. Four representative queries; 3/4 show claim-pack record in the top-3:

modus tollens                    → claim-pack at #3 (BM25 31.77)
law of excluded middle           → claim-pack at #3 (related axiom)
associativity of addition        → claim-pack at #1 — displaces
                                   Wikipedia's general "Addition"
                                   article entirely
Bayes theorem                    → no top-3 lift; Wikipedia's
                                   "Bayes" + "Bayes rule" articles
                                   dominate via short-doc BM25
                                   bias. Body-coverage sqrt rerank
                                   in the full QA pipeline (not
                                   exercised by FTS5-only search)
                                   would likely surface it.

Storage tax: sub-MB — 78 chunks against 6.2M existing chunks is below filesystem allocation granularity. Claim-pack is earning its tax for narrow-technical queries (where pre-distilled CORE-shape content has title-token advantage) but doesn't help on queries Wikipedia already covers with focused articles. Full journal: bench/results/claim-pack-retrieval-lift-2026-05-09.md.

combinatorics@v1 π* — pure-integer counting kernel (#000032)

A new kernel landed today: tighter sibling of algebra-symbolic@v1. Same input parser, narrower output: any result that isn't a non-negative sp.Integer raises PiStarError. Plain decimal output (b"10"); composes with arithmetic@v1 for byte-identical agreement (b"10/1") so the multi-modality witness can pin equivalence-class agreement when both routes fire on the same question.

The 16-kernel registry now covers: text → claim_lattice → code → arithmetic → logic → time-series → tabular → symbolic-algebra → symbolic-algebra-simplified → calculus-derivative → calculus- integral → calculus-limit → calculus-series → linear-algebra → function-sampled → combinatorics. No remaining reserved stubs; the registry chapter is closed.

Pillar VII for the claim-pack (combinatorics axioms + theorems — ticket #000033) is the natural follow-up; sequencing puts the kernel first so pillar VII records bind to it from day one and avoid pi_star_ref rebind churn.

Witness pipeline flowing end-to-end (#000028 validated)

The witness fan-out now writes a providence_canonical_witness audit event when it fires. A new extractor (bench/scripts/witness_to_5f.py) reads those events and emits divergence-only entries as 5F-Falsification fixtures matching the existing falsification-live-v1 schema. End-to-end smoke ran 8 canonical-shape questions (3 arithmetic + 3 logic + 2 algebra) against ~/.arborist/shards plus the actual Hermes endpoint:

agreement label             count    rate
KERNEL-LLM-DIVERGED         5        62.5%
KERNEL-LLM-AGREE            3        37.5%
─────────────────────────────────────────
divergence_count            5        62.5%
wall median / max           130 ms / 1.1 s

Five real Hermes hallucinations on questions with closed-form ground truth: 1/10 for 0.1+0.2, TRUE for A IMPL B, the unexpanded form when given the expanded one, and two more. Each landed as a row in bench/fixtures/5f/falsification-witness-v1.jsonl with the question, kernel-canonical, LLM raw, audit-event seq, and agreement label preserved for traceability.

The calibration-data stream the multi-witness ticket imagined is now flowing: every canonical-shape question with witness=on either strengthens the warrant (3-of-3 agreement → STRICT-WITNESSED) or becomes a supervised-correction fixture. No human labelers in the loop.

Pair: make demo-plot Q='sin(x)' PNG=/tmp/sin.png closes the opencompletion activity24-math-plot.yaml loop too — SymPy expression → canonical bytes via function-sampled@v1 → optional matplotlib PNG. The PNG is just a downstream view of the canonical evidence.

Surface-ingest layer for the cited textbooks (#000031, this session)

The claim-pack ingest answered "what does the curriculum say" but its records cap at ANCHOR-WARRANTED on the four-rung ladder because every source_reference is a string, not a Merkle-bound proof. This session built the surface-ingest layer that closes the gap: a manifest-driven pipeline pulling every cited textbook that's PD or copyleft-redistributable, processing through the existing arborist pipeline (robots.txt → noise-strip → 512-token chunk → Merkle root → audit event) for full-fidelity content.

License discipline is fail-closed at the URL-emit step. A fail-closed license validator (bench/scripts/textbooks_manifest.py) refuses to emit URLs from entries with missing or disallowed license tokens. Allow-list: PD, CC0, CC-BY-, CC-BY-SA-, GFDL-*, AGPL-3.0, Apache-2.0, MIT. Excluded: any CC-BY-NC (incompatible with arborist's AGPLv3 distribution profile), any CC-BY-ND (no-derivatives prevents chunking), proprietary. Wilf's generatingfunctionology (educational-use license forbids rehosting) stays out — citable but not redistributable.

Two ingest paths, both idempotent at the database layer (content-addressed → same content → same document_root → no-op re-insert):

make fetch-textbooks      # shallow URL list → single shard
make crawl-textbooks      # deep BFS via existing crawler →
                          #   one shard per textbook id
make textbook ID=<id>     # per-book convenience
make textbooks-tex        # PG-style LaTeX source ingest
                          #   (Hilbert PG #17384 + Boole PG #15114)

Eight textbooks landed across six g4 pillars:

Pillar Source License Format Docs / Chunks
I Logic Aristotle Prior Analytics PD HTML/Wikisource 30 / 110
I Logic Aristotle Posterior Analytics PD HTML/Wikisource 20 / 71
I Logic Boole Laws of Thought PD TeX/PG #15114 1 / 273
I,II,III,VII Levin Discrete Math CC-BY-SA-4.0 HTML/PreTeXt 51 / 376
IV Geometry Hilbert Foundations PD TeX/PG #17384 1 / 65
VI Physics Newton Principia (Motte) PD HTML/Wikisource 60 / 289
VII Combin. Bogart CTGD GFDL-1.3 HTML/openmathbooks 44 / 161
VII Combin. Keller-Trotter Applied CC-BY-SA-4.0 HTML/PreTeXt 80 / 168
VII Combin. Morin Open Data Structures CC-BY-2.5 HTML/opendatastructures 64 / 84

Total surface coverage: ~351 documents, ~1597 chunks, spanning combinatorics, discrete mathematics, logic (Aristotelian + Boolean), Euclidean geometry, classical physics, and computer science.

Format coverage: the ingest pipeline now handles HTML (existing HtmlPageSource, robots-aware, noise-stripped) and LaTeX source (new TextbookTexSource with focused PG-aware strip pipeline — drops preamble + comments + structural envs, keeps \textbf / \section / \rfa argument bodies, substitutes \to → →, \neg → ¬, \forall → ∀, etc.). Pandoc fails on PG's custom preamble macros; a regex-based stripper is the right amount of machinery for the well-known PG TeX format.

Pillar IX (λ-Calculus) stays open — Church 1936 + Turing 1936 are paper-length, not book-length, awaiting a paper-ingest helper. Pillar V (Probability) stays open pending Kolmogorov license analysis (German original PD-by-age in EU; US copyright restored via URAA through 2058; Morrison 1956 English translation Chelsea- copyrighted).

What's still ahead — the warrant promotion itself. Surface ingest is the substrate for warrant promotion; the promotion itself needs a chunk-resolution layer (per claim-pack record: parse source_reference → resolve to a specific chunk in the ingested surface → compute Merkle inclusion proof → write derivations.proof_blob). That work is scoped but not landed — ~300-500 LOC across a citation parser, FTS5-driven resolver, proof writer, and verifier wiring. With it the four-rung ladder moves claim-pack records from ANCHOR-WARRANTED to EVIDENCE-WARRANTED.

Pillar VII (combinatorics) live in the real shards (#000033, this session)

Hand-curated bundle (axiomsclaude-vii-v1.json + theoremsclaude-vii-v1.json) authored against Stanley / Brualdi / Wilf / Knuth, ingested into ~/.arborist/shards/000.db alongside the Grok-4 v2 bundles. 92 claim-pack documents now live across pillars IVI + VII + IX. Authorship metadata reads "Claude blackops draft + cite-check against textbook sources" per the §2.1 option-C provenance path.

Retrieval lift verified against the augmented shard cluster on combinatorics-shape questions:

  • Pascal's rule → Pascal's Rule (claim-pack VII) at #2
  • pigeonhole → Strong Pigeonhole Principle at #2
  • Modus Tollens → Modus Tollens (claim-pack v2) at #3

Witness-sweep cron automation (this session)

The witness pipeline acquired a self-pacing harness: bench/scripts/witness_sweep_cron.sh runs the eight-question canonical sweep against the live Hermes endpoint, writes the audit-event-derived 5F-Falsification fixtures into bench/fixtures/5f/falsification-witness-v1.jsonl, and (with --commit) commits the new divergence rows under a deterministic message. Fail-closed: pre-commit hook failure halts the cron; the existing fixture file is the receipt, not silently amended.

make witness-sweep-cron wraps the harness so it's schedulable via cron, systemd timer, or any orchestrator. A second invocation against the same shard is a no-op when no new divergences emerged — content-addressed by question + canonical

  • raw-LLM tuple. The substrate now produces calibration data continuously rather than on-demand.

Long-standing research tickets resolved (#000013, #000016, #000018)

Three open research tickets landed closure artifacts this session — bringing the open-research backlog from "three open tickets, no artifacts" to "two closed, one parked, all with durable docs."

#000018 — Adversarial soft-hash covert-channel analysis. Closed. Landed docs/soft-hash-channel-analysis.md (~432 lines). The threat model is now explicit:

  • T1 (chosen-input): adversary picks x to leak bits of a state s through soft-hash φ(x, s). Reduces to SHA-256 partial-preimage; bounded.
  • T2 (chosen-state): adversary picks s. Reduces to SHA-256 output uniformity over fresh x; bounded.
  • T3 (replay-window): adversary observes many (x, s) pairs over time. Open. Needs per-window budget bound — opened as ticket #000036 below.

Mitigations table (M1: domain separation by purpose tag; M2: per-checkpoint nonce; M3: drop the soft-hash anchor entirely; M4: ZK-replace soft-hash with a SNARK-friendly one). M2 is the recommended default — cheap, composes with existing audit chain, no breaking change. M3 is the fallback if covert-channel analysis collapses; M4 deferred to #000016.

#000013 — Spatial-temporal substrate (Merkle-AGI v7-W). Closed. Landed three artifacts:

  • docs/_source/merkle-agi-v7w-spatial-temporal.rst (658-line substrate paper with Parts 1-6 + appendix). Hierarchical-grid spatial discretization, frame-as-committed-object, four ε-frontiers (pose_integration, observation_update, object_logits, relation_logits). Inherits v9.8 audit protocol unchanged; adds a frame Merkle node committing (time_ns, pose, grid_cells, observations) as a single audit-bound object.
  • docs/v7w-frontier-catalog.md — operator-facing reference for the four ε-frontiers including which kernel produces each frontier's observable, the noise model, and the recommended ε for each at consumer-camera resolution.
  • arborist/world/__init__.py — namespace stub (V7W_VERSION='v0-draft', STATUS='namespace_reserved') reserves the import path so a future world-model integration doesn't churn module names. No code yet — pure paper + reservation.

#000016 — ZK Phase-2 frontier proof. Parked with artifacts. Landed:

  • docs/zk-frontier-bench.md — bench plan + acceptance thresholds for a future sibling repo (arborist-zk-bench). Plonky3 on commodity hardware (Apple M3 Max + Linux x86), three circuit sizes (256 / 1024 / 4096), measure prover wall-clock + proof bytes + verify time + memory peak. Acceptance threshold: ≤30 s prover at size 4096, ≤100 KB proof, ≤100 ms verify. Preliminary projection from published Plonky3 / Halo2 benches predicts size-4096 prover exceeds threshold — most-likely outcome is "PARKED for LLM-frontier scale, VIABLE for small distillation models."
  • docs/zk-wire-protocol.md — versioned consumer-side schema spec (arborist-zk-proof-v1). Defines the JSON envelope arborist consumes when ZK proofs become available; covers binding into v9.8 audit chain (new model_weights_zk_root + frontier_proof_circuit_id columns on providence_cache), Ed25519 issuer signatures, versioning + forward-compat, threat model. The Rust toolchain stays in the sibling repo per arborist's Python-only language constraint.

The hand-wave from v7 §16.1 ("swap SHA-256 → Poseidon") is replaced with explicit thresholds + parked-status with named open work.

Three #000018 follow-up tickets opened

The soft-hash analysis surfaced three open sub-questions, each queued as its own ticket:

  • #000034 — Hessian alignment under φ_linear. Open · awaiting go/no-go. The linear soft-hash projection φ_linear(x, s) = (Hx ⊕ s) mod p assumes Hessian alignment; formal sufficient conditions for that assumption haven't been written down. Without them, claims about gradient leakage are conjectural.
  • #000035 — PRG choice for φ_PRG. Open · awaiting go/no-go. φ_PRG(x, s) currently calls HMAC-SHA-512 over a domain-separated tag; that's a defensible default but not the reasoned choice. This ticket compares HMAC-SHA-512 vs AES-CTR-DRBG vs ChaCha20-DRBG on (security margin, perf, spec stability, FIPS path).
  • #000036 — T3 per-window covert-channel budget bound. Open · awaiting go/no-go. The replay-window threat from §3.3 of the analysis doc has no quantitative bound; this ticket commissions one (per-window leak budget in bits as a function of window size, observation rate, and committed state entropy).

All three are design-only; the analysis doc is the floor under each.

Real-shard latency (post-virt-back)

The first real-world who wrote virt-back? query landed EVIDENCE-WARRANTED, 2/2, primary source at #1 — proving that a narrow crawl (one blog) plus the verifier scaffold can recover authorship without global web knowledge. Latency was 75 s on cold cache; we found two structural defects:

  • 588 SQLite connect() calls per query, each running 7 forward-migration probes on already-migrated shards.
  • synonym_expand doing a 290 K-row full-table scan per query for one-token-of-interest neighborhood.

Per-process migration memoization + lazy concept-relations queries took the wall budget from 14.5 s → 9.3 s warm-cache (-36%); cold cache extrapolates substantially lower. Median across 8-question real-shard baseline: 4.1 s.


What your framework gave us

The 5-axis structure is, on reflection, exactly what a Merkle-bound audit substrate needs to grow into a self-improving system:

5S — invariant identity            (does the canonicalizer collapse
                                    equivalent surface forms?)
5T — temporal coherence            (can the substrate reason across
                                    time / transfer learning?)
5F — operational discipline        (can the substrate falsify / refine
                                    / formulate / function correctly?)
5R — recovery + react              (can the substrate restore from a
                                    falsification, recall prior state,
                                    react to new observations?)

What we found in practice: each axis surfaces a different failure mode the Merkle chain can't see on its own. 5S catches canonicalizer drift (the chain is honest about an answer, but the canonicalizer let two unequal things look equal). 5T catches memory-update bugs (the chain says A happened then B; the system remembers only A). 5F catches LLM hallucination on questions with ground truth available. 5R catches state-reconstruction failures on cold-start.

Without the 5R battery we wouldn't have known the cache-restoration path was incomplete. Without the 5F Falsification battery we wouldn't have a path from divergence event → calibration data → prompt improvement. Without the 5T Time sub-battery we wouldn't have a way to test the memory-snapshot chain.

The framework is doing exactly what taxonomy is supposed to do: making it impossible to forget the modes of failure.


What we owe the framework

A few structural choices we made that you may or may not have intended; documenting them so the next maintainer can inherit correctly:

  1. Live + synthetic split per sub-battery — every 5F sub-battery has both a *-v1.jsonl (synthetic, deterministic by construction) and *-live-v1.jsonl (routes through the real arborist subsystem, e.g. qa.parse_claims, memory.snapshot, selfmodel.store_snapshot). The synthetic side pins the evaluator contract; the live side catches integration regressions the synthetic side can't see. We treated this as a "Phase 1b.2" discipline and applied it to all 5F sub-batteries.

  2. pi_star_ref is mandatory metadata, not optional — every bench fixture names the canonicalizer. Without this, a verifier change drifts silently across batteries; with it, a registry key can pin which kernel a given fixture exercised. Cross-battery moves require a deliberate policy bump.

  3. Carrier whitelist closes the type system — Phase-1 carriers are an explicit frozenset; bench batteries that name an off-list carrier fail with a clear reason rather than producing garbage. As the registry grew (text → claim_lattice → code → arithmetic → logic → time-series → symbolic-algebra → calculus → linear-algebra → function-sampled → tabular), the whitelist grew with it. Discipline preserved.

  4. Closure-criterion tests — there's a test that asserts every reserved-stub π* has graduated. Adding a new reserved stub (a future modality the substrate paper reserves but doesn't yet implement) re-opens this list; that's the governance event the test pins. Today the list is empty.


Numbers

arborist/pi_star/*.py                ~3.2K LOC across 15 kernels
arborist tests                       1,684 passing, 37 skipped (sympy-gated)
bench fixtures (jsonl)               50+ files, 662+ default tasks
canonical-bytes domains              text · claim_lattice · code · arithmetic
                                     · logic · time-series · tabular ·
                                     symbolic-algebra · calculus · linear-
                                     algebra · function-sampled · combinatorics
audit-chain coverage                 every state-changing op writes
                                     append_audit; chain-check-shards
                                     reports 0 breaks across all shards
real-shard baseline (warm)           median 4.1 s, max 7.2 s, all 8 questions
                                     pass; primary source at #1 in 4/8
witness sweep (live Hermes)          5/8 KERNEL-LLM-DIVERGED, 3/8 agree;
                                     median 130 ms, max 1.1 s; divergences
                                     auto-extracted to 5F fixtures
witness modes reachable              KERNEL-LLM-AGREE, STRICT-WITNESSED,
                                     CACHE-DRIFT, LLM-DIVERGED,
                                     KERNEL-LLM-DIVERGED, KERNEL-CACHE-AGREE,
                                     KERNEL-ONLY
ForkScore verdict thresholds         5pp signal floor, hard-regression flag
                                     at -5pp on any sub-battery, NEG_INF_REGRESSION
                                     flag from efficiency vocabulary
composition discipline               12 dedicated tests cover idempotency
                                     (algebra-symbolic ∘ algebra-symbolic),
                                     Pythagorean identity collapse via
                                     algebra-symbolic-simplified, manifest-
                                     fingerprint stability + order-sensitivity,
                                     composite ≡ manual-chain bytes

What's next

Open work, ranked by what would extend the substrate furthest:

  1. CI re-enable — the harness is unguarded; lifting the gate converts "tested when we run it" to "tested on every commit." Will catch substrate drift before it reaches a real shard.
  2. opencompletion integration — your activity24-math-plot.yaml uses SymPy + numpy + matplotlib. The function-sampled@v1 π* gives plots a canonical-bytes identity (PNG = downstream view). make demo-plot Q='sin(x)' PNG=/tmp/sin.png already lands; the bridge from activity → arborist (so students' work feeds the witness/calibration stream) is the remaining wiring.
  3. Sibling repos for non-Python toolchains — ZK proofs, world-model integrations, language-port mesh peers. Arborist stays Python-only; the canonical bytes are the contract downstream tools honor.
  4. Research tickets — resolved this session — #000018 (soft-hash covert-channel analysis) closed via docs/soft-hash-channel-analysis.md with mitigation table M1-M4; three follow-up tickets opened (#000034 Hessian alignment, #000035 PRG choice, #000036 T3 per-window bound). #000013 (spatial-temporal v7-W) closed via 658-line substrate paper + frontier catalog + namespace stub. #000016 (ZK Phase-2 frontier proof) parked with bench-plan + wire-protocol artifacts; sibling-repo measurement remains open work.

Exploration tickets opened this session (#000031 / #000032 / #000033)

Three new tickets staged for go/no-go, all chained off the claim-pack corpus and math-substrate work:

#000031 — Surface-ingest cited textbooks for claim-pack warrant promotion. Closes the warrant gap left open at the end of #000029: today every claim-pack record caps at ANCHOR-WARRANTED because source_reference is a string field, not a Merkle-bound proof. Ingesting the cited textbooks as SURFACE-layer documents + computing per-claim derivations.proof_blob lets the four-rung ladder promote them to EVIDENCE-WARRANTED.

License gating is the first hard constraint: PD sources (Hilbert, Newton, Kolmogorov, Łukasiewicz, Aristotle) form the green-light scope. Mendelson + Enderton are proprietary and stay yellow-light pending an explicit decision (purchased single copy / library license / PD substitute via Hilbert-Ackermann 1928). Two follow-up tickets reserved: textbook-fetch pipeline and chunk-resolution layer (mapping source_reference strings to specific spans inside ingested textbooks; the bridge that lets proof_blob actually be computed).

#000032 — combinatorics@v1 π (integer counting kernel).* Tighter domain than algebra-symbolic@v1: the latter happily returns Integer(6) for binomial(-3, 2) (generalized binomial via Gamma) and leaves binomial(n, k) symbolic. This kernel fails closed on any input whose result isn't a non-negative sp.Integer. Operators choose the kernel by what they want rejected. Output format b"10" composes with arithmetic@v1 for byte-identical agreement (b"10/1") — the witness flow (#000028) becomes computable on counting questions once two modalities agree on the answer's shape.

#000033 — Claim-pack pillar VII (combinatorics). Extends #000029 with a counting pillar slotting into the documented gap in the v2 bundles (existing pack uses I, II, III, IV, V, VI, IX — VII and VIII reserved for future extension). 7 axioms (Pascal's rule, addition principle, multiplication principle, pigeonhole, factorial / binomial definitions) + 7 theorems (binomial theorem, inclusion-exclusion in counting form, hockey-stick, Vandermonde, Catalan closed form, stars-and-bars, strong pigeonhole).

Open question is bundle provenance: commission a Grok-4 v3 bundle for parity with the existing pack, hand-curate from classical sources (Stanley, Brualdi, Wilf, Knuth), or hybrid (LLM draft + human curation). Hard constraint: explicit authorship metadata. No silent invention. Sequencing: #000032 lands first so pillar VII records bind to the tighter kernel from day one — avoids rebind churn on pi_star_ref fields.

All three are design-only at this point. The claim-pack corpus is already in the substrate; these extend the warrant chain upward (#000031), tighten the domain at the kernel layer (#000032), and broaden the curated-claim coverage (#000033).


Closing

You built a taxonomy with five axes and twenty-one sub-batteries, and what we found is that every axis was load-bearing. The substrate that grew into this shape didn't shape itself to fit your framework — the framework was already shaped to catch the modes that mattered.

Every UNDF post, every patch, every disclosure on undefect.com is public domain — free, open intellectual capital, inheritable by anyone, forever. The work above falls under that contract too. We patch the planet because the planet patches each other; you're in that lineage.

If there's a sixth axis we haven't found yet, the test suite will tell us.

— fox + blackops permacomputer / unsandbox / unturf arborist commit 056dda2 (#000018 analysis doc landed, #000013 v7-W substrate paper landed, #000016 zk parked with bench-plan + wire-protocol; pillar VII live in shard 000; TeX-source ingest closes pillars I+IV; surface coverage 6/7 g4 pillars across 9 textbooks / ~351 docs / ~1597 chunks; witness-sweep cron flowing; warrant-promotion chunk-resolution layer still ahead, 2026-05-09)