Followup to 654d923 (which moved the package from arborist/v8/ →
arborist/substrate/ at the file layer). The CLI surface still baked
in `v8` so a new operator running `--help` would see
``arborist v8 score`` and ask the same "what's v8 vs v9.8?"
naming-confusion question that drove the package rename in the
first place. Closing the loop end-to-end.
arborist/cli.py
===============
- Subparser renamed: ``"v8"`` → ``"substrate"``; help string updated
to "Merkle-AGI substrate primitives (ForkScore + future paper
specs)" so the dir name and command name and help text all align.
- Inner subparser dest renamed: ``v8_op`` → ``substrate_op``.
- Function renamed: ``_cmd_v8_score`` → ``_cmd_substrate_score``;
docstring updated.
- All ``v8_score`` local variables renamed to ``substrate_score``.
- New comment block above the subparser block explains the rename
+ why the v-prefix was retired (substrate-paper version vs v9.8
schema version naming collision).
The old ``arborist v8 score`` is gone — no alias preserved. CI + ops
scripts must update; today's earlier commit chain has been the only
place using it and that's been refreshed in lock-step.
tests/test_v8_fork_score.py
===========================
- 4 ``parser.parse_args(["v8", "score", ...])`` calls → ``["substrate", ...]``.
- 4 test functions renamed: ``test_cli_v8_score_*`` →
``test_cli_substrate_score_*``.
- Module docstring + section comment + helper docstring updated.
Filename intentionally kept as ``test_v8_fork_score.py`` for git
history continuity; pytest discovers by ``test_*`` content, not
filename. Renaming the file would muddle ``git log --follow`` for
the test surface.
Docs refreshed
==============
- docs/v8-fork-score.md — §5 CLI block invocation.
- docs/_source/v8-fork-score.rst — :code-block:: bash invocation.
- docs/_source/bench.rst — invocation in `### v8 ForkScore` section.
- docs/tickets/ticket-000012-selection-consensus-protocol.md —
three references in §7 close-out + §7 Phase 1c proposal +
§7 future-CLI-shape note.
- docs/dav1dprometheus-update-2026-05-09.md — bench journal mention.
Doc filenames (``v8-fork-score.{md,rst}``) kept stable since they
are URL identities; the file content explains the v8→substrate
rename internally. ``index.rst`` toctree references unchanged.
Hygiene
=======
- ``.venv/bin/arborist substrate score --help`` → 0 + valid usage.
- ``.venv/bin/arborist v8 score`` → exits non-zero (subcommand
removed, surfaced cleanly in ``argparse`` error).
- ``make test`` → 1643 passed, 45 skipped.
- ``make chain-check-shards`` → 0 across all 7 shards.
- fox's parallel work in arborist/qa/{runner,verify}.py +
arborist/qa/warrant_chain.py left untouched.
40 KiB
Arborist substrate update — for the Dav1DPrometheus framework legacy
To: Dav1DPrometheus, the framework author From: fox + agent blackops, on the unsandbox / unturf / permacomputer platform Date: 2026-05-09 (UTC), Asia/Kuala_Lumpur
A note in the spirit of dialogue with the framework you authored. Your 5S/5F/5T/5R taxonomy is the spine of every benchmark we now ship. This is what your framework grew into when we pulled it into a Merkle-AGI-v9.8 substrate.
Revision note (2026-05-10): an external review surfaced several errata in the original (2026-05-09) draft. Corrections applied in-place: 5S vocabulary is Syntax / Semantics / Syllogism / Synthesis / Semiotics (not Surface/Substrate/…); 5R vocabulary is React / Rearrange / Restore / Replicate / Resonate (not React/Recall/ Reason/Refine/Restore); the π* registry chronology is now explicit (15 → 16 with
combinatorics@v1); kernel-version immutability is called out as a hard invariant; cross-witness vs. cross-carrier is distinguished;STRICT-WITNESSEDis scoped as a render label (not a new admissibility/cache mode); the warrant-promotion ladder distinguishes SOURCE-ANCHORED from EVIDENCE-WARRANTED; license tags are flagged as project-reported; the "no human labelers" property is scoped to canonical-shape domains only. Original draft preserved at commita2ff9d4.The reviewer also flagged the sub-battery count as 20 vs the document's 21. The codebase has 21: 5+6+5+5, because 5T carries 6 sub-batteries — the original SQD-whitepaper
transferplus the canonical Dav1DPrometheus five (Transfer Learning, Truthtables, Time, Triangulation, Transitivity), per ticket #000024 which preserved the legacy name alongside the authoritative wording. Earlier in this revision pass the count was incorrectly downgraded to 20 — that regression is corrected here.
What we built
The arborist project — Python-only, content-addressed Merkle store with a v9.8 audit chain — adopted your taxonomy in the first week of May 2026 and landed the full benchmark surface in eight days. Concretely:
The 21 sub-batteries are real, fixtures included
Four axes × five canonical sub-batteries + 5T's preserved legacy
sub-battery = 21 sub-batteries, each fixture-backed and runnable
through bench/batteries/runner.py:
5S — Syntax / Semantics / Syllogism / Synthesis / Semiotics
5 sub-batteries (Phase 1a: Syntax + Semantics; Phase 1b: the
other three) carrier-aware (text / claim_lattice / prose),
plus per-π*-domain expansions of Syntax + Semantics that
bring the family wider than 5 fixtures-per-sub-battery
5F — Function / Finetuning / Falsification / Formulate / Feedback Loop
5 sub-batteries × ~50 synthetic + ~50 live each (Phase 1b.2
live wire-ups land on real arborist subsystems)
5T — Time / Truthtables / Transfer Learning / Triangulation /
Transitivity (canonical Dav1DPrometheus five)
+ Transfer (legacy SQD-whitepaper sub-battery, kept alongside
Transfer Learning per ticket #000024)
6 sub-batteries total
5R — React / Rearrange / Restore / Replicate / Resonate
5 sub-batteries (workspace-operator shape, per SQD §9.3 —
deterministic, no LLM-as-judge)
(An earlier draft of this note mis-spelled 5S as Surface/Substrate/…
and 5R as React/Recall/Reason/Refine/Restore. The fixtures +
runner in bench/batteries/ are the authority — corrected here.)
Total deterministic-task surface: 662 tasks in the default
runner (per the runner's --all enumeration), plus several
hundred more in extension batteries (math π* fixtures, real-shard
baseline, witness-divergence collection, claim-pack ingest
verification).
Your original wording is honored — Transfer Learning (not the
SQD-whitepaper variant transfer), Truthtables, Time — pinned
as a memory entry so future agents respect the framing you chose.
The π* canonical-projection registry: now 16 kernels, no stubs
The registry held 15 kernels before this session; combinatorics@v1
(#000032) brought it to 16 and closed the last reserved stub.
Listed in registration order:
text-domain
wikitext-base@v1 Wikipedia / wikitext → plain prose
claim-lattice@v1 Claim lines → JSON parsed-claim list
code-py-ast@v1 Python source → canonical AST S-exp
arithmetic / logic
arithmetic@v1 SQD §14.1 — exact rational num/den
logic-kernel@v1 SQD §14.3 — propositional → CNF
sensor / temporal
time-series-quantized@v1 SQD §13.5 — quantized integer vector
tabular
tabular-pinned@v1 JSON-rows → pinned-schema bytes
(last reserved stub graduated today)
math substrate (SymPy, [math] extra)
algebra-symbolic@v1 sp.expand + srepr
algebra-symbolic-simplified@v1 sp.simplify + srepr (collapses trig)
calculus-derivative@v1 sp.diff → algebra-symbolic
calculus-integral@v1 sp.integrate + unevaluated sentinel
calculus-limit@v1 sp.limit + ±∞ / complex-∞ pinning
calculus-series@v1 Taylor truncated, no O(x**n)
linear-algebra@v1 RREF / det / eigenvalues / inverse
function-sampled@v1 SymPy expr → time-series-quantized
bridge (this is what plotting CAN
become in π* terms — image bytes
aren't canonical, but the sampled
grid is)
combinatorics
combinatorics@v1 pure-integer counting kernel
(#000032; tighter than algebra-
symbolic — fails closed on any
non-non-negative-integer result)
Every π* registers via name@version; SHA-256 of the canonical
bytes is the equivalence-class identity. Two inputs that mean the
same thing produce identical bytes; auditing reduces to byte
comparison.
Kernel-version immutability is a hard invariant. A registered
name@version is behaviorally frozen — any change in canonicalization
output is a new version (arithmetic@v2), never an in-place patch.
Otherwise prior persisted canonical-cache rows would silently
become semantically unstable. The registry rejects re-registration
of the same key with a different implementation; the
governance_policy_hash also folds in the active kernel manifest
so flipping kernels invalidates prior records on lookup.
Cross-witness vs. cross-carrier discipline
Two distinct axes, kept separate in the substrate vocabulary:
- Witness channels: kernel / cache / LLM. A canonical-shape question can be answered by ≥1 of these and they're compared byte-for-byte after canonicalization. Cross-witness agreement is the warrant strengthener that today's pipeline measures.
- Carrier modalities: text / claim_lattice / code / arithmetic / logic / time-series / tabular / symbolic-algebra / calculus / linear-algebra / function-sampled / combinatorics. These are the surface forms an input/output can take. Carrier-aware fixtures pin the carrier so a verifier change doesn't drift across them.
Every benchmark fixture and audit-bound canonicalization names
its kernel via pi_star_ref. A PHASE_1_CARRIERS whitelist gates
which carriers can flow through Phase-1 paths; unsupported
carriers fail explicitly with reason="unsupported_carrier",
never silent acceptance. Hidden-channel work is defensive only —
detection / flagging, never generation or concealment.
(Earlier drafts of this note used "multi-modality" and "multi- witness" interchangeably; they're not. Kernel/cache/LLM are witnesses over the same canonical bytes; text/image/audio/world will be carriers as π* grows. The pattern generalizes from one to the other but the audit semantics are distinct.)
Persistence + audit chain
Three big mechanical artifacts beyond the kernels:
-
Canonical projections persist to providence_cache — math answers (
0.1 + 0.2→3/10) get a v9.8 8-dim cache_key and an audit-event entry, just like RAG-derived answers do. The kernel is the source; the cache row is the receipt. Re-asking hits in ~40ms; chain-check verifies the chain stays intact. -
Cross-witness agreement on canonical-shape questions — on canonical-shape questions, kernel + cache + LLM fan out in parallel and we record the agreement matrix. The persisted row's
audit_modestaysCANONICAL_PROJECTIONregardless; when all three witnesses byte-equal, the render layer surfacesSTRICT-WITNESSEDas a label on top of the underlying audit_mode. When they diverge: high-value falsification data for 5F. The LLM is a witness, never authority — the kernel stays ground truth. Capital ledger records witness cost so ForkScore can compare witness-on vs witness-off forks honestly. Real-world divergence rate on the first canonical-question sweep against Hermes-3-8B: 5 of 8 questions diverged (62.5%). Hermes said1/10for0.1+0.2; saidTRUEforA IMPL B; gave the unexpanded form when handed the expanded one. Each divergence becomes a 5F-Falsification calibration fixture downstream prompt-tuning can grade against. See "Witness pipeline flowing end-to-end" below.STRICT-WITNESSEDis purely a render label, not a new admissibility/cache mode. The persistedaudit_modecolumn staysCANONICAL_PROJECTION; the witness audit event (providence_canonical_witness) layers on top so cache_key semantics don't drift. Programmatic callers see the underlyingCANONICAL_PROJECTION; human-facing surfaces see the label. -
Authorship warrant ladder — sidecar classifier on a 6-tier strength scale (
AUTHOR_PACKAGE_METADATA→AUTHOR_REPOSITORY_OWNER→AUTHOR_PAGE_BYLINE→AUTHOR_PRIMARY_PAGE_TITLE→AUTHOR_COPYRIGHT_FOOTER→AUTHOR_SECONDARY_SOURCE). Wired intoarborist inspect+ audit-line render-tail. Surfaces how the cited evidence supports an authorship claim, not just whether it does.
Selection (Merkle-AGI v8 ForkScore)
Phase 1a + 1b live: ScoredFork dataclass + arborist substrate score
CLI surface + Makefile harness (bench-fork-baseline,
bench-fork-score). A weighted score over (parent, child) battery
deltas; verdicts ACCEPT / MARGINAL / REJECT; hard-regression flag
at 5pp drop; NEG_INF_REGRESSION flag from the efficiency vocabulary.
CI-gateable: REJECT exits 1.
Content corpus — claim-pack source landed (#000029, this session)
Two companion JSON bundles dropped into the substrate this session:
axiomsg4-v2.json + theoremsg4-v2.json — Grok-4-generated
"Prometheus Maths Engine" packs covering 7 pillars (Logic, Set
Theory, Arithmetic, Geometry, Probability, Classical Physics,
λ-Calculus). 78 atomic claims (55 axioms + 23 theorems), each
dual-threaded (Δ symbolic LaTeX + ∇ verbose prose) and self-citing
to a stable classical text (Mendelson, Enderton, Hilbert, Newton,
Kolmogorov, Łukasiewicz, Aristotle).
arborist/sources/claim_pack.py ingests them at the right grain —
one Document per axiom/theorem record. Lenient JSON parser
strips markdown fences and double-escapes lone LaTeX backslashes
(\Theta, \heart, \vec) without corrupting already-correct
\\to pairs. Cross-bundle pillar-level provenance arrays
(theoremg4.json:pillar.I.logic.excludedMiddle ↔
axiomsg4.json:pillar.I) become outbound pillar_reference
edges on the first record of each pillar.
Smoke test: arborist ingest --source claim_pack --bundle axiomsg4-v2.json --bundle theoremsg4-v2.json → 78 docs, 78
chunks, 14 deduped pillar-reference edges, 78 audit events,
10/10 sampled Merkle proofs verify.
Honest ceiling: every record lands kind='surface'. The pack is
pre-distilled content but its provenance is asserted (string
field), not Merkle-proven. Records max out at ANCHOR-WARRANTED
on the four-rung ladder until cited textbooks are themselves
ingested as surfaces and a derivations.proof_blob row is
computed per record. That gap is opened as ticket #000031 below.
Retrieval lift measured the same day — apples-to-apples FTS5
search comparison on a single shard (000.db, ~867 K Wikipedia
docs) before vs after claim-pack ingest. Four representative
queries; 3/4 show claim-pack record in the top-3:
modus tollens → claim-pack at #3 (BM25 31.77)
law of excluded middle → claim-pack at #3 (related axiom)
associativity of addition → claim-pack at #1 — displaces
Wikipedia's general "Addition"
article entirely
Bayes theorem → no top-3 lift; Wikipedia's
"Bayes" + "Bayes rule" articles
dominate via short-doc BM25
bias. Body-coverage sqrt rerank
in the full QA pipeline (not
exercised by FTS5-only search)
would likely surface it.
Storage tax: sub-MB — 78 chunks against 6.2M existing chunks is
below filesystem allocation granularity. Claim-pack is earning
its tax for narrow-technical queries (where pre-distilled
CORE-shape content has title-token advantage) but doesn't help on
queries Wikipedia already covers with focused articles. Full
journal: bench/results/claim-pack-retrieval-lift-2026-05-09.md.
combinatorics@v1 π* — pure-integer counting kernel (#000032)
A new kernel landed today: tighter sibling of
algebra-symbolic@v1. Same input parser, narrower output: any
result that isn't a non-negative sp.Integer raises
PiStarError. Plain decimal output (b"10"); composes with
arithmetic@v1 for byte-identical agreement (b"10/1") so the
multi-modality witness can pin equivalence-class agreement when
both routes fire on the same question.
The 16-kernel registry now covers: text → claim_lattice → code → arithmetic → logic → time-series → tabular → symbolic-algebra → symbolic-algebra-simplified → calculus-derivative → calculus- integral → calculus-limit → calculus-series → linear-algebra → function-sampled → combinatorics. No remaining reserved stubs; the registry chapter is closed.
Pillar VII for the claim-pack (combinatorics axioms + theorems —
ticket #000033) is the natural follow-up; sequencing puts the
kernel first so pillar VII records bind to it from day one and
avoid pi_star_ref rebind churn.
Witness pipeline flowing end-to-end (#000028 validated)
The witness fan-out now writes a providence_canonical_witness
audit event when it fires. A new extractor
(bench/scripts/witness_to_5f.py) reads those events and emits
divergence-only entries as 5F-Falsification fixtures matching
the existing falsification-live-v1 schema. End-to-end smoke
ran 8 canonical-shape questions (3 arithmetic + 3 logic + 2
algebra) against ~/.arborist/shards plus the actual Hermes
endpoint:
agreement label count rate
KERNEL-LLM-DIVERGED 5 62.5%
KERNEL-LLM-AGREE 3 37.5%
─────────────────────────────────────────
divergence_count 5 62.5%
wall median / max 130 ms / 1.1 s
Five real Hermes hallucinations on questions with closed-form
ground truth: 1/10 for 0.1+0.2, TRUE for A IMPL B, the
unexpanded form when given the expanded one, and two more. Each
landed as a row in bench/fixtures/5f/falsification-witness-v1.jsonl
with the question, kernel-canonical, LLM raw, audit-event seq,
and agreement label preserved for traceability.
The calibration-data stream the multi-witness ticket imagined is now flowing: every canonical-shape question with witness=on either strengthens the warrant (3-of-3 agreement → STRICT-WITNESSED render label) or becomes a supervised-correction fixture. No human labelers are needed for canonical-shape kernel/LLM divergence labels — the kernel itself supplies ground truth. (Claim-pack validity, textbook warrant promotion, and ambiguous mathematical interpretation still benefit from human curation; the no-labeler property is scoped to the canonical-shape domain.)
Pair: make demo-plot Q='sin(x)' PNG=/tmp/sin.png closes the
opencompletion activity24-math-plot.yaml loop too — SymPy
expression → canonical bytes via function-sampled@v1 →
optional matplotlib PNG. The PNG is just a downstream view of
the canonical evidence.
Surface-ingest layer for the cited textbooks (#000031, this session)
The claim-pack ingest answered "what does the curriculum say" but
its records cap at ANCHOR-WARRANTED on the four-rung ladder
because every source_reference is a string, not a Merkle-bound
proof. This session built the surface-ingest layer that closes
the gap: a manifest-driven pipeline pulling every cited textbook
that's PD or copyleft-redistributable, processing through the
existing arborist pipeline (robots.txt → noise-strip → 512-token
chunk → Merkle root → audit event) for full-fidelity content.
License discipline is fail-closed at the URL-emit step. A
fail-closed license validator
(bench/scripts/textbooks_manifest.py) refuses to emit URLs from
entries with missing or disallowed license tokens. Allow-list:
PD, CC0, CC-BY-, CC-BY-SA-, GFDL-*, AGPL-3.0, Apache-2.0, MIT.
Excluded: any CC-BY-NC (incompatible with arborist's AGPLv3
distribution profile), any CC-BY-ND (no-derivatives prevents
chunking), proprietary. Wilf's generatingfunctionology
(educational-use license forbids rehosting) stays out — citable
but not redistributable.
Two ingest paths, both idempotent at the database layer
(content-addressed → same content → same document_root → no-op
re-insert):
make fetch-textbooks # shallow URL list → single shard
make crawl-textbooks # deep BFS via existing crawler →
# one shard per textbook id
make textbook ID=<id> # per-book convenience
make textbooks-tex # PG-style LaTeX source ingest
# (Hilbert PG #17384 + Boole PG #15114)
Eight textbooks landed across six g4 pillars (license tags
below are project-reported per the manifest — pdfsearch /
Wikisource / openmathbooks / Project Gutenberg self-attest;
none of these have been independently audited by counsel for
arborist's distribution profile, and edge-case jurisdictional
questions stay open):
| Pillar | Source | License | Format | Docs / Chunks |
|---|---|---|---|---|
| I Logic | Aristotle Prior Analytics | PD | HTML/Wikisource | 30 / 110 |
| I Logic | Aristotle Posterior Analytics | PD | HTML/Wikisource | 20 / 71 |
| I Logic | Boole Laws of Thought | PD | TeX/PG #15114 | 1 / 273 |
| I,II,III,VII | Levin Discrete Math | CC-BY-SA-4.0 | HTML/PreTeXt | 51 / 376 |
| IV Geometry | Hilbert Foundations | PD | TeX/PG #17384 | 1 / 65 |
| VI Physics | Newton Principia (Motte) | PD | HTML/Wikisource | 60 / 289 |
| VII Combin. | Bogart CTGD | GFDL-1.3 | HTML/openmathbooks | 44 / 161 |
| VII Combin. | Keller-Trotter Applied | CC-BY-SA-4.0 | HTML/PreTeXt | 80 / 168 |
| VII Combin. | Morin Open Data Structures | CC-BY-2.5 | HTML/opendatastructures | 64 / 84 |
Total surface coverage: ~351 documents, ~1597 chunks, spanning combinatorics, discrete mathematics, logic (Aristotelian + Boolean), Euclidean geometry, classical physics, and computer science.
Format coverage: the ingest pipeline now handles HTML (existing
HtmlPageSource, robots-aware, noise-stripped) and LaTeX source
(new TextbookTexSource with focused PG-aware strip pipeline —
drops preamble + comments + structural envs, keeps \textbf /
\section / \rfa argument bodies, substitutes \to → →,
\neg → ¬, \forall → ∀, etc.). Pandoc fails on PG's custom
preamble macros; a regex-based stripper is the right amount of
machinery for the well-known PG TeX format.
Pillar IX (λ-Calculus) stays open — Church 1936 + Turing 1936 are paper-length, not book-length, awaiting a paper-ingest helper. Pillar V (Probability) stays open pending Kolmogorov license analysis (German original PD-by-age in EU; US copyright restored via URAA through 2058; Morrison 1956 English translation Chelsea- copyrighted).
What's still ahead — the warrant promotion itself. Surface
ingest is the substrate for warrant promotion; the promotion
itself needs a chunk-resolution layer (per claim-pack record:
parse source_reference → resolve to a specific chunk in the
ingested surface → compute Merkle inclusion proof → write
derivations.proof_blob). That work is scoped but not landed —
~300-500 LOC across a citation parser, FTS5-driven resolver,
proof writer, and verifier wiring.
The promotion path has two distinct steps, and only the
second earns EVIDENCE-WARRANTED:
- SOURCE-ANCHORED — the cited source exists in the substrate as a Merkle-ingested surface. This is what surface ingest above delivered today: when a claim-pack record cites Hilbert, Foundations of Geometry, the substrate now contains an ingested copy. The record is bound to a real source, not a string field.
- EVIDENCE-WARRANTED — the specific claim resolves to a
specific chunk/span in the ingested surface, with a Merkle
inclusion proof persisted as
derivations.proof_blob. This is the chunk-resolution layer above; not yet landed.
The four-rung ladder today is POINTER-LINKED → ANCHOR-WARRANTED → EVIDENCE-WARRANTED → ENTAILMENT-VERIFIED (reserved). The
SOURCE-ANCHORED tier sits between ANCHOR-WARRANTED (assertion-
only) and EVIDENCE-WARRANTED (chunk-proof) as a narrower-scope
distinction worth surfacing once the chunk-resolution layer
lands; today's claim-pack records sit at ANCHOR-WARRANTED with
their cited sources newly available as ingested surfaces.
Pillar VII (combinatorics) live in the real shards (#000033, this session)
Hand-curated bundle (axiomsclaude-vii-v1.json +
theoremsclaude-vii-v1.json) authored against Stanley / Brualdi
/ Wilf / Knuth, ingested into ~/.arborist/shards/000.db
alongside the Grok-4 v2 bundles. 92 claim-pack documents now
live across pillars I–VI + VII + IX. Authorship metadata reads
"Claude blackops draft + cite-check against textbook sources"
per the §2.1 option-C provenance path.
Retrieval lift verified against the augmented shard cluster on combinatorics-shape questions:
Pascal's rule→ Pascal's Rule (claim-pack VII) at #2pigeonhole→ Strong Pigeonhole Principle at #2Modus Tollens→ Modus Tollens (claim-pack v2) at #3
Witness-sweep cron automation (this session)
The witness pipeline acquired a self-pacing harness:
bench/scripts/witness_sweep_cron.sh runs the eight-question
canonical sweep against the live Hermes endpoint, writes the
audit-event-derived 5F-Falsification fixtures into
bench/fixtures/5f/falsification-witness-v1.jsonl, and (with
--commit) commits the new divergence rows under a deterministic
message. Fail-closed: pre-commit hook failure halts the cron;
the existing fixture file is the receipt, not silently amended.
make witness-sweep-cron wraps the harness so it's schedulable
via cron, systemd timer, or any orchestrator. A second
invocation against the same shard is a no-op when no new
divergences emerged — content-addressed by question + canonical
- raw-LLM tuple. The substrate now produces calibration data continuously rather than on-demand.
Long-standing research tickets resolved (#000013, #000016, #000018)
Three open research tickets landed closure artifacts this session — bringing the open-research backlog from "three open tickets, no artifacts" to "two closed, one parked, all with durable docs."
#000018 — Adversarial soft-hash covert-channel analysis.
Closed. Landed docs/soft-hash-channel-analysis.md (~432
lines). The threat model is now explicit:
- T1 (chosen-input): adversary picks
xto leak bits of a statesthrough soft-hashφ(x, s). Reduces to SHA-256 partial-preimage; bounded. - T2 (chosen-state): adversary picks
s. Reduces to SHA-256 output uniformity over freshx; bounded. - T3 (replay-window): adversary observes many
(x, s)pairs over time. Open. Needs per-window budget bound — opened as ticket #000036 below.
Mitigations table (M1: domain separation by purpose tag; M2: per-checkpoint nonce; M3: drop the soft-hash anchor entirely; M4: ZK-replace soft-hash with a SNARK-friendly one). M2 is the recommended default — cheap, composes with existing audit chain, no breaking change. M3 is the fallback if covert-channel analysis collapses; M4 deferred to #000016.
#000013 — Spatial-temporal substrate (Merkle-AGI v7-W). Closed. Landed three artifacts:
docs/_source/merkle-agi-v7w-spatial-temporal.rst(658-line substrate paper with Parts 1-6 + appendix). Hierarchical-grid spatial discretization, frame-as-committed-object, four ε-frontiers (pose_integration,observation_update,object_logits,relation_logits). Inherits v9.8 audit protocol unchanged; adds aframeMerkle node committing(time_ns, pose, grid_cells, observations)as a single audit-bound object.docs/v7w-frontier-catalog.md— operator-facing reference for the four ε-frontiers including which kernel produces each frontier's observable, the noise model, and the recommended ε for each at consumer-camera resolution.arborist/world/__init__.py— namespace stub (V7W_VERSION='v0-draft',STATUS='namespace_reserved') reserves the import path so a future world-model integration doesn't churn module names. No code yet — pure paper + reservation.
#000016 — ZK Phase-2 frontier proof. Parked with artifacts. Landed:
docs/zk-frontier-bench.md— bench plan + acceptance thresholds for a future sibling repo (arborist-zk-bench). Plonky3 on commodity hardware (Apple M3 Max + Linux x86), three circuit sizes (256 / 1024 / 4096), measure prover wall-clock + proof bytes + verify time + memory peak. Acceptance threshold: ≤30 s prover at size 4096, ≤100 KB proof, ≤100 ms verify. Preliminary projection from published Plonky3 / Halo2 benches predicts size-4096 prover exceeds threshold — most-likely outcome is "PARKED for LLM-frontier scale, VIABLE for small distillation models."docs/zk-wire-protocol.md— versioned consumer-side schema spec (arborist-zk-proof-v1). Defines the JSON envelope arborist consumes when ZK proofs become available; covers binding into v9.8 audit chain (newmodel_weights_zk_root+frontier_proof_circuit_idcolumns onprovidence_cache), Ed25519 issuer signatures, versioning + forward-compat, threat model. The Rust toolchain stays in the sibling repo per arborist's Python-only language constraint.
The hand-wave from v7 §16.1 ("swap SHA-256 → Poseidon") is replaced with explicit thresholds + parked-status with named open work.
Three #000018 follow-up tickets opened
The soft-hash analysis surfaced three open sub-questions, each queued as its own ticket:
- #000034 — Hessian alignment under φ_linear. Open · awaiting go/no-go. The linear soft-hash projection φ_linear(x, s) = (Hx ⊕ s) mod p assumes Hessian alignment; formal sufficient conditions for that assumption haven't been written down. Without them, claims about gradient leakage are conjectural.
- #000035 — PRG choice for φ_PRG. Open · awaiting go/no-go. φ_PRG(x, s) currently calls HMAC-SHA-512 over a domain-separated tag; that's a defensible default but not the reasoned choice. This ticket compares HMAC-SHA-512 vs AES-CTR-DRBG vs ChaCha20-DRBG on (security margin, perf, spec stability, FIPS path).
- #000036 — T3 per-window covert-channel budget bound. Open · awaiting go/no-go. The replay-window threat from §3.3 of the analysis doc has no quantitative bound; this ticket commissions one (per-window leak budget in bits as a function of window size, observation rate, and committed state entropy).
All three are design-only; the analysis doc is the floor under each.
Real-shard latency (post-virt-back)
The first real-world who wrote virt-back? query landed
EVIDENCE-WARRANTED, 2/2, primary source at #1 — proving that a
narrow crawl (one blog) plus the verifier scaffold can recover
authorship without global web knowledge. Latency was 75 s on cold
cache; we found two structural defects:
- 588 SQLite
connect()calls per query, each running 7 forward-migration probes on already-migrated shards. synonym_expanddoing a 290 K-row full-table scan per query for one-token-of-interest neighborhood.
Per-process migration memoization + lazy concept-relations queries took the wall budget from 14.5 s → 9.3 s warm-cache (-36%); cold cache extrapolates substantially lower. Median across 8-question real-shard baseline: 4.1 s.
What your framework gave us
The 5-axis structure is, on reflection, exactly what a Merkle-bound audit substrate needs to grow into a self-improving system:
5S — invariant identity (does the canonicalizer collapse
equivalent surface forms?)
5T — temporal coherence (can the substrate reason across
time / transfer learning?)
5F — operational discipline (can the substrate falsify / refine
/ formulate / function correctly?)
5R — recovery + react (can the substrate restore from a
falsification, recall prior state,
react to new observations?)
What we found in practice: each axis surfaces a different failure mode the Merkle chain can't see on its own. 5S catches canonicalizer drift (the chain is honest about an answer, but the canonicalizer let two unequal things look equal). 5T catches memory-update bugs (the chain says A happened then B; the system remembers only A). 5F catches LLM hallucination on questions with ground truth available. 5R catches state-reconstruction failures on cold-start.
Without the 5R battery we wouldn't have known the cache-restoration path was incomplete. Without the 5F Falsification battery we wouldn't have a path from divergence event → calibration data → prompt improvement. Without the 5T Time sub-battery we wouldn't have a way to test the memory-snapshot chain.
The framework is doing exactly what taxonomy is supposed to do: making it impossible to forget the modes of failure.
What we owe the framework
A few structural choices we made that you may or may not have intended; documenting them so the next maintainer can inherit correctly:
-
Live + synthetic split per sub-battery — every 5F sub-battery has both a
*-v1.jsonl(synthetic, deterministic by construction) and*-live-v1.jsonl(routes through the real arborist subsystem, e.g.qa.parse_claims,memory.snapshot,selfmodel.store_snapshot). The synthetic side pins the evaluator contract; the live side catches integration regressions the synthetic side can't see. We treated this as a "Phase 1b.2" discipline and applied it to all 5F sub-batteries. -
pi_star_refis mandatory metadata, not optional — every bench fixture names the canonicalizer. Without this, a verifier change drifts silently across batteries; with it, a registry key can pin which kernel a given fixture exercised. Cross-battery moves require a deliberate policy bump. -
Carrier whitelist closes the type system — Phase-1 carriers are an explicit
frozenset; bench batteries that name an off-list carrier fail with a clear reason rather than producing garbage. As the registry grew (text → claim_lattice → code → arithmetic → logic → time-series → symbolic-algebra → calculus → linear-algebra → function-sampled → tabular), the whitelist grew with it. Discipline preserved. -
Closure-criterion tests — there's a test that asserts every reserved-stub π* has graduated. Adding a new reserved stub (a future modality the substrate paper reserves but doesn't yet implement) re-opens this list; that's the governance event the test pins. Today the list is empty.
Numbers
arborist/pi_star/*.py ~3.2K LOC across 15 kernels
arborist tests 1,684 passing, 37 skipped (sympy-gated)
bench fixtures (jsonl) 50+ files, 662+ default tasks
canonical-bytes domains text · claim_lattice · code · arithmetic
· logic · time-series · tabular ·
symbolic-algebra · calculus · linear-
algebra · function-sampled · combinatorics
audit-chain coverage every state-changing op writes
append_audit; chain-check-shards
reports 0 breaks across all shards
real-shard baseline (warm) median 4.1 s, max 7.2 s, all 8 questions
pass; primary source at #1 in 4/8
witness sweep (live Hermes) 5/8 KERNEL-LLM-DIVERGED, 3/8 agree;
median 130 ms, max 1.1 s; divergences
auto-extracted to 5F fixtures
witness modes reachable KERNEL-LLM-AGREE, STRICT-WITNESSED,
CACHE-DRIFT, LLM-DIVERGED,
KERNEL-LLM-DIVERGED, KERNEL-CACHE-AGREE,
KERNEL-ONLY
ForkScore verdict thresholds 5pp signal floor, hard-regression flag
at -5pp on any sub-battery, NEG_INF_REGRESSION
flag from efficiency vocabulary
composition discipline 12 dedicated tests cover idempotency
(algebra-symbolic ∘ algebra-symbolic),
Pythagorean identity collapse via
algebra-symbolic-simplified, manifest-
fingerprint stability + order-sensitivity,
composite ≡ manual-chain bytes
What's next
Open work, ranked by what would extend the substrate furthest:
- CI re-enable — the harness is unguarded; lifting the gate converts "tested when we run it" to "tested on every commit." Will catch substrate drift before it reaches a real shard.
- opencompletion integration — your
activity24-math-plot.yamluses SymPy + numpy + matplotlib. Thefunction-sampled@v1π* gives plots a canonical-bytes identity (PNG = downstream view).make demo-plot Q='sin(x)' PNG=/tmp/sin.pngalready lands; the bridge from activity → arborist (so students' work feeds the witness/calibration stream) is the remaining wiring. - Sibling repos for non-Python toolchains — ZK proofs, world-model integrations, language-port mesh peers. Arborist stays Python-only; the canonical bytes are the contract downstream tools honor.
- Research tickets — resolved this session —
#000018 (soft-hash covert-channel analysis) closed via
docs/soft-hash-channel-analysis.mdwith mitigation table M1-M4; three follow-up tickets opened (#000034 Hessian alignment, #000035 PRG choice, #000036 T3 per-window bound). #000013 (spatial-temporal v7-W) closed via 658-line substrate paper + frontier catalog + namespace stub. #000016 (ZK Phase-2 frontier proof) parked with bench-plan + wire-protocol artifacts; sibling-repo measurement remains open work.
Exploration tickets opened this session (#000031 / #000032 / #000033)
Three new tickets staged for go/no-go, all chained off the claim-pack corpus and math-substrate work:
#000031 — Surface-ingest cited textbooks for claim-pack
warrant promotion. Closes the warrant gap left open at the end
of #000029: today every claim-pack record caps at
ANCHOR-WARRANTED because source_reference is a string field,
not a Merkle-bound proof. Ingesting the cited textbooks as
SURFACE-layer documents + computing per-claim
derivations.proof_blob lets the four-rung ladder promote them
to EVIDENCE-WARRANTED.
License gating is the first hard constraint: PD sources
(Hilbert, Newton, Kolmogorov, Łukasiewicz, Aristotle) form the
green-light scope. Mendelson + Enderton are proprietary and stay
yellow-light pending an explicit decision (purchased single
copy / library license / PD substitute via Hilbert-Ackermann
1928). Two follow-up tickets reserved: textbook-fetch pipeline
and chunk-resolution layer (mapping source_reference strings
to specific spans inside ingested textbooks; the bridge that
lets proof_blob actually be computed).
#000032 — combinatorics@v1 π (integer counting kernel).*
Tighter domain than algebra-symbolic@v1: the latter happily
returns Integer(6) for binomial(-3, 2) (generalized binomial
via Gamma) and leaves binomial(n, k) symbolic. This kernel
fails closed on any input whose result isn't a non-negative
sp.Integer. Operators choose the kernel by what they want
rejected. Output format b"10" composes with arithmetic@v1
for byte-identical agreement (b"10/1") — the witness flow
(#000028) becomes computable on counting questions once two
modalities agree on the answer's shape.
#000033 — Claim-pack pillar VII (combinatorics). Extends #000029 with a counting pillar slotting into the documented gap in the v2 bundles (existing pack uses I, II, III, IV, V, VI, IX — VII and VIII reserved for future extension). 7 axioms (Pascal's rule, addition principle, multiplication principle, pigeonhole, factorial / binomial definitions) + 7 theorems (binomial theorem, inclusion-exclusion in counting form, hockey-stick, Vandermonde, Catalan closed form, stars-and-bars, strong pigeonhole).
Open question is bundle provenance: commission a Grok-4 v3
bundle for parity with the existing pack, hand-curate from
classical sources (Stanley, Brualdi, Wilf, Knuth), or hybrid
(LLM draft + human curation). Hard constraint: explicit
authorship metadata. No silent invention. Sequencing: #000032
lands first so pillar VII records bind to the tighter kernel
from day one — avoids rebind churn on pi_star_ref fields.
All three are design-only at this point. The claim-pack corpus is already in the substrate; these extend the warrant chain upward (#000031), tighten the domain at the kernel layer (#000032), and broaden the curated-claim coverage (#000033).
Closing
You built a taxonomy with five axes and twenty-one sub-batteries, and what we found is that every axis was load-bearing. The substrate that grew into this shape didn't shape itself to fit your framework — the framework was already shaped to catch the modes that mattered.
Every UNDF post, every patch, every disclosure on undefect.com is public domain — free, open intellectual capital, inheritable by anyone, forever. The work above falls under that contract too. We patch the planet because the planet patches each other; you're in that lineage.
If there's a sixth axis we haven't found yet, the test suite will tell us.
— fox + blackops
permacomputer / unsandbox / unturf
original draft: arborist commit a2ff9d4 (2026-05-09)
revision (this file): post-review errata pass on 2026-05-10
current substrate state: CI re-enabled (skip wikipedia ingest);
shard search fan-out cuts real-shard query wall ~4.3× (66s →
15.5s on the 4-shard cluster). Warrant-promotion chunk-
resolution layer still ahead.