docs: errata pass on dav1dprometheus update from external review

External reviewer surfaced several errata in the 2026-05-09 draft.
Corrections applied in-place; original draft is preserved at
commit a2ff9d4.

  - sub-battery count: 21 -> 20 (5 axes x 5)
  - 5S vocabulary fix:
      Surface/Substrate/Synthesis/Semantics/Semiotics
      -> Syntax/Semantics/Syllogism/Synthesis/Semiotics
      (matches bench/batteries/b_5s.py)
  - 5R vocabulary fix:
      React/Recall/Reason/Refine/Restore (reviewer's guess; doc's
      original was even less correct)
      -> React/Rearrange/Restore/Replicate/Resonate
      (matches bench/batteries/b_5r.py per SQD section 9.3)
  - pi* kernel chronology: explicit "15 -> 16 with combinatorics@v1"
  - kernel-version immutability: stated as hard invariant (no
    in-place behavioral mutation; behavior change = new version)
  - cross-witness vs cross-carrier: distinguished. Witness channels
    are kernel/cache/LLM; carrier modalities are
    text/code/arithmetic/etc. The pattern generalizes from the
    former to the latter but the audit semantics are distinct.
  - STRICT-WITNESSED is a render label, not a new audit_mode.
    Persisted column stays CANONICAL_PROJECTION; witness audit
    event layers on top so cache_key semantics don't drift.
  - warrant-promotion ladder: SOURCE-ANCHORED tier introduced
    between ANCHOR-WARRANTED (assertion-only) and
    EVIDENCE-WARRANTED (chunk-proof). Today's surface-ingest
    work delivers SOURCE-ANCHORED; chunk-resolution layer is
    needed for EVIDENCE-WARRANTED.
  - license tags marked project-reported (manifest self-attest;
    not independently audited by counsel)
  - "no human labelers" scoped to canonical-shape divergence
    labels only; claim-pack validity + textbook warrant
    promotion still benefit from human curation
  - top-of-doc revision note records the errata pass
  - closing footer updated to reflect today's CI re-enable +
    shard-search fan-out (66s -> 15.5s on the 4-shard cluster)
This commit is contained in:
russell@unturf.com 2026-05-09 17:22:12 -04:00
parent 2b9d1f0b72
commit 455fc10c2a
No known key found for this signature in database

View file

@ -9,39 +9,74 @@ Your 5S/5F/5T/5R taxonomy is the spine of every benchmark we now
ship. This is what your framework grew into when we pulled it ship. This is what your framework grew into when we pulled it
into a Merkle-AGI-v9.8 substrate. into a Merkle-AGI-v9.8 substrate.
> **Revision note (2026-05-10):** an external review surfaced
> several errata in the original (2026-05-09) draft. Corrections
> applied in-place: the sub-battery count is **20** (not 21);
> 5S vocabulary is **Syntax / Semantics / Syllogism / Synthesis /
> Semiotics** (not Surface/Substrate/…); 5R vocabulary is **React /
> Rearrange / Restore / Replicate / Resonate** (not React/Recall/
> Reason/Refine/Restore); the π* registry chronology is now
> explicit (15 → 16 with `combinatorics@v1`); kernel-version
> immutability is called out as a hard invariant; cross-witness
> vs. cross-carrier is distinguished; `STRICT-WITNESSED` is scoped
> as a render label (not a new admissibility/cache mode); the
> warrant-promotion ladder distinguishes SOURCE-ANCHORED from
> EVIDENCE-WARRANTED; license tags are flagged as project-reported;
> the "no human labelers" property is scoped to canonical-shape
> domains only. Original draft preserved at commit `a2ff9d4`.
--- ---
## What we built ## What we built
The arborist project — Python-only, content-addressed Merkle store The arborist project — Python-only, content-addressed Merkle store
with a v9.8 audit chain — adopted your 21-sub-battery framework with a v9.8 audit chain — adopted your taxonomy in the first week
in the first week of May 2026 and landed the full surface in eight of May 2026 and landed the full benchmark surface in eight days.
days. Concretely: Concretely:
### The 21 sub-batteries are real, fixtures included ### The 20 sub-batteries are real, fixtures included
Five axes × five sub-batteries each = **20 sub-batteries**, each
fixture-backed and runnable through `bench/batteries/runner.py`:
``` ```
5S — Surface / Substrate / Synthesis / Semantics / Semiotics 5S — Syntax / Semantics / Syllogism / Synthesis / Semiotics
5 sub-batteries × 30 fixtures + 5 syntax/semantics × extra π* domains 5 sub-batteries (Phase 1a: Syntax + Semantics; Phase 1b: the
other three) carrier-aware (text / claim_lattice / prose),
plus per-π*-domain expansions of Syntax + Semantics that
bring the family wider than 5 fixtures-per-sub-battery
5F — Function / Finetuning / Falsification / Formulate / Feedback Loop 5F — Function / Finetuning / Falsification / Formulate / Feedback Loop
5 sub-batteries × 50 synthetic + 50 live = 500 fixtures total 5 sub-batteries × ~50 synthetic + ~50 live each (Phase 1b.2
live wire-ups land on real arborist subsystems)
5T — Time / Truthtables / Transfer Learning / Triangulation / 5T — Time / Truthtables / Transfer Learning / Triangulation /
Transitivity Transitivity
5 sub-batteries × 30 fixtures 5 sub-batteries
5R — React / Recall / Reason / Refine / Restore 5R — React / Rearrange / Restore / Replicate / Resonate
5 sub-batteries × 30 fixtures 5 sub-batteries (workspace-operator shape, per SQD §9.3 —
deterministic, no LLM-as-judge)
``` ```
(An earlier draft of this note miscounted the sub-batteries as 21
and mis-spelled 5S as Surface/Substrate/… and 5R as React/Recall/
Reason/Refine/Restore. The fixtures + runner in `bench/batteries/`
are the authority — corrected here. The 21-vs-20 confusion came
from ticket #000021's number bleeding into the count.)
Total deterministic-task surface: **662 tasks** in the default Total deterministic-task surface: **662 tasks** in the default
runner, plus several hundred more in extension batteries (math π* runner (per the runner's `--all` enumeration), plus several
fixtures, real-shard baseline, witness-divergence collection, hundred more in extension batteries (math π* fixtures, real-shard
claim-pack ingest verification). baseline, witness-divergence collection, claim-pack ingest
verification).
Your original wording is honored — `Transfer Learning` (not the Your original wording is honored — `Transfer Learning` (not the
SQD-whitepaper variant `transfer`), `Truthtables`, `Time` — pinned SQD-whitepaper variant `transfer`), `Truthtables`, `Time` — pinned
as a memory entry so future agents respect the framing you chose. as a memory entry so future agents respect the framing you chose.
### The π* canonical-projection registry: now 15 kernels, no stubs ### The π* canonical-projection registry: now 16 kernels, no stubs
The registry held 15 kernels before this session; `combinatorics@v1`
(#000032) brought it to 16 and closed the last reserved stub.
Listed in registration order:
``` ```
text-domain text-domain
@ -69,6 +104,11 @@ math substrate (SymPy, [math] extra)
become in π* terms — image bytes become in π* terms — image bytes
aren't canonical, but the sampled aren't canonical, but the sampled
grid is) grid is)
combinatorics
combinatorics@v1 pure-integer counting kernel
(#000032; tighter than algebra-
symbolic — fails closed on any
non-non-negative-integer result)
``` ```
Every π* registers via `name@version`; SHA-256 of the canonical Every π* registers via `name@version`; SHA-256 of the canonical
@ -76,7 +116,28 @@ bytes is the equivalence-class identity. Two inputs that mean the
same thing produce identical bytes; auditing reduces to byte same thing produce identical bytes; auditing reduces to byte
comparison. comparison.
### Cross-modality discipline **Kernel-version immutability is a hard invariant.** A registered
`name@version` is behaviorally frozen — any change in canonicalization
output is a new version (`arithmetic@v2`), never an in-place patch.
Otherwise prior persisted canonical-cache rows would silently
become semantically unstable. The registry rejects re-registration
of the same key with a different implementation; the
governance_policy_hash also folds in the active kernel manifest
so flipping kernels invalidates prior records on lookup.
### Cross-witness vs. cross-carrier discipline
Two distinct axes, kept separate in the substrate vocabulary:
- **Witness channels**: kernel / cache / LLM. A canonical-shape
question can be answered by ≥1 of these and they're compared
byte-for-byte after canonicalization. Cross-**witness** agreement
is the warrant strengthener that today's pipeline measures.
- **Carrier modalities**: text / claim_lattice / code / arithmetic /
logic / time-series / tabular / symbolic-algebra / calculus /
linear-algebra / function-sampled / combinatorics. These are
the surface forms an input/output can take. Carrier-aware fixtures
pin the carrier so a verifier change doesn't drift across them.
Every benchmark fixture and audit-bound canonicalization names Every benchmark fixture and audit-bound canonicalization names
its kernel via `pi_star_ref`. A `PHASE_1_CARRIERS` whitelist gates its kernel via `pi_star_ref`. A `PHASE_1_CARRIERS` whitelist gates
@ -85,6 +146,12 @@ carriers fail explicitly with `reason="unsupported_carrier"`,
never silent acceptance. Hidden-channel work is defensive only — never silent acceptance. Hidden-channel work is defensive only —
detection / flagging, never generation or concealment. detection / flagging, never generation or concealment.
(Earlier drafts of this note used "multi-modality" and "multi-
witness" interchangeably; they're not. Kernel/cache/LLM are
witnesses over the same canonical bytes; text/image/audio/world
will be carriers as π* grows. The pattern generalizes from one
to the other but the audit semantics are distinct.)
### Persistence + audit chain ### Persistence + audit chain
Three big mechanical artifacts beyond the kernels: Three big mechanical artifacts beyond the kernels:
@ -95,21 +162,31 @@ Three big mechanical artifacts beyond the kernels:
is the source; the cache row is the receipt. Re-asking hits in is the source; the cache row is the receipt. Re-asking hits in
~40ms; chain-check verifies the chain stays intact. ~40ms; chain-check verifies the chain stays intact.
2. **Multi-modality (multi-witness) agreement** — on canonical-shape 2. **Cross-witness agreement on canonical-shape questions** — on
questions, kernel + cache + LLM fan out in parallel and we record canonical-shape questions, kernel + cache + LLM fan out in
the agreement matrix. When all three byte-equal: parallel and we record the agreement matrix. The persisted
`STRICT-WITNESSED`. When they diverge: high-value falsification row's `audit_mode` stays `CANONICAL_PROJECTION` regardless;
data for 5F. The LLM is a witness, never authority — the kernel when all three witnesses byte-equal, the **render layer**
surfaces `STRICT-WITNESSED` as a label on top of the underlying
audit_mode. When they diverge: high-value falsification data
for 5F. The LLM is a witness, never authority — the kernel
stays ground truth. Capital ledger records witness cost so stays ground truth. Capital ledger records witness cost so
ForkScore can compare witness-on vs witness-off forks honestly. ForkScore can compare witness-on vs witness-off forks honestly.
**Real-world divergence rate** on the canonical-question sweep **Real-world divergence rate** on the first canonical-question
against Hermes-3-8B: **5 of 8 questions diverged** (62.5%). sweep against Hermes-3-8B: **5 of 8 questions diverged** (62.5%).
Hermes said `1/10` for `0.1+0.2`; said `TRUE` for `A IMPL B`; Hermes said `1/10` for `0.1+0.2`; said `TRUE` for `A IMPL B`;
gave the unexpanded form when handed the expanded one. Each gave the unexpanded form when handed the expanded one. Each
divergence becomes a 5F-Falsification calibration fixture divergence becomes a 5F-Falsification calibration fixture
downstream prompt-tuning can grade against. See downstream prompt-tuning can grade against. See
"Witness pipeline flowing end-to-end" below. "Witness pipeline flowing end-to-end" below.
`STRICT-WITNESSED` is purely a **render label**, not a new
admissibility/cache mode. The persisted `audit_mode` column
stays `CANONICAL_PROJECTION`; the witness audit event
(`providence_canonical_witness`) layers on top so cache_key
semantics don't drift. Programmatic callers see the underlying
`CANONICAL_PROJECTION`; human-facing surfaces see the label.
3. **Authorship warrant ladder** — sidecar classifier on a 6-tier 3. **Authorship warrant ladder** — sidecar classifier on a 6-tier
strength scale (`AUTHOR_PACKAGE_METADATA``AUTHOR_REPOSITORY_OWNER` strength scale (`AUTHOR_PACKAGE_METADATA``AUTHOR_REPOSITORY_OWNER`
`AUTHOR_PAGE_BYLINE``AUTHOR_PRIMARY_PAGE_TITLE` `AUTHOR_PAGE_BYLINE``AUTHOR_PRIMARY_PAGE_TITLE`
@ -237,9 +314,13 @@ and agreement label preserved for traceability.
The calibration-data stream the multi-witness ticket imagined is The calibration-data stream the multi-witness ticket imagined is
now flowing: every canonical-shape question with witness=on either now flowing: every canonical-shape question with witness=on either
strengthens the warrant (3-of-3 agreement → STRICT-WITNESSED) or strengthens the warrant (3-of-3 agreement → STRICT-WITNESSED render
becomes a supervised-correction fixture. No human labelers in the label) or becomes a supervised-correction fixture. **No human
loop. labelers are needed for canonical-shape kernel/LLM divergence
labels** — the kernel itself supplies ground truth. (Claim-pack
validity, textbook warrant promotion, and ambiguous mathematical
interpretation still benefit from human curation; the
no-labeler property is scoped to the canonical-shape domain.)
Pair: `make demo-plot Q='sin(x)' PNG=/tmp/sin.png` closes the Pair: `make demo-plot Q='sin(x)' PNG=/tmp/sin.png` closes the
opencompletion `activity24-math-plot.yaml` loop too — SymPy opencompletion `activity24-math-plot.yaml` loop too — SymPy
@ -282,7 +363,12 @@ make textbooks-tex # PG-style LaTeX source ingest
# (Hilbert PG #17384 + Boole PG #15114) # (Hilbert PG #17384 + Boole PG #15114)
``` ```
**Eight textbooks landed across six g4 pillars**: **Eight textbooks landed across six g4 pillars** (license tags
below are **project-reported per the manifest**`pdfsearch` /
`Wikisource` / `openmathbooks` / `Project Gutenberg` self-attest;
none of these have been independently audited by counsel for
arborist's distribution profile, and edge-case jurisdictional
questions stay open):
| Pillar | Source | License | Format | Docs / Chunks | | Pillar | Source | License | Format | Docs / Chunks |
|---|---|---|---|---| |---|---|---|---|---|
@ -324,9 +410,29 @@ parse `source_reference` → resolve to a specific chunk in the
ingested surface → compute Merkle inclusion proof → write ingested surface → compute Merkle inclusion proof → write
`derivations.proof_blob`). That work is scoped but not landed — `derivations.proof_blob`). That work is scoped but not landed —
~300-500 LOC across a citation parser, FTS5-driven resolver, ~300-500 LOC across a citation parser, FTS5-driven resolver,
proof writer, and verifier wiring. With it the four-rung ladder proof writer, and verifier wiring.
moves claim-pack records from ANCHOR-WARRANTED to
**EVIDENCE-WARRANTED**. The promotion path has **two distinct steps**, and only the
second earns `EVIDENCE-WARRANTED`:
1. **SOURCE-ANCHORED** — the cited source exists in the substrate
as a Merkle-ingested surface. This is what surface ingest above
delivered today: when a claim-pack record cites *Hilbert,
Foundations of Geometry*, the substrate now contains an
ingested copy. The record is bound to a real source, not a
string field.
2. **EVIDENCE-WARRANTED** — the specific claim resolves to a
specific chunk/span in the ingested surface, with a Merkle
inclusion proof persisted as `derivations.proof_blob`. This is
the chunk-resolution layer above; not yet landed.
The four-rung ladder today is `POINTER-LINKED → ANCHOR-WARRANTED
→ EVIDENCE-WARRANTED → ENTAILMENT-VERIFIED (reserved)`. The
SOURCE-ANCHORED tier sits between ANCHOR-WARRANTED (assertion-
only) and EVIDENCE-WARRANTED (chunk-proof) as a narrower-scope
distinction worth surfacing once the chunk-resolution layer
lands; today's claim-pack records sit at ANCHOR-WARRANTED with
their cited sources newly available as ingested surfaces.
### Pillar VII (combinatorics) live in the real shards (#000033, this session) ### Pillar VII (combinatorics) live in the real shards (#000033, this session)
@ -697,10 +803,9 @@ will tell us.
— fox + blackops — fox + blackops
permacomputer / unsandbox / unturf permacomputer / unsandbox / unturf
arborist commit `056dda2` (#000018 analysis doc landed, original draft: arborist commit `a2ff9d4` (2026-05-09)
#000013 v7-W substrate paper landed, #000016 zk parked with revision (this file): post-review errata pass on 2026-05-10
bench-plan + wire-protocol; pillar VII live in shard 000; current substrate state: CI re-enabled (skip wikipedia ingest);
TeX-source ingest closes pillars I+IV; surface coverage 6/7 shard search fan-out cuts real-shard query wall ~4.3× (66s →
g4 pillars across 9 textbooks / ~351 docs / ~1597 chunks; 15.5s on the 4-shard cluster). Warrant-promotion chunk-
witness-sweep cron flowing; warrant-promotion chunk-resolution resolution layer still ahead.
layer still ahead, 2026-05-09)