pi_star: time-series-quantized@v1 graduates + Substrate docs section

Two items from the menu, fanned out:

1. time-series-quantized@v1 — last meaningful π* stub graduates.
   Sample-array carrier (sensor / temporal data) joins the registry
   alongside text · claim_lattice · code · arithmetic · logic.
   Quantizes to (dt, dv) grid, sorts by timestamp, dedupes
   collisions (last wins), serializes as integer-vector text:

       dt=1;dv=0.1;n=2;t0=0:10|20

   Equivalence classes preserved: timestamp jitter < Δ_t,
   value jitter < Δ_v/2 (banker's rounding), out-of-order samples,
   different JSON presentation. Distinct: any change to dt/dv grid,
   any quantized value or timestamp difference. Projective —
   canonical text is not valid JSON, so re-canonicalization raises.

   - 13 unit tests in tests/test_pi_star.py (jitter, dedupe, sort,
     fractional dv, error paths, idempotency-projective)
   - 10 syntax + 12 semantics fixtures under bench/fixtures/5s/
     (10/10 + 12/12 pass)
   - bench-5s-time-series Makefile target
   - time_series added to PHASE_1_CARRIERS whitelist
   - tabular-pinned@v1 is now the only remaining stub

2. Substrate docs — first formal coverage of the registry, bench
   harness, and v8 ForkScore at arborist.unturf.com:

   - docs/_source/pi-star.rst: registry overview, cross-modality
     discipline (carrier + pi_star_ref), math π* highlights
     (arithmetic + logic-kernel worked examples), composition
     algebra pointer, authoring checklist (8 steps).
   - docs/_source/bench.rst: 5S/5T/5F/5R structure, sub-batteries,
     phase-1 carriers, ForkScore integration, fixture format,
     reproducibility (runtime_digest, fixture_digest).
   - docs/_source/v8-fork-score.rst: formula, default weights,
     verdict thresholds (ACCEPT/MARGINAL/REJECT), hard-regression +
     NEG_INF_REGRESSION flags, CLI usage.
   - index.rst gets a "Substrate" toctree section above the existing
     module-reference autosummary.

   Sphinx build clean (3 new pages, no new warnings).

Test suite: 1269 passed, 36 skipped.
This commit is contained in:
russell@unturf.com 2026-05-08 09:33:19 -04:00
parent e4ecf89621
commit 508a25076d
No known key found for this signature in database
10 changed files with 739 additions and 12 deletions

140
docs/_source/bench.rst Normal file
View file

@ -0,0 +1,140 @@
Benchmark surface
=================
Arborist ships the complete **Dav1DPrometheus 5S/5T/5F/5R** evaluation
suite — 21 sub-batteries, ~660 deterministic fixtures — as
first-class infrastructure. Every benchmark is reproducible, no
LLM-as-judge, and many sub-batteries route through the actual
arborist surface (parser, verifier, audit chain, π* registry) rather
than synthetic gold output.
Quick reference
---------------
.. code-block:: bash
make bench-suite # complete 5S + 5T + 5F + 5R suite
make bench-5s # representational discipline
make bench-5t # temporal / cross-reasoning
make bench-5f # operational quality (Phase 1a embedded)
make bench-5f-live # 5F bridged to live arborist surfaces (Phase 1b.2)
make bench-5r # workspace operators
make bench-5s-math # arithmetic@v1 + logic-kernel@v1 fixtures
make bench-5s-code # code-py-ast@v1 fixtures
Each invocation emits a JSON :class:`bench.batteries.base.BatteryResult`
with per-task pass/fail, fixture digest, runtime digest, and
sub-battery-specific metrics.
The four batteries
------------------
5S — representation discipline
~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
Five sub-batteries testing what the system understands at the
sign / meaning / derivation level:
- **Syntax** — does the named π* parse the input without raising?
- **Semantics** — do two surface forms canonicalize to the same
bytes when they should (and not when they shouldn't)?
- **Syllogism** — does each step in a deductive chain validly
follow under the named rule (categorical_transitivity, chain_3,
invalid_converse, missing_premise)?
- **Synthesis** — does the system assemble cited facts into a
coherent derivation supported by the fact set?
- **Semiotics** — is meaning preserved under controlled label
swaps?
5T — temporal / cross-reasoning discipline
~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
Six sub-batteries (legacy ``transfer`` plus the canonical
Dav1DPrometheus five):
- **Transfer Learning** — does a learned pattern carry across
task / domain / carrier?
- **Triangulation** — do independent strategies (substring,
token_subset, token_overlap, entity_match) agree at threshold?
- **Truthtables** — exhaustive propositional coverage at N=2..4
variables.
- **Transitivity** — typed-relation chains under whitelist
(``implies``, ``subset_of``, ``ancestor_of``, ``before``,
``less_than``).
- **Time** — temporal-context preservation across memory_root
snapshots; integrates with #000017 surface.
5F — operational quality
~~~~~~~~~~~~~~~~~~~~~~~~
Five sub-batteries, all with **embedded** (Phase 1a synthetic) AND
**live** (Phase 1b.2, routes through real arborist surfaces) modes:
================ =============================================
Sub-battery Live surface
================ =============================================
Function ``arborist.qa.parse_claims.parse_pointer_claims``
Finetuning ``arborist.selfmodel.store_snapshot`` round-trip
Falsification ``arborist.qa.verify.verify_quotes``
Formulate ``arborist.qa.parse_claims.parse_pointer_claims``
Feedback Loop ``arborist.store.append_audit`` + ``memory.snapshot``
================ =============================================
Per-task ``detail.source`` reports ``"embedded"`` or ``"live"`` so
bench output distinguishes synthetic from production signal.
5R — workspace operators
~~~~~~~~~~~~~~~~~~~~~~~~
Five sub-batteries testing operators applied to a workspace
(SelfModel + memory_root + audit chain):
- **React** — observations integrate into downstream state.
- **Rearrange** — restructure without semantic shift.
- **Restore** — retrieve prior facts from history.
- **Replicate** — π*-determinism across N replicas.
- **Resonate** — variance-zero across N runs.
Cross-modality discipline
-------------------------
Every fixture carries:
- ``carrier`` — domain whitelist enforced by
:data:`bench.batteries.base.PHASE_1_CARRIERS`. Phase 1
domains: ``text``, ``claim_lattice``, ``memory_snapshot``,
``selfmodel_snapshot``, ``providence_record``, ``audit_event``,
``code``, ``arithmetic``, ``logic``, ``time_series``.
- ``domain`` — sub-domain qualifier (e.g.,
``rational``, ``propositional``, ``python_ast``).
- ``pi_star_ref`` — registry key naming the canonicalizer.
- ``loss_report_refs`` — optional projection-loss links.
- ``modality_notes`` — scope note.
Unsupported carriers fail explicitly with
``reason="unsupported_carrier"`` — never silently accepted. Hidden-
channel work is defensive only (detection / flagging, never
generation).
ForkScore consumes battery output
---------------------------------
The v8 ForkScore (see :doc:`v8-fork-score`) reads BatteryResult JSON
from a parent and child organism, computes a weighted scalar
verdict with ACCEPT / MARGINAL / REJECT classes:
.. code-block:: bash
make bench-suite # generates parent.json
# ... apply changes ...
make bench-suite # generates child.json
arborist v8 score --parent parent.json --child child.json
Authoring new fixtures
----------------------
See :file:`docs/spec-methodology.md` (per-author checklist) and the
existing fixture files under :file:`bench/fixtures/`. New
sub-batteries follow the protocol in :mod:`bench.batteries.base`
``Battery.run(fixtures_path) → BatteryResult``, deterministic, no
LLM-as-judge, carrier metadata mandatory.

View file

@ -13,6 +13,14 @@ Contents:
concepts
cookbook
.. toctree::
:maxdepth: 2
:caption: Substrate
pi-star
bench
v8-fork-score
.. toctree::
:maxdepth: 3
:caption: API Modules

134
docs/_source/pi-star.rst Normal file
View file

@ -0,0 +1,134 @@
π* domain library
=================
Arborist's canonical-projection registry — the substrate that
implements SQD whitepaper §3's invariant projection
``π*: Σ* → 𝓘 {⊥}``. Every modality the bench surface or audit
chain touches has (or reserves) a registered π* that maps surface
bytes to canonical bytes; the SHA-256 of the canonical bytes is
the equivalence-class identity.
Registry overview
-----------------
Lookup is by ``name@version`` key. Six concrete π*'s + one stub
ship today:
========================== =========== ===========================================================
Key Domain Status
========================== =========== ===========================================================
``wikitext-base@v1`` text Wikitext → plain prose (Phase 1a)
``claim-lattice@v1`` text Claim lines → JSON parsed-claim list (Phase 1a)
``code-py-ast@v1`` code Python source → canonical AST S-expression
``arithmetic@v1`` arithmetic Expression → exact rational ``num/den`` (SQD §14.1)
``logic-kernel@v1`` logic Boolean expression → canonical CNF (SQD §14.3)
``time-series-quantized@v1`` time-series JSON sample array → quantized integer vector
``tabular-pinned@v1`` tabular reserved (stub)
========================== =========== ===========================================================
Cross-modality discipline
-------------------------
Every bench fixture and audit-bound canonicalization names its
:class:`PiStar` via:
- ``carrier`` — a Phase-1 whitelist enforced by
:data:`bench.batteries.base.PHASE_1_CARRIERS`
- ``pi_star_ref`` — the registry key
Unsupported carriers fail or skip explicitly with
``reason="unsupported_carrier"`` — never silently accepted.
Hidden-channel work is defensive only (detection / flagging,
never generation or concealment).
The discipline gives every benchmark, every audit row, and every
selection score a stable answer to "what canonicalizer was used"
that survives schema migrations and is reproducible from the
canonical bytes alone.
Math π*'s — SQD §14
-------------------
Two of the most recently graduated π*'s implement the SQD
whitepaper's math substrate:
**arithmetic@v1** — closed-form rational arithmetic. Solves the
canonical SQD test exactly:
.. code-block:: python
from arborist.pi_star import get
ps = get("arithmetic@v1")
ps.canonicalize(b"0.1+0.2") # → b"3/10"
ps.canonicalize(b"0.3") # → b"3/10"
ps.canonicalize(b"1+2") # → b"3/1"
ps.canonicalize(b"6/4") # → b"3/2" (lowest terms)
No floating-point drift — ``Decimal(str(0.1))`` gives exact
``1/10``, then ``fractions.Fraction`` arithmetic stays in .
Identifiers, function calls, division by zero, and non-integer
exponents raise :class:`PiStarError`.
**logic-kernel@v1** — propositional Boolean expression →
Conjunctive Normal Form (CNF):
.. code-block:: python
ps = get("logic-kernel@v1")
ps.canonicalize(b"A AND B") # → b"A AND B"
ps.canonicalize(b"B AND A") # → b"A AND B" (commutativity)
ps.canonicalize(b"A IMPL B") # → b"(NOT A OR B)"
ps.canonicalize(b"(NOT B) IMPL (NOT A)") # → b"(NOT A OR B)" (contrapositive)
ps.canonicalize(b"NOT NOT A") # → b"A" (double negation)
ps.canonicalize(b"A OR NOT A") # → b"TRUE" (tautology)
Atom cap: 8 (CNF expansion is exponential; cap keeps
canonicalization deterministic in bounded time).
Equivalences preserved:
commutativity, associativity, IMPL/IFF/XOR rewrites, De Morgan,
double negation, distribution, idempotence, tautology collapse,
contrapositive.
Composition algebra
-------------------
Two π*'s can be composed into a third via
:func:`arborist.pi_star.compose`:
.. code-block:: python
from arborist.pi_star import compose
chain = compose("wikitext-base@v1", "claim-lattice@v1")
# Auto-registers as "wikitext-base-then-claim-lattice@v1"
chain.canonicalize(b"- The release date was July 3, 1985. [E1]")
Each composition is itself a registered π* with its own key. See
:file:`docs/pi-star-composition.md` for the algebra (type
compatibility, determinism preservation, equivalence-class
preservation, projective vs invertible compositions).
Authoring a new π*
------------------
1. Add ``arborist/pi_star/<name>.py`` with a frozen dataclass
declaring ``name``, ``version``, ``domain``, and a
``canonicalize(self, raw: bytes) -> bytes`` method.
2. Call :func:`arborist.pi_star.register` at module import time.
3. Import the new module in ``arborist/pi_star/__init__.py`` so
the registration fires at package load.
4. Write tests covering: idempotency on the canonical form (or
document projective behavior), equivalence classes preserved,
equivalence classes kept distinct, error paths (bad input
rejected explicitly).
5. Add bench fixtures under ``bench/fixtures/5s/{syntax,semantics}-<name>-v1.jsonl``
exercising the new carrier through the existing 5S Syntax and
Semantics runners.
6. Add the carrier name to
:data:`bench.batteries.base.PHASE_1_CARRIERS`.
7. Add ``make bench-5s-<name>`` Makefile target.
See :file:`docs/spec-methodology.md` for the full discipline.

View file

@ -0,0 +1,134 @@
v8 ForkScore
============
The Merkle-AGI v8 selection protocol's **scoring half**. Pure
function over a (parent, child) bench-result pair → a single
scalar with ACCEPT / MARGINAL / REJECT verdict. Phase 1a of
ticket ``#000012``.
The complementary canonicalization half (validator state machine,
acceptance protocol, slashing, fork-choice rule) is reserved for
the v8 paper itself; Phase 1a ships only the function ForkScore
without the consensus surface around it.
Formula
-------
.. code-block:: text
ForkScore =
α · Δ5S
+ β · Δ5T
+ γ · Δ5F (incl. efficiency-aware bonus)
+ δ · SelfModelCalibrationGain
+ ε · AuditCompleteness
+ ζ · ValidatorDiversity
- η · RegressionPenalty
- θ · CapitalCostPenalty
- ι · SecurityRiskPenalty (reserved; Phase 1a = 0)
- κ · ComplexityPenalty (reserved; Phase 1a = 0)
- λ · MemoryInvalidationPenalty
Each term consumes the metrics every 5S/5T/5F sub-battery emits in
its :class:`BatteryResult.metrics` dict. The Δ-rate per battery is
the mean of per-sub-battery rate deltas (child — parent).
Δ5F additionally consumes the inf-aware efficiency aggregations
landed under the 2026-05-08 ``fbd99a8`` review:
``adaptation_efficiency_mean_finite``,
``adaptation_efficiency_infinite_count``,
``adaptation_efficiency_neg_infinite_count``,
``feedback_efficiency_mean_finite``,
``feedback_efficiency_infinite_count``.
Verdict thresholds
------------------
========================== ================== ============
Score / flags Verdict CLI exit
========================== ================== ============
``score >= SIGNAL_FLOOR`` **ACCEPT** ``0``
``[0, SIGNAL_FLOOR)`` **MARGINAL** ``0``
``score < 0`` **REJECT** ``1``
hard-regression flag **REJECT** ``1``
``NEG_INF_REGRESSION`` flag **REJECT** ``1``
========================== ================== ============
``SIGNAL_FLOOR`` defaults to ``0.05`` (5pp; matches
:file:`docs/bench-maxing.md`'s noise floor).
Hard-regression flag fires if any single sub-battery rate drops by
``HARD_REGRESSION_FLOOR`` (default 5pp), regardless of net score.
``NEG_INF_REGRESSION`` flag fires when
``*_efficiency_neg_infinite_count`` increases parent → child:
*free regression* is unsafe regardless of other gains.
Default weights
---------------
.. code-block:: python
WeightSet(
alpha=1.0, # Δ5S
beta=1.0, # Δ5T
gamma=1.0, # Δ5F
delta=0.5, # SelfModelCalibrationGain
epsilon=0.3, # AuditCompleteness
zeta=0.0, # ValidatorDiversity (off in single-validator)
eta=2.0, # RegressionPenalty (heavy by design)
theta=0.5, # CapitalCostPenalty
iota=0.0, # SecurityRiskPenalty (reserved)
kappa=0.0, # ComplexityPenalty (reserved)
lambda_=0.5, # MemoryInvalidationPenalty
)
Override via JSON file passed to ``--weights``:
.. code-block:: json
{
"alpha": 1.5,
"eta": 5.0,
"lambda": 1.0
}
The JSON key ``"lambda"`` round-trips into ``WeightSet.lambda_``
because ``lambda`` is a Python reserved word.
CLI
---
.. code-block:: bash
arborist v8 score \
--parent parent-bench.json \
--child child-bench.json \
[--weights weights.json] \
[--capital-delta N] \
[--memory-invalidation-count N] \
[--audit-completeness 0..1] \
[--selfmodel-calibration-gain N]
``parent-bench.json`` and ``child-bench.json`` are the JSON output
of ``bench.batteries.runner --all``.
Output: :class:`arborist.v8.fork_score.ScoredFork` with
``score``, ``verdict``, per-term ``breakdown``, ``flags`` list,
and the active ``weights`` echoed back.
What's NOT in Phase 1a
----------------------
- Validator state machine (bonding / signing / slashing).
- Acceptance protocol (proposal / quorum / finalization).
- Challenge protocol (audit-replay disagreement).
- Fork-choice rule (which of two competing finalizations wins).
- Mesh wire format extensions for validator gossip.
- Stake mechanics + economic incentives.
- Cross-validator ZK proof exchange.
These are commissioned by the v8 paper itself
(``docs/merkle-agi-v8-consensus.rst``, still open under #000012).
Phase 1a's scoring function is the substrate the paper consumes;
landing it now lets the paper cite measured values instead of
stipulated ones.