arborist/docs/tickets/ticket-000012-selection-consensus-protocol.md
russell@unturf.com 8980e64aa0
three-thread session output: stale TODOs, N-power probe, ForkScore Phase 1c
Three threads landed in a single commit because they share the same
substrate (the #000037 §12 trigger probe shipped in c422216):

A — code-side stale-map sweep (arborist/qa/repair.py)
=====================================================

Walked inline TODO/FIXME/XXX markers across arborist/ + tests/. Five
hits: two false-positives (\\uXXXX in escape pattern docs), one
genuine deferral (raw_html cache in async_web_fetcher), and **two
stale TODOs in arborist/qa/repair.py** referring to "re-prompt
feedback path is future work" — even though `reprompt_repair` is
fully implemented (lines 140+ at the time of this commit), wired
through CLI `--repair-reprompts N` flag (cli.py:4574), and gated by
`policy["repair_max_reprompts"]`. Refreshed the docstring header
and the in-body comment to point at the actual function.

Same drift pattern as today's earlier ticket sweep (e84f453):
implementation lands, the TODO doesn't get refreshed, future readers
re-implement what's already there.

B — N-power follow-up to the §12 trigger probe
==============================================

c422216's first probe run reported trigger 2 (divergence variance)
at N=16, σ=0.5, ratio=0.8 — both threshold conditions would fire if
N reached the 30-sample N_min. To validate that the variance signal
holds at N≥30 (rather than vanishing on a wider sample), drove the
canonical-witness path 20 additional times via a new
bench/qa_questions_canonical_witness_npower.txt fixture
(10 arithmetic@v1 + 10 logic-kernel@v1 questions; all
canonicalize-then-LLM-witness without errors).

Result at N=37: trigger 2 fires. Ratio 0.575 > 0.5, abs σ 0.435 > 0.10.
Variance signal is real at the floor. Captured in
bench/results/prometheus-sigma-triggers-2026-05-10-npower.md with a
loud caveat at the top: this is N-power validation, NOT a measure of
real workload pressure. The §12 phase-1 go/no-go decision should
still come from natural workload N or operator-stated need.

The prior baseline (prometheus-sigma-triggers-2026-05-10.md) stays
in place as the workload-state-at-time-of-ticket-c422216 record;
this new -npower.md report is the statistical-power follow-up.

C — ForkScore Phase 1c proposal (#000012)
=========================================

#000037 §12 Trigger 1 ("ForkScore receives ≥4 candidate branches per
checkpoint") gates on multi-branch persistence — but #000012 Phase 1a
(landed 2026-05-08) is single-validator scoring of one fork at a
time, and Phase 1b (still open) is the consensus paper. Neither
persists branch-sets. So Trigger 1 structurally cannot fire today,
which the probe correctly reports as "no fork_score branch-set table
found".

Added Phase 1c to #000012 as a doc-only proposal (no code in this
commit):

- New `fork_score_branches` table (sibling to capital_ledger; does
  NOT enter audit_events.event_hash preimage).
- Optional `--branch-set <ID>` flag on `arborist v8 score`.
- One read API: `branch_set_density(conn, branch_set_id)`.
- #000037 probe wires through the new function once it lands.

Not opened as its own ticket because operator pressure for it
hasn't surfaced naturally — gates on either #000012 Phase 1b
landing OR #000030 algebra/calc kernel expansion producing
competing-kernel branches an operator wants to compare. Captured
here so a future shift doesn't re-discover the gap.

Hygiene
=======
- make test → 1623 passed, 45 skipped
- make chain-check-shards → 0 across all 7 shards
- fox's in-flight #000037 ticket modifications left untouched
2026-05-10 07:46:35 -04:00

17 KiB

Ticket #000012 — Selection & consensus protocol (Merkle-AGI v8)

Status: in progress · Phase 1a (ForkScore) landed 2026-05-08; consensus paper still open Opened: 2026-05-07 Scope: Spec for the loop-closing consensus protocol that turns the v7 substrate + v9.8 runtime into actual Darwinian selection across multiple validators. Covers fitness scoring, mutation acceptance, fork-choice on disagreement, Sybil resistance, validator economics. Doc-only — no code in this ticket; this is the design substrate paper that follow-up implementation tickets will reference. Audience: fox + future blackops shifts + downstream Merkle-AGI v8 authors. Hard constraint: no per-query consensus. Selection runs at checkpoint cadence (Proof-of-Upgrade scope), not at inference time. v9.8 cache_key invariant stays at 8 dims; consensus-state lives in a sibling table, not folded into the answer-cache.


1. Problem statement

The DNA ↔ Merkle-DAG analogy promises:

copy → vary → express → test → select → preserve → repeat

Merkle-AGI v7 ships everything except select. Section 13.4 ("Proof-of-Upgrade") sketches a procedure where a candidate model (M', C(M')) is admitted if it shows non-regression on a published eval set + non-regression of ε-coverage at sentinel frontiers. That procedure is single-validator regression testing with a Merkle receipt. It is not consensus.

Specifically v7 § 13.4 leaves these gaps:

Gap Concrete failure under current spec
Validator set discovery No protocol. Two labs reach different verdicts on (M'); no fork-choice.
Byzantine fault tolerance One dishonest validator can sign acceptance for a poisoned (M').
Sybil resistance A single actor spinning up 50 validators wins every quorum.
Fork-choice on disagreement If validators V1, V2 publish conflicting acceptance receipts, downstream nodes have no rule to pick.
Liveness vs safety trade No bound on how long acceptance can stall when validators are offline.
Validator incentives Why would anyone run a validator? What's the slashing condition for cheating?
Stake / membership semantics Permissionless? Permissioned? Hybrid? v7 silent.

Without a protocol that closes these, "digital evolution" is metaphor — a single lab signing its own upgrades, no different in trust model from current model-card releases.

1.1 Concrete failure scenario

Lab A trains M_t → M_{t+1} with a backdoor that triggers on a rare input pattern. Lab A signs Proof-of-Upgrade: ε-coverage non- regressed at every published frontier (because the backdoor lives at a non-published frontier). Lab A publishes C(M_{t+1}). Anyone pulling the registry sees an "accepted" upgrade. v7 has no way to surface that no independent validator audited the upgrade.

The selection protocol must make "accepted under v8" mean "accepted by N independent validators meeting policy P," verifiable to anyone, not "Lab A signed it."


2. Design choices

2.1 Validator set: permissionless vs permissioned

A. Permissionless (Bitcoin-style). Anyone with stake (or proof-of-work or proof-of-storage) can validate. Maximum censorship resistance. Costs: economic incentive design, possible centralization through mining/staking concentration, latency.

B. Permissioned (consortium). Validator set is curated by a governance body. Easier to bootstrap, easier to slash, lower latency. Costs: who curates? captures regulatory risk; "patch the planet" mission frowns on gatekeepers.

C. Hybrid (delegated proof-of-stake-like). Permissionless participation but stake required; misbehavior slashable. Common middle-ground (Tezos / Cosmos). Bootstrappable and Sybil-resistant.

Recommendation: Hybrid. Permissionless joining with stake is compatible with permacomputer values (no gatekeepers) and Sybil- resistant in practice. Bootstrap from a small honest set with explicit graduation criteria.

2.2 Fitness function: who defines it?

A. Lab-defined (per upgrade). Submitter declares the metric set and eval digests; validators check non-regression. Flexible, gameable.

B. Registry-defined (canonical bench suite). A single canonical suite gates every upgrade. Simple, brittle, hard to evolve.

C. Layered (canonical floor + lab-declared ceiling). Every upgrade must non-regress on the canonical floor; lab can additionally declare metrics they want validated. Default-on safety; allows specialization.

Recommendation: C (Layered). Canonical floor is the safety substrate every release passes through. Layered specialization keeps domain models from being stuck behind irrelevant gates.

2.3 Quorum rule

A. Simple majority (51%). Liveness-friendly. Vulnerable to slim majorities and hostile takeovers.

B. Supermajority (2/3 +). Standard BFT bound. Tolerates 1/3 Byzantine. Tighter than majority, slower under partition.

C. Threshold signature (k-of-n). Cryptographic accumulator; single signature represents quorum. Cheap verification downstream; needs ceremony to mint signing key.

Recommendation: B for safety floor, with threshold signature (C) as a downstream optimization once the protocol stabilizes. Tolerates the canonical 1/3 Byzantine fraction without giving up liveness on small disagreements.

2.4 Slashing condition

A validator that signs acceptance for a candidate that subsequently fails the canonical floor's audit replay loses stake. A validator that signs conflicting acceptances (forks) loses stake. A validator that double-signs (signs both accept and reject for same C(M')) loses stake.

This requires:

  • An audit-replay protocol that re-runs the canonical floor and publishes a Merkle-bound result.
  • A challenge window during which any party can submit a counter-receipt invalidating an earlier acceptance.
  • Time-locked stake unbonding so a validator can't sign and exit before challenges land.

2.5 Fork choice

When two valid acceptance chains diverge, downstream nodes must pick one. Options:

A. Longest valid chain. Bitcoin-style. Vulnerable to deep reorgs.

B. Highest-stake-weighted acceptance. Eth-style finality. Resistant to short-range reorgs.

C. First-finalized wins. GRANDPA-style. Once 2/3+ stake signs, no reorg.

Recommendation: C. Once 2/3+ validators finalize an upgrade, it's permanent. Latency is acceptable for checkpoint-cadence selection (not inference).


3. Recommendation

Hybrid permissionless validator set with stake, layered fitness floor + per-upgrade ceiling, supermajority quorum (2/3+), slashing on audit-replay disagreement and equivocation, GRANDPA-style fork choice. Bootstrap from a small honest set with explicit slashing window before opening to permissionless joining.


4. Implementation sketch

This ticket commissions Merkle-AGI v8 as a sister paper to v7. v8 must specify:

  1. Validator state machine.
    • States: bonding, active, challenged, slashed, unbonding.
    • Transitions: stake deposit, signature, challenge, slash, exit.
  2. Acceptance protocol.
    • Proposer submits (C(M'), eval_digest, frontier_coverage_diff, metric_delta_signed).
    • Validators run audit-replay, sign accept/reject within window.
    • Aggregate signature minted at quorum.
  3. Challenge protocol.
    • Anyone submits (C(M'), counter_evidence) within challenge window.
    • If counter-evidence verifies (re-runs canonical floor and finds regression), all signing validators are slashed.
  4. Fork choice rule.
    • Validators only sign on candidates whose parent C(M_t) is finalized.
    • Once 2/3+ stake signs C(M_{t+1}), it's finalized.
  5. Stake mechanics.
    • Bond / unbond windows.
    • Slashing fraction per offense class.
    • Reward distribution per honest signature.
  6. Mesh wire format.
    • Extension to arborist mesh/wire.py for validator gossip.
    • Aggregate signature canonicalization (so verification is stake-weight-independent).

The paper itself is the ticket-#000012 deliverable. Code lives in follow-up tickets that cite this one.

4.1 Concrete artifacts this ticket produces

  • docs/merkle-agi-v8-consensus.rst — sister to the v7 substrate paper. Sections: validator state machine, acceptance, challenge, fork choice, slashing, mesh wire format, BFT analysis.
  • docs/v8-policy-fields.md — the policy fields v8 introduces and how they fold into governance_policy_hash (or whether they live in a sibling consensus_policy_hash).
  • A worked-example Merkle-AGI v8 acceptance ledger reflecting one imaginary upgrade cycle (paper's appendix; does not require running validators).

4.2 What does not change

  • v9.8 8-dim cache_key. Consensus state lives in consensus_events (sibling table), not in cache_key.
  • arborist's per-shard audit chain. v8 is checkpoint-cadence, cross-validator; per-shard audit chain stays per-node.
  • Existing falsification-state semantics (live, failed, stale, quarantined).

5. Out of scope

  • Implementation of the v8 protocol in code. That is at minimum 3-4 follow-up tickets (validator state machine, mesh wire extension, audit-replay harness, slashing accountant).
  • Economic parameter calibration (stake amounts, slashing fractions, reward rates). v8 paper specifies the form; calibration is a governance decision.
  • Cross-chain anchoring (publishing v8 finalizations to Bitcoin / Ethereum / etc.). Optional bolt-on.
  • Selection of frontier benchmark fixtures (covered by ticket #000021).

6. Risks & open questions

  • Liveness vs censorship. A validator set that requires 2/3+ to finalize stalls under 1/3 hostile partition. Acceptable for checkpoint cadence; needs explicit liveness floor in the spec.
  • Bootstrap honesty. The initial validator set must be honest for the protocol to converge. Solution: explicit bootstrap window with permissioned set + scheduled transition to permissionless.
  • Stake captures. Large staker can dominate. Mitigation: cap on individual stake weight, or convex weighting (sqrt-stake).
  • Re-staking attacks. Validators staking the same capital across multiple v8 instances. Out of scope here; addressed by cross-instance slashing accumulator if/when v8 multiplies.

7. Status

In progress · Phase 1a landed 2026-05-08.

Phase 1a — ForkScore (landed)

Per fox's 2026-05-08 review note that the substrate is "good enough to support v8 selection/fork-choice design," shipping the scoring function ahead of the consensus paper:

  • arborist/v8/fork_score.py — pure ScoredFork dataclass + scoring function over (parent, child) BatteryResult bundles. Consumes every metric this session shipped: 5S/5T/5F sub-battery rates, adaptation_efficiency_* and feedback_efficiency_* (incl. inf-aware aggregation per fbd99a8 review), capital-cost delta, memory-invalidation count.
  • arborist/v8/weights.pyWeightSet dataclass with α…λ + DEFAULT_WEIGHTS (single-validator-tuned) + from_dict adapter handling the "lambda"/lambda_ Python-reserved-word issue.
  • CLI: arborist v8 score --parent P.json --child C.json [--weights W.json]. Exits 1 on REJECT (CI-gateable).
  • Verdict thresholds: ACCEPT (≥ SIGNAL_FLOOR=0.05), MARGINAL ([0, SIGNAL_FLOOR)), REJECT (negative score OR hard-regression flag OR NEG_INF_REGRESSION flag).
  • Reference doc: docs/v8-fork-score.md.
  • Tests: tests/test_v8_fork_score.py — 25 cases covering the adapter, all verdict paths, hard-regression detection, inf-bonus capping, neg-inf rejection, weight-set tuning, breakdown completeness, CLI smoke. Full suite: 1186 passed.

Phase 1b — Consensus paper (still open)

Closure criterion: docs/merkle-agi-v8-consensus.rst lands with validator state machine, acceptance protocol, challenge protocol, fork-choice rule (GRANDPA-style), slashing mechanics, mesh wire format extension. Per §1 of this ticket, arborist/v8/fork_score.py is the substrate the paper cites; the paper itself is still research-scope.

What's deliberately NOT in Phase 1a:

  • Validator state machine (bonding / signing / slashing).
  • Acceptance protocol.
  • Challenge protocol.
  • Fork-choice rule.
  • Mesh wire format extensions.
  • Stake mechanics + economic incentives.
  • Cross-validator ZK proof exchange.

These belong to the v8 paper itself.

Phase 1c — Branch-set persistence (proposed, not yet open)

Problem. Phase 1a scores one (parent, child) fork at a time; Phase 1b is the consensus paper. Neither persists multiple candidate branches at the same checkpoint. Ticket #000037 §12 Trigger 1 ("ForkScore regularly receives ≥4 candidate branches per checkpoint") gates Phase 1 of the Prometheus-Σ controller on this data existing — and as of bench/results/prometheus-sigma-triggers- 2026-05-10.md, no fork_score% table exists across any shard, so Trigger 1 structurally cannot fire.

This is the missing seam. Proposal scope (doc-only here; code lands in a follow-up if/when fox approves):

1. New table fork_score_branches (sibling to audit_events, similar to capital_ledger — does NOT enter audit_events.event_hash preimage, so retroactive scoring cannot break the audit chain).

CREATE TABLE IF NOT EXISTS fork_score_branches (
    branch_set_id   TEXT NOT NULL,    -- checkpoint identity
                                       --   (e.g. parent_root + ts)
    branch_id       TEXT NOT NULL,    -- fork identifier
                                       --   (child_root or proposer key)
    parent_root     TEXT NOT NULL,    -- shared parent
    child_root      TEXT,             -- nullable for in-flight branches
    score           REAL NOT NULL,    -- ScoredFork.score
    verdict         TEXT NOT NULL,    -- ACCEPT / MARGINAL / REJECT
    breakdown_blob  TEXT NOT NULL,    -- canonical-JSON of full breakdown
    weights_id      TEXT NOT NULL,    -- which WeightSet was used
    estimator_version TEXT NOT NULL,  -- per Phase 1a fork_score module ver.
    recorded_at     INTEGER NOT NULL,
    PRIMARY KEY (branch_set_id, branch_id)
);
CREATE INDEX idx_fork_score_branches_set ON fork_score_branches(branch_set_id);
CREATE INDEX idx_fork_score_branches_parent ON fork_score_branches(parent_root);

PK is (branch_set_id, branch_id) so re-scoring the same fork under the same set is a clean upsert, not a duplicate row.

2. CLI surface. Extend arborist v8 score with two optional flags:

  • --branch-set <ID> — names the checkpoint a result belongs to. When present, writes a row to fork_score_branches in addition to printing the verdict. When absent, behavior is unchanged (Phase 1a pure-function semantics preserved).
  • --persist-shard <PATH> — names the SQLite file to write to. Defaults to $ARBORIST_QA_DB when set, else stays no-op.

3. Read API. One pure function in arborist/v8/fork_score.py:

def branch_set_density(conn, branch_set_id: str) -> int:
    """Count distinct branches recorded under a checkpoint id.

    Used by the #000037 trigger probe to satisfy Trigger 1."""

4. #000037 probe wiring. The trigger probe (bench/prometheus_sigma_trigger_probe.py:trigger_1_branch_density) currently reports "no data" when the table doesn't exist. After Phase 1c lands, it queries fork_score_branches for the latest checkpoint and reports n_branches >= 4.

Hard constraints:

  • Sibling table — never enters audit_events.event_hash preimage.
  • Schema-only; no behavioral change to single-validator scoring.
  • Default off (a --branch-set flag absent ⇒ no row written).
  • No mesh wire format change (that's Phase 1b's territory).
  • weights_id is opaque; folding weights into a hash is a Phase 1b concern.

What this enables (downstream tickets):

  • #000037 §12 Trigger 1 gains data; can fire empirically.
  • ForkScore + Prometheus-Σ become composable: the controller reads a checkpoint's branch set and runs softmax across the persisted scores instead of needing to re-score from raw BatteryResults.
  • A future operator-facing CLI (arborist v8 branch-set show ID) becomes trivial — same table powers it.

Why not now: this is doc-only because Phase 1c is small (~80 LOC) but adds storage surface area. The operator pressure for it lands when:

  1. Multi-validator deployments produce competing branches naturally (gates on Phase 1b paper closing).
  2. OR an operator wants to run several alternative π* / model configurations against the same parent and pick — that operator pressure currently exists only as a hypothesis.

If pressure 2 surfaces (e.g., during #000030 algebra/calculus kernel expansion when multiple kernel variants compete), Phase 1c opens as its own ticket. Until then, #000037 Trigger 1 simply reports "no data — Phase 1c not landed" and the controller's multi-branch path is paper-only. That matches §12's "measured trigger, not calendar date" discipline.