Three threads landed in a single commit because they share the same substrate (the #000037 §12 trigger probe shipped inc422216): A — code-side stale-map sweep (arborist/qa/repair.py) ===================================================== Walked inline TODO/FIXME/XXX markers across arborist/ + tests/. Five hits: two false-positives (\\uXXXX in escape pattern docs), one genuine deferral (raw_html cache in async_web_fetcher), and **two stale TODOs in arborist/qa/repair.py** referring to "re-prompt feedback path is future work" — even though `reprompt_repair` is fully implemented (lines 140+ at the time of this commit), wired through CLI `--repair-reprompts N` flag (cli.py:4574), and gated by `policy["repair_max_reprompts"]`. Refreshed the docstring header and the in-body comment to point at the actual function. Same drift pattern as today's earlier ticket sweep (e84f453): implementation lands, the TODO doesn't get refreshed, future readers re-implement what's already there. B — N-power follow-up to the §12 trigger probe ============================================== c422216's first probe run reported trigger 2 (divergence variance) at N=16, σ=0.5, ratio=0.8 — both threshold conditions would fire if N reached the 30-sample N_min. To validate that the variance signal holds at N≥30 (rather than vanishing on a wider sample), drove the canonical-witness path 20 additional times via a new bench/qa_questions_canonical_witness_npower.txt fixture (10 arithmetic@v1 + 10 logic-kernel@v1 questions; all canonicalize-then-LLM-witness without errors). Result at N=37: trigger 2 fires. Ratio 0.575 > 0.5, abs σ 0.435 > 0.10. Variance signal is real at the floor. Captured in bench/results/prometheus-sigma-triggers-2026-05-10-npower.md with a loud caveat at the top: this is N-power validation, NOT a measure of real workload pressure. The §12 phase-1 go/no-go decision should still come from natural workload N or operator-stated need. The prior baseline (prometheus-sigma-triggers-2026-05-10.md) stays in place as the workload-state-at-time-of-ticket-c422216 record; this new -npower.md report is the statistical-power follow-up. C — ForkScore Phase 1c proposal (#000012) ========================================= #000037 §12 Trigger 1 ("ForkScore receives ≥4 candidate branches per checkpoint") gates on multi-branch persistence — but #000012 Phase 1a (landed 2026-05-08) is single-validator scoring of one fork at a time, and Phase 1b (still open) is the consensus paper. Neither persists branch-sets. So Trigger 1 structurally cannot fire today, which the probe correctly reports as "no fork_score branch-set table found". Added Phase 1c to #000012 as a doc-only proposal (no code in this commit): - New `fork_score_branches` table (sibling to capital_ledger; does NOT enter audit_events.event_hash preimage). - Optional `--branch-set <ID>` flag on `arborist v8 score`. - One read API: `branch_set_density(conn, branch_set_id)`. - #000037 probe wires through the new function once it lands. Not opened as its own ticket because operator pressure for it hasn't surfaced naturally — gates on either #000012 Phase 1b landing OR #000030 algebra/calc kernel expansion producing competing-kernel branches an operator wants to compare. Captured here so a future shift doesn't re-discover the gap. Hygiene ======= - make test → 1623 passed, 45 skipped - make chain-check-shards → 0 across all 7 shards - fox's in-flight #000037 ticket modifications left untouched
17 KiB
Ticket #000012 — Selection & consensus protocol (Merkle-AGI v8)
Status: in progress · Phase 1a (ForkScore) landed 2026-05-08; consensus paper still open Opened: 2026-05-07 Scope: Spec for the loop-closing consensus protocol that turns the v7 substrate + v9.8 runtime into actual Darwinian selection across multiple validators. Covers fitness scoring, mutation acceptance, fork-choice on disagreement, Sybil resistance, validator economics. Doc-only — no code in this ticket; this is the design substrate paper that follow-up implementation tickets will reference. Audience: fox + future blackops shifts + downstream Merkle-AGI v8 authors. Hard constraint: no per-query consensus. Selection runs at checkpoint cadence (Proof-of-Upgrade scope), not at inference time. v9.8 cache_key invariant stays at 8 dims; consensus-state lives in a sibling table, not folded into the answer-cache.
1. Problem statement
The DNA ↔ Merkle-DAG analogy promises:
copy → vary → express → test → select → preserve → repeat
Merkle-AGI v7 ships everything except select. Section 13.4
("Proof-of-Upgrade") sketches a procedure where a candidate model
(M', C(M')) is admitted if it shows non-regression on a published
eval set + non-regression of ε-coverage at sentinel frontiers. That
procedure is single-validator regression testing with a Merkle
receipt. It is not consensus.
Specifically v7 § 13.4 leaves these gaps:
| Gap | Concrete failure under current spec |
|---|---|
| Validator set discovery | No protocol. Two labs reach different verdicts on (M'); no fork-choice. |
| Byzantine fault tolerance | One dishonest validator can sign acceptance for a poisoned (M'). |
| Sybil resistance | A single actor spinning up 50 validators wins every quorum. |
| Fork-choice on disagreement | If validators V1, V2 publish conflicting acceptance receipts, downstream nodes have no rule to pick. |
| Liveness vs safety trade | No bound on how long acceptance can stall when validators are offline. |
| Validator incentives | Why would anyone run a validator? What's the slashing condition for cheating? |
| Stake / membership semantics | Permissionless? Permissioned? Hybrid? v7 silent. |
Without a protocol that closes these, "digital evolution" is metaphor — a single lab signing its own upgrades, no different in trust model from current model-card releases.
1.1 Concrete failure scenario
Lab A trains M_t → M_{t+1} with a backdoor that triggers on a rare input pattern. Lab A signs Proof-of-Upgrade: ε-coverage non- regressed at every published frontier (because the backdoor lives at a non-published frontier). Lab A publishes C(M_{t+1}). Anyone pulling the registry sees an "accepted" upgrade. v7 has no way to surface that no independent validator audited the upgrade.
The selection protocol must make "accepted under v8" mean "accepted by N independent validators meeting policy P," verifiable to anyone, not "Lab A signed it."
2. Design choices
2.1 Validator set: permissionless vs permissioned
A. Permissionless (Bitcoin-style). Anyone with stake (or proof-of-work or proof-of-storage) can validate. Maximum censorship resistance. Costs: economic incentive design, possible centralization through mining/staking concentration, latency.
B. Permissioned (consortium). Validator set is curated by a governance body. Easier to bootstrap, easier to slash, lower latency. Costs: who curates? captures regulatory risk; "patch the planet" mission frowns on gatekeepers.
C. Hybrid (delegated proof-of-stake-like). Permissionless participation but stake required; misbehavior slashable. Common middle-ground (Tezos / Cosmos). Bootstrappable and Sybil-resistant.
Recommendation: Hybrid. Permissionless joining with stake is compatible with permacomputer values (no gatekeepers) and Sybil- resistant in practice. Bootstrap from a small honest set with explicit graduation criteria.
2.2 Fitness function: who defines it?
A. Lab-defined (per upgrade). Submitter declares the metric set and eval digests; validators check non-regression. Flexible, gameable.
B. Registry-defined (canonical bench suite). A single canonical suite gates every upgrade. Simple, brittle, hard to evolve.
C. Layered (canonical floor + lab-declared ceiling). Every upgrade must non-regress on the canonical floor; lab can additionally declare metrics they want validated. Default-on safety; allows specialization.
Recommendation: C (Layered). Canonical floor is the safety substrate every release passes through. Layered specialization keeps domain models from being stuck behind irrelevant gates.
2.3 Quorum rule
A. Simple majority (51%). Liveness-friendly. Vulnerable to slim majorities and hostile takeovers.
B. Supermajority (2/3 +). Standard BFT bound. Tolerates 1/3 Byzantine. Tighter than majority, slower under partition.
C. Threshold signature (k-of-n). Cryptographic accumulator; single signature represents quorum. Cheap verification downstream; needs ceremony to mint signing key.
Recommendation: B for safety floor, with threshold signature (C) as a downstream optimization once the protocol stabilizes. Tolerates the canonical 1/3 Byzantine fraction without giving up liveness on small disagreements.
2.4 Slashing condition
A validator that signs acceptance for a candidate that subsequently fails the canonical floor's audit replay loses stake. A validator that signs conflicting acceptances (forks) loses stake. A validator that double-signs (signs both accept and reject for same C(M')) loses stake.
This requires:
- An audit-replay protocol that re-runs the canonical floor and publishes a Merkle-bound result.
- A challenge window during which any party can submit a counter-receipt invalidating an earlier acceptance.
- Time-locked stake unbonding so a validator can't sign and exit before challenges land.
2.5 Fork choice
When two valid acceptance chains diverge, downstream nodes must pick one. Options:
A. Longest valid chain. Bitcoin-style. Vulnerable to deep reorgs.
B. Highest-stake-weighted acceptance. Eth-style finality. Resistant to short-range reorgs.
C. First-finalized wins. GRANDPA-style. Once 2/3+ stake signs, no reorg.
Recommendation: C. Once 2/3+ validators finalize an upgrade, it's permanent. Latency is acceptable for checkpoint-cadence selection (not inference).
3. Recommendation
Hybrid permissionless validator set with stake, layered fitness floor + per-upgrade ceiling, supermajority quorum (2/3+), slashing on audit-replay disagreement and equivocation, GRANDPA-style fork choice. Bootstrap from a small honest set with explicit slashing window before opening to permissionless joining.
4. Implementation sketch
This ticket commissions Merkle-AGI v8 as a sister paper to v7. v8 must specify:
- Validator state machine.
- States:
bonding,active,challenged,slashed,unbonding. - Transitions: stake deposit, signature, challenge, slash, exit.
- States:
- Acceptance protocol.
- Proposer submits
(C(M'), eval_digest, frontier_coverage_diff, metric_delta_signed). - Validators run audit-replay, sign accept/reject within window.
- Aggregate signature minted at quorum.
- Proposer submits
- Challenge protocol.
- Anyone submits
(C(M'), counter_evidence)within challenge window. - If counter-evidence verifies (re-runs canonical floor and finds regression), all signing validators are slashed.
- Anyone submits
- Fork choice rule.
- Validators only sign on candidates whose parent C(M_t) is finalized.
- Once 2/3+ stake signs C(M_{t+1}), it's finalized.
- Stake mechanics.
- Bond / unbond windows.
- Slashing fraction per offense class.
- Reward distribution per honest signature.
- Mesh wire format.
- Extension to arborist
mesh/wire.pyfor validator gossip. - Aggregate signature canonicalization (so verification is stake-weight-independent).
- Extension to arborist
The paper itself is the ticket-#000012 deliverable. Code lives in follow-up tickets that cite this one.
4.1 Concrete artifacts this ticket produces
docs/merkle-agi-v8-consensus.rst— sister to the v7 substrate paper. Sections: validator state machine, acceptance, challenge, fork choice, slashing, mesh wire format, BFT analysis.docs/v8-policy-fields.md— the policy fields v8 introduces and how they fold intogovernance_policy_hash(or whether they live in a siblingconsensus_policy_hash).- A worked-example Merkle-AGI v8 acceptance ledger reflecting one imaginary upgrade cycle (paper's appendix; does not require running validators).
4.2 What does not change
- v9.8 8-dim cache_key. Consensus state lives in
consensus_events(sibling table), not incache_key. - arborist's per-shard audit chain. v8 is checkpoint-cadence, cross-validator; per-shard audit chain stays per-node.
- Existing falsification-state semantics (
live,failed,stale,quarantined).
5. Out of scope
- Implementation of the v8 protocol in code. That is at minimum 3-4 follow-up tickets (validator state machine, mesh wire extension, audit-replay harness, slashing accountant).
- Economic parameter calibration (stake amounts, slashing fractions, reward rates). v8 paper specifies the form; calibration is a governance decision.
- Cross-chain anchoring (publishing v8 finalizations to Bitcoin / Ethereum / etc.). Optional bolt-on.
- Selection of frontier benchmark fixtures (covered by ticket #000021).
6. Risks & open questions
- Liveness vs censorship. A validator set that requires 2/3+ to finalize stalls under 1/3 hostile partition. Acceptable for checkpoint cadence; needs explicit liveness floor in the spec.
- Bootstrap honesty. The initial validator set must be honest for the protocol to converge. Solution: explicit bootstrap window with permissioned set + scheduled transition to permissionless.
- Stake captures. Large staker can dominate. Mitigation: cap on individual stake weight, or convex weighting (sqrt-stake).
- Re-staking attacks. Validators staking the same capital across multiple v8 instances. Out of scope here; addressed by cross-instance slashing accumulator if/when v8 multiplies.
7. Status
In progress · Phase 1a landed 2026-05-08.
Phase 1a — ForkScore (landed)
Per fox's 2026-05-08 review note that the substrate is "good enough to support v8 selection/fork-choice design," shipping the scoring function ahead of the consensus paper:
arborist/v8/fork_score.py— pure ScoredFork dataclass + scoring function over (parent, child) BatteryResult bundles. Consumes every metric this session shipped: 5S/5T/5F sub-battery rates,adaptation_efficiency_*andfeedback_efficiency_*(incl. inf-aware aggregation perfbd99a8review), capital-cost delta, memory-invalidation count.arborist/v8/weights.py—WeightSetdataclass with α…λ +DEFAULT_WEIGHTS(single-validator-tuned) +from_dictadapter handling the"lambda"/lambda_Python-reserved-word issue.- CLI:
arborist v8 score --parent P.json --child C.json [--weights W.json]. Exits 1 on REJECT (CI-gateable). - Verdict thresholds: ACCEPT (≥ SIGNAL_FLOOR=0.05), MARGINAL
([0, SIGNAL_FLOOR)), REJECT (negative score OR hard-regression
flag OR
NEG_INF_REGRESSIONflag). - Reference doc:
docs/v8-fork-score.md. - Tests:
tests/test_v8_fork_score.py— 25 cases covering the adapter, all verdict paths, hard-regression detection, inf-bonus capping, neg-inf rejection, weight-set tuning, breakdown completeness, CLI smoke. Full suite: 1186 passed.
Phase 1b — Consensus paper (still open)
Closure criterion: docs/merkle-agi-v8-consensus.rst lands with
validator state machine, acceptance protocol, challenge protocol,
fork-choice rule (GRANDPA-style), slashing mechanics, mesh wire
format extension. Per §1 of this ticket, arborist/v8/fork_score.py
is the substrate the paper cites; the paper itself is still
research-scope.
What's deliberately NOT in Phase 1a:
- Validator state machine (bonding / signing / slashing).
- Acceptance protocol.
- Challenge protocol.
- Fork-choice rule.
- Mesh wire format extensions.
- Stake mechanics + economic incentives.
- Cross-validator ZK proof exchange.
These belong to the v8 paper itself.
Phase 1c — Branch-set persistence (proposed, not yet open)
Problem. Phase 1a scores one (parent, child) fork at a time;
Phase 1b is the consensus paper. Neither persists multiple
candidate branches at the same checkpoint. Ticket #000037 §12
Trigger 1 ("ForkScore regularly receives ≥4 candidate branches per
checkpoint") gates Phase 1 of the Prometheus-Σ controller on this
data existing — and as of bench/results/prometheus-sigma-triggers- 2026-05-10.md, no fork_score% table exists across any shard, so
Trigger 1 structurally cannot fire.
This is the missing seam. Proposal scope (doc-only here; code lands in a follow-up if/when fox approves):
1. New table fork_score_branches (sibling to audit_events,
similar to capital_ledger — does NOT enter audit_events.event_hash
preimage, so retroactive scoring cannot break the audit chain).
CREATE TABLE IF NOT EXISTS fork_score_branches (
branch_set_id TEXT NOT NULL, -- checkpoint identity
-- (e.g. parent_root + ts)
branch_id TEXT NOT NULL, -- fork identifier
-- (child_root or proposer key)
parent_root TEXT NOT NULL, -- shared parent
child_root TEXT, -- nullable for in-flight branches
score REAL NOT NULL, -- ScoredFork.score
verdict TEXT NOT NULL, -- ACCEPT / MARGINAL / REJECT
breakdown_blob TEXT NOT NULL, -- canonical-JSON of full breakdown
weights_id TEXT NOT NULL, -- which WeightSet was used
estimator_version TEXT NOT NULL, -- per Phase 1a fork_score module ver.
recorded_at INTEGER NOT NULL,
PRIMARY KEY (branch_set_id, branch_id)
);
CREATE INDEX idx_fork_score_branches_set ON fork_score_branches(branch_set_id);
CREATE INDEX idx_fork_score_branches_parent ON fork_score_branches(parent_root);
PK is (branch_set_id, branch_id) so re-scoring the same fork
under the same set is a clean upsert, not a duplicate row.
2. CLI surface. Extend arborist v8 score with two optional
flags:
--branch-set <ID>— names the checkpoint a result belongs to. When present, writes a row tofork_score_branchesin addition to printing the verdict. When absent, behavior is unchanged (Phase 1a pure-function semantics preserved).--persist-shard <PATH>— names the SQLite file to write to. Defaults to$ARBORIST_QA_DBwhen set, else stays no-op.
3. Read API. One pure function in arborist/v8/fork_score.py:
def branch_set_density(conn, branch_set_id: str) -> int:
"""Count distinct branches recorded under a checkpoint id.
Used by the #000037 trigger probe to satisfy Trigger 1."""
4. #000037 probe wiring. The trigger probe
(bench/prometheus_sigma_trigger_probe.py:trigger_1_branch_density)
currently reports "no data" when the table doesn't exist. After
Phase 1c lands, it queries fork_score_branches for the latest
checkpoint and reports n_branches >= 4.
Hard constraints:
- Sibling table — never enters
audit_events.event_hashpreimage. - Schema-only; no behavioral change to single-validator scoring.
- Default off (a
--branch-setflag absent ⇒ no row written). - No mesh wire format change (that's Phase 1b's territory).
weights_idis opaque; folding weights into a hash is a Phase 1b concern.
What this enables (downstream tickets):
- #000037 §12 Trigger 1 gains data; can fire empirically.
- ForkScore + Prometheus-Σ become composable: the controller reads a checkpoint's branch set and runs softmax across the persisted scores instead of needing to re-score from raw BatteryResults.
- A future operator-facing CLI (
arborist v8 branch-set show ID) becomes trivial — same table powers it.
Why not now: this is doc-only because Phase 1c is small (~80 LOC) but adds storage surface area. The operator pressure for it lands when:
- Multi-validator deployments produce competing branches naturally (gates on Phase 1b paper closing).
- OR an operator wants to run several alternative π* / model configurations against the same parent and pick — that operator pressure currently exists only as a hypothesis.
If pressure 2 surfaces (e.g., during #000030 algebra/calculus kernel expansion when multiple kernel variants compete), Phase 1c opens as its own ticket. Until then, #000037 Trigger 1 simply reports "no data — Phase 1c not landed" and the controller's multi-branch path is paper-only. That matches §12's "measured trigger, not calendar date" discipline.