Commit graph

3 commits

Author SHA1 Message Date
3ea27aa471
#000047 — close: delta_aggregator knob on ForkScore (Option D)
The #000025 §10.14 calibration showed _delta_5{s,t,f} mean over a
battery's 5 subs, so a single-sub gain weighs 1/5 of face value (the
5× dilution). #000047 ships the knob to pick the aggregation, default
unchanged.

WeightSet.delta_aggregator ∈ {"mean","max","sum"} (default "mean") —
a categorical field, validated in __post_init__ against
DELTA_AGGREGATORS; from_dict takes it as a string. Default unchanged →
ScoredFork output byte-identical → no fork_score.ESTIMATOR_VERSION
bump.

fork_score._aggregate(deltas, how): mean = arithmetic mean, max =
max(0.0, max_i Δ_i), sum = Σ Δ_i; empty → 0.0. _delta_5s/_delta_5t/
_delta_5f take an aggregator arg (default "mean"); the 5F efficiency
bonus is added after the aggregated base (aggregator-independent).
fork_score passes weights.delta_aggregator. The per-sub
HARD_REGRESSION_FLOOR flags are computed before aggregation, so a
single-sub regression still forces REJECT under max/sum. The chosen
aggregator is recorded in ScoredFork.weights["delta_aggregator"] (via
WeightSet.as_dict()); fork_score_branches traceability stays via the
opaque weights_id — no schema migration.

bench/scripts/fivef_threshold_calibration.py gained §5 — runs the
#000046 below-ceiling pack (5f/falsification at 0.333) and shows the
verdict / γ·Δ5f under each aggregator; bench/results/5f-threshold-
calibration-2026-05-11.md §5 is the captured record. Default stays
"mean" — the conservative, noise-robust, regression-symmetric choice
matching docs/bench-maxing.md's per-rate floor framing; v8 picks
max/sum per-deployment.

Tests: 8 new in tests/test_fork_score.py + 1 anchor in
tests/test_fivef_threshold_calibration.py; tests/test_weights.py
as_dict field-set test updated to include delta_aggregator;
test_fork_score.py AUTOCOUNT tags (#000012 §286, warrant-substrate-
cookbook.md ×2) bumped 23 → 31.

#000047 closed; #000012 §8 §3 + TICKETS.md row updated.
Full suite: 2330 passed, 28 skipped.
2026-05-11 08:27:38 -04:00
d53115efd7
#000012 Phase 1c: branch-set persistence — fork_score_branches table + CLI
Phase 1a scores one (parent, child) fork at a time; Phase 1b is the
consensus paper. Neither persists multiple candidate branches at the
same checkpoint — and #000037 §12 Trigger 1 ("ForkScore regularly
receives ≥4 candidate branches per checkpoint") gates the multi-
branch path of the Prometheus-Σ controller on this data existing.
Phase 1c lands the missing seam.

Schema (arborist/store.py): _migrate_fork_score_branches creates the
sibling table with PK (branch_set_id, branch_id) + indexes on
branch_set_id and parent_root. Sibling — never enters
audit_events.event_hash preimage, so re-scoring or back-filling
cannot break the audit chain.

Helpers (arborist/substrate/fork_score.py): persist_branch_score
upserts one row via ON CONFLICT (branch_set_id, branch_id) DO UPDATE
so re-scoring the same fork under the same checkpoint is a clean
overwrite, not a duplicate. branch_set_density(conn, branch_set_id)
returns the count of distinct branches recorded under a checkpoint
— the function the #000037 §12 Trigger 1 probe reads.
ESTIMATOR_VERSION = "fork-score-v1" pins the producer generation on
every persisted row.

CLI (arborist/cli.py): arborist substrate score gains six new flags
(--branch-set, --branch-id, --parent-root, --child-root,
--persist-shard, --weights-id). Default off — --branch-set absent
preserves Phase 1a pure-function semantics for every existing
caller. When present, requires --parent-root and either --branch-id
or --child-root; missing inputs return exit code 2.

Tests (tests/test_fork_score.py, count 18 → 23): migration creates
the table + both indexes; persist writes one row carrying
parent/child roots + verdict + weights_id + estimator_version;
upsert on the PK refreshes child_root + weights_id + recorded_at
without duplicating; branch_set_density counts per-checkpoint and
ignores cross-set rows; breakdown_blob round-trips as canonical
JSON whose values sum to the persisted score.

Status sync: #000012 §7 Phase 1c flipped from "proposed, not yet
open" to "landed 2026-05-10" with the original proposal preserved
below as design log. TICKETS row 117 mirror-updated. AUTOCOUNT
counters in #000012 + cookbook bumped 18 → 23 plus the cookbook's
fork_score.py LOC row refreshed (298 → 386 module, 403 → 609
tests, density 1.35 → 1.58).

End-to-end smoke verified: arborist substrate score writes a
fork_score_branches row with the expected schema (verdict / weights_id
/ estimator_version) and the row survives a clean SQLite read.
2026-05-10 20:12:30 -04:00
3d0f02e1ac
tests/fork_score: 18 tests for v8 ForkScore (#000012 Phase 1a — was zero coverage)
arborist/substrate/fork_score.py landed in #000012 Phase 1a but
shipped with no test file. 298 LOC of pure-function scoring +
verdict logic, exposed via the `arborist v8 score` CLI (now
substrate-rooted per ticket #000035 dir-rename).

Coverage:
  - bench_result_to_metrics adapter (BatteryResult JSON → nested
    {battery: {sub_battery: metrics}})
  - fork_score happy path: pure improvement → ACCEPT (score ≥
    SIGNAL_FLOOR=0.05)
  - marginal band: small improvement → MARGINAL (score in
    [0, SIGNAL_FLOOR))
  - zero parent + zero child → score 0 → MARGINAL
  - hard-reject paths: per-sub-battery HARD_REGRESSION_FLOOR
    (≥5pp drop on any 5S/5T/5F sub triggers REJECT regardless of
    overall positive score) + adaptation_efficiency_neg_infinite_count
    > 0 → NEG_INF_REGRESSION → REJECT
  - negative score → REJECT (separate path from hard-reject)
  - non-bench inputs: capital_delta penalty, audit_completeness
    bonus, security_risk subtracts WHEN iota>0 (default iota=0
    documented)
  - WeightSet customization flows through to output dict
  - ScoredFork.to_dict() JSON-serializable
  - score ≡ Σ breakdown.values() closure (no hidden term)
  - SIGNAL_FLOOR honored exactly (≥, not >) — score == 0.05 → ACCEPT

Fixed-point design discipline: tests use the constants from
arborist.substrate.fork_score directly (SIGNAL_FLOOR,
HARD_REGRESSION_FLOOR) so a bench-maxing PR that flips the floor
forces a tests-fail signal.

Default-iota=0 documented explicitly so future readers see "no,
you didn't break security_risk; it's deliberately opt-in."
2026-05-10 12:38:39 -04:00