#000037 §22 Finding 3: per-audit-mode τ_qa in dry-run
Per Finding 3, a uniform τ_qa=7d filtered out every recent CANONICAL_PROJECTION row (the π* graduations from #000027/#000030/ #000032 are all younger than 7d), so the sweep saw zero high-value kernel-only work. Splitting τ_qa by audit_mode lets the cheap kernel re-probe path (CP) run on a short cycle while the expensive LLM re-witness path (STRICT/HYBRID/UNGROUNDED) keeps the long cycle. bench/scripts/prometheus_sigma_sweep_dryrun.py: new build_tau_by_mode() helper + per-mode CASE in iter_target_a_candidates; sweep_target_a now takes the dict instead of a single seconds value. New CLI flag --tau-qa-cp-days (default 1d); --tau-qa-days now scopes to LLM-witness modes only (default 7d). Report renders the per-mode τ table in the header and marks Findings 2 and 3 RESOLVED with their landing commits. Makefile: PROMETHEUS_SWEEP_TAU_DAYS bumped to 7 (was 1, the prior Finding-3 workaround); new PROMETHEUS_SWEEP_TAU_CP_DAYS=1 makevar. bench/results/prometheus-sigma-sweep-dryrun-2026-05-10.md: regenerated under the new defaults — 9 CP rows surface alongside 1,904 LLM- witness candidates → 1,913 total Target A candidates → 1 ACCEPT, 2 MARGINAL, 205 REJECT, 271 DEFERRED chunks; 2 cache_drift vetoes preserved end-to-end. docs/tickets/ticket-000037 §22: Findings 2 + 3 marked RESOLVED with landing-commit references; total dry-run cost line updated to the new measurements (37.6 ms / 3,913 branches / 9.6 µs per branch).
This commit is contained in:
parent
43380b11b7
commit
1f882df22c
4 changed files with 174 additions and 87 deletions
7
Makefile
7
Makefile
|
|
@ -246,12 +246,17 @@ prometheus-trigger-probe: bootstrap ## #000037 §12 measured-pressure probe →
|
||||||
# the Phase 1 controller, reports decision distribution + Phase-3
|
# the Phase 1 controller, reports decision distribution + Phase-3
|
||||||
# design findings. No LLM calls, no mutations.
|
# design findings. No LLM calls, no mutations.
|
||||||
PROMETHEUS_SWEEP_DRYRUN_OUT ?= bench/results/prometheus-sigma-sweep-dryrun-$(shell date -u +%Y-%m-%d).md
|
PROMETHEUS_SWEEP_DRYRUN_OUT ?= bench/results/prometheus-sigma-sweep-dryrun-$(shell date -u +%Y-%m-%d).md
|
||||||
PROMETHEUS_SWEEP_TAU_DAYS ?= 1
|
# §22 Finding 3 fix: per-audit-mode τ_qa. CP rows are kernel-only
|
||||||
|
# (cheap re-probe → short τ); LLM-witness modes (STRICT/HYBRID/
|
||||||
|
# UNGROUNDED) re-witness via LLM (expensive → long τ).
|
||||||
|
PROMETHEUS_SWEEP_TAU_DAYS ?= 7
|
||||||
|
PROMETHEUS_SWEEP_TAU_CP_DAYS ?= 1
|
||||||
PROMETHEUS_SWEEP_B_SAMPLE ?= 500
|
PROMETHEUS_SWEEP_B_SAMPLE ?= 500
|
||||||
prometheus-sweep-dryrun: bootstrap ## #000037 Phase 3 sleep-sweep dry-run → markdown report
|
prometheus-sweep-dryrun: bootstrap ## #000037 Phase 3 sleep-sweep dry-run → markdown report
|
||||||
PYTHONUNBUFFERED=1 $(PY) bench/scripts/prometheus_sigma_sweep_dryrun.py \
|
PYTHONUNBUFFERED=1 $(PY) bench/scripts/prometheus_sigma_sweep_dryrun.py \
|
||||||
--shards-dir $(SHARDS_DIR) \
|
--shards-dir $(SHARDS_DIR) \
|
||||||
--tau-qa-days $(PROMETHEUS_SWEEP_TAU_DAYS) \
|
--tau-qa-days $(PROMETHEUS_SWEEP_TAU_DAYS) \
|
||||||
|
--tau-qa-cp-days $(PROMETHEUS_SWEEP_TAU_CP_DAYS) \
|
||||||
--sample-b $(PROMETHEUS_SWEEP_B_SAMPLE) \
|
--sample-b $(PROMETHEUS_SWEEP_B_SAMPLE) \
|
||||||
--out $(PROMETHEUS_SWEEP_DRYRUN_OUT)
|
--out $(PROMETHEUS_SWEEP_DRYRUN_OUT)
|
||||||
|
|
||||||
|
|
|
||||||
|
|
@ -1,39 +1,44 @@
|
||||||
# Prometheus-Σ Phase 3 sleep-sweep dry-run
|
# Prometheus-Σ Phase 3 sleep-sweep dry-run
|
||||||
|
|
||||||
Generated: 2026-05-10T20:54:12Z UTC
|
Generated: 2026-05-10T22:33:37Z UTC
|
||||||
Script: `bench/scripts/prometheus_sigma_sweep_dryrun.py`
|
Script: `bench/scripts/prometheus_sigma_sweep_dryrun.py`
|
||||||
Ticket: #000037 Phase 3 (read-only simulation)
|
Ticket: #000037 Phase 3 (read-only simulation)
|
||||||
|
|
||||||
## Parameters
|
## Parameters
|
||||||
|
|
||||||
- Shards dir: `/home/fox/.arborist/shards`
|
- Shards dir: `/home/fox/.arborist/shards`
|
||||||
- τ_qa: 1 days (86400 seconds)
|
- τ_qa per audit_mode (§22 Finding 3 fix):
|
||||||
|
- CANONICAL_PROJECTION: 1d (86400s)
|
||||||
|
- HYBRID: 7d (604800s)
|
||||||
|
- STRICT: 7d (604800s)
|
||||||
|
- UNGROUNDED: 7d (604800s)
|
||||||
- Target B sample/shard: 500
|
- Target B sample/shard: 500
|
||||||
- Controller budget (Hermes concurrency): 4
|
- Controller budget (Hermes concurrency): 4
|
||||||
- Weight profile: safe (§15.1)
|
- Weight profile: safe (§15.1)
|
||||||
|
|
||||||
## Target A — providence_cache sweep candidates
|
## Target A — providence_cache sweep candidates
|
||||||
|
|
||||||
- Total candidates (rows older than τ_qa): **2381**
|
- Total candidates (rows older than τ_qa): **1913**
|
||||||
- Sweep chunks (size 4): 596
|
- Sweep chunks (size 4): 479
|
||||||
- Controller runtime: 51.93 ms
|
- Controller runtime: 15.80 ms
|
||||||
|
|
||||||
### Audit-mode distribution of candidates
|
### Audit-mode distribution of candidates
|
||||||
|
|
||||||
| audit_mode | count |
|
| audit_mode | count |
|
||||||
|---|---|
|
|---|---|
|
||||||
| HYBRID | 873 |
|
| STRICT | 720 |
|
||||||
| STRICT | 864 |
|
| HYBRID | 713 |
|
||||||
| UNGROUNDED | 635 |
|
| UNGROUNDED | 471 |
|
||||||
| CANONICAL_PROJECTION | 9 |
|
| CANONICAL_PROJECTION | 9 |
|
||||||
|
|
||||||
### Controller decision distribution (per chunk)
|
### Controller decision distribution (per chunk)
|
||||||
|
|
||||||
| label | chunks |
|
| label | chunks |
|
||||||
|---|---|
|
|---|---|
|
||||||
| DEFERRED | 322 |
|
| DEFERRED | 271 |
|
||||||
| REJECT | 270 |
|
| REJECT | 205 |
|
||||||
| MARGINAL | 4 |
|
| MARGINAL | 2 |
|
||||||
|
| ACCEPT | 1 |
|
||||||
|
|
||||||
### Veto kinds observed
|
### Veto kinds observed
|
||||||
|
|
||||||
|
|
@ -45,14 +50,14 @@ Ticket: #000037 Phase 3 (read-only simulation)
|
||||||
|
|
||||||
- MemoryRoot proposals: 0
|
- MemoryRoot proposals: 0
|
||||||
- SelfModel proposals: 0
|
- SelfModel proposals: 0
|
||||||
- FalsificationFixture proposals (§13 Step 11 — high-divergence → 5F-fixture funnel): **576**
|
- FalsificationFixture proposals (§13 Step 11 — high-divergence → 5F-fixture funnel): **447**
|
||||||
- Advisory event entries (would write to `controller_events` under Phase 2): 1144
|
- Advisory event entries (would write to `controller_events` under Phase 2): 895
|
||||||
|
|
||||||
## Target B — document sweep sample
|
## Target B — document sweep sample
|
||||||
|
|
||||||
- Total sampled candidates: **2000**
|
- Total sampled candidates: **2000**
|
||||||
- Sweep chunks (size 4): 500
|
- Sweep chunks (size 4): 500
|
||||||
- Controller runtime: 43.37 ms
|
- Controller runtime: 21.82 ms
|
||||||
|
|
||||||
### Canonical-shape regex prefilter hits
|
### Canonical-shape regex prefilter hits
|
||||||
|
|
||||||
|
|
@ -86,9 +91,9 @@ Ticket: #000037 Phase 3 (read-only simulation)
|
||||||
|
|
||||||
## Total simulation cost
|
## Total simulation cost
|
||||||
|
|
||||||
- Total branches scored: 4381
|
- Total branches scored: 3913
|
||||||
- Total controller runtime: 95.30 ms
|
- Total controller runtime: 37.62 ms
|
||||||
- Mean per-branch latency: 21.75 µs
|
- Mean per-branch latency: 9.61 µs
|
||||||
|
|
||||||
## Findings & fixes (dry-run iteration log)
|
## Findings & fixes (dry-run iteration log)
|
||||||
|
|
||||||
|
|
@ -96,9 +101,9 @@ Three iterations on the dry-run heuristic surfaced four real Phase-3-design find
|
||||||
|
|
||||||
**Finding 1 — chunk-size dominates Kelly threshold.** Iteration 1 (chunk_size=64) returned 100% DEFERRED. Kelly's `f_i = max(0, (p_i·b − q_i)/b)` requires `p_i > 0.5`; with 64 branches in a softmax, no single branch reaches that mass. Phase 3's scheduler must size chunks to the budget (Hermes concurrency = 4), not to the candidate pool. Dropping to chunk_size=4 split the distribution into DEFERRED (no positive-U branch) vs REJECT (positive-p but selected_u ≤ 0) vs ACCEPT/MARGINAL (positive-U winner).
|
**Finding 1 — chunk-size dominates Kelly threshold.** Iteration 1 (chunk_size=64) returned 100% DEFERRED. Kelly's `f_i = max(0, (p_i·b − q_i)/b)` requires `p_i > 0.5`; with 64 branches in a softmax, no single branch reaches that mass. Phase 3's scheduler must size chunks to the budget (Hermes concurrency = 4), not to the candidate pool. Dropping to chunk_size=4 split the distribution into DEFERRED (no positive-U branch) vs REJECT (positive-p but selected_u ≤ 0) vs ACCEPT/MARGINAL (positive-U winner).
|
||||||
|
|
||||||
**Finding 2 — flat capital_cost = no allocation.** Iteration 1 also assigned `capital_cost=1.0` to every branch (one full Hermes call). Combined with small audit-mode-based Δ5F deltas (±0.05–0.10), every utility was negative. Split kernel-only re-probe cost from full LLM-witness cost: `CANONICAL_PROJECTION=0.05`, `UNGROUNDED=0.4`, `HYBRID=0.8`, `STRICT=1.0` for Target A; `0.02` (HEAD-only freshness) vs `0.05` (canonical-shape probe) for Target B. Recommendation for Phase 3: model `capital_cost` as the expected cost given which witness paths fire (kernel / lexical / LLM), not a flat per-call estimate.
|
**Finding 2 — flat capital_cost = no allocation. RESOLVED in Phase 1.c (commit `4b85a0a`).** Iteration 1 assigned `capital_cost=1.0` to every branch (one full Hermes call). Combined with small audit-mode-based Δ5F deltas (±0.05–0.10), every utility was negative. The dry-run since splits kernel-only re-probe cost from full LLM-witness cost: `CANONICAL_PROJECTION=0.05`, `UNGROUNDED=0.4`, `HYBRID=0.8`, `STRICT=1.0` for Target A; `0.02` (HEAD-only freshness) vs `0.05` (canonical-shape probe) for Target B. Phase 1.c then promoted the split into the controller's input contract — `ControllerBranch` now exposes `kernel_cost` + `llm_cost` fields and a back-compat `effective_cost` property (legacy `capital_cost` callers continue to work).
|
||||||
|
|
||||||
**Finding 3 — τ_qa=7d filters out every CANONICAL_PROJECTION row.** All 29 CP rows in `qa.db` are ≤ 7 days old (they're the recent π* graduations from tickets #000027/#000030/#000032). Sweep with τ_qa=7d shows zero CP candidates → zero ACCEPT chunks → zero high-value sleep work surfaced. Phase 3 should split τ by audit_mode: kernel-only modes (CP) get τ_qa=1d (cheap to re-probe, high value when kernel-LLM divergence surfaces); lexical/quote modes (STRICT/HYBRID/UNGROUNDED) stay at τ_qa=7d (expensive LLM calls).
|
**Finding 3 — τ_qa=7d filters out every CANONICAL_PROJECTION row. RESOLVED in this dry-run iteration.** All 29 CP rows in `qa.db` were ≤ 7 days old (they're the recent π* graduations from tickets #000027/#000030/#000032); a uniform τ_qa=7d returned zero CP candidates → zero ACCEPT chunks → zero high-value sleep work surfaced. The dry-run now splits τ by audit_mode via `build_tau_by_mode()` + per-mode `CASE` in `iter_target_a_candidates`: kernel-only modes (CP) default to τ_qa=1d (cheap to re-probe, high value when kernel-LLM divergence surfaces); LLM-witness modes (STRICT/HYBRID/UNGROUNDED) stay at τ_qa=7d (expensive LLM calls). Tunable via `--tau-qa-cp-days` + `--tau-qa-days` CLI flags. Phase 3 lifts both into governance parameters.
|
||||||
|
|
||||||
**Finding 4 — canonical-shape Target B is the real headline.** 4.40% of sampled documents (sample n=2000) contain math/logic/time-series canonical shapes. Extrapolated to ~152K candidates across 3.5M docs in shards. The controller correctly returns MARGINAL on chunks containing these (high entropy = uncertain = queue for sleep). At Hermes concurrency=4 a full Target B sweep would take ~38K Hermes-call rounds even with the regex prefilter — so the prefilter is doing real load-shaping and Phase 3 must still cap the per-window budget.
|
**Finding 4 — canonical-shape Target B is the real headline.** 4.40% of sampled documents (sample n=2000) contain math/logic/time-series canonical shapes. Extrapolated to ~152K candidates across 3.5M docs in shards. The controller correctly returns MARGINAL on chunks containing these (high entropy = uncertain = queue for sleep). At Hermes concurrency=4 a full Target B sweep would take ~38K Hermes-call rounds even with the regex prefilter — so the prefilter is doing real load-shaping and Phase 3 must still cap the per-window budget.
|
||||||
|
|
||||||
|
|
@ -106,8 +111,8 @@ Three iterations on the dry-run heuristic surfaced four real Phase-3-design find
|
||||||
|
|
||||||
**Recommendations for the eventual Phase 3 scheduler:**
|
**Recommendations for the eventual Phase 3 scheduler:**
|
||||||
|
|
||||||
1. Per-audit-mode τ + per-audit-mode cost class (split kernel-cost from LLM-cost on the input contract — consider adding `kernel_cost` + `llm_cost` fields to `ControllerBranch` in a v2 dataclass).
|
1. ~~Per-audit-mode τ + per-audit-mode cost class.~~ **Landed.** Cost-class split landed in Phase 1.c (commit `4b85a0a`); per-audit-mode τ landed in this dry-run iteration. Phase 3 promotes both into governance parameters (cache-key hash inputs).
|
||||||
2. Chunk size = Hermes concurrency (4 today, governance param going forward).
|
2. Chunk size = Hermes concurrency (4 today, governance param going forward).
|
||||||
3. Sweep-specific weight profile (gamma_5f bumped, lambda_capital_cost dropped) since sweep work is deliberately accepting capital cost in exchange for falsification discovery.
|
3. ~~Sweep-specific weight profile (gamma_5f bumped, lambda_capital_cost dropped).~~ **Landed in Phase 1.c (commit `4b85a0a`)** — `arborist.substrate.prometheus.sweep_weights()` returns the tuned profile (gamma_5f=1.5, lambda_capital_cost=0.25, nu_witness_divergence=0.5). Phase 3's scheduler picks it up via `WEIGHT_PROFILES['sweep']`.
|
||||||
4. MARGINAL queue from Target B becomes the funnel for 5F falsification-fixture mining (§3: "Divergence → candidate falsification fixture").
|
4. MARGINAL queue from Target B becomes the funnel for 5F falsification-fixture mining (§3: "Divergence → candidate falsification fixture").
|
||||||
5. The 152K extrapolated canonical-shape doc count is a real-corpus pressure signal — Phase 3 needs a sustained-throughput floor, not a one-shot burst design.
|
5. The 152K extrapolated canonical-shape doc count is a real-corpus pressure signal — Phase 3 needs a sustained-throughput floor, not a one-shot burst design.
|
||||||
|
|
|
||||||
|
|
@ -71,11 +71,31 @@ from arborist.substrate.prometheus import (
|
||||||
|
|
||||||
|
|
||||||
DEFAULT_SHARDS_DIR = Path.home() / ".arborist" / "shards"
|
DEFAULT_SHARDS_DIR = Path.home() / ".arborist" / "shards"
|
||||||
DEFAULT_TAU_QA_DAYS = 7
|
DEFAULT_TAU_QA_DAYS = 7 # LLM-witness modes (STRICT / HYBRID / UNGROUNDED)
|
||||||
|
DEFAULT_TAU_QA_CP_DAYS = 1 # kernel-only mode (CANONICAL_PROJECTION)
|
||||||
DEFAULT_TARGET_B_SAMPLE_PER_SHARD = 1000
|
DEFAULT_TARGET_B_SAMPLE_PER_SHARD = 1000
|
||||||
DEFAULT_BUDGET = 4 # Hermes concurrent-request ceiling (§11)
|
DEFAULT_BUDGET = 4 # Hermes concurrent-request ceiling (§11)
|
||||||
|
|
||||||
|
|
||||||
|
def build_tau_by_mode(
|
||||||
|
cp_days: int = DEFAULT_TAU_QA_CP_DAYS,
|
||||||
|
llm_days: int = DEFAULT_TAU_QA_DAYS,
|
||||||
|
) -> dict[str, int]:
|
||||||
|
"""Per §22 Finding 3 (RESOLVED): split τ_qa by audit_mode.
|
||||||
|
|
||||||
|
Kernel-only modes (CANONICAL_PROJECTION) take a short τ — cheap
|
||||||
|
to re-probe, high value when kernel-LLM divergence surfaces.
|
||||||
|
LLM-witness modes (STRICT / HYBRID / UNGROUNDED) take a longer τ
|
||||||
|
— re-witness is expensive. Returns seconds per mode.
|
||||||
|
"""
|
||||||
|
return {
|
||||||
|
"CANONICAL_PROJECTION": cp_days * 86400,
|
||||||
|
"STRICT": llm_days * 86400,
|
||||||
|
"HYBRID": llm_days * 86400,
|
||||||
|
"UNGROUNDED": llm_days * 86400,
|
||||||
|
}
|
||||||
|
|
||||||
|
|
||||||
# ---------------------------------------------------------------------
|
# ---------------------------------------------------------------------
|
||||||
# Target A — providence_cache classification + synthesis
|
# Target A — providence_cache classification + synthesis
|
||||||
# ---------------------------------------------------------------------
|
# ---------------------------------------------------------------------
|
||||||
|
|
@ -90,18 +110,31 @@ AUDIT_MODE_TO_DELTA_5F: dict[str, float] = {
|
||||||
|
|
||||||
|
|
||||||
def iter_target_a_candidates(
|
def iter_target_a_candidates(
|
||||||
conn: sqlite3.Connection, tau_qa_seconds: int, now: int
|
conn: sqlite3.Connection, tau_by_mode: dict[str, int], now: int
|
||||||
) -> Iterator[sqlite3.Row]:
|
) -> Iterator[sqlite3.Row]:
|
||||||
"""Enumerate providence_cache rows older than τ_qa."""
|
"""Enumerate providence_cache rows older than per-mode τ_qa.
|
||||||
|
|
||||||
|
``tau_by_mode`` maps audit_mode → seconds; rows whose audit_mode
|
||||||
|
is absent from the dict fall through to ``fallback_seconds`` (the
|
||||||
|
longest declared τ — never filter MORE aggressively than declared).
|
||||||
|
See :func:`build_tau_by_mode` for the §22 Finding 3 rationale.
|
||||||
|
"""
|
||||||
conn.row_factory = sqlite3.Row
|
conn.row_factory = sqlite3.Row
|
||||||
cur = conn.execute(
|
fallback_seconds = max(tau_by_mode.values(), default=7 * 86400)
|
||||||
|
case_clauses: list[str] = []
|
||||||
|
case_params: list[int] = []
|
||||||
|
for mode in sorted(tau_by_mode):
|
||||||
|
case_clauses.append(f"WHEN audit_mode = '{mode}' THEN ?")
|
||||||
|
case_params.append(tau_by_mode[mode])
|
||||||
|
sql = (
|
||||||
"SELECT cache_key, audit_mode, falsification_state, n_quotes,"
|
"SELECT cache_key, audit_mode, falsification_state, n_quotes,"
|
||||||
" n_verified, unverified_quotes, hit_count, created_at,"
|
" n_verified, unverified_quotes, hit_count, created_at,"
|
||||||
" question_text"
|
" question_text"
|
||||||
" FROM providence_cache"
|
" FROM providence_cache"
|
||||||
" WHERE ? - created_at >= ?",
|
f" WHERE ? - created_at >= (CASE {' '.join(case_clauses)}"
|
||||||
(now, tau_qa_seconds),
|
f" ELSE ? END)"
|
||||||
)
|
)
|
||||||
|
cur = conn.execute(sql, [now, *case_params, fallback_seconds])
|
||||||
yield from cur
|
yield from cur
|
||||||
|
|
||||||
|
|
||||||
|
|
@ -290,7 +323,12 @@ def target_b_branch(
|
||||||
# ---------------------------------------------------------------------
|
# ---------------------------------------------------------------------
|
||||||
|
|
||||||
|
|
||||||
def sweep_target_a(shards_dir: Path, tau_qa_seconds: int, now: int, chunk_size: int = 4) -> dict:
|
def sweep_target_a(
|
||||||
|
shards_dir: Path,
|
||||||
|
tau_by_mode: dict[str, int],
|
||||||
|
now: int,
|
||||||
|
chunk_size: int = 4,
|
||||||
|
) -> dict:
|
||||||
"""Iterate Target A candidates across all shards; controller-decide per chunk.
|
"""Iterate Target A candidates across all shards; controller-decide per chunk.
|
||||||
|
|
||||||
Default chunk_size = 4 mirrors the Hermes concurrency budget (§11)
|
Default chunk_size = 4 mirrors the Hermes concurrency budget (§11)
|
||||||
|
|
@ -335,7 +373,7 @@ def sweep_target_a(shards_dir: Path, tau_qa_seconds: int, now: int, chunk_size:
|
||||||
continue
|
continue
|
||||||
|
|
||||||
batch: list[ControllerBranch] = []
|
batch: list[ControllerBranch] = []
|
||||||
for row in iter_target_a_candidates(conn, tau_qa_seconds, now):
|
for row in iter_target_a_candidates(conn, tau_by_mode, now):
|
||||||
candidates_total += 1
|
candidates_total += 1
|
||||||
audit_mode_seen[(row["audit_mode"] or "UNKNOWN").upper()] += 1
|
audit_mode_seen[(row["audit_mode"] or "UNKNOWN").upper()] += 1
|
||||||
batch.append(target_a_branch(row))
|
batch.append(target_a_branch(row))
|
||||||
|
|
@ -526,8 +564,9 @@ def render_markdown(a_results: dict, b_results: dict, opts: dict) -> str:
|
||||||
lines.append("## Parameters")
|
lines.append("## Parameters")
|
||||||
lines.append("")
|
lines.append("")
|
||||||
lines.append(f"- Shards dir: `{opts['shards_dir']}`")
|
lines.append(f"- Shards dir: `{opts['shards_dir']}`")
|
||||||
lines.append(f"- τ_qa: {opts['tau_qa_days']} days "
|
lines.append(f"- τ_qa per audit_mode (§22 Finding 3 fix):")
|
||||||
f"({opts['tau_qa_seconds']} seconds)")
|
for mode, secs in sorted(opts["tau_by_mode"].items()):
|
||||||
|
lines.append(f" - {mode}: {secs // 86400}d ({secs}s)")
|
||||||
lines.append(f"- Target B sample/shard: {opts['sample_b']}")
|
lines.append(f"- Target B sample/shard: {opts['sample_b']}")
|
||||||
lines.append(f"- Controller budget (Hermes concurrency): "
|
lines.append(f"- Controller budget (Hermes concurrency): "
|
||||||
f"{DEFAULT_BUDGET}")
|
f"{DEFAULT_BUDGET}")
|
||||||
|
|
@ -681,31 +720,39 @@ def render_markdown(a_results: dict, b_results: dict, opts: dict) -> str:
|
||||||
)
|
)
|
||||||
lines.append("")
|
lines.append("")
|
||||||
lines.append(
|
lines.append(
|
||||||
"**Finding 2 — flat capital_cost = no allocation.** "
|
"**Finding 2 — flat capital_cost = no allocation. "
|
||||||
"Iteration 1 also assigned `capital_cost=1.0` to every "
|
"RESOLVED in Phase 1.c (commit `4b85a0a`).** "
|
||||||
"branch (one full Hermes call). Combined with small "
|
"Iteration 1 assigned `capital_cost=1.0` to every branch "
|
||||||
"audit-mode-based Δ5F deltas (±0.05–0.10), every utility "
|
"(one full Hermes call). Combined with small audit-mode-"
|
||||||
"was negative. Split kernel-only re-probe cost from full "
|
"based Δ5F deltas (±0.05–0.10), every utility was "
|
||||||
"LLM-witness cost: `CANONICAL_PROJECTION=0.05`, "
|
"negative. The dry-run since splits kernel-only re-probe "
|
||||||
"`UNGROUNDED=0.4`, `HYBRID=0.8`, `STRICT=1.0` for Target A; "
|
"cost from full LLM-witness cost: "
|
||||||
"`0.02` (HEAD-only freshness) vs `0.05` (canonical-shape "
|
"`CANONICAL_PROJECTION=0.05`, `UNGROUNDED=0.4`, "
|
||||||
"probe) for Target B. Recommendation for Phase 3: model "
|
"`HYBRID=0.8`, `STRICT=1.0` for Target A; `0.02` "
|
||||||
"`capital_cost` as the expected cost given which witness "
|
"(HEAD-only freshness) vs `0.05` (canonical-shape probe) "
|
||||||
"paths fire (kernel / lexical / LLM), not a flat per-call "
|
"for Target B. Phase 1.c then promoted the split into the "
|
||||||
"estimate."
|
"controller's input contract — `ControllerBranch` now "
|
||||||
|
"exposes `kernel_cost` + `llm_cost` fields and a "
|
||||||
|
"back-compat `effective_cost` property (legacy "
|
||||||
|
"`capital_cost` callers continue to work)."
|
||||||
)
|
)
|
||||||
lines.append("")
|
lines.append("")
|
||||||
lines.append(
|
lines.append(
|
||||||
"**Finding 3 — τ_qa=7d filters out every "
|
"**Finding 3 — τ_qa=7d filters out every "
|
||||||
"CANONICAL_PROJECTION row.** All 29 CP rows in `qa.db` are "
|
"CANONICAL_PROJECTION row. RESOLVED in this dry-run "
|
||||||
"≤ 7 days old (they're the recent π* graduations from "
|
"iteration.** All 29 CP rows in `qa.db` were ≤ 7 days old "
|
||||||
"tickets #000027/#000030/#000032). Sweep with τ_qa=7d "
|
"(they're the recent π* graduations from tickets "
|
||||||
"shows zero CP candidates → zero ACCEPT chunks → zero "
|
"#000027/#000030/#000032); a uniform τ_qa=7d returned zero "
|
||||||
"high-value sleep work surfaced. Phase 3 should split τ "
|
"CP candidates → zero ACCEPT chunks → zero high-value "
|
||||||
"by audit_mode: kernel-only modes (CP) get τ_qa=1d (cheap "
|
"sleep work surfaced. The dry-run now splits τ by "
|
||||||
"to re-probe, high value when kernel-LLM divergence "
|
"audit_mode via `build_tau_by_mode()` + per-mode `CASE` "
|
||||||
"surfaces); lexical/quote modes (STRICT/HYBRID/UNGROUNDED) "
|
"in `iter_target_a_candidates`: kernel-only modes (CP) "
|
||||||
"stay at τ_qa=7d (expensive LLM calls)."
|
"default to τ_qa=1d (cheap to re-probe, high value when "
|
||||||
|
"kernel-LLM divergence surfaces); LLM-witness modes "
|
||||||
|
"(STRICT/HYBRID/UNGROUNDED) stay at τ_qa=7d (expensive "
|
||||||
|
"LLM calls). Tunable via `--tau-qa-cp-days` + "
|
||||||
|
"`--tau-qa-days` CLI flags. Phase 3 lifts both into "
|
||||||
|
"governance parameters."
|
||||||
)
|
)
|
||||||
lines.append("")
|
lines.append("")
|
||||||
lines.append(
|
lines.append(
|
||||||
|
|
@ -734,20 +781,24 @@ def render_markdown(a_results: dict, b_results: dict, opts: dict) -> str:
|
||||||
)
|
)
|
||||||
lines.append("")
|
lines.append("")
|
||||||
lines.append(
|
lines.append(
|
||||||
"1. Per-audit-mode τ + per-audit-mode cost class (split "
|
"1. ~~Per-audit-mode τ + per-audit-mode cost class.~~ "
|
||||||
"kernel-cost from LLM-cost on the input contract — "
|
"**Landed.** Cost-class split landed in Phase 1.c (commit "
|
||||||
"consider adding `kernel_cost` + `llm_cost` fields to "
|
"`4b85a0a`); per-audit-mode τ landed in this dry-run "
|
||||||
"`ControllerBranch` in a v2 dataclass)."
|
"iteration. Phase 3 promotes both into governance "
|
||||||
|
"parameters (cache-key hash inputs)."
|
||||||
)
|
)
|
||||||
lines.append(
|
lines.append(
|
||||||
"2. Chunk size = Hermes concurrency (4 today, governance "
|
"2. Chunk size = Hermes concurrency (4 today, governance "
|
||||||
"param going forward)."
|
"param going forward)."
|
||||||
)
|
)
|
||||||
lines.append(
|
lines.append(
|
||||||
"3. Sweep-specific weight profile (gamma_5f bumped, "
|
"3. ~~Sweep-specific weight profile (gamma_5f bumped, "
|
||||||
"lambda_capital_cost dropped) since sweep work is "
|
"lambda_capital_cost dropped).~~ **Landed in Phase 1.c "
|
||||||
"deliberately accepting capital cost in exchange for "
|
"(commit `4b85a0a`)** — `arborist.substrate.prometheus."
|
||||||
"falsification discovery."
|
"sweep_weights()` returns the tuned profile "
|
||||||
|
"(gamma_5f=1.5, lambda_capital_cost=0.25, "
|
||||||
|
"nu_witness_divergence=0.5). Phase 3's scheduler picks it "
|
||||||
|
"up via `WEIGHT_PROFILES['sweep']`."
|
||||||
)
|
)
|
||||||
lines.append(
|
lines.append(
|
||||||
"4. MARGINAL queue from Target B becomes the funnel for "
|
"4. MARGINAL queue from Target B becomes the funnel for "
|
||||||
|
|
@ -783,7 +834,13 @@ def main(argv: list[str] | None = None) -> int:
|
||||||
"--tau-qa-days",
|
"--tau-qa-days",
|
||||||
type=int,
|
type=int,
|
||||||
default=DEFAULT_TAU_QA_DAYS,
|
default=DEFAULT_TAU_QA_DAYS,
|
||||||
help="re-witness providence_cache rows older than this many days",
|
help="τ_qa for LLM-witness modes (STRICT/HYBRID/UNGROUNDED), days",
|
||||||
|
)
|
||||||
|
p.add_argument(
|
||||||
|
"--tau-qa-cp-days",
|
||||||
|
type=int,
|
||||||
|
default=DEFAULT_TAU_QA_CP_DAYS,
|
||||||
|
help="τ_qa for kernel-only mode (CANONICAL_PROJECTION), days",
|
||||||
)
|
)
|
||||||
p.add_argument(
|
p.add_argument(
|
||||||
"--sample-b",
|
"--sample-b",
|
||||||
|
|
@ -809,10 +866,12 @@ def main(argv: list[str] | None = None) -> int:
|
||||||
return 1
|
return 1
|
||||||
|
|
||||||
now = int(time.time())
|
now = int(time.time())
|
||||||
tau_qa_seconds = args.tau_qa_days * 86400
|
tau_by_mode = build_tau_by_mode(
|
||||||
|
cp_days=args.tau_qa_cp_days, llm_days=args.tau_qa_days
|
||||||
|
)
|
||||||
|
|
||||||
sys.stderr.write(f"Target A sweep over {args.shards_dir}...\n")
|
sys.stderr.write(f"Target A sweep over {args.shards_dir}...\n")
|
||||||
a_results = sweep_target_a(args.shards_dir, tau_qa_seconds, now)
|
a_results = sweep_target_a(args.shards_dir, tau_by_mode, now)
|
||||||
sys.stderr.write(
|
sys.stderr.write(
|
||||||
f" {a_results['candidates_total']} candidates, "
|
f" {a_results['candidates_total']} candidates, "
|
||||||
f"{a_results['chunks_total']} chunks, "
|
f"{a_results['chunks_total']} chunks, "
|
||||||
|
|
@ -833,7 +892,8 @@ def main(argv: list[str] | None = None) -> int:
|
||||||
),
|
),
|
||||||
"shards_dir": str(args.shards_dir),
|
"shards_dir": str(args.shards_dir),
|
||||||
"tau_qa_days": args.tau_qa_days,
|
"tau_qa_days": args.tau_qa_days,
|
||||||
"tau_qa_seconds": tau_qa_seconds,
|
"tau_qa_cp_days": args.tau_qa_cp_days,
|
||||||
|
"tau_by_mode": tau_by_mode,
|
||||||
"sample_b": args.sample_b,
|
"sample_b": args.sample_b,
|
||||||
}
|
}
|
||||||
|
|
||||||
|
|
|
||||||
|
|
@ -1246,27 +1246,38 @@ branches never gives any single branch that much mass. **Fix for
|
||||||
Phase 3:** chunk size = Hermes concurrency (4 today; governance
|
Phase 3:** chunk size = Hermes concurrency (4 today; governance
|
||||||
parameter going forward), not the candidate pool size.
|
parameter going forward), not the candidate pool size.
|
||||||
|
|
||||||
**Finding 2 — flat `capital_cost` blocks every allocation.** Iteration
|
**Finding 2 — flat `capital_cost` blocks every allocation. RESOLVED
|
||||||
1 also assigned `capital_cost=1.0` to every branch. Combined with
|
in Phase 1.c (commit `4b85a0a`).** Iteration 1 assigned
|
||||||
small audit-mode-based Δ5F deltas (±0.05–0.10), every utility came
|
`capital_cost=1.0` to every branch; combined with small audit-mode-
|
||||||
out negative — DEFERRED for chunks where Kelly's guard never fired,
|
based Δ5F deltas (±0.05–0.10), every utility came out negative —
|
||||||
REJECT for chunks where it did. Iteration 2 split the cost class
|
DEFERRED for chunks where Kelly's guard never fired, REJECT for chunks
|
||||||
|
where it did. The dry-run since splits the cost class
|
||||||
(`CANONICAL_PROJECTION=0.05`, `UNGROUNDED=0.4`, `HYBRID=0.8`,
|
(`CANONICAL_PROJECTION=0.05`, `UNGROUNDED=0.4`, `HYBRID=0.8`,
|
||||||
`STRICT=1.0` for Target A; `0.02` HEAD-only vs `0.05` canonical-shape
|
`STRICT=1.0` for Target A; `0.02` HEAD-only vs `0.05` canonical-shape
|
||||||
for Target B) and surfaced real ACCEPT/MARGINAL signal. **Fix for
|
for Target B) and surfaces real ACCEPT/MARGINAL signal. Phase 1.c
|
||||||
Phase 3:** consider splitting `capital_cost` into `kernel_cost` +
|
then promoted the split into the controller's input contract:
|
||||||
`llm_cost` on the input contract in a v2 dataclass, or compute
|
`ControllerBranch` now exposes `kernel_cost` + `llm_cost` fields and
|
||||||
`capital_cost` as the expected cost given which witness paths fire.
|
a back-compat `effective_cost` property (legacy `capital_cost`
|
||||||
|
callers continue to work, byte-identical utility values when only
|
||||||
|
one path is populated). The matching `sweep_weights()` profile
|
||||||
|
(γ_5f=1.5, λ_capital_cost=0.25, ν_witness_divergence=0.5) ships in
|
||||||
|
the same commit.
|
||||||
|
|
||||||
**Finding 3 — τ_qa=7d filters out every `CANONICAL_PROJECTION` row.**
|
**Finding 3 — τ_qa=7d filters out every `CANONICAL_PROJECTION` row.
|
||||||
All 29 CP rows in `qa.db` are ≤7 days old (they're the recent π*
|
RESOLVED in dry-run iteration 2026-05-10T22:33Z.** All 29 CP rows in
|
||||||
graduations from #000027 / #000030 / #000032). At τ_qa=7d, zero CP
|
`qa.db` are ≤7 days old (recent π* graduations from #000027 /
|
||||||
rows surface → zero ACCEPT chunks → zero high-value sleep work
|
#000030 / #000032); a uniform τ_qa=7d returned zero CP candidates →
|
||||||
discovered. At τ_qa=1d, 9 CP rows surface and 4 chunks return
|
zero ACCEPT chunks → zero high-value sleep work surfaced. The dry-
|
||||||
MARGINAL. **Fix for Phase 3:** split τ_qa by audit_mode. Kernel-only
|
run now splits τ_qa by audit_mode via `build_tau_by_mode()` + per-
|
||||||
modes (CP) take a short τ (1d default — they're cheap to re-probe);
|
mode `CASE` in `iter_target_a_candidates`: kernel-only modes (CP)
|
||||||
LLM-witness modes (STRICT/HYBRID/UNGROUNDED) take a longer τ (7d
|
default to τ_qa=1d (cheap re-probe, high value when kernel-LLM
|
||||||
default — re-witness is expensive).
|
divergence surfaces); LLM-witness modes (STRICT/HYBRID/UNGROUNDED)
|
||||||
|
stay at τ_qa=7d. Tunable via `--tau-qa-cp-days` + `--tau-qa-days`
|
||||||
|
CLI flags + `PROMETHEUS_SWEEP_TAU_*` make-variables. Result with
|
||||||
|
defaults: 9 CP rows surface alongside the 1,904 LLM-witness
|
||||||
|
candidates → 1,913 total Target A candidates → 1 ACCEPT, 2 MARGINAL,
|
||||||
|
205 REJECT, 271 DEFERRED chunks (chunk_size=4). Phase 3 lifts the
|
||||||
|
per-mode τ table into governance parameters (cache-key hash inputs).
|
||||||
|
|
||||||
**Finding 4 — Target B canonical-shape detection is the real
|
**Finding 4 — Target B canonical-shape detection is the real
|
||||||
headline.** 4.40% of sampled documents (n=2000) contain
|
headline.** 4.40% of sampled documents (n=2000) contain
|
||||||
|
|
@ -1287,11 +1298,17 @@ burst design.
|
||||||
path exercises end-to-end against real corpus data — no fixture-only
|
path exercises end-to-end against real corpus data — no fixture-only
|
||||||
mocking.
|
mocking.
|
||||||
|
|
||||||
Total dry-run cost: **~55 ms** to score 4,379 branches across all
|
Total dry-run cost (per-mode τ_qa, defaults CP=1d / LLM=7d, run
|
||||||
|
2026-05-10T22:33Z): **~37.6 ms** to score 3,913 branches across all
|
||||||
sweep targets and both shards' worth of providence_cache + a 2,000-
|
sweep targets and both shards' worth of providence_cache + a 2,000-
|
||||||
doc sample of documents — **12.5 µs per branch** at chunk_size=4 in
|
doc sample of documents — **9.6 µs per branch** at chunk_size=4 in
|
||||||
pure Python. Phase 3 latency budget for the controller itself is
|
pure Python. (Prior run with uniform τ_qa=1d: ~55 ms / 4,379 branches
|
||||||
non-binding; the witness fan-out (Hermes calls) is the dominant cost.
|
/ 12.5 µs per branch — fewer LLM-witness candidates surface under
|
||||||
|
the per-mode default, so the total branch count drops; the
|
||||||
|
per-branch latency delta is within measurement noise and not
|
||||||
|
attributable to a specific code change.) Phase 3 latency budget for
|
||||||
|
the controller itself remains non-binding; the witness fan-out
|
||||||
|
(Hermes calls) is the dominant cost.
|
||||||
|
|
||||||
Detailed numbers + iteration log per shard:
|
Detailed numbers + iteration log per shard:
|
||||||
`bench/results/prometheus-sigma-sweep-dryrun-YYYY-MM-DD.md`
|
`bench/results/prometheus-sigma-sweep-dryrun-YYYY-MM-DD.md`
|
||||||
|
|
|
||||||
Loading…
Add table
Add a link
Reference in a new issue