fox asked: JIT-trained, gossip-pulled, or repo-embedded?
Added §2.4 — three shapes, two dead ends:
- JIT-trained per deployment: ❌ — GPU-hours over MNLI/SNLI/ANLI, not
a per-deployment step; non-deterministic across runs → different
audit_mode → breaks the verifier's reproducibility invariant.
"Training" is a one-time act by someone; the distributed thing is
the resulting fixed blob, not the recipe.
- Embedded in the git repo as data: ❌ — a few-hundred-MB binary in
.git blows the standing "fresh clone = python3.12 + venv + sqlite3"
invariant; Git LFS adds a dep and still bloats; the repo is source
+ tiny fixtures + docs, not a model registry.
- One-time obtained, content-addressed, fetched on demand, mesh-
distributable: ✓ — the repo carries a tiny manifest
(arborist/qa/nli/manifest.json: {checkpoint_sha256, source_url,
license, nli_model_version}); the weights live under
~/.arborist/models/nli/<hash>/, fetched on first use (make
fetch-nli or auto-fetch on first --enable-nli query) with the
sha256 verified fail-closed against the manifest; nli_model_version
folds into governance_policy_hash (a swap stales the cache, like
chunking_version / canonicalization_version), keeping audit_mode a
deterministic function of (answer, source, policy, pinned hash);
and once arborist's mesh grows blob-sync, a deployment with peers
pulls the checkpoint from a peer (AXFR-style, like ClouDNS slaves
pulling a zone, like a shard rehydrating from snapshot.db) instead
of the origin URL — content-addressing makes peer-pull and
origin-pull interchangeable, the mesh is the optimization not the
canonical source, cold start falls back to origin. Source: a
published off-the-shelf NLI checkpoint (MIT/Apache, recorded) or a
one-time fox/blackops-trained-and-published one.
Bottom line: not JIT-trained, not repo-embedded weights — a
content-addressed hash-pinned checkpoint, manifest in the repo,
weights under ~/.arborist/, fetched on demand (origin or peer), hash
in governance_policy_hash. Precedents cited: the textbook manifest
(pointers + licenses not texts), the [vec] extra (fastembed's
bge-small fetched not committed — #000039), ~/.arborist/ operator-
state, docs/mesh.md + snapshot rehydrate. All presupposes the §2.2/§5
discipline call came back "yes" (is a fixed NLI model allowed in the
audit_mode path?); if "no", none of it builds. Also wired §2.2 con
(b)/(c), §3, §5, §6 to point at §2.4. Doc-only.
bge-small-en-v1.5 batched on a 4090 ≈ 10^3-10^4 chunks/s → full
6.24M-chunk corpus in minutes. Drop a CUDA Embedder (fastembed
CUDAExecutionProvider, or sentence-transformers device=cuda) into
default_embedder()'s pluggable callable; int8 quant stays CPU-side
post-embed; CUDA stack lives only on the producer box. Determinism
note: GPU backfill isn't byte-identical to CPU — irrelevant, soft
data, pin model revision + recipe not output bytes.
#000051: gossip the embedding backfill over mesh — backfill once on any
CPU box, publish a vecpack (leaf_hash-keyed, soft data, cheap structural
gate, never proof path), peers pull + bulk-load. The laptop never runs
the transformer; this is the mechanism behind whitepaper §1's "the
embedding pass runs off the device". Supplies #000050's backfill prereq
as a distributable artifact.
#000050: fold the Dav1dPrometheus-review high-value items into Phase-2
scope — accept-path-5 (vec hits clear the title gate via span-level
warrant, not similarity score, so the gate doesn't drop the semantic
candidates vec exists for), six vec config fields into
governance_policy_hash + a cache-write guard until wired, run-DAG
records the vec stage, four-condition bench (A/B/C/D, C-beats-D) on
semantic-allusion + curated + adversarial-semantic-neighbor fixtures.
Backfill prereq now routes through #000051.
Next ID 000051 -> 000052.
The #000012 top-of-file Status line still said "Phase 1c (branch-set
persistence) remains proposed-not-opened" — but Phase 1c landed
2026-05-10 (d53115e: fork_score_branches table + persist_branch_score
/ branch_set_density + the default-off CLI flags), the #000037 §12
Trigger 1 probe was wired to it 2026-05-11 (dbe8249), and the ticket's
own §7 Phase 1c section + the TICKETS.md index row already say
"landed". Only the header was stale.
Refreshed: Phase 1c landed (with the probe wire — measurable, not yet
fired); §8 = the ForkScore threshold-calibration handoff from #000025
§10.14 (delta_aggregator knob #000047; harder-fixture-tier line
#000046 closed / #000048 closed / #000049 open); Phase 2 (the v8
multi-validator consensus implementation) noted as gating on a
multi-validator deployment, like #000045. Doc-only.
make bench-emergent EMERGENT_N=30 — 30 fresh 3-word-triangulation
cycles appended to bench/emergent_log.jsonl.
Aggregate: 29/30 UNGROUNDED, 1/30 HYBRID, 0 STRICT — the expected
healthy null result. The triangulator generates absurd questions
("How does the depth of one's music-making abilities on a kazoo …")
with no real corpus support; the verifier ladder honestly returns
UNGROUNDED rather than hallucinating-then-STRICTing. 0 false-positive
STRICT on garbage input.
The one HYBRID (perversely/succeeds/disagreeably) is a re-run of
failure shape E ("total deflection to unrelated trivia"): the common
words pulled Hirabah / Original-righteousness chunks via FTS5; the
model quoted them verbatim (3/4 claims verify) but the answer is
total topic-drift → DEFLECTION_DETECTED + CITATION_MISMATCH → capped
at HYBRID, not STRICT. The deflection sidecar + verbatim verifier
holding on adversarial-by-construction input — nothing new, logged as
a confirmation.
No new failure shape or correctness concern. #000006 amend records
it; the 30 entries join the --print-pending queue if a future
teacher-review pass wants a closer look. Doc + log only.
Capture the positioning fox articulated: arborist's per-document ingest
is ~10-100x cheaper than building a vector-DB representation — same
SQLite substrate, different retrieval philosophy — which is the
difference between "ingest + search runs on a phone" and "the NPU is
now a hand-warmer."
New docs/lexical-first-rationale.md (positioning/architecture
reference, not a ticket): the cost asymmetry with the measured numbers
(FTS5 + SHA-256 leaf + Merkle commit + sqlite + zstd pipeline << 1 ms/
chunk vs bge-small ONNX inference ~5-30 ms/chunk, worse contended; the
query side too — a vec query embeds the query string first, an FTS5
query is B-tree lookups); the same-SQLite-different-philosophy table
(inverted index vs dense vectors + ANN; build cost; query cost;
matching; proof-bearing); the deep version of the point — arborist IS
the Merkle Providence model and that model is cheap by construction,
embeddings are a soft signal (CLAUDE.md "soft hash vs hard hash") that
never enter a proof and are the expensive bolt-on; the edge/mobile
consequence; the honest caveat (lexical-first trades the semantic
allusion gap — which is why vec is opt-in/additive, never the default,
and the embed pass is lazy/out-of-band so the heavy transformer work
runs off-device/off-peak; int8 keeps the storage tax at +6%).
Wired in: TICKETS.md "Distinction from other docs" reference list gains
the doc; #000039 §14.6 gains a "Strategic framing" pointer to it.
(A possible follow-up: fold the mobile-viability argument into the
Merkle Providence Reverse RAG whitepaper proper — noted in the doc's
references; not done here, that's a deliberate cross-repo paper edit.)
Doc-only.
fox 2026-05-12: the attribution-aware path (#000048's deferred §2.3 —
closing the 2 recombination over-grounds in falsification-hard) is its
own ticket, not a #000048 phase. So:
#000048 → closed (at 2.1 + 2.4). Steps 2.1 + 2.4 landed 2026-05-11
(12 of 16 residual items: 4 HYBRID_ENTITY over-grounds + 8 Formulate
mis-segments → formulate-hard 12/12, falsification-hard 10/12; each
bench-gated, no STRICT-rate regression). Step 2.2 (single-clause-
containment paraphrase check) attempted + reverted — recombination-
vs-summary isn't lexical (§5 "What we learned"). The 2 residual
falsification-hard fixtures (hard-003 Mercury, hard-005 Einstein)
stand as a documented marker of where the lexical verifier stops.
Header + §5 Closure + §2.3 updated; cross-refs in #000046 / #000012
§8 / TICKETS.md repointed from "#000048 §2.3" to "#000049".
#000049 opened (doc-only, awaiting go/no-go) — "Attribution-aware
grounding check (the recombination boundary)". The recombination
class needs an attribution / dependency-parse or mini-NLI check
(distinguishing "Mercury is the largest" against "Jupiter is the
largest; Mercury is the smallest" from a legit cross-sentence
summary). Options: 2.1 hand-rolled dependency-attribution heuristic
(no model, brittle — same threshold-can't-separate problem one rung
up); 2.2 small purpose-built NLI model ([nli] extra, policy-gated,
off-by-default, bench-gated — the right capability, but forces the
"is a fixed NLI model an LLM-judge?" discipline call + a model
dependency + a non-determinism surface to pin); 2.3 do nothing (the
2 fixtures are a boundary marker, no observed real-traffic harm).
Recommends 2.3 until real-traffic recombination over-grounds show up,
then 2.2 *if* fox rules a fixed NLI model is acceptable in the
proof-adjacent path; the first decision the ticket needs is that
discipline question. Next ID 000049 → 000050. #000048 follow-up.
Doc-only — no code change.
New §14.6 captures the 2026-05-11 findings:
- Measured embed rate: ~4.3 chunks/s on the contended dev box (load
~11 on 8 cores; the embed process got ~28% of one core). One wiki
shard (~1.56M chunks) at that rate ≈ 100h ≈ 4+ days; all four ≈
~16 days. The full int8 backfill was abandoned as not feasible to
brute-force there.
- The 54K-chunk partial on 000.db confirmed ~409 B/chunk apparent
→ 384 B amortized → ~2.4 GB for the full 6.24M-chunk corpus at
int8 — the deterministic number a full backfill would only
re-confirm, so finishing it bought nothing.
- On an idle healthy box (batching + all cores) bge-small does
~50-200 chunks/s → full corpus ≈ ~9-35h (the ticket's earlier
"~17h" is the optimistic end).
- Per-chunk cost breakdown: bge-small ONNX inference ~5-30 ms/chunk
dominates; the existing arborist ingest steps (chunker + SHA-256
leaf + Merkle commit + sqlite INSERTs + zstd + FTS5) are well
under 1 ms total. So adding vec multiplies ingest by ~10-100×,
entirely in the ONNX matmuls — the non-vec ingest path is
unchanged and still runs at hundreds of chunks/s.
- Implication for §14.2: this is *why* lazy-out-of-band is the
default and `arborist ingest --embed` is the opt-in. Production
guidance: a corpus-wide backfill is a one-time batch job (hours
on idle / days on contended), best run off-peak or on a dedicated
box; it does not slow ongoing ingest (which never embeds unless
--embed is passed); a GPU/accelerated embedder is a drop-in via
the pluggable Embedder callable if backfill latency matters.
Status line updated to point at §14.6. Doc-only.
What we learned from the (reverted) step-2.2 attempt, stated as a
general principle in #000048 §5 "What we learned":
A recombination ("Mercury is the largest planet …" reusing the
source's "largest planet …" with its "Mercury") and a legitimate
cross-sentence summary ("Batman, who is the alias of Bruce Wayne,
lives in Gotham City." reusing two adjacent source sentences) are
the same SHAPE to any lexical signal — both scatter the answer's
content tokens across source clauses, and the recombination's
best-single-clause coverage (4/5 = 0.8) sits ABOVE the legit
summary's (4/6 = 0.67), so no token-coverage / clause-containment /
bigram threshold separates them in the safe direction. The
discriminating thing is *attribution* — in the source, are these
tokens attached to the same subject/predicate the answer attaches
them to? — which is a dependency / NLI question, not a string
metric. That's the boundary of the deterministic, no-LLM-judge
lexical verifier: absence signals (#000046 numeric gate, #000048
step 2.1 entity gate) and structure-of-the-model's-own-output
signals (step 2.4 segmenter) are lexical and work; "the source
contradicts this pairing" is not, and proxying it with a coverage
cut trades a small contrived-fixture win for honest demotions of
real summaries — a net loss against bench-maxing's 5-pp floor.
Updated: #000048 §2.2 (the attempted idea kept as design log + the
no-threshold-separates finding), §2.3 (now framed as the only path to
the last 2 — attribution / mini-NLI, its own ticket if ever), §3
(original plan annotated with the LANDED / ATTEMPTED+REVERTED / NOT
DONE outcome), §5 (the step-2.2 receipt + the "What we learned"
subsection + the closure recommendation). Stale cross-refs fixed:
TICKETS.md #000048 + #000046 rows, #000046 ticket Headroom section,
#000012 §8 #4 — all of which said "#000048 step 2.2 closes the last
2", now corrected to "step 2.2 attempted + reverted; the 2 recombination
fixtures stand as documented residue; #000048 §2.3 is the path if
ever wanted".
Recommendation unchanged: close#000048 at 2.1+2.4 (12 of 16 residual
items closed — formulate-hard 12/12, falsification-hard 10/12). Doc-
only — no code change.
Tried a single-clause-containment check on the paraphrase path (and
the entity path's weakest slot): a span is grounded only if a single
source clause (;/. -delimited) covers >= paraphrase_coverage of its
content tokens — a recombination spreads them across clauses. On the
contrived fixtures it works (hard-003 Mercury, hard-005 Einstein →
UNGROUNDED → falsification-hard 10/12 → 12/12). But it ALSO rejects
legitimate cross-sentence summaries — the established fox-2026-04-29
Batman case ("Batman, who is the alias of Bruce Wayne, lives in
Gotham City." against "Batman is the alias of Bruce Wayne. Batman
lives in Gotham City.": spans two clauses, best clause covers 4/6 =
0.67 → demoted, but it's a correct paraphrase). And there's no
threshold that separates the two directions — the recombination
("Mercury …": best clause covers 4/5 = 0.8) sits ABOVE the legit
summary's 0.67, so any cut that rejects Mercury also rejects Batman.
Distinguishing "recombined into a different statement" from
"summarized two adjacent sentences" needs attribution / dependency
parsing or a mini-NLI model (the §2.3 territory), not a conservative
lexical proxy.
Reverted — verify.py stays at its post-step-2.4 state; falsification-
hard stays at 10/12 with hard-003 + hard-005 as documented residue.
#000048 §5 records the attempt + the no-threshold-separates argument;
recommendation: close#000048 at 2.1+2.4 (12 of 16 residual items
closed), with the 2 recombination fixtures a marker for a future
attribution-aware verifier (its own ticket). Awaiting fox's call.
Doc-only — no code change in this commit.
Closes the 8 mis-segments #000046 left in formulate-hard-v1.jsonl.
The parser was line/bullet-only — one line ⇒ one claim — so a line
that crammed several pointered claims onto one row ("Water is wet
[E1]; fire is hot [E2]", "X happened [E1]. Y followed [E2]") became
one monolithic claim with all the pointers, and a wrapped bullet
became two.
arborist/qa/parse_claims.py: _SEGMENT_SEP_RE splits a line on ';',
sentence boundaries ('. '/'! '/'? ' then a Capital), spaced dashes
(' - '/' — '/' – '), ' and '/' or '/' because '/' although '/' since
'/' while ', inline '(N)' enumeration markers, and commas — with
'(?![^\[]*\])' so a comma inside a [E1, E2] bracket never splits it.
_segment_line keeps the split ONLY IF every resulting non-empty
segment is a well-pointered claim — a legit single claim ("The cat
is black and white [E1].", "The cast: A, B, C [E1].") is never
broken because splitting it would manufacture pointer-less prose
fragments → guard rejects; a leading colon-terminated header with no
pointer ("Two facts:", "Key points:") is allowed and dropped. Plus a
wrapped-bullet join: a continuation line (leading whitespace then a
lowercase letter, no bullet glyph) folds its text + pointers into the
previous claim.
Effect: formulate-hard rate 4/12 → 12/12 (the pack is now at ceiling
— a harder Formulate tier would re-open below-ceiling headroom; a
#000046 follow-up). Remaining #000048 headroom: 2 STRICT_PARAPHRASE
recombinations in falsification-hard (Mercury, Einstein — step 2.2).
Bench gate: make bench-qa (n=3 × 75 × 3 = 675 cells; parse_pointer_claims
feeds the 450 claim_lattice_pointer + claim_lattice cells) after
(bench/qa_results/2026-05-11T20-26-37Z) vs the pre-step-2.4 baseline
(...T17-12-41Z = HEAD's parse_claims.py). STRICT-rate quote 0.54→0.55,
pointer 0.22→0.22, lattice 0.43→0.45 — all within the 5-pp noise
floor. Per-row diff: the segmenter changed the parsed-claim count on
the SAME answer text for 7 of the 450 lattice cells (0 in
claim_lattice, 7 in claim_lattice_pointer); of those, 2 caused an
audit_mode change — both correct: a wrap-join recovered an answer's
intended structure (4 claims, 2 pointer-less wrap-fragments → HYBRID)
into 2 well-pointered claims → STRICT; and a crammed-one-line blob (1
monolithic claim, all pointers → STRICT) split into 8 claims, some
not individually verifying → HYBRID (the honest verdict — false-
positive STRICT was the corruption). Every other lattice/quote delta
is LLM re-answer variance. No regression — the segmenter's only
visible effects on real traffic are honest improvements. Summarized
in qa-modes-bench.md Addendum 7 + ticket-000048 §5 step 2.4.
Tests: 8 new in test_claim_lattice.py (semicolon/sentence/conjunction
splits; pointerless-fragment + cast-list guards; leading-colon-header
drop; wrapped-bullet join; pointer-order/multi-pointer); existing
parse_pointer_claims tests pass untouched; test_5f_formulate_hard_pack
re-pinned 4/12 → 12/12. make test 2358 passed, 28 skipped.
#000048 → steps 2.1 + 2.4 landed; #000046 / #000012 §8 / TICKETS.md /
Makefile / fixture _meta + notes updated.
Brought the Merkle-AGI v7 formal substrate spec into the repo as
docs/_source/merkle-agi-dag-v7.rst (previously referenced only as the
un-version-controlled ~/Downloads/merkle-agi-dag_v7.txt). Section
structure converted to reStructuredText; inline math kept in the
source's literal notation; added to the docs/_source/index.rst
"Substrate" toctree (also added the pre-existing merkle-agi-v8-consensus
entry that was missing from it).
Folded ticket #000035's § 9.10 + § 9.10.1 (anchor PRG map φ_PRG;
dav1d-reviewed-final, little-endian, HMAC-SHA-512 / 32-byte seed) in at
their numbered positions, after § 9.9, with a .. note:: citing the
reference implementation (arborist/substrate/anchor_prg.py). #000035 ->
closed (Phase 1 + Phase 2 both landed); #000018 §9.2 ("which PRG?")
resolves to HMAC-SHA-512 with a 32-byte committed seed. Full upstream
v7 spec revision stays exogenous; this lands the amendment into the
tracked in-repo copy where future amendments also go.
(docs/TICKETS.md also carries the in-flight #000048 index-row update
from a concurrent session.)
Closes the 4 HYBRID_ENTITY over-grounds #000046 left in
falsification-hard-v1.jsonl. The entity strategy grants HYBRID when a
multi-word proper noun matches the source — but "Insulin was
discovered by Alexander Fleming" against "Penicillin was discovered by
Alexander Fleming" matches on the shared "Alexander Fleming" while the
swapped subject "Insulin" (the falsehood) is ignored.
arborist/qa/verify.py: _entity_salient_disagrees(answer_text, norm_ctx)
flags a >4-char Capitalized content token (stopword-filtered) or a
digit-number in the answer absent from the source.
_is_single_sentence(text) — no internal '. '/'! '/'? ' break. Gated in
verify_quotes' entity branch (proximity policy) in the weakest-grounding
slot only: not cluster AND len(verified) <= 1 AND _is_single_sentence
AND _entity_salient_disagrees → UNGROUNDED. The narrow caller-gate is
what keeps a structured multi-claim summary untouched — the Matrix cast
list (many entities, a tight cluster) and the TMNT answer (a numbered
list with parenthetical nicknames the source omits): model-added
accurate detail in a real summary isn't a contradiction, only the
single-sentence-one-weak-match shape is. The Matrix/TMNT/hybrid
entity-path regression tests still pass, pinned untouched.
Effect: falsification-hard rate 6/12 → 10/12 = 0.833 (Insulin / Berlin
/ 1889 / Pacific now correctly UNGROUNDED). The 2 live-pack fixtures it
newly demotes — 5f-fal-live-003 (the exact gap #000046 built its hard
pack around) and 5f-fal-live-028 — had expected_reason updated
HYBRID_ENTITY → UNGROUNDED (the live pack records what verify_quotes
actually does). Remaining hard-pack headroom: 2 STRICT_PARAPHRASE
recombinations (Mercury, Einstein — step 2.2) + 8 Formulate
mis-segments (step 2.4).
Bench gate: make bench-qa (n=3 × 75 × 3 = 675 cells) after
(bench/qa_results/2026-05-11T17-12-41Z) vs the pre-step-2.1 baseline
(...T14-19-51Z = HEAD's verify.py). STRICT-rate quote 0.50→0.54,
pointer 0.25→0.22, lattice 0.45→0.43 — all within the 5-pp noise
floor. Per-row diff (675 common cells, 30 quote-mode rows changed
audit_mode): 0 quote-mode rows demoted to UNGROUNDED from the entity
path — the gate fired on 0 legitimate QA answers in the whole bench.
Every transition was LLM re-answer variance (verifier quote→quote with
the verdict flipping); pointer/lattice deltas are noise too (the gate
is in verify_quotes / quote mode, not the claim-lattice verifier). No
regression — the gate is provably narrow on real traffic. Summarized in
qa-modes-bench.md Addendum 6 + ticket-000048 §5 step 2.1.
Tests: 4 new in test_verify.py (_is_single_sentence helper,
_entity_salient_disagrees helper, swapped-subject → UNGROUNDED,
gate-narrow-on-multi-claim); test_5f_falsification_hard_pack_below_ceiling
re-pinned 6/12 → 10/12; test_fork_score_positive_gamma_5f_... updated
(positive γ·Δ5f on the real lift — possibly MARGINAL given the ÷5
dilution; ACCEPT via a degraded-parent sub-scenario).
make test 2343 passed, 28 skipped.
#000048 → step 2.1 landed; #000046 / #000012 §8 / TICKETS.md /
Makefile / fixture _meta + notes / baseline JSON updated.
Wire vector quantization (the §3.1 production knob): arborist embed
--quant int8 [--rebuild]. The chunk_vecs vec0 column becomes int8[384]
vs float[384] per quant; the quant folds into VEC_BACKEND_VERSION
(...-384int8-... / ...-384float32-...) and vec_meta records it per
shard. Switching quant on an existing chunk_vecs requires --rebuild
(the vec0 element type can't be altered in place — embed_documents
raises ValueError telling you to --rebuild).
int8 serialization: scale each bge component by 127 (theoretical
[-1,1] range), clamp to [-127,127], round, serialize_int8. The same
scaling on the query vector → distances comparable; cosine is
scale-invariant so the uniform x127 cancels in the ranking.
sqlite-vec v0.1.9 quirk worked around: a bare blob inserted into a
vec0 column is interpreted as float32 regardless of the column's
declared type — int8 vectors MUST be wrapped in vec_int8(...). So
the INSERT and the MATCH now wrap the blob in vec_f32(?) (float32)
or vec_int8(?) (int8) — constructor name from a fixed dict, no
injection surface. (Discovered the hard way: a bare int8 blob into
an int8[384] column → "expected int8, but a float32 vector was
provided".)
VecBackend reads the quant from the existing chunk_vecs schema (or
defaults to float32) so search uses the matching wrapper. New module
exports: QUANTS, EMBED_QUANT, vec_backend_version(quant), existing_quant.
int8 head-to-head on crawl_appliedcombinatorics_org.db (168 chunks):
- storage: float32 1,597,440 B -> int8 417,792 B = 3.8x smaller
(~4x at corpus scale where the 1024-vector blocks fill; the
~28 KB of vec0 metadata doesn't quarter, hence 3.8 not 4.0).
- recall vs the float32 baseline:
Q "how many ways to choose k things from n":
identical top-5 (Combinations, Permutations, Exercises,
Derangements, Graph Coloring).
Q "pigeonhole principle counting":
identical top-2 (Graph Coloring, Exercises); ranks 3-4 swap
Derangements <-> Permutations at Δdistance 0.002 — sub-noise.
- embed speed unchanged (~3.9 chunks/s — model-load-dominated).
Conclusion: int8 is the obvious production config (§3.1's +6%-tax
recommendation confirmed empirically). v1 default stays float32 for
max fidelity; flipping the default to int8 is a fox call.
CLI (arborist/cli.py): arborist embed --quant {float32,int8}; output
JSON gains "quant"; embed_documents ValueError → exit 2 with the
"--rebuild" hint.
tests/test_search_vec.py (9 -> 16): test_int8_quant_roundtrips
(int8[384] schema, vec_meta version, search round-trip, quant
inferred by VecBackend), test_quant_mismatch_requires_rebuild,
test_invalid_quant_rejected.
#000039 status updated. Full suite: 2343 passed, 28 skipped.
(Unrelated parallel-clone work in the tree — Makefile, arborist/qa/
verify.py, bench/fixtures/5f/*, tests/test_bench_batteries.py,
tests/test_verify.py — is #000046's hard-fixture tier, not touched.)
Tracks the headroom #000046 (closed) left: the 6 over-grounds still
in falsification-hard-v1.jsonl (4 HYBRID_ENTITY where the
entity-proximity strategy matches a shared proper noun while the
answer's other salient term is wrong; 2 STRICT_PARAPHRASE where the
false claim recombines source tokens into a different true statement —
"Mercury is the largest" — and lexical token-coverage can't tell
recombination from grounding) + the 8 mis-segments in
formulate-hard-v1.jsonl (line/bullet-only parse_pointer_claims merges
multi-claim lines / splits wrapped bullets).
Options + recommended order: 2.1 verify_quotes entity
salient-token-disagreement gate (the direct analogue of #000046's
numeric gate — narrow, lexical, bench-safe; catches the 4
HYBRID_ENTITY) → 2.4 parse_pointer_claims sentence/clause
segmentation (closes the Formulate pack) → 2.2 sequence-aware
paraphrase match (conservative threshold; catches the 2
recombinations); defer 2.3 mini-NLI (heavier; only if 2.1+2.2 leave a
residue worth a model dep). Each landing bench-gated — verify_quotes
changes run `make bench-qa` before/after, parse_pointer_claims changes
run the Formulate fixtures + a QA smoke. Doc-only proposal; no code in
this ticket.
Next ID 000048 → 000049; #000046 "Headroom" section + TICKETS.md
index updated to point at #000048.
Closes the harder-fixture-tier ticket: a real verify_quotes
tightening lifts the falsification-hard rate 4/12 → 6/12, bench-gated,
ForkScore's bench-Δ goes positive on it — the loop is closed
end-to-end.
arborist/qa/verify.py: _numeric_signature(text) extracts comma-
stripped digit-runs ('8,849' and '8849' collapse; '300' stays
distinct from '300000' ← '300,000'). _check_each_with_paraphrase
gains a gate: a span that token-covers the source ≥ paraphrase_coverage
but asserts a digit-number the source lacks (modulo thousands-comma)
is no longer paraphrase-grounded — it goes to unverified. Catches the
near-miss the lexical coverage check is blind to ("Water boils at 50
degrees" against a source saying 100 token-covers 100% because
'50'/'100' aren't >4-char content tokens). Narrow by construction:
fires only on the paraphrase fallback (verbatim/span/entity/claim-
lattice paths untouched), only on a digit-number. A rounding-
paraphrase demoting here is the honest verdict — it isn't a verbatim
grounding.
Hard pack: lifts 5f-fal-hard-004 (50 vs 100) and -007 (300 vs
300,000) to UNGROUNDED → falsification-hard rate 4/12 → 6/12 = 0.5.
The other 6 over-grounds (4 HYBRID_ENTITY + 2 recombined-no-number
STRICT_PARAPHRASE) and the Formulate hard pack are unaffected —
headroom for a bigger order/dependency-aware verifier upgrade, an
optional follow-up, not a #000046 blocker.
Bench gate: make bench-qa (n=3 × 75 questions × 3 modes = 675 cells)
before / after. STRICT-rate quote 0.53→0.50, pointer 0.23→0.25,
lattice 0.44→0.45 — all within the 5-pp noise floor. Per-row diff
(675 common cells, 74 changed audit_mode): the only clearly
gate-attributable QA shift was the fictional "our cold fusion
breakthrough" year-claim demoting STRICT→HYBRID (×3 samples) — a
correct demotion; every other transition was quote→quote /
claim_lattice→claim_lattice LLM re-answer variance. No regression on
legit answers. Artifacts: bench/qa_results/2026-05-11T13-42-38Z (before)
and ...T14-19-51Z (after) — gitignored; summarized in qa-modes-bench.md
Addendum 5 + ticket-000046 §5 Phase 3.
Worked example: fork_score on the real change — parent {5f/
falsification: 4/12} → child {5f/falsification: 6/12} → γ·Δ5f =
(1/6)/5 ≈ +0.033 > 0 (positive; a single improvement of this size is
MARGINAL by the ÷5 dilution, the test pins the full-lift-to-1.0 case
at ACCEPT).
Tests: 5 new in tests/test_verify.py (numeric_signature normalization
+ subset-matches-comma-variant + disagreement-rejected +
gate-is-narrow + number-present-still-verifies);
test_5f_falsification_hard_pack_below_ceiling re-pinned 4/12 → 6/12.
make test 2339 passed, 28 skipped.
#000046 → closed; #000012 §8 §4 + TICKETS.md row + falsification-hard
_meta / hard-004,007 notes + Makefile 8/12→6/12 comments updated.
(Makefile also carries an uncommitted chain-check-SQL improvement from
the concurrent session — NOT in this commit; staged only the
#000046-comment hunks.)
Fox: "i like both" — keep the lazy out-of-band pass as the default
AND add the eager opt-in. Plus the Phase-1 gap fix (incremental embed).
Key insight folded into the design (new ticket section 14): a chunk_id's
content is immutable in arborist — same content → same chunk_id;
different content → a NEW chunk_id (re-ingest makes a new doc_root +
new chunk_ids linked by supersedes; a chunker bump re-chunks → new
chunk_ids). So a chunk, once embedded, never needs re-embedding — the
ONLY re-embed trigger is the embedder changing (VEC_BACKEND_VERSION
bump). That makes the idempotency story clean.
arborist/search/vec.py — embed_documents() now:
- incremental=True (default): embed only chunk_ids NOT already in
chunk_vecs (chunk_id NOT IN (SELECT chunk_id FROM chunk_vecs)).
This is the after-ingest / cron / Prometheus-Sigma-sweep path —
it picks up exactly the newly-ingested chunks; re-running is a
cheap no-op once everything's embedded.
- incremental=False: re-embed every chunk with content (delete-then-
insert all) — the embedder-changed case.
- rebuild=True: DROP + recreate chunk_vecs first, then a full pass —
the clean VEC_BACKEND_VERSION-bump path (a search mid-rebuild never
mixes old- and new-model embeddings: the recreated table starts
empty and grows new-model as the pass runs). Implies non-incremental.
- Cold-evicted chunks (content NULL) still skipped; vec rows persist
and stay valid (content is identical on rehydrate).
CLI (arborist/cli.py):
- arborist embed --rebuild — the DROP+recreate+full-re-embed path
(default is incremental). Output JSON now reports "mode".
- arborist ingest --embed — eager opt-in: after the chunk+Merkle-
commit pass, incremental-embed this run's new chunks. Default
ingest does NOT embed. ingest output gains "chunks_embedded" when
--embed is set. Only surfaced when the [vec] extra is installed.
- Hoisted the _vec_ok check up to the top of build_parser so both
the ingest --embed flag and the search --backend / embed subcommand
can gate on it.
tests/test_search_vec.py (7 -> 9): test_embed_incremental_only_embeds
_new_chunks (second pass after a follow-up ingest embeds only the new
chunk; third pass is a no-op), test_embed_rebuild_re_embeds_all
(DROP+recreate+full pass; vec_meta still records the version).
Verified on crawl_appliedcombinatorics_org.db: incremental on an
already-embedded shard reports chunks_embedded=0 in ~1.8s; --rebuild
re-embeds all 168 in ~61s; semantic search after rebuild still returns
topically-correct hits ("pigeonhole principle counting" -> "AC Graph
Coloring" chunk containing "Generalized Pigeon Hole Principle").
Ticket section 14 added: the idempotency table (re-ingest / chunker
bump / cold eviction / embedder bump / superseded docs), the two
integration models (lazy default + eager opt-in; the lazy pass's
natural home is a Prometheus-Sigma unconscious-sweep task per #000037
section 3.1), the command matrix, concurrency notes, versioning.
Status line updated.
Full suite: 2339 passed, 28 skipped.
(Unrelated parallel-clone work in the working tree — Makefile,
arborist/qa/verify.py, bench/fixtures/5f/*, tests/test_bench_batteries.py,
tests/test_verify.py — is #000046's hard-fixture tier, not touched here.)
Extends the #000046 hard tier to the Formulate sub-battery.
bench/fixtures/5f/formulate-hard-v1.jsonl — 12 prose inputs that
arborist.qa.parse_claims.parse_pointer_claims SHOULD segment into a
particular claim lattice (recorded in expected_lattice). The parser
is line/bullet-based — one line ⇒ one claim, [E#] tokens attach to
it — so 8 of 12 it mis-segments: merges and-/semicolon-/dash-joined
or (1)(2)-enumerated multi-claim lines into one claim with all the
pointers, or splits a wrapped bullet into two. Those 8 fail at HEAD
on claim-count mismatch; the other 4 are well-formed bullet/numbered
lists / single claims the parser handles right. Rate at HEAD = 4/12
= 0.333, stable (parse_pointer_claims is deterministic). A
claim-lattice parser that does sentence/clause segmentation (split on
'. ', ';', subordinating conjunctions, inline enumerations) + joins
wrapped bullets lifts the rate toward 1.0 → positive γ·Δ5f for that
child fork.
make bench-5f-formulate-hard runs it (|| true past the runner's
nonzero-on-failures exit). Not in `make bench-5f` / `runner --all`.
tests/test_bench_batteries.py — test_5f_formulate_hard_pack_below_ceiling
(pins rate 4/12, source=live, the 8 fails are claim-count
mis-segments).
#000046 → "Phase 1 + Phase 2 landed"; two below-ceiling 5F subs now
exist (falsification, formulate). Closure still pending an actual
surface improvement (verify_quotes tightening — bench-gated — or
parse_pointer_claims segmentation) that lifts a rate. §5 + §6 +
TICKETS.md row updated.
(Makefile also carries an uncommitted chain-check-SQL improvement
from the concurrent #000039 session — NOT included in this commit;
staged only the bench-5f-formulate-hard hunk + the .PHONY line.)
The #000025 §10.14 calibration showed _delta_5{s,t,f} mean over a
battery's 5 subs, so a single-sub gain weighs 1/5 of face value (the
5× dilution). #000047 ships the knob to pick the aggregation, default
unchanged.
WeightSet.delta_aggregator ∈ {"mean","max","sum"} (default "mean") —
a categorical field, validated in __post_init__ against
DELTA_AGGREGATORS; from_dict takes it as a string. Default unchanged →
ScoredFork output byte-identical → no fork_score.ESTIMATOR_VERSION
bump.
fork_score._aggregate(deltas, how): mean = arithmetic mean, max =
max(0.0, max_i Δ_i), sum = Σ Δ_i; empty → 0.0. _delta_5s/_delta_5t/
_delta_5f take an aggregator arg (default "mean"); the 5F efficiency
bonus is added after the aggregated base (aggregator-independent).
fork_score passes weights.delta_aggregator. The per-sub
HARD_REGRESSION_FLOOR flags are computed before aggregation, so a
single-sub regression still forces REJECT under max/sum. The chosen
aggregator is recorded in ScoredFork.weights["delta_aggregator"] (via
WeightSet.as_dict()); fork_score_branches traceability stays via the
opaque weights_id — no schema migration.
bench/scripts/fivef_threshold_calibration.py gained §5 — runs the
#000046 below-ceiling pack (5f/falsification at 0.333) and shows the
verdict / γ·Δ5f under each aggregator; bench/results/5f-threshold-
calibration-2026-05-11.md §5 is the captured record. Default stays
"mean" — the conservative, noise-robust, regression-symmetric choice
matching docs/bench-maxing.md's per-rate floor framing; v8 picks
max/sum per-deployment.
Tests: 8 new in tests/test_fork_score.py + 1 anchor in
tests/test_fivef_threshold_calibration.py; tests/test_weights.py
as_dict field-set test updated to include delta_aggregator;
test_fork_score.py AUTOCOUNT tags (#000012 §286, warrant-substrate-
cookbook.md ×2) bumped 23 → 31.
#000047 closed; #000012 §8 §3 + TICKETS.md row updated.
Full suite: 2330 passed, 28 skipped.
Implements the optional vec backend from the #000039 doc, with the
"obvious" v1 tuning, and demonstrates it on a real corpus shard.
arborist/search/vec.py (new):
- VecBackend(SearchBackend) — ANN over chunk_vecs, UNGROUNDED hits
(same as FTS5; vec changes recall, never warrant — embeddings are
soft signal, never in the proof path).
- chunk_vecs vec0 virtual table + vec_meta — sibling tables, additive,
don't touch chunks/documents/the audit chain.
- embed_documents() — batched ingest; delete-then-insert per chunk_id
(vec0 doesn't honor INSERT-OR-REPLACE — re-inserting an existing PK
is a hard UNIQUE error), so re-runs are idempotent and content-
changed → re-embed works. Skips cold-evicted chunks (content NULL).
- Pluggable Embedder callable; default = fastembed bge-small-en-v1.5
(~130 MB ONNX, downloads on first use). load_vec_extension(conn)
toggles enable_load_extension + sqlite_vec.load.
- v1 hyperparams (VEC_BACKEND_VERSION = vec-v1-bge-small-en-v1.5-
384float32-cosine-flat): model bge-small-en-v1.5, dim 384, quant
float32 (int8/binary = the production storage knob per §3.1, not
wired in v1), metric cosine (bge outputs L2-normalized, so cosine
ranking ≡ L2 ranking), ANN flat (vec0 default), top_k 20. These
five fold into governance_policy_hash in a later phase (§6).
CLI (arborist/cli.py):
- — populate chunk_vecs
for --db; prints progress + timing.
- — semantic ANN search (errors with
an install/embed hint if [vec] missing or chunk_vecs empty).
- Both surfaced only when sqlite_vec imports (mirrors the [html] /
selectolax pattern).
pyproject.toml: [vec] optional extra (sqlite-vec>=0.1.9, fastembed>=0.4);
added to [dev]. Note: sentence-transformers is the heavier "official"
embedder path §5 names; fastembed is the lightweight ONNX one.
tests/test_search_vec.py (7 tests, skip-if-no-[vec]): deterministic
stub embedder (hash → unit vector) so the suite exercises the
sqlite-vec plumbing — ext load, schema, ingest, KNN, JOIN, Hit shape,
limit, idempotent re-embed, --limit cap, empty/unpopulated — without
the heavy fastembed model. Semantic quality is demonstrated on a
shard, not unit-tested.
Demonstrated on ~/.arborist/shards/crawl_appliedcombinatorics_org.db:
168 chunks embedded in ~37 s (mostly model load); semantic queries
return topically-correct hits — "how many ways to choose k things
from n" → top hit "AC Combinations", "binomial coefficient counting"
→ "AC Introduction" (integer-solution counting) + "AC Combinatorial
Proofs". None of the query tokens need stem-match the chunk — the
semantic-allusion-gap closure the ticket promised. chain-check on
that shard reports 0 after embedding (chunk_vecs is a sibling table).
#000039 status flipped to "in progress · Phase 1 landed"; Phase 2
(RRF hybrid fusion in query.py) gated on a ≥5pp recall-lift
measurement with no STRICT-rate regression (§8).
(Unrelated: tests/test_weights.py::test_as_dict_returns_all_eleven_fields
fails in the working tree — that's a parallel-clone in-flight change
to arborist/substrate/weights.py + its test, not touched here.)
Closes the "everything is at rate 1.0 so fork_score's bench-Δ terms
are inert" gap the #000025 §10.14 calibration surfaced — at least on
the 5F/falsification axis.
bench/fixtures/5f/falsification-hard-v1.jsonl — 12 near-misses, each a
FALSE/unsupported claim whose correct verdict is UNGROUNDED (recorded
in expected_reason). 8 of 12 are over-grounded by
arborist.qa.verify.verify_quotes at HEAD — its paraphrase
token-coverage strategy returns STRICT_PARAPHRASE, its entity-proximity
strategy returns HYBRID_ENTITY, both matching on incidental overlap
(shared entities/numbers, the same key terms stated in the opposite
direction) — so those tasks fail by design; the other 4 the verifier
handles correctly. Rate at HEAD = 4/12 = 0.333, stable (verify_quotes
is pure-lexical / deterministic). Built around the pre-documented gap
5f-fal-live-003.
make bench-5f-falsification-hard runs the pack; make
bench-fork-baseline-hard pins it to
bench/results/baseline-falsification-hard.json. Both targets `|| true`
past the runner's nonzero-on-failures exit (8 fixtures fail by design;
the JSON is still written).
tests/test_bench_batteries.py — test_5f_falsification_hard_pack_below_ceiling
(pins rate 4/12, source=live, every fixture asserts UNGROUNDED, the 8
fails are over-grounds not abstentions) +
test_fork_score_positive_gamma_5f_on_hard_falsification_improvement
(the worked example: fork_score(parent={5f/falsification: 1/3},
child={5f/falsification: 1.0}) → gamma*Delta5f ≈ +0.133 > 0, verdict
ACCEPT, no regression flags — the bench Δ-rate carrying signal it
can't carry while every canonical pack is at ceiling).
NOT in `make bench-5f` / `make bench-5s5t5f` / `make
bench-fork-baseline` / `runner --all` — the hard pack is a separate,
deliberately-failing artifact pinned on its own.
#000046 flipped to "in progress · Phase 1 landed"; closure pending an
actual verify_quotes tightening that lifts the rate (a separate,
larger task). ticket-000012 §8 §4 + TICKETS.md row updated.
Full suite: 2314 passed, 28 skipped.
The 5F threshold calibration (#000025 §10.14) surfaced two loose ends
that #000012's v8 acceptance protocol owns; opening them as proper
proposals so they don't get lost.
#000046 — Harder 5S/5T/5F fixture tier (below-ceiling baselines).
Every canonical 5S/5T/5F pack runs at rate 1.0 at HEAD, so fork_score's
α·Δ5s+β·Δ5t+γ·Δ5f terms are inert (a child can only move them ≤ 0).
The embedded packs can't be made "harder" (their evaluator is the
gold); a harder tier is necessarily a live-path tier where a real,
imperfect arborist surface produces wrong/partial output. Recommends
Option A scoped narrow: a `falsification-hard-v1.jsonl` built around
the verify_quotes entity-strategy gap already documented by
5f-fal-live-003, confirm the rate sits below 1.0 and is stable, pin it
via `make bench-fork-baseline-hard`, demonstrate a planted surface fix
lifting the rate. Status: open · awaiting go/no-go.
#000047 — ForkScore _delta_* aggregator (mean vs max vs sum). _delta_*
means over a battery's 5 sub-batteries, so a single-sub gain is worth
a fifth of face value (the 5× dilution from the §10.14 calibration:
+0.06 on one 5F sub → MARGINAL, same +0.06 on all five → ACCEPT).
Recommends Option D: parameterize delta_aggregator on WeightSet,
default "mean", record it in the breakdown + fork_score_branches; bump
ESTIMATOR_VERSION only if the default changes. Defer implementation
until #000046 lands a below-ceiling baseline so the aggregator choice
is benchable, not guessed. Status: open · awaiting go/no-go; parks
behind #000046.
TICKETS.md index rows + Next ID bump (000046 → 000048) landed in
8599ce3 (swept in by a concurrent commit; content is correct).
ticket-000012 §8 §3/§4 updated to name #000047 / #000046.
"One more iteration then close" (fox): added committed KAT-regeneration
scripts for both the T3 calculator and φ_PRG — the regen step was a
throwaway temp script before; now it's reproducible and the phi_prg
test's skipif reason ("run scripts/generate_phi_prg_kat.py") points at
a file that exists. Then closed#000036.
New scripts:
- scripts/generate_t3_bound_kat.py — regenerates
bench/fixtures/t3-bound/known-answer-tests.jsonl from a fixed 12-config
list (the §7 worked examples under max_envelope + non-default-C_B*
+ g=0 edge + explicit-b1_model pins for the other three models).
- scripts/generate_phi_prg_kat.py — regenerates
bench/fixtures/phi-prg/known-answer-tests.jsonl from a fixed 10-entry
list (placeholder/random seeds, one-bit-flip variants, block-boundary
dim_h=16/17, 4096 counter-rollover stress).
- Both verified to reproduce the committed fixture data lines byte-
for-byte (only the header comments changed, to reference the script).
Each docstring states: run after any algorithm change, then bump the
module version (CALCULATOR_VERSION / PHI_PRG_VERSION) so the fixture's
version field changes too.
Doc/test:
- test_t3_bound_calculator.py skipif reason now references the regen
script (matches the phi_prg test pattern).
- #000035 §3.3 + t3-bound.md §10.1 reference the regen scripts.
Closure (#000036):
- Status → closed · 2026-05-11 in the ticket file + TICKETS.md row.
Phase 1 + dav1d Tier-1/Tier-2 (Option B in v1) + KAT-regen tooling
all landed; all §5 acceptance criteria met; both dav1d closure
blockers cleared. Continuation: empirical C_B1/C_B2/C_B3 tightening
under #000043 (parks on v7 deployment data); landing the bound's
framing into a v7 plastic-training spec parks on that spec gaining
a deployment target; R2's architectural integrations (Merkle audit-
event commitment, SQD canonicalization, CTI clause-lattice, 5F
trigger, ForkScore security-risk) are separate tickets if wanted.
- t3-bound.md header flipped to "closed 2026-05-11".
Full suite: 2312 passed, 28 skipped.
v7's canonical integer byte-order was confirmed little-endian by
inspecting merkle-agi-dag_v7.txt §A1 — every to_bytes / astype in the
TLV encoding is little-endian (TLV length prefixes to_bytes(4,'little'),
enc_int to_bytes(8,'little'), quantized tensors '<i8'); no big-endian
anywhere. Per dav1d's 2026-05-11 review rule ("if v7 TLV canonical
integer encoding is little-endian, flip §3.4 to little-endian before
KAT freeze"), flip done — this is the -le variant.
Implementation (arborist/substrate/anchor_prg.py):
- PHI_PRG_VERSION → "phi-prg-v1-hmac-sha512-le" (still "v1";
the -le suffix records the endianness; future re-flip MUST bump).
- _expand: counter.to_bytes(4, 'big') → 'little'.
- _bytes_to_floats: int.from_bytes(..., 'big') → 'little' (the
uint32-word interpretation, for full consistency with v7).
- Module + function docstrings updated: little-endian throughout,
with the merkle-agi-dag_v7.txt §A1 verification note.
- Note: at counter=0 the bytes are identical regardless of
endianness, so 5 of the 10 KAT entries (dim_h ≤ 16, single block)
keep the same output_sha256; the 5 multi-block entries (dim_h 17/
32/64×3/4096) change.
KAT fixture (bench/fixtures/phi-prg/known-answer-tests.jsonl):
- Regenerated under the little-endian counter. Each entry now also
carries a "version" field (phi-prg-v1-hmac-sha512-le). Header
comment updated.
Tests (tests/test_anchor_prg.py, 30 → 31):
- test_module_exports_version_string: assert the -le suffix.
- test_bytes_to_floats_midpoint_maps_to_zero: 2^31 is b'\x00\x00\x00\x80'
in little-endian, not b'\x80\x00\x00\x00'.
- New test_bytes_to_floats_reads_little_endian: pins the byte-order
so an accidental re-flip is caught.
- test_phi_prg_first_block_matches_direct_hmac: uint32-word reads
little-endian (counter=0 bytes unchanged either way).
- test_phi_prg_known_answer_tests: assert kat['version'] == module
version when present.
Spec text (#000035 §3.4): folded the little-endian variant of
dav1d's §9.10 wording — counter_le32, uint32_le word reads, an
"all integers little-endian, matching v7 TLV §A1" preamble, and an
"Endianness — RESOLVED 2026-05-11" note replacing the open
big-vs-little question. soft-hash-channel-analysis.md §9.2/§11 +
#000035 status + TICKETS.md row updated. AUTOCOUNT for
test_anchor_prg.py bumped 30 → 31; PHI_PRG_VERSION refs in docs
bumped to -le.
Full suite: 2312 passed, 28 skipped.
Closes the three open Phase-1b items of #000025; every §10 closure
criterion is now met, so the ticket flips to closed.
§10.14 — ForkScore threshold-calibration handoff to #000012.
bench/scripts/fivef_threshold_calibration.py (make bench-5f-threshold-
calibration) runs the canonical 5S/5T/5F packs + the 5F live packs and
reports baseline rates, observability granularity (1/n), and fork_score
verdicts on the parent vs synthetic child perturbations →
bench/results/5f-threshold-calibration-2026-05-11.md. Findings written
into ticket-000012 §8: keep SIGNAL_FLOOR / HARD_REGRESSION_FLOOR at
0.05; the small 5S packs (syntax n=10, semantics n=8) are coarser than
the floors so any regression there trips hard-reject (intended zero-
tolerance); the 5x averaging dilution in _delta_*; ceiling saturation
(every pack at 1.0 -> delta-rate terms <= 0). No constant change
shipped. 6 tests in tests/test_fivef_threshold_calibration.py.
§10.13 — feedback latency / efficiency on real workload.
run_feedback_loop now computes feedback_latency (listed in §5.5 since
Phase 1a, never implemented) — wall-clock seconds to apply a live
chain against its temp shard, surfaced per-task
(feedback_latency_seconds) + battery (feedback_latency_mean_seconds,
feedback_live_task_count). For live chains feedback_efficiency's cost
denominator switched from len(chain) (count of requested ops) to the
persisted footprint _persisted_cost = audit-event rows the chain
actually wrote + their body bytes / 1e6. Embedded chains keep
len(chain) and report feedback_latency_seconds = None. Latency is a
wall-clock field (run-to-run variable, like BatteryResult.timestamp)
and is not a fork_score input. 3 tests in tests/test_bench_batteries.py.
§10.11 — real selfmodel finetuning chains.
bench/scripts/selfmodel_chain_snapshot.py (make bench-5f-selfmodel-
snapshot) appends one chained SelfModel snapshot per run to a
persistent shard (~/.arborist/shards/selfmodel-chain.db, override via
ARBORIST_SELFMODEL_CHAIN_DB) with one CapabilityClaim per sub-battery
(metric = "5S-syntax" etc., measured_value = that pack's rate,
eval_digest = the pack's fixture digest, threshold = SIGNAL_FLOOR).
snapshot() auto-parents, so each snapshot is a distinct root and the
lineage grows by one per run. run_finetuning gains a third dispatch
mode — shard-chain (gated on a task's selfmodel_shard key) — via
_chain_finetuning_measure: reads the two most-recent snapshots
(latest() = child, its parent_selfmodel_root = parent) and measures
improvement on target_capability between them. This is the real
lineage replacing Phase-1a's synthetic parent->child pairs; the
chained delta reflects genuine cross-run drift (0.0 today — the
embedded packs are at ceiling). Operator pack
bench/fixtures/5f/finetuning-shardchain-v1.jsonl (6 tasks) + make
bench-5f-finetuning-shardchain; not in `make bench-5f`, `make test`,
or a fresh checkout (a missing/too-short chain fails honestly). The
real chain shard was bootstrapped 2-deep on 2026-05-11; make
chain-check-shards reports 0 breaks on it (and all other shards).
10 tests in tests/test_selfmodel_chain.py.
Full suite: 2311 passed, 28 skipped.
dav1d returned the §3.4 φ_PRG anchor-map review with a decision set:
HMAC-SHA-512 / 32-byte seed / uint32-be counter from 0 / SHALL-replace
all LOCKED; manifest field renamed; float-map prose corrected; two
ADDs (exhaustion guard + seed-independence rule); M1-policy separation.
Spec text (#000035 §3.4):
- Folded dav1d's full corrected §9.10 wording (RESPONSE_1 §1).
- Manifest field phi_prg_seed → anchor_prg_seed (purpose-scoped, not
implementation-scoped; phi_prg_seed kept only as a code-local alias;
phi_seed / m1_anchor_seed rejected as too vague / too policy-tied).
- Float map 2·(u32/2^32)−1 unchanged (KAT compat) but the prose now
says "uniform over a 2^32-point grid in [-1, 1) with negligible
finite-grid mean −2^−32" — NOT "unbiased". -1.0 reachable, +1.0
not. If exact zero-mean is ever needed → midpoint map x =
2·((u32+0.5)/2^32)−1 with a PHI_PRG_VERSION bump + new KATs, never
a silent change.
- Added dim_h ≤ 16·2^32 exhaustion guard (4-byte counter ceiling).
- Added seed-independence + single-purpose-seed requirements (seed
must be generated independently of model/data, not adversary-
selected, not reused for other PRG domains — no domain-separation
tag in v1).
- Added §9.10.1: M1 enablement is a mitigation-selection-policy
decision (e.g. skippable under #000034 NO_ALIGNMENT), not a §9.10
function-definition question; "MUST NOT claim M1 while still using
embed_hard_to_vec" prevents fake-M1 deployments.
- Added an endianness-confirmation note: big-endian is pinned to the
impl + KATs; flip only if v7 TLV convention turns out little-endian
(would need a PHI_PRG_VERSION bump).
- SHALL-replace wording kept (RFC-2119 strong mandate inside M1).
Implementation (arborist/substrate/anchor_prg.py):
- New dim_h > 16·2^32 → ValueError guard (clean message naming the
ceiling rather than overflowing the counter deep in _expand).
- bool dim_h now rejected explicitly (isinstance(True, int) is True).
- Module + function docstrings updated: manifest field is
anchor_prg_seed; seed-independence / single-purpose rules; corrected
float-map distribution wording (negligible mean −2^−32, not exactly
zero); endianness note.
Tests (tests/test_anchor_prg.py, 27 → 30):
- test_phi_prg_rejects_bool_dim_h (True/False params).
- test_phi_prg_rejects_dim_h_above_counter_ceiling.
Doc cross-refs: soft-hash-channel-analysis.md §9.2 + §11 status note
the dav1d-reviewed §9.10 wording + anchor_prg_seed field name.
#000035 ticket status + TICKETS.md row updated. AUTOCOUNT markers
for test_anchor_prg.py bumped 27 → 30 across 5 doc files.
Full suite: 2291 passed, 28 skipped.
Per fox: apply the conservative max_envelope B1 model by changing the
v1 calculator's default — NOT by forking a v2. CALCULATOR_VERSION stays
"t3-bound-v1-bottou-refinement" (the descriptor names the unchanged B3
term); b1_model is echoed in the output AND the inputs dict so KAT
replays are unambiguous about which model produced a row.
Calculator (bench/scripts/t3_bound_calculator.py):
- New b1_model kwarg + --b1-model CLI flag, choices:
max_envelope (default) max(fraction_channels, aggregate_bias)
fraction_channels g · W · log₂(1 + G/σ)
aggregate_bias W · log₂(1 + g·G/σ)
effective_control_v1 g · W · log₂(1 + g·G/σ) (old non-worst-case)
- Default is now max_envelope — genuinely upper-bounding across both
interpretations of g (dav1d review §3 closure blocker, RESOLVED).
- Every output reports all three concrete B1 variants
(B1_fraction_channels / B1_aggregate_bias / B1_effective_control_v1),
b1_selected, and both SNR readings (snr_grad = g·G/σ,
snr_per_channel = G/σ) regardless of which b1_model was requested.
- model_assumptions[] now carries f"B1_model_{b1_model}".
- inputs echo now includes c_b1/c_b2/c_b3/b1_model (replay-complete).
- Invalid b1_model rejected with a ValueError naming the field.
- Baseline I_window: 625.8716 (effective_control_v1) → 6183.0154
(max_envelope: B1=aggregate_bias 5849.63 dominates fraction_channels
1729.72), certification_status NOT_CERTIFIED_BY_BOUND at W=10000.
KAT fixture (bench/fixtures/t3-bound/known-answer-tests.jsonl):
- Regenerated 2026-05-11 — 12 entries: the 8 §7-derived configs under
the new max_envelope default, a g=0 edge case, plus explicit-mode
pins for effective_control_v1 / fraction_channels / aggregate_bias.
- Each entry carries b1_model, expected_b1_selected,
expected_b1_{fraction_channels,aggregate_bias,effective_control_v1},
expected_snr_per_channel, expected_certification_status.
Tests (tests/test_t3_bound_calculator.py, 75 → 83):
- test_t3_bound_known_answer_tests no longer skips (fixture active);
pins b1_model, b1_selected, certification_status + numbers, tolerates
optional new fields on older fixtures.
- New: test_b1_max_envelope_exact_formula, test_invalid_b1_model_rejected,
test_cli_b1_model_flag (effective_control_v1 / fraction_channels /
aggregate_bias). test_b1_exact_formula renamed
test_b1_effective_control_v1_exact_formula and now passes the explicit
model. Updated baseline / below-256 / CLI tests for the new numbers.
Doc (docs/soft-hash-channel-t3-bound.md):
- Header + §0 + §3.1 + §6 + §7 (worked examples) + §8 (operator
guidance W-solving) + §10 (closure blockers RESOLVED) + §10.1 +
§11 (calculator schema) + §12 all updated for the max_envelope
default. §8: target-256 W drops from ~4196 to ~415 steps under the
conservative model — the ~10× cost of not assuming which g-reading
holds; operators who can measure effective-control applies can use
--b1-model effective_control_v1 for the looser W (a calibration
claim they must justify, not a default).
Status (#000036 ticket + TICKETS.md): both prior dav1d closure
blockers cleared (B1 worst-case model + active KAT fixture); remaining
= fox's final close-or-iterate call.
AUTOCOUNT markers bumped 75 → 83. Full suite: 2288 passed, 28 skipped.
trigger_1_branch_density in bench/prometheus_sigma_trigger_probe.py read
the fork_score_branches table via branch_set_density() instead of the
"density check not yet implemented" stub. It groups rows by branch_set_id,
fires when the most-recently-recorded checkpoint carries >= 4 branches
(BRANCH_DENSITY_FLOOR), and surfaces n_checkpoints / latest_density /
max_density / n_checkpoints_clearing_floor in the markdown report so
section 12's "regularly" qualifier stays visible. Density sums across
shards per branch_set_id.
With no branch sets persisted yet the probe reports "table present but
empty across shards" (data_available True, fires False) rather than a
false negative. Re-ran the probe against the live shards:
bench/results/prometheus-sigma-triggers-2026-05-11.md.
6 new tests in tests/test_prometheus_trigger_probe.py (probe loaded via
importlib): no-table to no-data, empty-table to data-available-no-fire,
latest-checkpoint->=4 to fires, earlier-dense-but-latest-sparse to no-fire,
density-sums-across-shards, report-renders-density-lines.
Updated #000012 Phase 1c landing receipt, #000037 section 12 Trigger 1
note, and the TICKETS.md index rows for both. Pure measurement: no
mutation, no LLM call, no schema change.
dav1d's review (RESPONSE_1 + RESPONSE_2) returned 2026-05-11. This
lands the Tier-1 items — everything that doesn't change numeric
outputs or invalidate the KAT discipline. The Tier-2 B1 conservative-
envelope (v2 calculator) is a separate decision and stays a closure
blocker.
Calculator (bench/scripts/t3_bound_calculator.py):
- Recommendation wording: "M2's single-window guarantee is broken"
→ "this conservative bound CANNOT CERTIFY M2's residual". An upper
bound exceeding 256 bits means we cannot certify, NOT that the
adversary can steer 256 bits — the prior wording overclaimed.
- New structured output fields: b1_model ("effective_control_v1"),
certification_status ∈ {CERTIFIED_BY_BOUND, NOT_CERTIFIED_BY_BOUND},
certification_threshold_bits (256), model_assumptions[]. Callers
read a machine-readable status, not just prose.
- Input validation hardening: _require_finite_float / _require_positive_int
helpers reject bools (isinstance(True, int) is True in Python — a
real leak risk for a security calculator) and NaN / ±inf for every
numeric input and constant.
- gradient_fraction = 0 now accepted (no T2 surface; B1 = 0; T3's
LR + batch-order channels still contribute) — improves component
isolation. CLI help + module docstring updated accordingly.
- Numeric outputs UNCHANGED: baseline still 625.8716 / 292.4813 /
300.0 / 33.3904; b1_model stays effective_control_v1; KAT discipline
intact.
Tests (tests/test_t3_bound_calculator.py, 53 → 75):
- Hard-coded cwd="/home/fox/git/arborist" → pathlib.Path(__file__).
resolve().parents[1] so the suite runs on any checkout.
- New: test_gradient_fraction_zero_accepted, test_bool_rejected_for_int_fields,
test_bool_rejected_for_float_fields, test_nonfinite_numbers_rejected,
test_output_carries_b1_model_and_certification_fields,
test_certification_status_certified_below_threshold.
- test_recommendation_exceeds_sha256 now also asserts "CANNOT CERTIFY"
+ certification_status == NOT_CERTIFIED_BY_BOUND.
Doc (docs/soft-hash-channel-t3-bound.md):
- §0 reworked into a reviewer brief recording dav1d's findings
(§2 accepted, §4 accepted, §5 accepted as model-bound, §3 = closure
blocker, wording/validation = applied).
- New §3.1: the B1-double-g issue spelled out — effective_control_v1
vs fraction_channels vs aggregate_bias vs max_envelope, with the
baseline-spread table (292 / 1730 / 5850 / 5850 bits); v2 path
described.
- §5: "B3 is a model-bound, not a directly-quoted theorem" note.
- §10: items 1-2 are now the closure blockers (B1 envelope v2; active
KAT fixture); items 3-7 are tightening paths (#000043). New §10.1
records what the 2026-05-11 hardening pass already landed.
- §11: calculator-output example updated to show the new fields +
corrected recommendation wording.
- §12: references add the dav1d review + clarify Bottou-Bousquet
"inspires" (not "underlies") the §5 model-bound.
Status (#000036 ticket + TICKETS.md row): review-returned + Tier-1-
applied; closure blockers = B1 v2 envelope (awaits fox go/no-go) +
active KAT fixture. R2's architectural integrations (Merkle audit-
event commitment, SQD canonicalization, CTI clause-lattice, 5F
trigger, ForkScore security-risk) noted as out-of-scope (separate
tickets if wanted).
AUTOCOUNT markers in docs/calculator-test-patterns.md +
docs/warrant-substrate-cookbook.md bumped 53 → 75.
Full suite: 2264 passed, 28 skipped.
Phase 1a scores one (parent, child) fork at a time; Phase 1b is the
consensus paper. Neither persists multiple candidate branches at the
same checkpoint — and #000037 §12 Trigger 1 ("ForkScore regularly
receives ≥4 candidate branches per checkpoint") gates the multi-
branch path of the Prometheus-Σ controller on this data existing.
Phase 1c lands the missing seam.
Schema (arborist/store.py): _migrate_fork_score_branches creates the
sibling table with PK (branch_set_id, branch_id) + indexes on
branch_set_id and parent_root. Sibling — never enters
audit_events.event_hash preimage, so re-scoring or back-filling
cannot break the audit chain.
Helpers (arborist/substrate/fork_score.py): persist_branch_score
upserts one row via ON CONFLICT (branch_set_id, branch_id) DO UPDATE
so re-scoring the same fork under the same checkpoint is a clean
overwrite, not a duplicate. branch_set_density(conn, branch_set_id)
returns the count of distinct branches recorded under a checkpoint
— the function the #000037 §12 Trigger 1 probe reads.
ESTIMATOR_VERSION = "fork-score-v1" pins the producer generation on
every persisted row.
CLI (arborist/cli.py): arborist substrate score gains six new flags
(--branch-set, --branch-id, --parent-root, --child-root,
--persist-shard, --weights-id). Default off — --branch-set absent
preserves Phase 1a pure-function semantics for every existing
caller. When present, requires --parent-root and either --branch-id
or --child-root; missing inputs return exit code 2.
Tests (tests/test_fork_score.py, count 18 → 23): migration creates
the table + both indexes; persist writes one row carrying
parent/child roots + verdict + weights_id + estimator_version;
upsert on the PK refreshes child_root + weights_id + recorded_at
without duplicating; branch_set_density counts per-checkpoint and
ignores cross-set rows; breakdown_blob round-trips as canonical
JSON whose values sum to the persisted score.
Status sync: #000012 §7 Phase 1c flipped from "proposed, not yet
open" to "landed 2026-05-10" with the original proposal preserved
below as design log. TICKETS row 117 mirror-updated. AUTOCOUNT
counters in #000012 + cookbook bumped 18 → 23 plus the cookbook's
fork_score.py LOC row refreshed (298 → 386 module, 403 → 609
tests, density 1.35 → 1.58).
End-to-end smoke verified: arborist substrate score writes a
fork_score_branches row with the expected schema (verdict / weights_id
/ estimator_version) and the row survives a clean SQLite read.
Per fox's just-codified §-status drift discipline (#000044 commit
4e41c73), close the doc loop on this evening's commits before
moving on. Three surfaces synced to truth:
(1) Header status line — was "Phases 0 + 1 + 1.b + 2 landed",
silent on Phase 1.c, the 4th event kind, the inspector, the live-
harvest pipeline, the §22 Findings 2 + 3 resolution, and #000045.
Now mentions all of them in one tight paragraph.
(2) §20 Status detail rows — Phase 1 LOC count refreshed
893 → 944 + test count 36 → 42; new Phase 1.c bullet (commit
4b85a0a) describes the kernel_cost/llm_cost split + effective_cost
back-compat property + sweep_weights() profile + §15.4
documentation; Phase 2 row collapsed multi-commit history
(a786d6d + 43380b1 + cc72784 + 70c2184) into a single bullet
covering all four event kinds, the QA-runner wiring, the live
inspector subcommand, and the live-harvest pipeline; LOC 200 → 239,
tests 14 → 25.
(3) TICKETS.md row 103 — mirror of (1) at index granularity. Now
includes Phase 1.c, the 4 event kinds, both downstream consumers
(inspector + harvest), and the #000045 gating-ticket pointer.
No code changes — pure status-drift cleanup. 91/91 tests still
green (test_directives + test_prometheus + test_prometheus_audit).
Two doc-only fixes from #000045 walk-through:
== #000045 §4 — `sweep_weights()` swap marked resolved ==
§4 "Out of scope" included:
"**No `sweep_weights()` swap in the dry-run.** ... A later
change can swap it ... that is a follow-up inside #000037,
not this ticket."
But fox's commit `6734f80` (2026-05-10 19:22 EDT) already did
the swap — landed ~3 minutes before #000045 was opened. The §4
item was stale at the moment it was written; the follow-up
already happened. Strike-through marker + resolved note added,
explaining that the dry-run report numbers were intentionally
allowed to drift from prior-bench comparability since sweep
profile is operationally correct for sleep-sweep economics
(γ_5f=1.5 / λ_capital_cost=0.25 / ν_witness_divergence=0.5),
not the safe profile.
This is a third-order drift: a ticket's "deferred follow-up"
marker becoming stale because the follow-up landed concurrently.
Caught by reading #000045 §4 line-by-line during the walk.
== Cookbook prometheus_audit count 22 → 25 ==
Third same-day AUTOCOUNT drift catch since the harness landed:
- 36 → 42 (catch 1, during cookbook write — fox in-flight work)
- 17 → 22 (catch 2, after fox's `cc72784` inspector CLI landed)
- 22 → 25 (catch 3, fox's in-flight further additions during
#000045 walk)
Each refresh has been the same shape — file+line+claim/live diff
fires, 60-second turnaround. The discipline is structurally
holding under heavy commit pressure.
fox's 6 currently-uncommitted files (cli.py + prometheus_audit
+ harvest_falsification_proposals + test_bench_batteries +
test_prometheus_audit + the touched docs) preserved in working
tree — my doc-only commit stages only these two doc files
explicitly.
Verification:
$ pytest tests/test_doc_counts.py
3 passed in 2.81s
The dry-run was still using safe_weights(); Phase 1.c shipped
sweep_weights() (γ_5f=1.5, λ_capital_cost=0.25,
ν_witness_divergence=0.5) precisely for sleep-sweep economics. Swap
both call sites in prometheus_sigma_sweep_dryrun.py and label the
report Weight profile: sweep (§15.4).
§15.4 adds the sweep profile alongside §15.1 safe / §15.2
conservative / §15.3 exploratory so the named-profile registry has
a single doc source of truth and the dry-run report's §15.4
reference resolves.
§22 dry-run-runs table (3 rows: initial uniform-τ-1d under safe;
per-mode τ under safe; per-mode τ under sweep) replaces the prior
inline narrative count. Includes a Phase-3-design observation: the
sweep profile flattens the softmax (DEFERRED 271 → 369; REJECT
205 → 110; ACCEPT 1 → 0; MARGINAL 2 → 3) because reducing
λ_capital_cost + ν_witness_divergence shrinks the gap between
high-Δ5F and low-Δ5F branches, so fewer branches reach Kelly's
p_i > 0.5 floor. Falsification-fixture proposal count stayed flat
(447 → 449) — the §13 step 11 emission path is upstream of decision
labeling, so sweep mode preserves information-gathering value while
concentrating action-taking on the high-confidence tail. Phase 3
ticket #000045 §2 deliberately leaves softmax temperature un-pinned
because of this trade-off.
Today's audit wave (820409b + 952abc5 + 3b30126) surfaced a
sibling drift class to AUTOCOUNT: ticket body §-status sections
freeze at landing time while file headers update inline. 5 of 6
in-progress tickets audited had this drift. Codify the two
distinct rules (in-progress: refresh in place; closed: archival)
as new §10 of #000044 so future blackops shifts have a structured
reference instead of having to re-derive the discipline.
== New §10 — Sibling discipline — §-status section drift ==
Three sub-sections:
§10.1 In-progress tickets — body §-status must match file
header. Walk pattern when amending: (1) update header, (2) walk
body §-status + refresh, (3) walk TICKETS.md index row + refresh.
The header is the load-bearing surface; lagging body sections
are pure drift.
§10.2 Closed tickets — body §-status is archival. Do NOT
rewrite when later phases land. Phase 2/2.5/B-1/B-2 of a closed
ticket (e.g. #000031) go in:
- The file header (consolidated)
- A NEW section appended below ("Subsequent phases" /
"Phase B follow-ups")
Rewriting §8 of #000031 with Phase 2 content would erase the
Phase 1 landing record — the design log.
§10.3 The rule in one line:
**In-progress §-status must match header. Closed §-status is
archival and stays frozen.**
Table at the top shows the 6 in-progress tickets audited today:
- #000034 §7 lead stale → refreshed
- #000035 §7 lead CLEAN
- #000036 §7 lead + closure(c) stale → refreshed
- #000037 §20 + §17.1 stale → refreshed
- #000025 §11 lead stale → refreshed
- #000012 §7 lead + Phase 1a tests stale → refreshed (file rename
caught: test_v8_fork_score.py →
test_substrate_fork_score.py per a4058a4)
Rule is NOT machine-checked. The audit pattern is manual +
periodic. Sweep-triggers documented:
- Any in-progress → closed flip (final § refresh before freezing)
- Any in-progress gains a new phase landing (refresh §)
- Quarterly housekeeping pass
Long-term, could become a tagged claim (an AUTOCOUNT-like metric
asserting §-status text matches header text up to whitespace
normalization). Deferred until the audit pattern recurs enough
to justify the surface. Speculative; currently below threshold.
== Numbering ==
References renumbered §10 → §11 (one section pushed; new §10
inserted between §9 Scope boundaries and §10 References → §11).
No other content moved.
== Why land this in #000044 specifically ==
#000044 already documents doc-drift discipline (the AUTOCOUNT
mechanism). §-status drift is a parallel manifestation of the
same root cause (point-in-time snapshots freezing while the
load-bearing surface updates inline). Both deserve a single
home. The alternative — a separate #000046 ticket for §-status
drift — would split the discipline into two designs that share
99% of the rationale.
== Verification ==
$ pytest tests/test_doc_counts.py
3 passed in 2.73s
No new tags. No code change. Doc-only.
§7 lead "In progress · Phase 1a landed 2026-05-08" was stale —
file header (current) says "Phase 1a (ForkScore) landed
2026-05-08; Phase 1b (consensus paper) landed 2026-05-10; Phase
1c (branch-set persistence) remains proposed-not-opened."
Refresh to match file header so the index claim and the body
section are saying the same thing.
§7 Phase 1a body referenced "tests/test_v8_fork_score.py — 25
cases" on two counts of drift:
1. **Filename stale**: file was renamed to
`tests/test_substrate_fork_score.py` in `a4058a4` 2026-05-10
under the v-prefix-retirement convention (substrate-paper
version vs schema-version disambiguation). The original path
no longer resolves.
2. **Count + scope conflated**: 25 cases at the time, but the
surface has since split into two test files:
- `tests/test_fork_score.py` (18 tests) pins the pure
ScoredFork dataclass + scoring contract.
- `tests/test_substrate_fork_score.py` (27 tests) pins the
`arborist substrate score` CLI surface.
Refresh: name both files explicitly + AUTOCOUNT-tag each count
so future drift fires (#000044 discipline). Note the rename
provenance for future readers walking the design log.
Six in-progress tickets audited for §-status drift in today's
sweep wave:
#000034 §7 lead stale (now refreshed in 952abc5)
#000035 §7 lead clean
#000036 §7 + closure stale (refreshed in 952abc5)
#000037 §20 + §17.1 stale (refreshed in 3b30126 + 952abc5)
#000025 §11 lead stale (refreshed in 3b30126)
#000012 §7 lead stale (THIS COMMIT)
5 of 6 had body-section drift. Pattern is consistent: file
header is the load-bearing surface fox updates; body sections
are point-in-time prose snapshots that freeze at landing time.
Sweep-on-amend is the right discipline; this commit closes the
last in-progress ticket §-status drift surface for today.
#000006 (rolling research log) and #000044 (closed today,
retroactive design log) excluded from sweep — different shapes.
Closed tickets prior to today excluded entirely per #000044 §5
discipline ("closed-ticket point-in-time snapshots stay
UNtagged / refreshing them changes archival meaning").
Verification:
$ pytest tests/test_doc_counts.py
3 passed in 2.48s
Tags by metric after this commit: 48 tests + 6 fixture-rows +
3 db-rows + 3 db-where = 60 active claims (was 58; +2 from this
commit's two new tests-metric tags on fork_score test files).
Two parallel landings in one commit:
== §-status sweep on in-progress tickets ==
Checked #000034 / #000035 / #000036 §7 Status sections against
their file-header status. Same pattern as earlier sweep on
#000037 / #000025 / #000006 — body-section prose freezes at
earlier snapshots while file headers stay current. 2 of 3 had
drift.
**#000034 §7** lead: was "Open · awaiting go/no-go" + "Phase 1a
proposed below" + body subsection "Phase 1a (landed 2026-05-10)".
The lead line was a 2-versions-old prose snapshot contradicting
the same section's own subsection. Refresh: "In progress · Phase
1a landed 2026-05-10 (commit 1dfb8b9; test backfill a4b3056).
Phase 1b parks until v7 reference checkpoint..."
**#000035 §7**: already says "In progress · Phase 1 landed
2026-05-10". No drift; sample-confirmed clean.
**#000036 §7** lead: was "Awaits fox + cryptographer review of
constants." Refresh to current state — "pre-review polish pass
in 8916bf3; math review in flight with dav1d (forwarded
2026-05-10 Asia/Kuala_Lumpur as Tier 2 bundle: t3-bound.md +
soft-hash-analysis.md + t3_bound_calculator.py +
test_t3_bound_calculator.py + this ticket). Empirical tightening
tracked separately under #000043." Matches file header verbatim.
**#000036 closure criterion (c)** was marked "[pending —
docs-only commit]" but soft-hash-channel-analysis.md §9.3 line
475 already reads:
"§9.3 closed 2026-05-10 via the T3 per-window bound at
docs/soft-hash-channel-t3-bound.md (under #000036)"
So criterion (c) is done. Refresh: "[done — §9.3 closure landed
in line 475 of that doc]". Also clarify criterion (d) gates on
the dav1d review verdict (was ambiguous as just "explicitly
accepted by fox").
== #000037 §20 — reference #000045 follow-on ==
fox opened #000045 (Prometheus-Σ Phase 3 sleep-sweep scaffold)
in commit 0379e4c. My §20 refresh from 3b30126 said "Phase 3
gates on a renewed §12 trigger plus a Phase 3 ticket" — now
that Phase 3 ticket is named: #000045. Refresh to point at it
explicitly so future readers don't have to grep TICKETS.md to
find the follow-on.
== CLAUDE.md — AUTOCOUNT discipline rule ==
Added a paragraph to "Operational rules" pointing future
blackops shifts at the AUTOCOUNT discipline (#000044 + the
harness at tests/test_doc_counts.py). Without this, the harness
is discoverable only by accidentally hitting a test failure or
reading commit messages. With it, "tag at write time" is a
documented operational rule alongside Python-only, PYTHONUNBUFFERED,
and fail-closed.
Rule names: format, 4 metrics (tests / fixture-rows / db-rows /
db-where), code-fence-skip behavior, closed-ticket-stay-untagged
discipline, pointer to #000044 + the harness module.
== Verification ==
$ pytest tests/test_doc_counts.py
3 passed in 2.80s
No drift introduced. fox's parallel 4 commits (4b85a0a /
43380b1 / 1f882df / 0379e4c) landed independently; my
modifications stack on top cleanly. fox's commits added 155 +
89 lines to test_prometheus.py + test_prometheus_audit.py but
collected counts stayed at 42 + 17 — those landings modified
existing tests (parametrize additions / refactors), didn't add
new test functions. AUTOCOUNT claims from 3b30126 still match.
Phase 3 of #000037 (the actual sleep-sweep scheduler that runs the
Phase 1 controller on a cadence over real shards) gates on a
measured retrigger, not a calendar date. This ticket is the gate.
§2 commits 8 governance parameters that fold into governance_policy_hash
when Phase 3 lands: chunk_size (Hermes concurrency), per-mode τ_qa
seconds (CP/LLM split from Finding 3), weight profile (default
"sweep" from Phase 1.c), active sweep targets, per-window budget
cap, scheduling cadence, quarantined-row policy.
§3 names 4 retrigger gates: ≥1000 advisory rows from Phase 2 wiring
showing reproducible REJECT/DEFERRED structure (Retrigger 1); three
consecutive weekly dry-runs with sustained ACCEPT/MARGINAL on Target
A (Retrigger 2); 5F-fixture funnel demand from #000025 plateauing
on Target A's stream and needing Target B's larger candidate pool
(Retrigger 3); operator mission need (Retrigger 4, mirrors #000037
§12 Trigger 4).
§4 explicitly excludes implementation, schema migration,
governance_policy_hash bump, dry-run sweep_weights swap, and
Hermes-call planner — all deferred to the implementation ticket
that this ticket gates.
Includes TICKETS.md index row + Next ID bump 000045 → 000046.
Per Finding 3, a uniform τ_qa=7d filtered out every recent
CANONICAL_PROJECTION row (the π* graduations from #000027/#000030/
#000032 are all younger than 7d), so the sweep saw zero high-value
kernel-only work. Splitting τ_qa by audit_mode lets the cheap kernel
re-probe path (CP) run on a short cycle while the expensive LLM
re-witness path (STRICT/HYBRID/UNGROUNDED) keeps the long cycle.
bench/scripts/prometheus_sigma_sweep_dryrun.py: new build_tau_by_mode()
helper + per-mode CASE in iter_target_a_candidates; sweep_target_a now
takes the dict instead of a single seconds value. New CLI flag
--tau-qa-cp-days (default 1d); --tau-qa-days now scopes to LLM-witness
modes only (default 7d). Report renders the per-mode τ table in the
header and marks Findings 2 and 3 RESOLVED with their landing commits.
Makefile: PROMETHEUS_SWEEP_TAU_DAYS bumped to 7 (was 1, the prior
Finding-3 workaround); new PROMETHEUS_SWEEP_TAU_CP_DAYS=1 makevar.
bench/results/prometheus-sigma-sweep-dryrun-2026-05-10.md: regenerated
under the new defaults — 9 CP rows surface alongside 1,904 LLM-
witness candidates → 1,913 total Target A candidates → 1 ACCEPT,
2 MARGINAL, 205 REJECT, 271 DEFERRED chunks; 2 cache_drift vetoes
preserved end-to-end.
docs/tickets/ticket-000037 §22: Findings 2 + 3 marked RESOLVED with
landing-commit references; total dry-run cost line updated to the
new measurements (37.6 ms / 3,913 branches / 9.6 µs per branch).
Lock the AUTOCOUNT regression-test pattern as the design log
canonical record. Previously declined when surface was 1-metric
+ 29 tags; now mature enough (4 metrics + 58 tags + 1 same-day
drift-catch since landing) to formalize.
== Ticket content ==
10 sections covering:
1. Why this exists — the 4-drift-day baseline (6cbbf95 / 14bcb99 /
5c21e83 / 30a9488) that motivated mechanization. Five-step
walk through justifying each choice (Step 5 last).
2. Format — `<!--AUTOCOUNT:metric:path-->N<!--/AUTOCOUNT-->`.
3. Four supported metrics with examples + skip semantics:
`tests`, `fixture-rows`, `db-rows`, `db-where`.
4. Skip-on-absence — operator state (shards, qa.db) absence is a
logged skip, not a fail. Smoke verified 2026-05-10 with
HOME=/tmp/empty.
5. What NOT to tag — closed-ticket point-in-time snapshots,
aggregate floors ("2000+"), historical journey arcs.
6. Install discipline at write time + at refresh time.
7. Future metrics deferred (file-lines, gh-pr-comments-count,
module-loc, commit-hash-exists) with the "add a metric"
recipe.
8. Empirical baseline at landing (3 test functions, 58 active
tagged claims across 8 doc files, harness runtime 2-4s).
9. Scope boundaries — does NOT auto-rewrite, does NOT validate
prose quality, does NOT scan docstrings, does NOT lock
values, does NOT add deps.
10. References — every landing commit + sister doc.
Closed at landing (status quo since fc5ba50 2026-05-10 morning;
this ticket is retroactive design log per the convention "every
ticket flips to `closed · landed in commit <sha>` when the work
ships").
== Code-fence parser fix ==
Adding the ticket itself surfaced an oversight: my AUTOCOUNT
examples in §3.3 + §3.4 used literal tag pairs in ``` fenced
code blocks. The parser was reading them as live claims and
firing on the illustrative `db-rows:002.db:concept_relations`
claim (compared 1234 vs live 72576 — both meaningless because
it's an example).
Fix: `_strip_fenced_code_blocks` substitutes the body of every
triple-backtick block with newlines before regex scanning. Line
numbers stay aligned (newline-preserving substitution); tags
inside fences are skipped because their parent text no longer
matches the regex.
Both helper functions (`_iter_claims` and the well-formed-tags
test) walk through the stripped text, so the strip discipline
is consistent across all three test functions.
== TICKETS.md index ==
Added #000044 row marked closed with the 5-commit landing trail.
Bumped Next ID 000044 → 000045.
== Verification ==
$ pytest tests/test_doc_counts.py
3 passed in 2.80s
$ pytest tests/ -q
2337 passed, 37 skipped in 108.29s
Hygiene: fox's in-flight changes to arborist/qa/runner.py +
arborist/substrate/prometheus.py + tests/test_prometheus*.py
left untouched in working tree.
Three coordinated doc landings tying together fox's evening
#000037 → #000025 closed-loop work (commits f625cac through
ff1752c + 8999b55):
== #000037 §20 Status — stale prose refresh ==
§20 said "open · awaiting go/no-go" but file header + TICKETS.md
both say "in progress" with phases 0/1/1.b/2 all landed. Refresh
§20 to reflect actual state with per-phase commit anchors:
- Phase 0 (doc): landed; David review applied per §21
- Phase 1 (pure-function controller): f625cac
(arborist/substrate/prometheus.py + test_prometheus.py)
- Phase 1.b (gap-close): f9f5ae4 (§14 row 4 Hermes-saturation
guard, §13 step 11 falsification-fixture proposal, §15 entropy
+ memory gates weight-tunable, ESCALATE > QUARANTINE > REJECT
priority cascade)
- Phase 2 (advisory audit writes): a786d6d
(prometheus_audit.py + controller_events sibling table; does
NOT enter audit_events.event_hash preimage)
- §12 Trigger 2 fired 2026-05-10: divergence variance ratio
0.575 > 0.5 with N=37 — Phase 1 opening is now empirically
gate-satisfied per 8999b55
- Phase 3: deliberately NOT landed; replaced with read-only
dry-run simulator surfacing five design findings (see §22)
- Closed-loop signal: 40 corpus-derived 5F fixtures harvested
per ff1752c
== #000037 §17.1 — "78 atomic claim-pack records" → 92 ==
§17.1 said "#000031 Phase 2 has 78 atomic claim-pack records
that max out at ANCHOR-WARRANTED". #000031 closed at 92 records
(78 was an interim count during Phase 2). Refresh with the
journey (78 → 92) + AUTOCOUNT-tagged via db-where metric so
future drift fires immediately. Note the resolution context: all
92 now resolve via 74 citation-aliases + 13 term-aliases under
#000031 Phase 2.5 + B-1 + B-2. The unconscious sweep drains the
ANCHOR-WARRANTED → EVIDENCE-WARRANTED promotion backlog
(derivations.proof_blob rows still need computation even on
resolved chains).
== #000025 §11 Status — Phase 1f closure note ==
#000025 file header lists Phase 1a/1b.2/1c/1d/1e but the §11
Status section was frozen at "Open · awaiting go/no-go" — a
two-versions-old prose snapshot. Refresh with the per-phase
landing trail + add a Phase 1f section for the corpus-derived
falsification harvest that ff1752c shipped:
- Phase 1f closes the controller → 5F battery loop fox designed
in #000037 §3 ("Divergence → candidate falsification fixture")
- bench/scripts/harvest_falsification_proposals.py reads qa.db,
stratifies top-20-by-cache_key per audit_mode (HYBRID +
UNGROUNDED), writes the 41-line JSONL pack (1 _meta + 40
fixtures, AUTOCOUNT-tagged via fixture-rows)
- Every fixture row carries _harvest_meta with cache_key,
witness_divergence at harvest time, audit_mode_at_harvest,
harvest_threshold, source_ticket: "#000037 §13 step 11"
- 5F battery exercises them every test run;
test_5f_falsification_harvested_pack_runs_clean asserts
error_detection_rate == 1.0 by construction (every harvested
row IS a falsification)
- Self-amplifying — coverage grows with corpus, not with
hand-curation
Listed open items (§10.11 / 10.13 / 10.14) preserved verbatim
from file header so the index claim "still open" stays in sync.
== Cookbook: new "Adjacent: live-corpus → bench-fixture
harvest" section ==
New section in warrant-substrate-cookbook.md between "Re-running
the substrate build" and "References" documenting the harvest
pattern. Three discipline patterns reused from the textbook
substrate noted explicitly:
1. Attribution metadata on every derived artifact (same shape as
derivations.proof_blob carrying inclusion proofs back to
source chunks)
2. Determinism via sort-and-cap (same shape as citation-alias
cascade's "top 5 AND-join then top 3 OR-join" stratification)
3. Pin the metadata contract in tests (same shape as the
cookbook appendix's discipline pins — silent regression
becomes loud test failure)
Plus the reusable recipe for any controller emitting Proposal
records: define the dataclass, write a harvester filtering +
stratifying, pin metadata in tests, wire a make target.
References section gets four new entries pointing to #000037,
#000025, the harvest script, and the fixture pack.
== Drift caught + refreshed during this commit ==
While editing the cookbook, the AUTOCOUNT regression test
(from fc5ba50 / 6c6defb) caught two stale counts from fox's
in-flight prometheus work:
- test_prometheus.py: 36 → 42 (fox's uncommitted +6 for
Phase 1.c sweep weight profile)
- test_prometheus_audit.py: 14 → 17 (fox's uncommitted +3)
Refreshed both inline + in the test/code-density table
(prometheus row test LOC also bumped 804 → 955 to match wc -l).
This is the harness firing exactly as designed — fox's
uncommitted tests changed live state and my doc claims went
stale within minutes. The test message named the file + line
+ claimed-vs-live, refresh was a 60-second turnaround.
== Verification ==
$ pytest tests/test_doc_counts.py
3 passed in 2.92s
$ pytest tests/ -q
2337 passed, 37 skipped in 107.43s
Hygiene: only docs/ paths staged. fox's in-flight changes to
arborist/qa/runner.py + arborist/substrate/prometheus.py +
tests/test_prometheus.py + tests/test_prometheus_audit.py
remain in their working tree, untouched by this commit.
Re-running `make prometheus-trigger-probe` after today's controller
landings shows the divergence-variance trigger has crossed both
thresholds:
Trigger 2 — divergence variance
Sample count: 37 (N_min = 30 ✓)
Mean: 0.7568, σ: 0.435
σ/mean ratio: 0.5748 (> 0.5 threshold)
Absolute σ: 0.435 (> 0.1 threshold)
Same-day morning probe (commit baseline) had only 16 samples and
did not fire; the additional witness-sweep / dry-run / harvest
activity through the afternoon brought sample count above N_min.
Agreement-label distribution across all shards:
KERNEL-LLM-DIVERGED 22
KERNEL-LLM-AGREE 6
LLM-DIVERGED 6
STRICT-WITNESSED 3
Trigger 1 (branch density) and Trigger 3 (witness cost share) did
NOT fire. Per §12 a single trigger firing is sufficient for Phase 1
gating — and Phase 1 has already landed. This commit captures the
empirical evidence that Phase 1 was on the right side of the gate.
Trigger 1 remains structurally blocked on #000012 Phase 1c
(fork_score_branches sibling table); that's the natural next move
if anyone wants to surface multi-branch consensus signals.
Two coupled doc updates capturing today's session state:
1. #000036 status pin — math review in flight with dav1d
- Ticket status line: 'awaits fox math review' → 'pre-review
polish pass 8916bf3; math review in flight with dav1d
(forwarded 2026-05-10 — Tier 2 bundle)'
- TICKETS.md index row mirrors same change
- Future shifts can now see review is live, not blocked on fox.
2. #000006 rolling research log — 2026-05-10b amend
- Fourth qualitatively different experimental shape:
Prometheus-Σ dry-run simulator (joining random-word,
witness-sweep, warrant-chain)
- Captures the five scheduler-calibration findings (F1-F5)
from bench/scripts/prometheus_sigma_sweep_dryrun.py:
- F1: chunk_size = Hermes concurrency, not pool size
- F2: capital_cost must split by audit_mode (CP=0.05 vs
STRICT=1.0); flat-1.0 blocks every allocation
- F3: τ_qa must split by audit_mode (1d for kernel-only,
7d for LLM-witness); single-τ hides CP-rows
- F4: Target B headline = 4.40 percent of docs are
canonical-shape candidates (~152K across the corpus)
- F5: quarantined-row veto exercises end-to-end on
real-corpus data, no fixture-only mocking
- Updates the distinct-signal table to four rows
- Cross-references #000037 §22 for the full per-shard log
Phase 3 scheduler (when it ships) inherits F1-F5 as known-good
defaults — the dry-run is the calibration substrate the eventual
implementation will reference for choice justification.
Doc-only updates; no schema, no governance hash, no code change.
Phases 1 (controller) and 2 (sibling-table audit writes) landed in
prior commits. This commit adds the Phase 3 dry-run simulator
instead of the actual sleep-sweep scheduler, since Phase 3's value
is mostly in what we'd learn from running it — and the dry-run
captures those findings without committing to a scheduler design
prematurely.
bench/scripts/prometheus_sigma_sweep_dryrun.py — read-only
simulator that classifies §3 Target A (providence_cache) + Target B
(documents) sweep candidates, synthesizes ControllerBranches from
real shard data, runs the Phase 1 controller, reports decision
distribution + Phase-3-design findings. No LLM calls, no
mutations.
make prometheus-sweep-dryrun — produces a dated markdown report
at bench/results/prometheus-sigma-sweep-dryrun-YYYY-MM-DD.md.
Five findings surfaced by three dry-run iterations against the
live ~/.arborist/shards corpus (3.5M docs + 2839 providence_cache
rows) — captured in ticket §22:
1. chunk-size dominates Kelly threshold (must = Hermes
concurrency, not candidate pool)
2. flat capital_cost blocks every allocation (split kernel-cost
vs LLM-cost on the contract)
3. τ_qa=7d filters every CANONICAL_PROJECTION row (all 29 are
<7d old; need per-audit-mode τ)
4. Target B canonical-shape detection is the real headline
(~152K candidates extrapolated; controller correctly returns
MARGINAL on shape-match chunks)
5. quarantined rows correctly veto via cache_drift hard-veto
Mean per-branch controller latency in dry-run: 12.5 µs at
chunk_size=4. Phase 3's actual bottleneck is the witness fan-out
(Hermes calls), not the controller itself.
Ticket #000037 status flipped to in-progress with Phases 0+1+2
landed; Phase 3 scheduler remains future work but is informed by
the five findings.
Fan-out follow-up to ``fc5ba50``. Two thrusts in one commit since
they exercise the same surface:
== Task 3: extend AUTOCOUNT with db-rows metric ==
New metric ``db-rows`` for tagging live SQLite row counts (alias
tables, claim-pack records, etc — operator state that drifted on
``30a9488`` and earlier). Target syntax::
<!--AUTOCOUNT:db-rows:citation_aliases-->74<!--/AUTOCOUNT-->
<!--AUTOCOUNT:db-rows:002.db:concept_relations-->1234<!--/AUTOCOUNT-->
Default shard: ``~/.arborist/shards/000.db`` (where the alias
tables live per ``arborist.cli._aliases_db_path``). Operator state
is graceful-skip semantics: when DB or table is absent (CI, fresh
checkout, sibling repo), the claim is logged as skipped and the
test still passes. Drift only fires when the DB IS present and
the count diverged.
Sentinel returns:
- ``_DB_MISSING`` (-2): shards dir not present → skip
- ``_TABLE_MISSING`` (-3): DB present but table absent → skip
- ``_DB_ERROR`` (-4): malformed table name or sqlite error → skip
Table name validated against ``[A-Za-z_][A-Za-z0-9_]*`` regex
before string-interpolating into ``SELECT COUNT(*) FROM <table>``;
this is belt-and-suspenders since AUTOCOUNT tags are author-
controlled, but the dynamic SQL surface deserves a bouncer.
Smoke verified under HOME redirect to ``/tmp/<empty>``: 3 db-rows
claims gracefully skip with informative line-numbered messages,
suite still passes.
== Task 2: backfill 15 tags ==
Cookbook test/code-density table (lines 569-579, 10 rows) — every
``(N tests)`` cell now machine-checked:
| aliases.py | 512 | 469 (28 tests) | 0.92 |
→
| aliases.py | 512 | 469 (<!--AUTOCOUNT:tests:tests/test_aliases.py-->28<!--/AUTOCOUNT--> tests) | 0.92 |
Markdown renderers strip HTML comments — table cells display
``28 tests`` unchanged. The ``warrant_resolver.py`` row stays
untagged because its test count is split across two test files
(verifier + parser) and the cell encodes a combined "~430"
instead of one collected count.
Cookbook alias-count surfaces (3 db-rows tags):
- L364 ``citation_aliases (74 rows live as of 2026-05-10)``
- L437 ``#000041 — citation-aliases table + 74 live rows``
- L438 ``#000042 — term-aliases table + 13 live rows``
Ticket #000035 (in progress, line 274) — refresh ``20 tests``
→ ``27 tests`` for ``test_anchor_prg.py`` + tag. Same drift
pattern as ``5c21e83``: ticket prose was written before the
``de997f7`` 2026-05-10 pattern backfill that added 7 tests
(prefix-extension closure, hand-formula, parametrized
invalid-input cones). Also tagged ``L279``'s 10-vector KAT
fixture claim with ``fixture-rows`` metric.
== Closed-ticket counts deliberately not tagged ==
#000028, #000030, #000042, #000031, #000004, #000026, #000009,
#000032, #000008 all carry historical "N tests pass" snapshots
from their landing date. Those are point-in-time records, not
live claims — drifting from current state is BY DESIGN. Tagging
them would fire the test on every successive change to the
codebase. Closed tickets are the design log; we don't backfill
them.
== Coverage summary ==
Total tags after this commit: 44 (was 29; +15)
Tags by metric:
tests: 39
fixture-rows: 2
db-rows: 3
Files with tags:
docs/warrant-substrate-cookbook.md 27 (was 14)
docs/soft-hash-channel-analysis.md 5
docs/tickets/ticket-000006-bench-emergent... 4
docs/seven-point-program.md 3
docs/calculator-test-patterns.md 3
docs/tickets/ticket-000035-prg-choice-phi-prg.md 2 (new)
== Verification ==
$ .venv/bin/pytest tests/test_doc_counts.py -v
3 passed in 4.32s
$ .venv/bin/pytest -q
2276 passed, 54 skipped in 168.34s
$ HOME=/tmp/empty pytest tests/test_doc_counts.py -v -s
3 db-rows AUTOCOUNT claim(s) skipped:
docs/warrant-substrate-cookbook.md:364 db-rows:citation_aliases skipped — /tmp/empty/.arborist/shards not present (CI / fresh checkout)
docs/warrant-substrate-cookbook.md:437 db-rows:citation_aliases skipped — /tmp/empty/.arborist/shards not present (CI / fresh checkout)
docs/warrant-substrate-cookbook.md:438 db-rows:term_aliases skipped — /tmp/empty/.arborist/shards not present (CI / fresh checkout)
3 passed in 4.78s
No new dependencies. No schema changes.
The doc-drift pattern recurred four times today on 2026-05-10
(commits 6cbbf95, 14bcb99, 5c21e83, 30a9488). Each fix was the
same shape: walk a doc, find a count that drifted from live truth
during the hours after the doc was written, refresh it. Cost: ~5
min per drift × 4 = 20 min of manual catching, with no guarantee
the next drift gets caught before someone external reads it.
Per fox's selection: regression test that makes drift loud at
test time instead of relying on visual catching.
== Mechanism ==
`tests/test_doc_counts.py` scans `docs/**/*.md` for AUTOCOUNT
tags of the form:
<!--AUTOCOUNT:metric:path-->N<!--/AUTOCOUNT-->
Two metrics supported:
- `tests` — pytest collected count for path. Batches every
tagged path into one `pytest --collect-only` subprocess
(~0.5s total).
- `fixture-rows` — non-blank-non-comment line count in a JSONL
fixture.
GitHub and most markdown renderers strip HTML comments, so
readers see only `N`. The tags are invisible in rendered output
but make the claim machine-checkable. Three tests in the file:
1. `test_doc_autocount_claims_match_live` — the core invariant
2. `test_autocount_tags_are_well_formed` — open/close balance
3. `test_autocount_metric_names_are_documented` — fail-closed on
undocumented metrics (catches typos)
Failure message names the doc file, line number, and the
claimed-vs-live diff. Example:
`docs/foo.md:42 AUTOCOUNT(tests:tests/test_x.py) claims 23, live is 27`
== 29 tags installed across 5 docs ==
While installing tags I had to read the surrounding prose, which
surfaced six stale counts that had drifted same-day:
`docs/soft-hash-channel-analysis.md`:
- L392 14 → 23 tests for phi_alignment_probe
- L417 20 → 27 tests for anchor_prg
- L463 14 → 23 tests for phi_alignment_probe (status section)
`docs/seven-point-program.md`:
- L77 68 → 58 tests for metacognition (drift -10; the file
shed tests during a refactor and the doc didn't catch up)
- L78 9 tests for `test_dag.py::test_preflight_*` — removed
count entirely; pytest selector subsets aren't currently
supported by the AUTOCOUNT metric set (would need a
`tests-matching` metric; not worth the surface for one claim).
- L110 24 → 33 tests for test_dag.py
`docs/calculator-test-patterns.md`:
- L35 33 → 23 tests for warrant_resolver
- L35 10 → 9 tests for warrant_chain
- L16, L265 51 → 53 tests for t3_bound_calculator (kept
initial-shipment provenance in prose)
== Coverage installed ==
calculator-test-patterns.md 3 tagged claims
soft-hash-channel-analysis.md 5 tagged claims
warrant-substrate-cookbook.md 14 tagged claims
seven-point-program.md 3 tagged claims
tickets/ticket-000006-bench-... 4 tagged claims
---
29 tagged claims
Every count that drifted today is now tagged. Future drift
fires the regression test at the next pytest run instead of
waiting for human catching.
== Discipline pattern ==
Walk this pattern for any new doc that names a count:
1. Surround the number with the tag pair:
`<!--AUTOCOUNT:tests:tests/test_foo.py-->N<!--/AUTOCOUNT-->`
2. Run `pytest tests/test_doc_counts.py` (~3.5s)
3. If it passes, the claim is now machine-verified
Aim to tag counts on first authorship. Retrofitting is cheap
but only catches drift after the fact.
== Out of scope ==
Test counts inside source code (docstrings, CLI --help) are not
scanned — would expand the test surface significantly and the
drift pattern hasn't manifested there. Add `**/*.py` scope when
that pattern surfaces.
Alias-row counts and claim-pack-record counts could be tagged
with new `db-rows:<table>` and `db-where:<sql>` metrics; deferred
until the next drift on those numbers (none caught today after
30a9488's cookbook refresh).
== Verification ==
$ .venv/bin/pytest tests/test_doc_counts.py -v
3 passed in 3.89s
$ .venv/bin/pytest -q
2276 passed, 54 skipped in 153.21s
No new dependencies. No schema changes. No source-code changes.
`docs/_source/merkle-agi-v8-consensus.rst` (834 lines, RST sister
to the v7-W substrate paper at the same path). Closes Phase 1b of
ticket #000012 — the loop-closing consensus protocol that turns
single-validator Proof-of-Upgrade into Darwinian selection across
an open validator set.
11 parts:
Part 1 Introduction & motivation — gap table from v7 § 13.4,
concrete backdoor-attack scenario, paper IS/IS-NOT
scope.
Part 2 Substrate definition — SQD A1/A2/A3 inheritance,
consensus_events row schema, consensus_policy_hash
sibling (never enters cache_key).
Part 3 Validator state machine — bonding/active/challenged/
slashed/unbonding with full transition graph + invariants.
Part 4 Acceptance protocol — proposer submission, layered
fitness floor (canonical + lab-declared ceiling),
audit-replay procedure, 2/3-stake quorum + GRANDPA-
style finalization, liveness floor.
Part 5 Challenge protocol — counter-evidence shape,
adjudication, challenger reward, frivolous-challenge
bond.
Part 6 Stake mechanics — bond/unbond/challenge window
recommendations, offense-class slashing schedule,
reward distribution, optional stake cap + sqrt-weighting.
Part 7 Fork choice rule — GRANDPA-style finality, pre-finality
constraints, liveness recovery.
Part 8 Mesh wire format extension — three new message kinds,
BLS-or-concat aggregate signatures, bandwidth profile.
Part 9 BFT analysis — safety, liveness, Sybil resistance,
bootstrap honesty, re-staking attacks.
Part 10 Worked example — 7-validator deployment, one upgrade
cycle with successful challenge against one fraudulent
validator.
Part 11 Out of scope — implementation, calibration, cross-chain
anchoring, fixture selection, bootstrap-set membership,
cross-instance slashing accumulator, branch-set
persistence.
Closure §: open questions tracked separately (initial validator
set composition, threshold-key ceremony, ZK-replay, policy-hash
transition mechanics).
Ticket #000012 status updated; Phase 1c (branch-set persistence)
remains proposed-not-opened. Implementation follow-up tickets that
cite this paper land later — one per validator-state-machine,
mesh-wire-format extension, audit-replay harness, slashing
accountant.