arborist/CLAUDE.md
russell@unturf.com 0ddff43142
docs: capture Wikipedia 2010 ingest origin (the 100-min, 4-way parallel run)
The original ingest of Wikipedia 2010 into fox's 4 production shards
was 2026-04-27 21:33-23:13 UTC — ~100 minutes wall, 4-way parallel,
3.47M docs / 14.12M chunks. Captured here because the audit chain is
the only durable record but querying 3.47M rows to recover the
headline number is friction; one line in a doc removes it.

Also informs #000065 reshard planning: ingest rate ceiling on real
XML workload is ~2,350 chunks/sec aggregate (4-way), vs the
~6,400 chunks/sec the M-sweep bench measured on the 2003 cur dump
(which skips XML parsing). The teleport-style reshard should beat
both ceilings because it's just SQLite INSERT throughput, no XML
parse + canonicalize + edge extraction.

New file: docs/corpus-history.md. Indexed in CLAUDE.md docs section.
Append-only convention; future migrations + cold-pack runs add
entries here so the operator log isn't only in the audit chain.

Derivation query (sqlite3 audit_events) embedded in the entry so
future re-derivation is one copy-paste.
2026-05-26 12:48:12 -04:00

682 lines
38 KiB
Markdown
Raw Permalink Blame History

This file contains ambiguous Unicode characters

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

# Agent Blackops — arborist
This repo is operated by **agent blackops** for fox/timehexon on the
unsandbox / unturf / permacomputer platform.
Identity shard: `~/git/unsandbox.com/blackops/BLACKOPS.md`.
## What arborist is
A content-addressed, Merkle-committed document store. Implements the
runtime spec from **Merkle Providence Reverse RAG** (April 2026
whitepaper) scaled up to the **Merkle-AGI v9.8** admissibility ledger.
Three layers stacked on one SQLite file:
1. **Surface** — ingested documents (Wikipedia dumps, HTML pages,
anything with a URI). Chunked, Merkle-rooted, FTS5-indexed.
2. **Core** — distilled documents Merkle-bound back to surfaces via
per-chunk inclusion proofs in `derivations.proof_blob`. Recursive.
3. **Providence cache** — Q&A records keyed on the v9.8 8-dim
invariant. Each record carries `audit_mode` (STRICT / HYBRID /
UNGROUNDED) decided by the verifier in `arborist/qa/verify.py`.
For claim-lattice modes the renderer maps that token to a four-
rung ladder (POINTER-LINKED → ANCHOR-WARRANTED → EVIDENCE-WARRANTED;
ENTAILMENT-VERIFIED reserved); UNGROUNDED below all rungs. The
schema column stays unchanged — programmatic callers see the
trichotomy, human-facing surfaces see the ladder. See
`arborist/cli.py:_render_audit_label`.
## Source papers
- `~/git/unfirehose-nextjs-logger/whitepaper/merkle-providence-reverse-rag-whitepaper.rst`
— canonical whitepaper (rst → PDF). Edit here, not the PDF.
- `~/Downloads/merkle-agi-dag_v7.txt` — formal substrate (TLV
encoding A1, public quantization A2, collision-resistant hash A3,
theorems T1T5).
- `~/git/proxy.unturf.com/pkg/verified/merkle.go` — fox's existing
Go merkle. Arborist Python ports its conventions exactly.
- `~/git/proxy.unturf.com/docs/merkle-tree.md` — convention reference.
## Architecture
```
arborist/
├── merkle.py # Python port of proxy.unturf.com Go conventions
├── store.py # v9.8 SQLite schema, audit-chain helpers
├── document.py # Document, Edge, Chunker (TokenChunker default)
├── source.py # Source ABC: iter_documents()
├── ingest.py # batched normalize → chunk → merkle → upsert
├── evict.py # hot↔cold + rehydrate (v9.8 falsification on drift)
├── compress.py # chunk pack/unpack (zstd dictionary trained on corpus)
├── snapshot.py # snapshot.db creation + load
├── journal.py # NDJSON unfirehose journal sink
├── wikitext.py # to_base(): wikitext → plain prose (BASE_VERSION-pinned)
├── search/ # SearchBackend ABC + AuditMode + FTS5
├── sources/ # one file per corpus (wikipedia, html_page,
│ # claim_pack, textbook_tex, grok, vcs, …)
├── distill/ # surface → core distillation (tfidf, first_sentence)
├── qa/ # Q&A: 8-dim cache_key + Merkle-bound answers
│ ├── client.py # ChatClient + StubClient + OpenAICompat
│ ├── keys.py # cache_key, question_hash (pure functions)
│ ├── verify.py # layered verifier + claim-lattice verifiers (G0)
│ ├── evidence.py # EvidenceObject + spotlight excerpt
│ ├── inspect.py # read-only sidecars (deflection, title-relevance)
│ ├── dag.py # per-run Merkle-DAG (7/8 quote · 9/10 CTI · 3 reject)
│ ├── quantifier.py # broad-quantifier classifier (#000008 P1)
│ ├── model_profiles.py # per-model claim-cap profiles (#000008 P2)
│ ├── quantifier_reminder.py # broad-query reminder text (#000008 P3)
│ ├── canonical_cache.py # canonical-projection persistence (#000027)
│ ├── witness.py # multi-witness fan-out (#000028)
│ ├── warrant_resolver.py # claim-pack warrant chain resolution (#000031)
│ ├── crosslang.py # cross-lang guard: signal + es stoppack (#000001 §7 P0)
│ ├── mt/ # Operation Sandwich MT edges (#000056)
│ │ # opus-mt es/fr/ru↔en + entity_mask
│ └── runner.py # ask(): cache → infer → verify → write
├── concepts/ # corpus-derived synonym + rivalry layer (#000018 sib.)
├── pi_star/ # canonical projection π* registry (#000015)
│ # arithmetic@v1, logic-kernel@v1, algebra-symbolic@v1, …
├── memory/ # MemoryRoot lifelong-learning audit chain (#000017)
├── selfmodel/ # SelfModel snapshot + falsify (#000014)
├── capital/ # 8-form capital ledger (#000020)
├── substrate/ # Merkle-AGI substrate primitives (paper-spec'd)
│ ├── anchor_prg.py # φ_PRG HMAC-SHA-512 (v7 §9.10; #000035)
│ ├── fork_score.py # ScoredFork decision fn (v8; #000012 P1a)
│ └── weights.py # ForkScore weight set
├── world/ # v7-W spatial-temporal substrate reservation (#000013;
│ # namespace stub; future kernels under world/pi_star/,
│ # world/frontier/, world/adapters/)
├── mesh/ # mesh wire format + group-key state machine
└── cli.py # ingest / search / verify / stats / distill /
# evict / rehydrate / ask / providence / emergent /
# reclassify / inspect / analyze / canon / sweep /
# warrant-resolve / alias / capital / selfmodel
```
**Dir naming convention.** Topic-named, never version-prefixed. The
substrate-paper version (v7 plastic-training, v8 selection/consensus,
v9 falsification controller, …) and the live SQLite schema version
(v9.8) are two unrelated numbering schemes that share decimals;
version-prefixed dirs (``arborist/v7/``, ``arborist/v8/``) were tried
2026-05-10 and retired the same day because readers asked "is this
schema-v7 or paper-v7?". Substrate-paper-spec'd primitives now live
under ``arborist/substrate/``; topic dirs (``capital/``, ``memory/``,
``selfmodel/``, ``concepts/``, ``pi_star/``) hold cross-version
mechanisms.
## Build, test, run
Every workflow is a `make` target. Bare python is not the user
interface. See the Makefile for the full list.
```
make bootstrap # venv + editable install with [dev] extras
make test # pytest -q
make all # bootstrap + fetch-cur + ingest-cur + verify + stats
make verify-shards # round-trip Merkle proofs (cross-shard sample)
make analyze-shards # cross-shard compression + audit integrity
make chain-check-shards # audit-chain break count per shard (0 = intact)
make query Q="..." [JSON=1 BURN=1 K="extra retrieval keywords" ANSWER_MODE=…]
make bench-qa # QA-quality sweep (live LLM)
# Textbook surface-ingest layer (#000031 — PD/open-licensed
# math/logic/CS textbooks for warrant promotion):
make textbooks-summary # license + URL counts per manifest entry
make crawl-textbooks # BFS-crawl every entry with crawl_url
# → ~/.arborist/crawl/textbook_<id>.db
make textbooks-tex # PG LaTeX-source ingest (Hilbert, Boole)
make textbook ID=<id> # ingest one textbook by manifest id
# (idempotent at DB layer)
make textbooks-base-knowledge # bulk: Cantor + De Morgan + Russell
# IMP + Judson (the four 2026-05-09
# base-knowledge additions)
# Per-textbook convenience targets (one per active manifest id):
make textbook-bogart textbook-keller-trotter textbook-levin
make textbook-aristotle-prior textbook-aristotle-posterior
make textbook-newton textbook-morin
make textbook-judson textbook-cantor textbook-demorgan
make textbook-russell-imp textbook-russell-pom
make textbook-laplace textbook-pm textbook-grinstead-snell
make textbook-hilbert textbook-boole textbook-peano textbook-dedekind
make textbook-plfa textbook-sf-lf
# Claim-pack warrant-chain resolver (#000031 Phase 2 + 2.5):
make sweep TARGET=warrants # warrant-resolve --use-aliases --write
arborist warrant-resolve --use-aliases --write
# → 92 / 92 (100%) coverage as of 2026-05-10:
# 18 textbook substrates + 74 fox-decided
# citation-aliases + 13 term-aliases.
# (alias counts grow as fox adds substitutions
# — `arborist alias citation list | jq length`
# for live count). Per-pillar 13/13 · 10/10 ·
# 13/13 · 18/18 · 5/5 · 5/5 · 14/14 · 14/14.
arborist alias citation list # see substitutions
arborist alias term list # see vocabulary aliases
```
**Hygiene after any state-changing op** (rebuild, reclassify, hash
bump, mass falsify): `make chain-check-shards` first (every shard
should report `0`), then `make analyze-shards` for the spectrum +
chain audit. Chain breaks are the loudest possible signal.
## Schema invariants (do not break)
- **v9.8 8-dim cache_key** (+ optional legible 9th): `source_root |
question_hash | model_profile_hash | conversation_hash |
governance_policy_hash | schema_version | canonicalization_version |
chunking_version`. Bumping any one invalidates prior records on
lookup. `cache_key()` also accepts an **optional 9th dimension**,
`verifier_policy_hash` (`keys.py`), for *audit legibility* — it does
NOT add correctness coverage, because the verifier-policy fields are
a subset of the policy dict and so already fold into
`governance_policy_hash` (a verifier-rule change already changes the
cache_key today; the 9th just makes "did the *verifier* rules
change?" answerable from one hash diff). 8-dim is the default write
form; the 9-dim form is opt-in. Mandatory-vs-legible is #000058.
- **`falsification_state ∈ {live, failed, stale, quarantined}`**.
Cache lookups filter on `state='live'`. Drift → `stale`.
- **Audit chain**: every state-changing op writes one row in
`audit_events` with `event_hash = sha256(prev || canonical(body))`.
Verified by `make chain-check-shards`. Use `arborist.store.append_audit`
— never insert into `audit_events` directly.
- **Cores never evict.** `evict_to_cold` only touches `kind='surface'`.
- **Idempotent re-ingest.** Same content → same `document_root` →
no-op insert. Same URI + different content → new doc + `supersedes`
edge (lossless history).
## Conventions (do not silently change)
Each rule below has full rationale in the named source file. Don't
revert without reading why. When in doubt, walk the
[Five-step algorithm](#five-step-algorithm) first.
- **Merkle conventions**: non-commutative `HashCombine` prefix `0x03`,
leaves `0x00`, odd-element rule = self-duplicate (NOT zero-pad).
`MerkleProof.siblings` carries `is_left` flag — never sort lexically.
See `arborist/merkle.py`.
- **Versioned defaults**: `tok-512-v1` (chunker), `norm-v1`
(canonicalization), `v9.8.0` (schema), `wikitext-base-v1` (prose).
Changing any default stales every prior cache record. Add a new
`name` instead.
- **`question_hash` is dedup-mode-aware** (`strict` |
`equivalence_class`); folds into `governance_policy_hash`. JIT
`fidelity` parameter on `query()`/`ask()` decouples lookup tolerance
from write policy. See `arborist/qa/keys.py`.
- **`audit_mode` is decided by the verifier, never asserted.** Four
layered strategies tried in order, first to find evidence
classifies: **quote** (sequential pair-matching, NOT regex —
prevents phantom inter-pair captures), **span** (verbatim line
match), **entity** (proximity-clustered proper nouns;
`entity_policy ∈ {strict, hybrid, drop, proximity}`), **paraphrase**
(token-coverage, prose-shaped only; `verifier_method='paraphrase'`).
Trichotomy: STRICT = every unit verifies, HYBRID = mixed,
UNGROUNDED = none. Never overclaim. See `arborist/qa/verify.py`.
- **Verifier stays binary; falsifications carry soft signal.** No
per-quote diagnosis fields on hard verifier output. Sidecars
(`arborist.qa.inspect.diagnose_*`, `arborist inspect --cache-key X`)
classify unverified spans, deflection, title-relevance — never
write to `providence_cache` or `audit_events`.
- **Trailing-citation strip**: `_strip_trailing_citation` peels one
trailing parenthetical at end-of-span (gated on a citation cue or
URL) before substring testing. See `arborist/qa/verify.py`.
- **Soft hash vs hard hash**: hard = SHA-256 (commitments, proofs,
cache_key); soft = embeddings/TF-IDF/similarity (training, ranking,
distillation). Soft never enters proof path.
- **Cross-language = the sandwich, MT on the edges only**: translate
query in (retrieval + LLM prompt) → English answer → the
**byte-identical verifier** grounds English-vs-English → translate
the verified answer out as **display-only** (`display_*`, banner-
labelled, zero grounding — the `_render_audit_label` projection
discipline). Translation NEVER re-enters the verifier (that's the
#000049 model-in-proof-path cage). Invariants: `question_hash` =
the user's original question (untranslated); `verifier_policy_hash`
unchanged; MT engine identity binds into `RetrievalPlan.mt_*` (run-
DAG), NOT `governance_policy_hash` (the *flag* moves it like any
policy flag — correct cache partition — but that hash covers the
whole policy, `keys.py:182`; don't mistake "no new governance
field" for "governance untouched"). Engine = local `[mt]` opus-mt,
hash-pinned, never Hermes-for-translation, never an API. Default
OFF (`crosslang_guard_enabled` P0; `crosslang_translate_enabled`
Sandwich; `crosslang_entity_mask` **default-OFF** — measured
net-negative, kept only behind the flag). Measured net win over
the real "nothing" baseline (raw es → noise/UNGROUNDED): es ≈0 % →
71 % grounded; the 14 pp vs *English* is the cost of a new
capability, not a regression (compare to no-cross-lang, never to
native English). The recall lever (entity-preservation) is **still
open** — `entity_mask` v1 failed at bench scale; corpus-title
anchoring is the untried idea. See `arborist/qa/crosslang.py`,
`arborist/qa/mt/`, #000001 §7, #000056 §9.
- **Three answer modes**: `policy["answer_mode"] ∈ {"quote",
"claim_lattice_pointer", "claim_lattice"}`, default `"quote"`.
Bench 2026-05-02T15:07Z on Hermes-3-8B (post-Sprint-1b/2, n=3
× 71 questions, sample-shuffled @ c=4): quote 0.54 strict-rate,
pointer 0.20, JSON 0.42. Quote leads on raw lexical grounding;
JSON leads among lattice modes. Per-mode peak buckets: quote
8-16KB (0.58), pointer 16-32KB (0.20), JSON 32-64KB (0.48) —
these drive `max_context_chars_by_mode`. 99% directive coverage
(D2/D3/D4/D6/D7) on lattice modes. Both lattice modes share
`verifier_method="claim_lattice"`; `answer_mode` on the run-DAG
+ `json_fixups` disambiguate. Each mode folds into
`governance_policy_hash`. See `arborist/qa/verify.py`,
`docs/qa-modes-bench.md`.
- **Four-rung ladder (lattice-mode display layer)**: POINTER-LINKED
→ ANCHOR-WARRANTED → EVIDENCE-WARRANTED → (ENTAILMENT-VERIFIED
reserved); UNGROUNDED below. Each rung names a strictly stronger
property the lexical verifier could confirm. WARRANT_MISSING drops
to POINTER-LINKED; soft-demote violations (LAZY_ANCHOR_DEMOTED,
POINTER_OVERFLOW_TRIMMED, TOO_MANY_CLAIMS, BARE_NAME_CLAIM,
TITLE_MISMATCH) cap at ANCHOR-WARRANTED. Quote / span / entity /
paraphrase modes keep their original audit_mode tokens (those
verify against pinned spans, not synthesis). Pure render-layer —
cache_key, governance_policy_hash, & all programmatic callers see
the underlying audit_mode unchanged. See
`arborist/cli.py:_render_audit_label`.
- **Claim-lattice-pointer mode (G0 / CTI)**: runtime mints
`pointer_id` (E1, E2, … — what the model sees) and content-addressed
`evidence_id` (what the cache & run-DAG store). Renderer interpolates
literal source spans via `_spotlight_excerpt`. Synthetic-elision-by-
construction-impossible — model never types the quote string. 9-stage
run-DAG. See `arborist/qa/evidence.py`, `docs/cti-architecture.md`.
- **Claim-count ceiling (`TOO_MANY_CLAIMS`)**: default 12 per answer.
Catches "tell me all there is to know about X" runaway. Demotes
STRICT → HYBRID without truncating. Folds into
`governance_policy_hash`. See `arborist/qa/verify.py`.
- **Broad-quantifier preflight guard (Ticket #000008)**: pure lexical
classifier (`arborist/qa/quantifier.py`) maps a question onto a
10-rung intensity ladder (ABSENT → SINGULAR → … → ALL →
COMPREHENSIVE). Per-model `arborist/qa/model_profiles.py` PROFILES
dict picks a per-call claim-cap from the (intensity, model) pair;
`arborist/qa/quantifier_reminder.py` synthesizes a one-line user-turn
reminder for broad questions. Six-level disable hierarchy
(per-test, per-call CLI, per-phase policy, per-mode, per-model,
master-via-governance-hash). 7 policy fields fold into
`governance_policy_hash` so flipping any of them invalidates prior
records. Defaults preserve dry-run discipline:
`quantifier_guard_apply_caps=False`,
`quantifier_reminder_enabled=False`,
`quantifier_reject_broad=False`. CLI flags on `arborist query`:
`--no-quantifier-guard`, `--allow-broad`, `--reject-broad`,
`--apply-quantifier-caps`. Bench A/B (2026-05-03, n=3 × 9 broad
questions × 3 modes): reminder default-on supported (FORMAT_COLLAPSED
100%, NO_EVIDENCE_POINTER 33%, JSON UNGROUNDED 22pp); cap default-on
for JSON only (+14pp STRICT-rate, no gain on pointer); cap+reminder
best on pointer mean ratio (0.684) but not strictly best on JSON
STRICT-rate. See `docs/tickets/ticket-000008-broad-quantifier-preflight-guard.md`
§12 for the four-cell A/B data.
- **Wikitext base prose**: `arborist/wikitext.py:to_base()` runs
before the LLM call AND inside `verify_quotes` so model and verifier
see the same prose. Optional dep — graceful fallback when
`mwparserfromhell` is missing.
- **Deflection sidecar**: `diagnose_deflection(question, answer)`
detects topic-shift via subject-anchor heuristic (LAST content
token in question must appear in answer). Suppressed for date /
count / cause shapes ("when", "why", "how many"). See
`arborist/qa/inspect.py`.
- **Coherence sidecar** (#000052 §3.1): `diagnose_coherence(answer)`
in `arborist/qa/inspect.py` — per-sentence lexical shape check,
**no model**. Emits `kind ∈ {phrase_component_reuse, circular,
vacuous, ok, empty}`: `circular` = subject content-tokens ⊆
predicate's and the predicate leads with a subject token ("X is
X"); `phrase_component_reuse` = the subject quotes a phrase & the
predicate reuses one of that phrase's own tokens as a bare
`the/a/an <token>` referent (the 2026-05-12 "the phrase 'Zionist
entity' is used as the entity" field case — a token collision the
verifier + deflection + title-relevance all pass and NLI returns
*neutral* on); `vacuous` = predicate is only placeholder hypernyms
+ filler ("X is a thing"). Surfaced in `inspect_cache_key` + the
`arborist inspect` human view (`· incoherent: <kind>`). Advisory
sidecar — **never** writes `providence_cache` / `audit_events` /
`run_dag_root`; a demote-only verifier hook is possible but
deliberately not wired (would fold into `governance_policy_hash`).
- **Title-relevance hard check (Rule 8)**: `_claim_title_overlap`
in `arborist/qa/verify.py`. For each claim that resolved, at least
one cited evidence's source title must share ≥1 stemmed content
token with the claim text. When NO cited title overlaps, record
a `TITLE_MISMATCH` violation & demote `STRICT → HYBRID`. Catches
retrieval-driven hallucinations where the cited chunk's SOURCE
is structurally unrelated to the claim's subject (2026-05-02
spin-glass case: claim about spin glass cited to a chunk from
*Quantum chromodynamics*; span had incidental physics-vocab
overlap, but the source title shared zero stems with the claim).
Renderer surfaces a `· title mismatch` tail on the audit-line
label alongside `· warrant missing`.
- **Title-relevance sidecar (legacy diagnostic)**:
`diagnose_title_relevance(claim, cited_titles)` in
`arborist.qa.inspect` returns the same signal in dict form for
per-cache-key inspection. Sidecar; never enters proof path.
Pre-dates the Rule 8 promotion (2026-05-02).
## Live endpoints
- LLM: `https://hermes.ai.unturf.com/v1` (Hermes-3 Llama-3.1-8B-FP8-
Dynamic on vLLM, 82K ctx, no auth). `uncloseai.com` is marketing
only. Override via `--endpoint` or `ARBORIST_LLM_ENDPOINT`.
- Wikipedia dumps: two corpora available — the 2003-05-16 cur snapshot
(`https://dumps.wikimedia.org/archive/2003/2003-05-16/en/`) and a
**2010** snapshot. **The 2010 corpus is what is currently loaded** in
`~/.arborist/shards` (866K docs/shard; contains post-2003 articles like
Barack Obama / YouTube — the discriminator). Fabrication-bait fixtures
(`qa_questions_stale_map`) target post-**2010** events.
`robots.txt` returned 404 → no rules.
### Budget discipline — Hermes / Qwen first, Opus deferred
2026-05-19: the huge-N `#000057` control-arm sweep
(`f63b00d` → `9dc02e4`) burned our Opus quota. Until fox explicitly
re-adds it, sweep + `bench-qa` LLM calls route to **Hermes-3-8B**
(`hermes.ai.unturf.com/v1`) and **Qwen via uncloseai.com**; Opus drops
out of control-arm grids. This is a budget rule, not a quality claim —
bench results already on disk under `bench/results/` stay valid; just
don't expand them with new Opus tokens without a go. Translation at
scale is still NOT a Hermes/Qwen 8B job (sandwich MT uses local pinned
`opus-mt`, see crosslang rule). Grok candidacy noted from ajax synthetic-
data-distillation benchmarks.
<!-- Grok API rejoins the grid once fox approves credit-card spend.
When unflagged, add a `grok-*` arm to `bench/control_sweep.py`
and document model id + endpoint here. -->
### Per-call model selection (general agent work, not sweep arms)
Within a 4-hour Max-plan window, haiku/sonnet/opus draw equally from
quota — prefer the largest model that fits (fewer retries = better
quota efficiency). During **paid overflow** (beyond Max), pay-per-token
applies: start haiku, escalate sonnet then opus only on failure.
Hermes/Qwen via `uncloseai.com` serve classification & code-adjacent
work, **never human-language translation** (8B quality too low; that's
the sandwich-MT `opus-mt` edge).
## Retrieval pipeline (`arborist/qa/query.py`)
Multi-stage. Each stage exists because something earlier wasn't
enough; revert at your peril. Order:
1. **Four parallel FTS5 search routes per shard, merged** — body BM25,
title-LIKE, core-keyword (TF-IDF cores), and **phrase-pattern**
(verbatim n=5/n=6 sequences from the question). Phrase route
closes the allusion gap (Orwell case: "always been at war" verbatim
matches the 1984 article whose title shares zero tokens with the
query).
2. **Body-coverage `sqrt` rerank** — counters BM25's short-doc bias.
3. **Title-token boost** — `boost × overlap` on title-token-matching hits.
4. **`_filter_by_title_relevance` — four accept paths**: title-token
overlap, TF-IDF core match, body density, **phrase match**
(accept-path 4 lets phrase-route hits with no title overlap survive).
5. **Rivalry exclusion + synonym expansion** (`arborist/concepts/`) —
Intel-titled docs drop from AMD queries; reverse holds. Backed
by the per-shard `concept_relations` SQLite table (corpus-derived,
not hand-curated). 1.6% storage tax measured at backfill on 6 GB
wiki — kept flat, no further compaction. (Storage choice rationale
lived in `docs/concept-relations-design.md` prior to its deletion
in `bb6a89c` — the design is now documented inline in
`arborist/concepts/extract.py` + the data is self-describing.)
6. **Stem-aware token matching** — possessive / plural collapse
(`superman's → supermans → superman`).
7. **Per-source context cap** — `max_context_chars / top_k`. Prevents
one huge doc from monopolizing the budget.
8. **Wikitext base prose** runs on assembled context BEFORE the LLM.
9. **Template-phrase stopwords** (`_FTS5_STOPWORDS` and
`_TITLE_STOPWORDS` must stay in sync) — strips `tell show describe
explain summarize say give list find make please all there know
everything anything something` so "tell me all there is to know
about X" doesn't dilute query tokens.
`--retrieval-keywords` (CLI: `K="..."`) lets an operator augment
retrieval-side tokens without changing what the LLM sees as its
question. Provenance gap on this is tracked in
[Ticket #000001](docs/tickets/ticket-000001-retrieval-keywords-audit-gap.md).
## Hot path / gotchas
- **Hand-rolled wikitext parser** (`arborist/sources/wikipedia.py`):
char-position state machine, escape-aware, 4× faster than char-
by-char loops via `str.find` + slicing. cProfile any change.
- **`PRAGMA synchronous=NORMAL`** per-connection in `store.connect()`.
Safe under WAL. Don't downgrade without measured reason (~5× cost).
- **HTML source has optional deps**: `pip install '.[html]'` for
`selectolax`. CLI surfaces `--source html` only if import succeeds.
- **Background ingest/distill processes**: stdout is buffered. Use
`export PYTHONUNBUFFERED=1` or `python -u`.
- **Disk pressure**: full cur ~2 GB; full old ~58 GB. `df -h` first.
## Operational rules
- I propose, fox decides. Unsure = ask. Can't ask = stop.
- **Never offer "stop here / take a break / fresh eyes tomorrow" as a
next-action option.** Fox keeps going. When proposing options at a
decision point, list the actual moves (high-payoff vs low-risk
cleanup, etc.); don't pad the menu with a no-op "we're done for
today" choice. If a task is genuinely complete, say so flat —
don't dress it up as a third option.
- No autonomous destructive ops (`clean-data`, `clean-db`, force-push,
DB drops) without explicit instruction.
- **Never add `Co-Authored-By` or "Generated with Claude" lines to
commits.** Code speaks for itself.
- **Python only in arborist.** No Rust, C, JS, or other languages
inside this repo. arborist is the source-of-truth implementation;
forks and downstream clients/servers in any language follow our
schemas, canonical encodings, and audit protocols. Optional
toolchains for ZK/world-model/etc. live in sibling repos
(`arborist-zk-bench`, `arborist-world`, etc.) so a fresh checkout
needs only `python3.12 + venv + sqlite3`.
- **Always `export PYTHONUNBUFFERED=1`** for long-running processes.
- **AUTOCOUNT discipline on numeric claims in `docs/`.** Any
numeric claim added to a doc (test count, fixture-row count,
SQLite row count, filtered-row count) should be wrapped in an
AUTOCOUNT tag at write time so future drift fires the
regression test. Format:
``<!--AUTOCOUNT:metric:path-->N<!--/AUTOCOUNT-->`` where
``metric`` is one of ``tests`` / ``fixture-rows`` / ``db-rows``
/ ``db-where``. Tags are invisible in rendered markdown
(GitHub strips HTML comments). Tags inside ```` ``` ```` fenced
code blocks are auto-skipped (illustrative examples, not live
claims). Closed-ticket "N tests pass" snapshots stay UNtagged
(they're point-in-time historical records). Full discipline +
4-metric reference + future-metric recipe in
``docs/tickets/ticket-000044-autocount-doc-drift-discipline.md``;
harness at ``tests/test_doc_counts.py``.
- Fail-closed. Cleanup crew, not demolition.
- DRY in context — single source of truth, no sprawl.
- Never say "AI" — always say "machine learning."
- Prefer "defect" over "bug."
- Check `robots.txt` before any web fetch the user didn't authorize.
## Five-step algorithm
When proposing or evaluating change, walk these in sequence. Skipping
a step makes the next ones expensive and the system worse.
1. **Make the requirements less dumb.** Every requirement gets a
person's name, not a department. If you can't name who asked or
which defect closed, the requirement is suspect.
2. **Delete the part or the process.** If you aren't putting back at
least 10% of what you delete, you aren't deleting hard enough.
Verifier-stays-binary and "no soft signals in hard chain" are
deletion-first guardrails.
3. **Simplify and optimize.** Only after 1 + 2. Don't optimize a
process that shouldn't exist.
4. **Accelerate cycle time.** Only after simplifying. "Don't dig the
grave faster."
5. **Automate.** Last, not first. Hand-rolled before scripted,
scripted before declarative, declarative before generated.
When in doubt, ask "have we tried deleting it?" before reaching for
steps 3-5.
## Bench-maxing — measure deltas, not opinions
Full discipline + worked examples in `docs/bench-maxing.md`. Headlines:
- Bench before AND after every change (n=3, signal under 5pp is noise).
- **When a lever's failure class is sub-noise-floor, fix the
instrument, not just the lever.** The curated n=3 audit_mode bench
can't resolve a class that is ≤3-5 of 75 q (four 2026-05-18
hypotheses died there — round-trip, entity-mask, disambiguation).
Mine **ground-truth-carrying** questions from corpus titles
(`bench/mine_questions.py`) and grade by **deterministic retrieval
recall@k** via `query --dry-run` (`bench/recall_at_k.py`): no LLM,
no verifier, no n=3 noise, no 5pp floor, scalable to the corpus.
Resolves a single lever to ±1 question (worked example: numeral-
fold recall 55→75%, `a3ac653`). Caveats: measures *retrieval
surfacing* — necessary-not-sufficient for STRICT; mined fixtures
are answerable-by-construction so they **complement, never
replace**, the curated adversarial set (the verifier-honesty/trap
gate). Conflating the two is itself a bench-maxing error.
- **Report recall@1/@3/@k, not one lenient k — a coarse k hides a
rank-only lift.** `recall_at_k.py` returns the target's rank, so
recall at every k is free from one retrieval. Measured 2026-05-18:
accent-fold looked inert at recall@8 (95→98, noise) but the OFF
baseline was recall@1 55% vs @8 95% — a too-lenient k flattered
it to a near-ceiling and nearly got a real lever wrongly reverted.
recall@1/@3 is the resolution that matters (primary-source
selection keys on rank, not mere top-k presence). Prevalence ≠
miss-rate either: the corpus survey ranked accent #1 at 8.1% of
*titles*, but the measured miss-rate (recall@1) is what decides —
measure headroom, never rank candidates by raw prevalence.
- **Fan out independent measurements; serial-by-caution is halting
in disguise.** Mined recall is deterministic per query (read-only,
no LLM, no shared state) — concurrency cannot change which sources
rank; the only risk is the per-probe timeout, and dry-run
retrieval (~2 s) has huge margin under the 120 s cap. Run the whole
fold-search backlog (accent / hyphen / honorific / …) as parallel
background sweeps; only same-fixture A/B that toggles `query.py`
state needs git-stash serialisation.
- Avoid negation in prompts (Hermes-3-8B inverts under attention).
- Bench is the scoreboard; live fixtures are the gates.
- Self-heal beats retry (preserve partial output, never fabricate).
- Honest verdicts beat optimistic ones (false-positive STRICT is
corruption).
- Old maps vs runtime maps — when the model and runtime disagree,
the runtime wins. Pointer IDs, runtime-interpolated spans, evidence
maps, policy hashes, hard verifier checks all move authority OUT
of the model's prior and INTO runtime artifacts.
- **When researching a config/model decision: sweep wide, sweep
deep, on real-traffic-sized samples — don't guess, and don't trust
small denominators.** Three times running (ticket #000049 §7
#18→#24) the answer flipped: a clean synthetic eval said one thing,
bench-qa pipeline output said another; the *default* config said one
thing, a `{model × hyperparams}` grid sweep said another; an n=1 FP
sample said "this config passes", an n=3 FP sample said "that one
collapses — a *different* config passes". So: (1) gate numbers come
from bench-qa pipeline output, never contrived fixtures; (2) the
sweep must be wide enough — multiple checkpoints, the full
hyperparam grid; there is no "bigger model is better" law, the
*specific checkpoint* and the *score-shape* (e.g. a single-threshold
margin vs a two-threshold rule) dominate — wide enough to include
the config that survives; (3) every gate number is provisional until
the denominators are big enough — bump `BENCH_QA_N` and re-confirm
before promoting anything. The GPU box (`ai`, the 4090 — `make
bootstrap-nli-only`, then `ARBORIST_NLI_DEVICE` auto-detects cuda)
makes this cheap: a 7-model × full-grid × ~300-record sweep is
~3.5 min. **That electricity is well spent — burn it; a guessed
config that ships is far more expensive than a sweep that doesn't.**
Heavy off-device passes (NLI sweeps, embedding backfills) run on the
GPU producer, never in arborist's `python+sqlite3` core (cf. the
#000051 vecpack pattern).
## Docs index
North-star:
- `docs/seven-point-program.md` — the architectural directive distilled
2026-05-01. Every new ticket / feature / prompt edit walks past this.
Bench harness reports per-mode directive coverage.
- `docs/warrant-substrate-cookbook.md` — per-pillar map of the 18
open-licensed textbook substrates that back every claim-pack
record. Ingest patterns (HTML / textbook_tex / PDF→localhost),
alias discipline, cascade tuning, honest tier breakdown.
Read first when extending substrate coverage.
Architecture / ongoing work:
- `docs/cti-architecture.md` — CTI Clause Tree Intelligence.
- `docs/mesh.md` — mesh wire format + group-key state machine.
- `docs/crawler.md` — web crawler: BFS discovery, robots/feed
handling, polite vs `--fast`, and the content-addressed
diagnostics (dedupe by `document_root`, orphan finding).
- `docs/embedding.md` — embedding arborist as a library in another
Python app (`arborist.embed`): produce `Document`s → ingest →
dedup + FTS5 + audit chain. The neopig-backend seam.
- `docs/corpus-history.md` — durable note of state changes at scale
that aren't otherwise captured outside the audit chain (initial
Wikipedia 2010 ingest, future migrations, cold-pack runs). Append-
only; one entry per event. Surface for the headline numbers + a
pointer to the audit-chain query that derived them.
- `docs/cold-object-store.md` — cold-pack distribution tier (#000061):
serialize the corpus into `tar.zst` packs and ship them via S3-compatible
buckets (DO Spaces / AWS S3 / R2 / B2 / GCS / MinIO via boto3) and/or
burn to DVD-R via `--local-dir` + growisofs. Packs are point-in-time
snapshots; each pack pins the `snapshot_root` it covers so falsifications
between repacks produce new pack_hashes. ≤4.4 GB safe-fit per pack (DVD-R
with ~6.5 % buffer below the 4.7 GB marketing capacity).
- `docs/benchmarks.md` — orientation: harnesses, fixtures,
signal floor, make targets, bench-row schema, addenda index.
Read first when running a bench.
- `docs/qa-modes-bench.md` — bench journal (rolling addenda).
Headlines + cross-references to per-ticket bench data.
- `docs/bench-maxing.md` — bench discipline (5pp signal floor etc).
- `docs/soft-hash-channel-analysis.md` — #000018 closure; analysis
of soft-hash covert channel risks under M0/M1/M2 threat models.
- `docs/soft-hash-channel-t3-bound.md` — #000036 closed-form
per-window budget bound; pairs with t3_bound_calculator.
- `docs/onnx-vendor-capture-immunity.md` — why the model-in-proof-
path cage (#000049) makes the inference engine (ORT / torch /
tinygrad) an interchangeable sidecar, never a trust dependency;
public-domain positioning capital.
- `docs/calculator-test-patterns.md` — sister discipline for
testing calculator-style code (12 patterns; see #000044 for
numeric drift discipline).
- `docs/spec-methodology.md` — #000019 specification methodology
for π* canonical projections.
- `docs/v7w-frontier-catalog.md` — #000013 v7-W spatial-temporal
frontier catalog (4 ε-frontiers).
- `docs/v8-fork-score.md` — #000012 ForkScore Phase 1a reference.
- `docs/pi-star-composition.md` — π* cross-domain composition
reference (#000015 deliverable).
Tickets: `docs/TICKETS.md` is the authoritative index with `Next
ID`. Closed tickets stay in place as the design log. New tickets
bump `Next ID` atomically. **Close tickets when the work lands —
flip Status to `closed · landed in commit <sha>` (or
`closed · YYYY-MM-DD`) in the ticket file AND in the index row, in
the same commit as the implementation.** An open ticket whose code
already shipped is a stale map.
**Default: extend an existing ticket. Don't proliferate.** When
follow-up work surfaces during implementation (a consumer-side fix
the new code needs to actually take effect; a bench-gated tuning
step; a small downstream tweak), **add it to the open / most-related
ticket**: reopen `closed` → `in progress` if the prior closure was
premature, extend the scope, list the new sub-items, keep the design
log linear. fox tracks the index by skimming a small set of threads;
spawning #000054.1 / #000055 / #000056 / … for every follow-up turns
the index into a wall of micro-tickets he can't easily keep state on,
and dilutes the design log instead of concentrating it. **Split only
when the new piece needs a *distinct audience*** — specifically, when
it's a self-contained design decision a Dav1d de-novo review (the
external code-review thread that runs against one ticket at a time)
needs to read independently. Architectural inflection points, large
scope changes, or fundamental discipline questions (NLI may touch
`audit_mode`? — #000049 split from #000048) cross that bar. "The
extractor needs a small consumer-side tweak" does not. When in doubt,
ask before opening a new ticket — a two-sentence "fold into #X or
spawn a sibling?" is cheap; an unwanted ticket is friction. Closing
on the implementation commit is good practice **only** when the work
is actually complete end-to-end; pattern-match closure on the
user-visible outcome, not on the commit.
## Orientation protocol
```bash
date -u
pwd
git log --oneline -5
git status
make test
make chain-check-shards # 0 per shard = intact
.venv/bin/arborist --shards-dir ~/.arborist/shards stats
.venv/bin/arborist --shards-dir ~/.arborist/shards analyze --gravity-top 5
```
Then ask fox what the mission is.