docs(#000057): scaffold — minimal deterministic recursive-drift A/B (Hamming de-novo review)

The 2026-05-19 GPT-5.5 Hamming-framed review's ONE arborist-scoped,
ticket-worthy nugget: prove the Merkle-Providence-Reverse-RAG
whitepaper's headline claim (untracked evidence loss -> unbounded
recursive drift; witness-preserving state bounds it). Scaffold only,
awaiting fox go/no-go on scope.

Discipline encoded from the 2026-05-18 precedent (a grand target is
unfalsifiable until the instrument can resolve it — four hypotheses
died, only the deterministic mined-recall instrument broke it):
instrument-before-experiment, ONE task not eight, minimal ON/OFF
A/B, non-claims pinned (necessary substrate, NOT AGI).

Everything else in the review (rename, corpus hierarchy, IQ/talent,
ToE/Riemann/identity/geopolitics) deliberately NOT ticketed —
narrative/positioning, not arborist engineering; don't-proliferate.
Exactly one ticket. Next ID 000057 -> 000058 (same commit).
This commit is contained in:
russell@unturf.com 2026-05-19 08:14:04 -04:00
parent c00639ed1d
commit 7100f7277b
No known key found for this signature in database
2 changed files with 91 additions and 1 deletions

View file

@ -111,6 +111,7 @@ Newest first. Update on every open/close.
| ID | Title | Status | Opened | Directive |
|----------|------------------------------------------------|-----------------------|------------|-----------|
| #000057 | Witness-preserving vs ordinary recursive loop: minimal deterministic drift A/B | **open · awaiting go/no-go · doc-only scaffold** (2026-05-19; fox relaying a Hamming-framed GPT-5.5 de-novo review). The review's one ticket-worthy nugget: prove the whitepaper's headline claim — *untracked evidence loss → unbounded recursive drift; witness-preserving state bounds it (detectable+reversible)*. Everything else in the review (rename, corpus hierarchy, IQ, ToE/Riemann/identity) **deliberately NOT ticketed** — narrative, not arborist engineering; don't-proliferate. Hard discipline encoded from the 2026-05-18 precedent: **instrument before experiment** (deterministic, no-LLM-judge, ground-truth-carrying, noise-resolvable — the `recall_at_k` discipline), **one task not eight** (recursive stale-source-invalidation *or* contradiction-repair — the falsification-state-exercising ones), minimal A/B (witness-binding ON vs OFF, N iterations, deterministic surviving-unsupported-claim count), non-claims pinned (necessary substrate, NOT AGI). Scaffold only until fox rules scope (task + metric + spend). One ticket, not ten. | 2026-05-19 | — |
| #000056 | Operation Sandwich — cross-language grounding via query+display MT | **implemented & landed 2026-05-17 · default-OFF** (fox: "call it operation sandwich, create a new ticket and finish it"). Mechanism + bright line **verified live end-to-end** (real Hermes + real opus-mt: es query → English answer+verifier → es display; `answer_text` English, `display_answer` Spanish additive, `question_hash`/`verifier_policy_hash` invariant; 6 tests + full suite 2477 passed, 0 regressions). **Fan-out measured (§9, n=1, 75 q):** EN baseline 85% → es+sandwich 71% = **14pp cost**; transitions PRESERVED 31 / DOWNGRADE 16 / LOST 17 / N/A 11. A deterministic round-trip predictor was tried and **refuted** (12/17 LOST round-tripped CLEAN; another instance of the codified CLAUDE.md bench-maxing lesson — not re-added). LOST taxonomy from the artifact: ≈7 entity-translate (`Boltzmann`→"perntzmann", `Tarsus`→"Tarso"), ≈3 broad-enum, ≈several n=1 noise. **Lever built+validated:** `arborist/qa/mt/entity_mask.py` mask/restore (real opus-mt: `who is Paul of Tarsus?` "Pablo de Tarso"→**"Paul of Tarsus"**); default-ON within the default-OFF sandwich. Caveat: bench is lowercased so cap-detector lift is a **lower bound** (corpus-title anchor = v2). Also: fan-out caught + fixed an 88%-engine-error concurrency defect (per-call model load → memoised singleton + lazy per-pair); French + Russian breads added (manifest, `crosslang_source_lang`). **Lift measured 2026-05-17 (comparator corrected, fox):** the true baseline is the pre-ticket ≈0% (raw es query → song-title noise, UNGROUNDED, 10.4s) — NOT native English. Against that: **the sandwich is a large net win (≈0% → 71% es grounded, proof core never corrupted); the 14pp vs English is the cost of a new capability, not a regression — calling it a "fail" was a comparator error.** Firmed 2026-05-18 (n=1, 0 engine-err): EN 85% · es-nomask 71% · fr-nomask **61%** (NOT 47% — that was the mask artifact; honest fr correction). The genuine negative is the **entity-mask lever**: net-negative in *both* languages (es 71→65 borderline, fr 61→47 = 14pp unambiguous; isolated Paul-of-Tarsus win didn't replicate — 3rd bench-maxing-lesson instance), now **default-OFF** (`crosslang_entity_mask=False`); no-mask sandwich is the keeper. Recommendation flipped: **worth continuing (minus the mask)**, not park. Remaining: n=3, fr no-mask, corpus-title anchoring (only untried lowercase-capable detector). CLAUDE.md updated with the durable cross-lang *convention* only. Tasks #14#17. Phase 1 of the #000001 §7 family; new ticket clears don't-proliferate (fox-directed + distinct Dav1d audience + architectural inflection: a model dependency `[mt]` + a presentation-translation layer — anticipated by #000001 §7's "split the `[mt]` model-distribution work like `[nli]`/vecpack"). **Sandwich:** translate query es→en (retrieval-side, == `--retrieval-keywords`, binds into `retrieval_plan_hash`, NOT `question_hash`) → English answer through the **byte-for-byte untouched verifier** → translate the verified English `answer_text` en→es into a NEW `display_answer` field, banner-labelled, zero grounding (the `_render_audit_label` render-projection pattern). Engine: local `[mt]` extra, Helsinki-NLP `opus-mt-es-en`/`-en-es`, Apache-2.0, hash-pinned, off-repo `~/.arborist/models/mt/`, optional dep, graceful-degrade — mirrors `[nli]`/`ShadowNLI` (#000049) + vecpack (#000051) verbatim; never Hermes-3-8B; not an external API (reproducibility + zero egress + es↔en is the best-resourced pair). Default OFF (`crosslang_translate_enabled`, gated under Phase-0 `crosslang_guard_enabled`); `--crosslang-translate` / `XLANG_MT=1`. Hash invariants (corrected 2026-05-17 — `governance_policy_hash` is sha256 of the *whole* policy, keys.py:182): `question_hash` + `verifier_policy_hash` untouched (user question preserved, verifier byte-identical); `governance_policy_hash` moves like every policy flag → correct cache partitioning by config (not a leak); MT engine identity binds into `RetrievalPlan.mt_*` (run-DAG), not a policy hash. | 2026-05-17 | — |
| #000055 | Windows quickstart without `make` (`tasks.py` + `make.bat`) | **in progress** — opened 2026-05-16 (fox: "bat files or some shit … avoid needing makefile for windows … we will test the quickstart on windows"). Pure-stdlib `tasks.py` runner mirroring the **quickstart subset** of the Makefile (bootstrap / fetch-cur / ingest-cur-attached / distill ×2 / query / inspect / falsify / burn / bootstrap-crawler / crawl-ingest / stats / verify / search / clean) + a ~10-line `make.bat` shim so `make <target>` works in Windows cmd and `.\make.bat <target>` in PowerShell. Audited the artifact (not the docs): the 2003 dump is opened via stdlib `bz2` (no external `bzip2`); only `fetch-cur` used `curl` (→ stdlib `urllib`); bash `for…&wait``subprocess.Popen` fan-out; `arborist` console-script lands at `.venv\Scripts\arborist.exe`. Net: quickstart needs only **Python 3.10+ + sqlite3** — the repo's existing ethos, now true on native Windows. Same `KEY=VALUE` make-style args so documented commands translate 1:1 (one doc form). Makefile untouched, still canonical on POSIX ("keep it as an option"). Found + fixed a README/Makefile discrepancy: README claimed `[dev,html]` bootstrap extras, Makefile installs `.[dev]` — artifact wins. Drift-pinned by `tests/test_tasks_runner.py`. | 2026-05-16 | — |
| #000054 | Acronym-parens concept extractor (closes the abbreviation→expansion retrieval gap) | **in progress** — Phase 1 (extractor + 481K edges) landed `58027e9`; Phase 2 (consumer-side surfacing — `synonym_expand` rank-and-truncate over the per-token cap, FTS5-`bm25` ordering in `_search_titles`, expanded `accept_tokens` in title-search + core-keyword + title-rerank, `synonym_expand_strict()` for the multiplicative title-purity rerank to exclude noisy `link_reciprocity` edges, tightened extractor regex to `[A-Z]{3,6}` purging 2-letter homonym edges) landed `ce855db`. **End-to-end verified:** `what is a CPU?` → Central processing unit at #1; `what is a GPU?` → Graphics processing unit at #1 EVIDENCE-WARRANTED 1/1; Mount Kilimanjaro / Soviet Union queries unchanged. **bench-qa n=3 limit=5** (2026-05-13T14:24Z): 30/45 STRICT (67%), zero regressions on basics (mona lisa / capital of france / new london bridge each 9/9 STRICT). 2026-05-13 — `arborist/concepts/extract.py:acronym_parens_synonym` lands as a new corpus-agnostic extractor in `EXTRACTORS` (`evidence_kind="acronym_parens"`). Scans each doc's lead chunk for `<Multi-Word Phrase> (ACRO)` where the all-caps acronym's letters match the content-word initials of the phrase in order; emits bidirectional synonym edges between the lowercased acronym and each ≥3-char content token of the phrase. Conservative (strict 1:1 initials, function words filtered, repeated definitions deduped per doc). Closes the *retrieval-side* abbreviation gap (`CPU↔central processing unit`, `GPU↔graphics processing unit`, `RAM↔random access memory`, `FBI↔federal bureau of investigation`, `WHO↔world health organization`, …) that `link_reciprocity_synonym` can't reach because the relation lives in body text, not the wiki link graph (Wikipedia represents abbreviation→expansion as a *redirect* — not an edge). Per-shard like all `concept_relations` data; corpus-agnostic so HTML/blogs/textbooks benefit equally. Retrieval-side only — never proof-path. 8 new tests; full suite green. Closes #000050 §2a's CPU/GPU fixture rows *upstream* of vec; the Orwell-shape conceptual-allusion row remains the genuine #000050 justification. Operational follow-up (not code): `arborist concepts derive --extractor acronym_parens` on each shard. | 2026-05-13 | — |
@ -170,4 +171,4 @@ Newest first. Update on every open/close.
## Next ID
`000057`
`000058`

View file

@ -0,0 +1,89 @@
# Ticket #000057 — Witness-preserving vs ordinary recursive loop: minimal deterministic drift A/B
**Status:** open · awaiting go/no-go · doc-only scaffold (2026-05-19)
**Opened:** 2026-05-19
**Asked by:** fox, relaying a Hamming-framed de-novo review (GPT-5.5
Thinking, 2026-05-19 19:13 Asia/Kuala_Lumpur).
**Scope:** ONE ticket, scaffold only. Define the *smallest*
deterministic experiment that would prove the whitepaper's headline
empirical claim. NOT a build. NOT the 8-task/8-metric flagship.
**Audience:** fox + a Dav1d de-novo review of the experiment design +
the Merkle-Providence-Reverse-RAG whitepaper (this is its missing
empirical headline; the paper currently *asserts* the architecture).
---
## 1. The one claim worth proving (the review's §8.2, compressed)
> In a recursive loop where outputs feed future context, **untracked
> evidence loss → unbounded epistemic drift**. A witness-preserving
> state transition system **bounds** drift by making unsupported
> transitions detectable and reversible.
Everything else in the review (renaming, corpus hierarchy, IQ/talent,
ToE/Riemann/identity/geopolitics) is **deliberately not ticketed**
positioning/whitepaper-narrative, not arborist engineering;
don't-proliferate + this session is arborist + whitepapers only.
## 2. The discipline this ticket exists to enforce
The blunt review line is *"make the **smallest** irreversible
proof."* The 2026-05-18 session is the cautionary precedent: a grand
target (STRICT-rate) was unmovable for four hypotheses **because the
instrument couldn't resolve it**; the win came only from a
deterministic, noise-free, ground-truth-carrying instrument
(mined recall@k). A "prove recursive trustworthy intelligence"
experiment is the *same trap at larger scale* — it will produce
unfalsifiable mush unless the instrument is built first.
**Therefore, hard scope gates (do not skip in order):**
1. **Instrument before experiment.** Define ONE deterministic drift
metric — replayable, no-LLM-as-judge, ground-truth-carrying,
resolvable above a stated noise floor — exactly the
`mine_questions.py`/`recall_at_k.py` discipline. If the metric
isn't deterministic and ground-truthed, stop; the experiment is
not yet attackable.
2. **One task, not eight.** Pick the single task that most directly
exercises arborist's *distinguishing* mechanism — falsification
state — under recursion: **recursive stale-source invalidation**
*or* **recursive contradiction repair**. The other five §8.1
tasks are explicitly OUT until the one-task A/B clears its floor.
3. **Minimal A/B.** Same task, N recursive iterations, the only
difference being witness-binding/falsification-state ON vs OFF.
Metric = a deterministic count of unsupported claims that survive
K iterations (and whether they are reversible). Before vs after,
like every fold in #000001-family work.
4. **Non-claims, pinned.** NOT an AGI claim. The provable statement
is "a *necessary substrate* for recursive trustworthy
retrieval-grounded answering, demonstrated on one task." The
review agrees (§9 AGI). No STRICT/global-QA/multilingual claim.
## 3. What this is NOT (recorded so it isn't re-litigated)
- Not the multi-task flagship matrix (§8.1's 6 tasks × 8 metrics) —
that is the *grandness* the review itself warns against; it earns
consideration only after gate 3 clears with measured signal.
- Not a rename (`Witness-Preserving Recursive AI` etc.) — narrative,
not engineering; whitepaper editorial, separate from this ticket.
- Not corpus-triage / physics / Riemann / identity — out of arborist
scope entirely.
- Not autonomous-executable: this is a scope/strategy decision (I
propose, fox decides) — unlike the folds, which were measured,
bounded, and self-evidently shippable.
## 4. Why a ticket at all (don't-proliferate justification)
Crosses the split bar: a self-contained, Dav1d-reviewable
architectural-inflection design decision (the whitepaper's headline
experiment), distinct audience, not a sub-tweak of an existing
ticket. Exactly **one** ticket — the rest of the review is
explicitly refused above.
## 5. Decision required (go/no-go, fox)
Scaffold only until fox rules on **scope**: (a) which single task
(stale-source-invalidation vs contradiction-repair), (b) the exact
deterministic drift metric, (c) whether the minimal A/B is worth the
spend now vs parked. No design past §2 gate 1 until that ruling —
building the experiment before the metric is the precedent error.