ticket #000039: close — Phase 0 + Phase 1 landed; Phase 2 split to #000050

#000039 (the optional sqlite-vec retrieval backend) closed 2026-05-12.
Shipped: Phase 0 doc + Phase 1 — arborist/search/vec.py (VecBackend,
chunk_vecs vec0 + vec_meta sibling tables, embed_documents
incremental/--rebuild, pluggable Embedder w/ fastembed bge-small-en-v1.5
default), CLI (arborist embed [--limit/--batch-size/--quant/--rebuild]
+ search --backend vec + ingest --embed eager opt-in), [vec] extra,
--quant {float32,int8} with the int8 head-to-head (3.8-4x smaller,
recall ~= float32 — int8 is the production config), ingest integration
+ idempotency (§14), embed-throughput measurement (§14.6 — ~4 chunks/s
contended, ~2.4 GB int8 full-corpus, full backfill abandoned as a
days-long batch job, non-vec ingest unchanged), 16 vec tests.
Demonstrated on crawl_appliedcombinatorics_org.db + a 54K-chunk partial
on wiki shard 000. UNGROUNDED hits, never proof path; vec config folds
into governance_policy_hash (noted in Phase 1, wired in Phase 2).

Phase 2 (RRF hybrid fusion in query.py — wire VecBackend as a 5th
retrieval route alongside the 4 FTS5 routes) → new ticket #000050,
gated on (a) a corpus backfill and (b) a recall bench clearing the
5pp floor. New doc-only scaffold docs/tickets/ticket-000050-vec-rrf-
hybrid-fusion.md (the design already lives in #000039 §4.2 + §8;
#000050 is the tracked continuation). Next ID 000050 -> 000051.
TICKETS.md: #000039 row -> closed, #000050 row added.

(Working tree also has parallel-clone work — ticket-000048-*.md
modified, ticket-000049-*.md untracked — not touched here.)
This commit is contained in:
russell@unturf.com 2026-05-12 08:33:45 -04:00
parent 20f6061f83
commit 906606072d
No known key found for this signature in database
3 changed files with 135 additions and 3 deletions

View file

@ -92,6 +92,7 @@ Newest first. Update on every open/close.
| ID | Title | Status | Opened | Directive |
|----------|------------------------------------------------|-----------------------|------------|-----------|
| #000050 | Vec RRF hybrid fusion (#000039 Phase 2) | open · awaiting go/no-go · doc-only scaffold; design in #000039 §4.2 (RRF) + §8 (the gate). Wire `VecBackend` as a 5th retrieval route in `query.py`, RRF-merged with the 4 FTS5 routes; UNGROUNDED hits, additive not replacement; folds the 5 vec hyperparams into `governance_policy_hash`. **Gated** on (a) a corpus backfill (`arborist embed --quant int8` — a one-time batch job, hours-on-idle/days-on-contended per #000039 §14.6) AND (b) a recall bench clearing the 5pp floor (vec-only beats FTS5-only ≥5pp on ≥1 semantic-allusion fixture; hybrid lifts STRICT-rate ≥5pp, no UNGROUNDED regression — else park, vec stays opt-in `--backend vec`). #000039 follow-up | 2026-05-12 | — |
| #000049 | Attribution-aware grounding check (the recombination boundary) | open · awaiting go/no-go · doc-only; the home for #000048's deferred §2.3 — closing the 2 recombination over-grounds in `falsification-hard` (hard-003 Mercury / hard-005 Einstein) needs an attribution / dependency-parse or mini-NLI check, which is *not lexical* (#000048 §5). First decision = the discipline question: may a small fixed purpose-built NLI/entailment *model* influence `audit_mode`? (vs the "no LLM-as-judge" rule). Recommends Option 2.3 (do nothing — the 2 fixtures are a boundary marker) until real-traffic recombination-over-grounds show up, then Option 2.2 (`[nli]` extra, policy-gated, off-by-default, bench-gated) *if* fox rules a fixed NLI model is acceptable. #000048 follow-up | 2026-05-12 | — |
| #000048 | Verifier upgrade — recombination-aware grounding + clause segmentation | **closed · 2026-05-12** — steps 2.1 + 2.4 landed 2026-05-11 (12 of 16 residual items: 4 HYBRID_ENTITY over-grounds + 8 Formulate mis-segments → `formulate-hard` 12/12, `falsification-hard` 10/12; each bench-gated, no STRICT-rate regression — 2.1's gate fired on 0 QA answers, 2.4's segmenter touched 7 of 450 lattice cells both verdict changes correct). Step 2.2 (single-clause-containment paraphrase check) attempted + reverted — catches the 2 recombination fixtures but also rejects legit cross-sentence summaries with no threshold separating the two; recombination-vs-summary isn't lexical (§5 "What we learned"). The attribution-aware path moved to **#000049** (fox 2026-05-12). 2 live-pack `expected_reason` updated HYBRID_ENTITY→UNGROUNDED; 12+ tests; `make bench-5f-falsification-hard` / `bench-5f-formulate-hard` / `bench-fork-baseline-hard`. #000046 follow-up; #000047 closed | 2026-05-11 | — |
| #000047 | ForkScore `_delta_*` aggregator (mean vs max vs sum) | **closed · 2026-05-11** — Option D: `WeightSet.delta_aggregator` ∈ {`mean`,`max`,`sum`} (default `mean` unchanged → no `ESTIMATOR_VERSION` bump), `fork_score._delta_5{s,t,f}` dispatch via `_aggregate`, recorded in `ScoredFork.weights`, per-sub `HARD_REGRESSION_FLOOR` flags aggregator-independent; bench data behind keeping `mean` in `5f-threshold-calibration-2026-05-11.md` §5; 8+1 tests. #000012-revision / #000025 §10.14 follow-up | 2026-05-11 | — |
@ -102,7 +103,7 @@ Newest first. Update on every open/close.
| #000042 | Term-aliases table (vocabulary-mismatch bridge) | closed · 13 rows live across geometry + classical-physics + arithmetic domains by 2026-05-10 | 2026-05-09 | — |
| #000041 | Citation-aliases table (PD substitutes for proprietary cites) | closed · 74 rows live as of 2026-05-10 (count grew 40 → 54 → 74; Goldstein/Newton, Mendelson/Enderton/Jech/Landau/Gödel→{Russell IMP, Russell PoM, De Morgan, Boole, Cantor, Peano, Dedekind, SF-LF}, Stanley/Brualdi/Knuth → Bogart+Levin+Keller-Trotter, Dummit-Foote/Barendregt/Böhm-Jacopini → Judson/PLFA/SF, Kolmogorov → Grinstead-Snell+Laplace) | 2026-05-09 | — |
| #000040 | Phase 5 resolver fix — phrase + content-token cascade (Hilbert terminology mismatch surfaced) | closed · cascade landed 2026-05-09; lift blocked by 1902-vs-modern vocab; follow-up #000042 | 2026-05-09 | — |
| #000039 | Optional `sqlite-vec` retrieval backend (A/B vs FTS5, hybrid not replacement) | in progress · Phase 0 doc + **Phase 1 landed 2026-05-11** (`arborist/search/vec.py``VecBackend` + `chunk_vecs` vec0 + `embed_documents` + `arborist embed` / `search --backend vec`; `[vec]` extra = sqlite-vec + fastembed; v1 tuning: bge-small-en-v1.5 / 384 / float32 / cosine / flat / top_k 20; 7 tests; demonstrated on `crawl_appliedcombinatorics_org.db` — 168 chunks, semantic hits topically correct, chain-check 0). Phase 2 (RRF hybrid fusion) gated on ≥5pp recall lift | 2026-05-09 | — |
| #000039 | Optional `sqlite-vec` retrieval backend (A/B vs FTS5, hybrid not replacement) | **closed · 2026-05-12** — Phase 0 (doc) + Phase 1 landed 2026-05-11: `arborist/search/vec.py` (`VecBackend`, `chunk_vecs` vec0 + `vec_meta` sibling tables, `embed_documents` incremental/`--rebuild`, pluggable `Embedder` w/ fastembed `bge-small-en-v1.5` default), CLI `arborist embed [--limit/--batch-size/--quant/--rebuild]` + `search --backend vec` + `ingest --embed` (eager opt-in), `[vec]` extra; `--quant {float32,int8}` with int8 head-to-head (3.8-4× smaller, recall ≈ float32 — int8 is the production config); 16 vec tests; ingest integration + idempotency (§14); embed-throughput measured (§14.6 — ~4/s contended, ~2.4 GB int8 full-corpus, full backfill abandoned as a days-long batch job, non-vec ingest unchanged). UNGROUNDED hits, never proof path; vec config folds into `governance_policy_hash` (noted, wired in Phase 2). **Phase 2** (RRF hybrid fusion in `query.py`) → **#000050** (gated on a corpus backfill + a ≥5pp recall bench). | 2026-05-09 | — |
| #000038 | Phase 4 content acquisition — proprietary textbook license decisions for warrant coverage | closed · obviated 2026-05-10 by alias-substitution sprint under #000031 (74 rows in #000041 + 13 rows in #000042); 92/92 records now resolve. Residue (multilingual PD, Hilbert-Ackermann OCR, Knuth permission, personal-copy path B) preserved as design log §8 | 2026-05-09 | — |
| #000037 | Prometheus-Σ recursive falsification controller (bicameral substrate) | in progress · Phases 0 + 1 + 1.b + 1.c + 2 landed 2026-05-10; **§12 Trigger 2 fired** (divergence variance 0.575 / N=37); §22 Findings 2 + 3 RESOLVED (kernel/llm cost split + sweep_weights §15.4 + per-mode τ_qa); `controller_events` carries 4 event kinds (decision · difficulty · budget_allocation · falsification_proposal) feeding `arborist controller-events` inspector + live-harvest third bucket in `bench/scripts/harvest_falsification_proposals.py`; §12 Trigger 1 probe wired 2026-05-11 (`trigger_1_branch_density` reads `fork_score_branches` — measurable, not yet fired); Phase 3 sleep-sweep scheduler tracked under #000045 (gating ticket) | 2026-05-09 | — |
| #000036 | T3 per-window covert-channel budget bound | **closed · 2026-05-11** · Phase 1 + dav1d review → Tier-1 + Tier-2 (Option B = `b1_model=max_envelope` default, in v1, no v2 fork) + KAT-regen tooling (`scripts/generate_t3_bound_kat.py`) all landed 2026-05-11; baseline 625.87 → 6183.02 (max_envelope), `NOT_CERTIFIED_BY_BOUND` at W=10000; 53 → 83 tests; 12-entry active KAT; both dav1d closure blockers cleared, all §5 acceptance criteria met. Continuation: empirical C_B* tightening under #000043 (parks on v7) | 2026-05-09 | — |
@ -144,4 +145,4 @@ Newest first. Update on every open/close.
## Next ID
`000050`
`000051`

View file

@ -1,6 +1,6 @@
# Ticket #000039 — Optional `sqlite-vec` retrieval backend (A/B vs FTS5, hybrid not replacement)
**Status:** in progress · Phase 0 (doc) + **Phase 1 landed 2026-05-11**. Phase 1: `arborist/search/vec.py``VecBackend(SearchBackend)` (UNGROUNDED hits, never in proof path), `chunk_vecs` vec0 virtual table + `vec_meta` (sibling tables — don't touch chunks/documents/audit chain), `embed_documents()` ingest (delete-then-insert idempotent; vec0 doesn't honor INSERT-OR-REPLACE), pluggable `Embedder` callable with a fastembed `bge-small-en-v1.5` default. CLI: `arborist embed [--limit] [--batch-size]` + `arborist search --backend vec`. `[vec]` optional extra (sqlite-vec + fastembed). **Obvious v1 tuning** (`VEC_BACKEND_VERSION = vec-v1-bge-small-en-v1.5-384float32-cosine-flat`): embedder `BAAI/bge-small-en-v1.5`, dim 384, quant float32 (int8/binary = the production storage knob per §3.1, not wired in v1), metric cosine (bge outputs L2-normalized so cosine ≡ L2 ranking), ANN flat (vec0 default), top_k 20. 7 tests (`tests/test_search_vec.py`, stub embedder — plumbing only; semantic quality demonstrated on a real shard). **Demonstrated on `crawl_appliedcombinatorics_org.db`** (168 chunks embedded in ~37s incl. model load; semantic queries return topically-correct hits — "how many ways to choose k things from n" → top hit "AC Combinations"; chain-check on that shard reports 0 after embedding). The 5 hyperparams fold into `governance_policy_hash` in a later phase (§6 — not wired yet). **Ingest integration landed 2026-05-11** (see §14): `embed_documents()` is now **incremental by default** (embeds only `chunk_id`s not already in `chunk_vecs` — re-runs are cheap no-ops); `arborist embed --rebuild` does the DROP+recreate+full-re-embed for a `VEC_BACKEND_VERSION` bump; `arborist ingest --embed` is the **eager opt-in** (embed this run's new chunks after the chunk+Merkle-commit pass; default ingest does NOT embed — the lazy `arborist embed` pass / cron / Prometheus-Σ unconscious sweep is the usual path). **`--quant {float32,int8}` landed 2026-05-11** (`arborist embed --quant int8 [--rebuild]`; the `chunk_vecs` vec0 column is `int8[384]` vs `float[384]` per quant; int8 blobs are `vec_int8(?)`-wrapped — sqlite-vec v0.1.9 treats a bare blob as float32; quant change on an existing table requires `--rebuild` since the vec0 element type can't be altered in place; the quant folds into `VEC_BACKEND_VERSION``...-384int8-...`, recorded in `vec_meta`). **int8 head-to-head on `crawl_appliedcombinatorics_org.db`**: storage 1.60 MB → **0.42 MB (3.8× smaller; ~4× at corpus scale where blocks fill)**; recall ≈ float32 — Q2 "how many ways to choose k things from n" identical top-5, Q1 "pigeonhole principle counting" identical top-2 with a sub-noise rank-3/4 swap (Δdistance 0.002). So int8 is the obvious production config (§3.1's +6%-tax recommendation confirmed) — v1 default stays float32 for max fidelity; switching the default to int8 is a fox call. 16 vec tests. **Embed throughput measured 2026-05-11** (§14.6): ~4.3 chunks/s on the contended dev box (~50-200/s idle); a 54K-chunk wiki backfill confirmed ~409 B/chunk → ~2.4 GB full-corpus at int8 (the deterministic number), then was **abandoned** — full corpus is hours-on-idle / days-on-contended, a one-time off-peak/dedicated-box batch job, not something to brute-force inline. The non-vec ingest path is unchanged (hundreds of chunks/s); adding vec multiplies ingest by ~10-100× (all in the ONNX matmuls) — which is exactly why lazy-out-of-band is the default and `ingest --embed` is the opt-in. **Phase 2** (RRF hybrid fusion in `query.py`) gated on a ≥5pp recall-lift measurement on bench fixtures with no STRICT-rate regression (§8).
**Status:** **closed · 2026-05-12** — Phase 0 (doc) + Phase 1 (the optional vec backend + ingest path + `--quant`) landed; **Phase 2 (RRF hybrid fusion in `query.py`) split to #000050**, gated on a corpus backfill + a ≥5pp recall bench (the corpus-wide `arborist embed` is a one-time hours-on-idle / days-on-contended batch job per §14.6 — not done here). The vec layer ships demonstrated on `crawl_appliedcombinatorics_org.db` (semantic hits topically correct) + a 54K-chunk partial on wiki shard 000 (confirming ~409 B/chunk → ~2.4 GB full-corpus at int8). 16 vec tests; full suite green. Detail below ↓. Phase 0 (doc) + **Phase 1 landed 2026-05-11**. Phase 1: `arborist/search/vec.py``VecBackend(SearchBackend)` (UNGROUNDED hits, never in proof path), `chunk_vecs` vec0 virtual table + `vec_meta` (sibling tables — don't touch chunks/documents/audit chain), `embed_documents()` ingest (delete-then-insert idempotent; vec0 doesn't honor INSERT-OR-REPLACE), pluggable `Embedder` callable with a fastembed `bge-small-en-v1.5` default. CLI: `arborist embed [--limit] [--batch-size]` + `arborist search --backend vec`. `[vec]` optional extra (sqlite-vec + fastembed). **Obvious v1 tuning** (`VEC_BACKEND_VERSION = vec-v1-bge-small-en-v1.5-384float32-cosine-flat`): embedder `BAAI/bge-small-en-v1.5`, dim 384, quant float32 (int8/binary = the production storage knob per §3.1, not wired in v1), metric cosine (bge outputs L2-normalized so cosine ≡ L2 ranking), ANN flat (vec0 default), top_k 20. 7 tests (`tests/test_search_vec.py`, stub embedder — plumbing only; semantic quality demonstrated on a real shard). **Demonstrated on `crawl_appliedcombinatorics_org.db`** (168 chunks embedded in ~37s incl. model load; semantic queries return topically-correct hits — "how many ways to choose k things from n" → top hit "AC Combinations"; chain-check on that shard reports 0 after embedding). The 5 hyperparams fold into `governance_policy_hash` in a later phase (§6 — not wired yet). **Ingest integration landed 2026-05-11** (see §14): `embed_documents()` is now **incremental by default** (embeds only `chunk_id`s not already in `chunk_vecs` — re-runs are cheap no-ops); `arborist embed --rebuild` does the DROP+recreate+full-re-embed for a `VEC_BACKEND_VERSION` bump; `arborist ingest --embed` is the **eager opt-in** (embed this run's new chunks after the chunk+Merkle-commit pass; default ingest does NOT embed — the lazy `arborist embed` pass / cron / Prometheus-Σ unconscious sweep is the usual path). **`--quant {float32,int8}` landed 2026-05-11** (`arborist embed --quant int8 [--rebuild]`; the `chunk_vecs` vec0 column is `int8[384]` vs `float[384]` per quant; int8 blobs are `vec_int8(?)`-wrapped — sqlite-vec v0.1.9 treats a bare blob as float32; quant change on an existing table requires `--rebuild` since the vec0 element type can't be altered in place; the quant folds into `VEC_BACKEND_VERSION``...-384int8-...`, recorded in `vec_meta`). **int8 head-to-head on `crawl_appliedcombinatorics_org.db`**: storage 1.60 MB → **0.42 MB (3.8× smaller; ~4× at corpus scale where blocks fill)**; recall ≈ float32 — Q2 "how many ways to choose k things from n" identical top-5, Q1 "pigeonhole principle counting" identical top-2 with a sub-noise rank-3/4 swap (Δdistance 0.002). So int8 is the obvious production config (§3.1's +6%-tax recommendation confirmed) — v1 default stays float32 for max fidelity; switching the default to int8 is a fox call. 16 vec tests. **Embed throughput measured 2026-05-11** (§14.6): ~4.3 chunks/s on the contended dev box (~50-200/s idle); a 54K-chunk wiki backfill confirmed ~409 B/chunk → ~2.4 GB full-corpus at int8 (the deterministic number), then was **abandoned** — full corpus is hours-on-idle / days-on-contended, a one-time off-peak/dedicated-box batch job, not something to brute-force inline. The non-vec ingest path is unchanged (hundreds of chunks/s); adding vec multiplies ingest by ~10-100× (all in the ONNX matmuls) — which is exactly why lazy-out-of-band is the default and `ingest --embed` is the opt-in. **Phase 2** (RRF hybrid fusion in `query.py`) gated on a ≥5pp recall-lift measurement on bench fixtures with no STRICT-rate regression (§8).
**Opened:** 2026-05-09
**Scope:** Spec an optional `sqlite-vec` backend that runs **alongside** the
existing FTS5 retrieval pipeline (never replacing it), with phased gates

View file

@ -0,0 +1,131 @@
# Ticket #000050 — Vec RRF hybrid fusion (the #000039 Phase 2)
**Status:** open · awaiting go/no-go — gated on (a) a corpus backfill
and (b) a bench measurement (see "Gate" below). Doc-only scaffold;
the design already lives in `docs/tickets/ticket-000039-sqlite-vec-optional-backend.md`
§4.2 (RRF) + §8 (the measurement gate). #000039 closed 2026-05-12
with Phase 0 (doc) + Phase 1 (`arborist/search/vec.py``VecBackend`,
`chunk_vecs`, `embed_documents`, `arborist embed` / `search --backend
vec`, `[vec]` extra, ingest integration, `--quant {float32,int8}`)
landed; this ticket carries the remaining Phase 2.
**Opened:** 2026-05-12
**Scope:** Wire `VecBackend` as a *fifth* retrieval route in
`arborist/qa/query.py`, merged with the four existing FTS5 routes
(body BM25, title-LIKE, core-keyword TF-IDF, phrase-pattern) via
**Reciprocal Rank Fusion** (RRF — `Σ_routes 1/(k+rank)`, `k=60`),
so the assembled context for a Q&A run draws from both lexical and
semantic candidates. Vec hits stay `UNGROUNDED` (soft signal, never
proof path); the post-merge filter chain (rivalry exclusion,
four-accept-paths title-relevance, stem-aware token matching,
per-source context cap, Rule-8 title-mismatch) runs on vec hits
exactly as on FTS5 hits. **Not** a replacement for any FTS5 route —
vec is additive (#000039 §4).
**Audience:** fox + future blackops shifts + whoever runs the
recall bench.
**Hard constraint:** vec retrieval **never enters the proof path**
(hits land `audit_mode=UNGROUNDED`). The 8-dim `cache_key` invariant
is untouched — the vec hyperparams fold into `governance_policy_hash`
(#000039 §6), which Phase 2 must wire (Phase 1 left it noted-only).
RRF (rank-based) is the merge, not raw-score addition — BM25 scores
and cosine distances aren't comparable scales. No change to the four
FTS5 routes' behavior; hybrid is opt-in / governance-gated until the
bench clears it.
---
## 1. Why this is a separate ticket
#000039's phased plan: Phase 0 (doc) → Phase 1 (the backend +
ingest path) → **Phase 2 (hybrid fusion), which opens only when
Phase 1 measures ≥ 5pp recall lift on the bench fixtures with no
STRICT-rate regression** (#000039 §8). Phase 1 shipped 2026-05-11
and #000039 closed 2026-05-12; Phase 2 is real, scoped, and
remaining, so it gets its own ticket per the repo convention rather
than leaving #000039 open indefinitely.
## 2. Prerequisites (the gate)
Phase 2 implementation does **not** open until both:
1. **A corpus backfill exists.** `arborist embed --quant int8` over
the shard set. Per #000039 §14.6 this is a one-time batch job —
hours on an idle box, days on a contended one (~4 chunks/s
measured on the dev box; ~50-200/s idle). Best run off-peak or
on a dedicated box; ~2.4 GB total at int8 (+6 % over the 38 GB
shards). The lazy `arborist embed` pass / a cron / a Prometheus-Σ
unconscious-sweep task (#000037 §3.1) is the home for this.
2. **A recall bench clears the gate.** Run the bench fixtures
(bench-emergent semantic-allusion shapes + the curated set)
with `--backend vec` vs `--backend fts5` vs the planned hybrid;
require **vec-only beats FTS5-only by ≥ 5pp on at least one
semantic-allusion fixture** AND **hybrid lifts STRICT-rate ≥ 5pp
over FTS5-only with no UNGROUNDED-rate regression** (the 5pp
signal floor per `docs/bench-maxing.md`). If hybrid lift is
sub-floor: park here — vec stays opt-in per-call (`--backend
vec`), hybrid does NOT become the default, and the ~2.4 GB tax
isn't paid for nothing.
## 3. Implementation sketch (when the gate clears)
- `arborist/qa/query.py`: after the four FTS5 routes produce ranked
hit lists and before the rerank/filter chain, add a fifth list
from `VecBackend.search(query, limit=K)`. Merge all five by RRF:
`score(doc) = Σ_route 1/(60 + rank_route(doc))`. The merged
ranking then feeds the existing body-coverage `sqrt` rerank,
title-token boost, `_filter_by_title_relevance`, rivalry
exclusion, stem-aware matching, per-source cap, wikitext-base
prose — unchanged. Vec hits with no title overlap survive via the
same accept-path-4 (phrase-match) logic the phrase route uses, or
add a fifth accept path (semantic-similarity above a threshold) —
decide at implementation time from the bench.
- Skip the vec route cleanly when `chunk_vecs` is absent/empty on a
shard (mixed-fleet: some shards backfilled, some not) — falls back
to FTS5-only for that shard.
- Governance: fold the 5 vec hyperparams (embedder name+version,
dim, quant, metric, ann) into `governance_policy_hash` (#000039
§6). A Q&A answer whose context drew on vec hits then carries the
vec config in its cache_key's `governance_policy_hash` dim;
bumping the embedder stales prior vec-consulting cache rows on
lookup. The 8-dim invariant is untouched (existing dim, not a new
one).
- `arborist query` / the bench harness: a `--retrieval-backend
{fts5,hybrid}` knob (default `fts5` until the bench clears
`hybrid`, then flip the default — a fox call + a `governance_policy_hash`
bump).
## 4. Out of scope
- The corpus backfill itself (a batch op, not code — #000039 §14.6;
prerequisite #1 above).
- A GPU/accelerated embedder (drop-in via `default_embedder()`'s
pluggable `Embedder` callable; orthogonal — #000039 §14.6).
- int8 vs float32 default (settled: int8 is the production config
per #000039's head-to-head; switching the v1 default is a
separate `VEC_BACKEND_VERSION` bump, not Phase 2's call).
- binary quantization + two-stage re-rank (#000039 §3.1's
storage-emergency option; its own future ticket if corpus scale
ever makes +0.8 % vs +6 % matter).
## 5. Acceptance criteria
1. `arborist/qa/query.py` merges a fifth (vec) route via RRF, behind
a `--retrieval-backend` knob; FTS5-only behavior unchanged when
the knob says `fts5` or `chunk_vecs` is absent.
2. The 5 vec hyperparams fold into `governance_policy_hash`.
3. Bench data recorded showing the §2 gate cleared (or, if it
didn't, this ticket parks with the bench result documented and
`hybrid` left non-default).
4. Tests: hybrid-merge unit tests (RRF math; vec-absent fallback;
filter chain runs on vec hits); a bench-row in the QA-modes
journal.
## 6. References
- #000039 — the parent (Phase 0 + Phase 1, closed 2026-05-12);
§4 (additive not replacement), §4.2 (RRF design), §8 (the gate),
§14 (ingest integration), §14.6 (embed throughput).
- `docs/bench-maxing.md` — the 5pp signal floor + bench discipline.
- `docs/benchmarks.md` — bench harness orientation.
- #000037 §3.1 — the Prometheus-Σ unconscious-sweep task that's the
natural home for the lazy embed pass that backfills `chunk_vecs`.
- CLAUDE.md "Retrieval pipeline" — the four FTS5 routes vec joins.