docs: lexical-first-rationale.md — why the cheap retrieval path is the default

Capture the positioning fox articulated: arborist's per-document ingest
is ~10-100x cheaper than building a vector-DB representation — same
SQLite substrate, different retrieval philosophy — which is the
difference between "ingest + search runs on a phone" and "the NPU is
now a hand-warmer."

New docs/lexical-first-rationale.md (positioning/architecture
reference, not a ticket): the cost asymmetry with the measured numbers
(FTS5 + SHA-256 leaf + Merkle commit + sqlite + zstd pipeline << 1 ms/
chunk vs bge-small ONNX inference ~5-30 ms/chunk, worse contended; the
query side too — a vec query embeds the query string first, an FTS5
query is B-tree lookups); the same-SQLite-different-philosophy table
(inverted index vs dense vectors + ANN; build cost; query cost;
matching; proof-bearing); the deep version of the point — arborist IS
the Merkle Providence model and that model is cheap by construction,
embeddings are a soft signal (CLAUDE.md "soft hash vs hard hash") that
never enter a proof and are the expensive bolt-on; the edge/mobile
consequence; the honest caveat (lexical-first trades the semantic
allusion gap — which is why vec is opt-in/additive, never the default,
and the embed pass is lazy/out-of-band so the heavy transformer work
runs off-device/off-peak; int8 keeps the storage tax at +6%).

Wired in: TICKETS.md "Distinction from other docs" reference list gains
the doc; #000039 §14.6 gains a "Strategic framing" pointer to it.

(A possible follow-up: fold the mobile-viability argument into the
Merkle Providence Reverse RAG whitepaper proper — noted in the doc's
references; not done here, that's a deliberate cross-repo paper edit.)

Doc-only.
This commit is contained in:
russell@unturf.com 2026-05-12 09:23:50 -04:00
parent 906606072d
commit 1d4d7c549c
No known key found for this signature in database
3 changed files with 131 additions and 0 deletions

View file

@ -82,6 +82,15 @@ proposing change. Examples that stay un-numbered:
exemplars (`pi_star_arithmetic` / `_logic` / `_code` /
`_time_series`). Explicit "doesn't apply to verifier-style
code" caveat.)
- `lexical-first-rationale.md` (positioning/architecture reference
for why lexical-first FTS5 + content-addressed Merkle proofs is
arborist's *default* retrieval and dense-vector semantic search is
an opt-in additive layer: arborist's per-document ingest is
~10-100× cheaper than building a vector representation — same
SQLite substrate, different philosophy — which is the difference
between "ingest + search runs on a phone" and an NPU hand-warmer.
Generalizes the "soft hash vs hard hash" discipline; the numbers
come from #000039 §14.6. Written 2026-05-12.)
If a doc proposes change AND awaits a decision AND has scoped
implementation cost, it's a ticket. Otherwise it's reference.

View file

@ -0,0 +1,114 @@
# Lexical-first retrieval — why the cheap path is the default
**What this is:** the rationale for arborist's retrieval philosophy —
lexical-first (FTS5 BM25 + content-addressed Merkle proofs) as the
default, with dense-vector semantic retrieval as an *optional,
additive* layer (#000039). Not a ticket; a positioning/architecture
reference. The headline: arborist's ingest is **10-100× cheaper per
document** than building a vector-DB representation, on the same
SQLite substrate — which is the difference between "ingest + search
runs on a phone" and "the NPU is now a hand-warmer."
---
## 1. The cost asymmetry, measured
| step | per-chunk cost (healthy CPU core) |
|---|---|
| **bge-small-en-v1.5 ONNX inference** (a vector-DB-style embed) | ~5-30 ms — chunk-length-dependent, 512-token cap; *worse* under load (~250 ms/chunk ≈ 4 chunks/s measured on a contended dev box at load 11/8 cores) |
| arborist's ingest pipeline: chunker (`tok-512-v1`) + SHA-256 leaf hash + Merkle commit + SQLite INSERTs (`documents` / `chunks` / `edges` / `audit_events`) + zstd content compression + FTS5 inverted-index update | **well under 1 ms total** |
So adding embeddings multiplies ingest time by **~10-100×**, entirely
in the ONNX matmuls. The non-vec ingest path runs at hundreds of
chunks/s; a corpus-wide embed of arborist's 6.24 M-chunk wiki shard
set is a one-time hours-on-idle / days-on-contended *batch job*
(#000039 §14.6) — not something you'd do on a phone, ever.
**And the query side is asymmetric too.** An FTS5 query: tokenize,
look up posting lists in the inverted-index B-tree per term, BM25-rank
— a handful of B-tree traversals, microseconds to low-milliseconds. A
vector query: tokenize, run the embedder (**one transformer forward
pass — the same ~5-30 ms as embedding a document chunk**), *then* the
ANN scan. On a phone, every search would pay a transformer forward
pass before it even touches the index.
## 2. Same SQLite file, different philosophy
Both approaches put a `.db` file on disk. The split is in *what's in
it* and *how retrieval works*:
| | lexical-first (arborist) | embedding-first (a vector DB) |
|---|---|---|
| index | FTS5 inverted index (term → posting list) | dense vectors + ANN index (IVF / HNSW / flat) |
| build cost | ~free (sort tokens into B-trees) | a transformer forward pass per chunk |
| query cost | B-tree lookups + BM25 | embed the query (transformer) + ANN scan |
| matching | exact term / phrase / stem | semantic similarity |
| proof-bearing? | **yes** — chunks are content-addressed (`leaf_hash`), Merkle-rooted, the retrieved span is inclusion-provable against `C(M)` | **no** — embeddings are lossy, non-canonical; they can't enter a proof path (CLAUDE.md "soft hash vs hard hash") |
That last row is the deep version of the point. arborist *is* the
Merkle Providence model — content-addressed chunks, Merkle inclusion
proofs, the audit chain — and that model is **cheap by construction**:
the proof-bearing layer (hard hashes, FTS5-retrievable chunks) costs
~nothing to build. Embeddings are a *soft* signal — they help recall,
they never enter a proof — and they're the expensive bolt-on. The
"soft hash vs hard hash" discipline (hard = SHA-256 commitments /
proofs / cache_key; soft = embeddings / TF-IDF / similarity, never in
the proof path) maps directly onto "cheap-and-mandatory vs
expensive-and-optional."
## 3. The consequence: edge / mobile viability
A phone (or a Raspberry Pi, or a permacomputer node) can:
- **Ingest** documents at arborist's full rate — sorting tokens into
B-trees, hashing, writing SQLite rows is what mobile CPUs do well.
- **Query** the lexical index essentially for free — B-tree lookups,
no neural net in the path.
- **Use** a precomputed int8 `chunk_vecs` someone else built
off-device, if a semantic layer is wanted — querying it still costs
a query-embedding forward pass per search, but the *build* cost
(the days-long batch job) was paid elsewhere.
A vector-DB-first design can't do the first two cheaply: ingest pegs
the CPU/NPU per chunk, and even browsing pays a forward pass per
query. On battery, that's the difference between a usable offline
knowledge store and a hand-warmer that's flat by lunch.
## 4. The honest caveat
Lexical-first **trades** the *semantic allusion gap* — the case where
a query and the target chunk share zero stems but mean the same thing
("what did Orwell call the country at war with Oceania?" → an article
that says "Eastasia" / "Eurasia" but never the query's words). FTS5's
four routes (body BM25, title-LIKE, core-keyword TF-IDF, phrase-
pattern) close the *lexical* allusion gap, not the semantic one. Dense
vectors close the semantic gap.
That's exactly why arborist's answer isn't "we don't do vectors" —
it's: **vectors are an opt-in, additive 5th retrieval route, never the
default**, and the embed pass is lazy / out-of-band (a cron, or a
Prometheus-Σ unconscious-sweep task per #000037 §3.1) so the heavy
transformer work runs off-device / off-peak. The cheap, mobile-viable,
proof-bearing path is what ships by default; the semantic layer is
there when you want it and can afford it, and it doesn't compromise
the rest. (`int8` quantization keeps the storage tax at +6% over the
shards — vs +25% at float32 — and the embeddings still never enter the
proof path.)
## 5. References
- `docs/tickets/ticket-000039-sqlite-vec-optional-backend.md` — the
optional vec backend; §3 (storage), §4 (additive not replacement),
§14 (ingest integration + idempotency), §14.6 (the embed-throughput
measurement these numbers come from).
- `docs/tickets/ticket-000050-vec-rrf-hybrid-fusion.md` — Phase 2:
wiring vec as a 5th retrieval route (gated on a corpus backfill + a
recall bench).
- CLAUDE.md "Retrieval pipeline" — the four FTS5 routes; "soft hash vs
hard hash" — the discipline this rationale generalizes.
- `~/git/unfirehose-nextjs-logger/whitepaper/merkle-providence-reverse-rag-whitepaper.rst`
— the canonical Merkle Providence Reverse RAG paper (this doc is the
arborist-side positioning note; folding the mobile-viability argument
into the paper proper is a possible follow-up).
- #000037 §3.1 — the Prometheus-Σ unconscious-sweep task that's the
natural home for the lazy embed pass.

View file

@ -940,3 +940,11 @@ If the backfill latency matters, options are: run it off-peak; use a
dedicated box; or wire GPU/accelerated embedding (out of scope for
this ticket — `default_embedder()` is a pluggable `Embedder` callable,
so a GPU/onnxruntime-gpu/external-endpoint variant is a drop-in).
**Strategic framing:** this ~10-100× ingest-cost asymmetry — and the
matching query-side asymmetry (a vec query embeds the query string
first; an FTS5 query is B-tree lookups) — is *why* arborist is
lexical-first by default and dense-vector is an opt-in additive layer:
the cheap, proof-bearing, mobile-viable path is what ships; the
expensive semantic layer is there when you want it and can afford it.
See `docs/lexical-first-rationale.md`.