arborist/docs/cold-object-store.md
russell@unturf.com d43714a503
cold pack: --jit-blobs mode for online JIT consumer flow
Replaces the batched chunk-pack phase with per-chunk content-addressed
blob uploads to `blobs/<hash[:2]>/<hash[2:]>`. The metadata pack still
ships (small, fast to restore), but consumers no longer have to pull
multi-GB chunk packs to get queryable: `cold unpack --mode just-enough`
+ `ARBORIST_JIT_CHUNKS=1` fetches single chunks on cache miss.

Producer (`_stream_jit_blobs` in evict.py):
- ThreadPoolExecutor with bounded queue (workers*4) keeps memory flat
  across millions of chunks
- HEAD-checks object_size for idempotent re-upload
- Mutually exclusive with chunk packs — manifest's `chunk_pack_hashes`
  is empty in JIT mode (consumer reads that as "JIT-only")

Consumer (`hydrate_doc_jit` in cold_clone.py + `_maybe_jit_hydrate` in
qa/query.py):
- Detects both content shapes that need JIT: NULL (Tier B raw-clone) and
  zeroblob placeholders (just-enough pack restore, per #53). Discriminator
  is first-byte = NUL — zstd-framed bodies start with 0x28, plain UTF-8
  prose never has leading NUL.
- Same placeholder filter applied to chunk-read sites in qa/query.py so
  partial hydrate doesn't surface zero-bytes content into the LLM context.

Test (`TestJitBlobsPackMode` in tests/test_cold_unpack_routed.py):
- End-to-end push → just-enough hydrate → JIT-fetch → content matches
  original byte-for-byte through `unpack_chunk`.

Docs (cold-object-store.md):
- Hard-invariant #1 updated: bucket holds packs by default; `blobs/`
  and `clones/` are opt-in prefixes for the JIT and Tier-A flows.
- New "Three consumer modes" section: full-pack vs JIT-blobs vs raw-clone
  comparison table + operator decision tree.
2026-05-30 07:19:01 -04:00

411 lines
21 KiB
Markdown
Raw Permalink Blame History

This file contains ambiguous Unicode characters

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

# Cold-pack distribution tier (ticket #000061)
A point-in-time corpus distribution mechanism. arborist serializes its
local chunks into `tar.zst` packs, ships them to an S3-compatible bucket
(and/or to local disk for DVD-burning), and any new peer hydrates by
downloading those packs from the bucket's CDN edge and unpacking them
into a fresh shard.
## What this is, and what it is not
**Is:** a backup-and-distribution unit. Pack bytes are content-addressed.
Same chunk set on two writers → same `pack_hash`. The bucket is a
delivery medium for a *delayed* snapshot of the corpus — repackaging
after falsifications produces a new pack with a new hash.
**Is not:** a live mirror. Packs do not see falsifications that happen
*after* the pack was built. They do not see ingests after the pack was
built. They are frozen artifacts, identified by `snapshot_root` of the
corpus state at pack time.
**Is not:** an individual-chunk fetch tier. There is no per-chunk URL in
the bucket — corpus chunks live exclusively inside packs. New peers and
backup consumers download packs whole.
## Hard invariants
1. **Bucket holds packs by default; `blobs/` is opt-in for JIT.** Layout:
```
<bucket>/packs/<pack_hash>.tar.zst # pack body
<bucket>/packs/<pack_hash>.manifest.ndjson # pack contents sidecar
<bucket>/blobs/<hash[:2]>/<hash[2:]> # per-chunk body — opt-in,
# populated only by
# `cold pack --jit-blobs`
# (online-JIT consumer flow)
<bucket>/clones/<snapshot_id>/<NNN>.db # raw shard clones — opt-in,
# populated only by
# `cold stream-snapshot`
# (Tier A raw-clone flow)
```
The classic pack flow (DVD-burn channel) keeps one pack ↔ one disc ↔
one bucket object. The two opt-in prefixes (`blobs/`, `clones/`) light
up additional consumer flows; see "Three consumer modes" below.
Pack contents (v2 format, self-sufficient for new-peer hydration):
```
manifest.ndjson # chunk leaf_hash + size catalog
tables/documents.jsonl # array-per-line, sorted
tables/chunks.jsonl # without content column
tables/merkle_nodes.jsonl
tables/edges.jsonl # FAN-IN restructured
tables/audit_events.jsonl
tables/derivations.jsonl
tables/concept_relations.jsonl
tables/concept_token_idf.jsonl
tables/providence_cache.jsonl
tables/citation_aliases.jsonl
tables/term_aliases.jsonl
tables/snapshots.jsonl
tables/document_http_meta.jsonl
blobs/<hash[:2]>/<hash[2:]> # raw UTF-8 chunk bodies
```
Not shipped (per-peer or rebuildable):
`mesh_*`, `selfmodel_*`, `capital_ledger`, `controller_events`,
`fork_score_branches`, `memory_*`, `adapter_loss_reports`,
`falsifications`, `schema_meta`, `meta`, `chunks_fts*`,
`documents_fts*`.
The FTS5 shadow tables (`chunks_fts*` / `documents_fts*`) are
**derived** from chunk content, so `cold pack` defaults to `--no-fts`
and the consumer regenerates them on hydrate. They CAN be shipped
(`--with-fts`) and restored faithfully (~24 s/shard, no re-tokenize)
for constrained consumers — see "FTS: restore or rebuild".
2. **`pack_hash = hash_leaf(manifest_bytes)`.** The manifest is sorted
by `leaf_hash` and deduped before hashing, so input order and
accidental duplicates don't move the hash. Two writers producing the
same chunk set produce the same `pack_hash` — bucket upload is
idempotent, DVD burns at two sites are byte-identical.
3. **`hash_leaf(chunk_body).hex() == leaf_hash`** is verified on every
`open_pack` member. The pack's tar member name is `blobs/<hash[:2]>/<hash[2:]>` —
that's a within-tar convention, not a bucket layout. Tampering with
pack bytes is caught at unpack time, never reaches the local DB.
4. **Every pack pins a `snapshot_root`.** Pack creation reads the
corpus's current snapshot root (`arborist/snapshot.py:compute_snapshot_root`)
and records it in:
- the audit row (`cold_pack_pushed.body.snapshot_root`)
- the `push_pack` return body
- the local-dir filenames implicitly (pack_hash itself encodes the
manifest, which encodes the chunk set, which encodes that snapshot's
content)
Consumers can run `arborist snapshot verify <root>` after unpack to
detect drift between the pack and the corpus state on the consuming
node.
5. **Cores never evict** (CLAUDE.md rule). Packs include cores AND
surfaces — cores carry the distillation derivations a new peer needs
to bootstrap the v9.8 chain.
6. **No credentials in audit body.** Backend identity is endpoint URL +
bucket name only. Credentials live in env vars / `~/.aws/credentials`
via standard boto3 discovery — Operation Voyeur.
## Delayed snapshots and falsifications
Packs are not live. Between two pack runs, three things can happen:
1. **New ingest.** `ingest_source` adds new documents. They aren't in
the old pack; they show up in the next pack. The old pack stays a
valid snapshot of *its* state.
2. **Falsification.** Drift detection, `arborist falsify`, or
`rehydrate_drift` flips a `providence_cache` row to
`falsification_state='stale'` and/or marks a document for
re-derivation. Chunk content does NOT change (chunks are immutable;
content-addressed). A new pack covers the same chunk *bytes* but with
a different `providence_cache` view.
3. **Re-pack.** A new pack run reads the current corpus and produces a
pack with a new `pack_hash` (because the manifest covers a different
chunk set — newly ingested, possibly with the same hashes minus any
superseded ones).
Three operational consequences:
- **Stale packs accumulate.** Old `pack_hash`es stay in the bucket
until explicitly garbage-collected. They're still valid snapshots
of past corpus states. There's no automatic cleanup; that's a future
ticket.
- **A peer hydrated from an old pack is honestly old.** It has the
corpus state from the pack's `snapshot_root`. To catch up, it
follows the same path any live peer does — ingest new sources,
receive falsification events on the mesh, re-derive cores.
- **The bucket is eventually consistent with intent**, not with the
live corpus. Re-pack cadence (daily? weekly? per-event?) is an
operational policy, not a code property.
## Two distribution channels — same artifact
The same `.tar.zst` file serves two channels:
| Channel | Transport | Default cap |
|---------------------|----------------------------|---------------|
| **Bucket + CDN** | `S3CompatibleBackend.put_pack` → public-read DO Spaces / R2 / S3, CDN edge serves consumers | 4.4 GB / pack |
| **DVD-R archival** | `--local-dir DIR` → `growisofs -dvd-compat -Z /dev/sr0=<pack>` | 4.4 GB / pack |
Pack files are byte-identical between channels. A DVD burned from one
local-dir pack and a CDN-fetched pack of the same content collide on
`sha256sum`.
## Three consumer modes
A fresh peer has three ways to hydrate from the bucket; producer mode
decides which the bucket supports.
| mode | producer | consumer | local disk | bucket bytes pulled | first-query latency |
|------|---------------------|-------------------------------------------|-------------------|---------------------|---------------------|
| **full pack** | `cold pack` (default) | `cold unpack --mode full` | full corpus (~35 GB) | metadata + every chunk pack | seconds (all data local) |
| **JIT-blobs** | `cold pack --jit-blobs` | `cold unpack --mode just-enough` + queries with `ARBORIST_JIT_CHUNKS=1` | metadata-only (~27 GB at FTS+schema floor) | metadata + N×blob per query | one round-trip per chunk on first hit, cached after |
| **raw clone (Tier A)** | `cold stream-snapshot` | `cold clone` | full corpus (~35 GB) | raw `.db` files via SQLite Backup API | seconds (all data local) |
**full-pack mode** is the canonical path: hardened, fully tested, FTS
restorable or rebuildable. Use for genesis recovery of a self-hosting
peer, DVD-archival workflows, and any consumer that wants the full
corpus offline-queryable.
**JIT-blobs mode** trades steady-state download for fastest time-to-
queryable. The metadata pack is small; chunk bodies stream in as
queries access them. Good for ephemeral nodes, demo VMs, edge servers
that only serve a subset of the corpus. The `blobs/<hash[:2]>/<hash[2:]>`
prefix is content-addressed: a chunk that exists in two snapshots costs
one bucket object, and producer reuploads are free idempotent no-ops.
Failure mode is graceful — `hydrate_doc_jit` skips missing blobs and
hash-mismatches rather than crashing the query.
**raw-clone mode** is the simplest. `cold stream-snapshot` runs the
SQLite Backup API over each shard to a tempfile and PUTs the bytes; the
consumer DOWNLOAD the `.db` files and is queryable immediately. No
compression, no restore phase. Larger bucket footprint than the pack
flow but skips the slow metadata-restore step entirely.
```bash
# Producer (JIT-blobs mode):
arborist cold pack --jit-blobs --shards-dir ~/.arborist/shards
# Consumer:
arborist cold unpack --mode just-enough <metadata_pack_hash> \
--hydrate-shards-dir ~/.arborist/shards --hydrate-M 4
export ARBORIST_JIT_CHUNKS=1
arborist query --shards-dir ~/.arborist/shards "what is X?"
# → first hit on each chunk fetches blobs/<hash> from the bucket and
# caches it back into the local row; subsequent reads are free.
```
## DO Spaces quickstart
```bash
# 1. Install the optional backend.
make bootstrap-object-store
# 2. Set boto3 standard env vars (never hard-code in scripts).
export AWS_ACCESS_KEY_ID=<your-spaces-key>
export AWS_SECRET_ACCESS_KEY=<your-spaces-secret>
# 3. Set bucket config.
export ARBORIST_COLD_ENDPOINT_URL=https://nyc3.digitaloceanspaces.com
export ARBORIST_COLD_BUCKET=arborist-corpus
# 4. Build packs and push them. Default cap = 4.4 GB / pack (DVD-R safe-
# fit). One shard typically yields 1-3 packs.
make cold-pack
# 5. Confirm what's in the bucket.
make cold-stats
```
Same flow works on AWS S3 (`endpoint_url=https://s3.<region>.amazonaws.com`),
Cloudflare R2, Backblaze B2, GCS S3-interop, MinIO.
## Hydrating a new peer from CDN
Genesis recovery is **M-aware** — every document routes to
`shard_for_document(root, M)`, the same deterministic function the
producer used — and is driven by one make target:
```bash
# On the fresh node: install arborist + the [object-store] extra, set
# the bucket env (ARBORIST_COLD_BUCKET / _ENDPOINT_URL) + boto3 creds,
# then hydrate the full corpus into M target shards:
make cold-hydrate HYDRATE_DIR=~/.arborist/shards HYDRATE_M=4
```
What it does, in order:
1. **Pull + restore metadata and chunk content — serially.** Each
metadata pack routes its rows into *all M shared target shards*, so
parallel workers contend on the same files (and an unbounded restore
transaction blows up RAM). The hydrate runs **serial by default**
(`COLD_PACK_JOBS=1`; override at your own risk). The restore is
bulk-tuned — 512 MB page cache applied *before* the heavy edge
fan-out, plus bounded incremental commits — so it runs in
~820 min/pack, not the hours an untuned restore took (the edge
fan-out into an indexed table thrashes a default ~2 MB cache).
2. **Get a searchable FTS index.** If the packs shipped FTS, it was
restored during step 1 (fast); otherwise rebuild from content in
parallel (`cold rebuild-fts`, one process per shard). This is
`cold verify`-gated, so the rebuild only runs when the index isn't
already searchable. See "FTS: restore or rebuild" below.
3. **Self-verify** (`cold verify`). Every non-empty shard must have
materialized chunk content (not zero-filled placeholders) **and** a
searchable FTS index. A bad recovery exits non-zero here — it fails
**loudly** instead of silently serving empty results. Run it any
time: `arborist cold verify --shards-dir ~/.arborist/shards`.
The publisher's pinned `snapshot_root` is in each pack's audit row; the
consumer recomputes it after unpack and they match on a clean restore.
### FTS: restore or rebuild (both work)
FTS is a deterministic function of chunk content, so a fresh peer has
two ways to get a searchable index after hydrate. Measured 2026-05-29 on
the live 4-shard / 6.2M-chunk corpus:
| path | extra download | consumer work | result |
|----------------------------|----------------|-------------------------|---------------|
| restore prebuilt FTS pack | +4.76 GB | ~24 s/shard | working index |
| rebuild from content | none | ~5 min (4-way parallel) | working index |
Both produce a byte-identical, queryable index (identical `MATCH`
counts). Restore skips re-tokenization (~12× less consumer CPU) but adds
~20 % to the download (FTS packs were 4.76 GB of a 23.8 GB bucket). So:
- **Bandwidth-sensitive / capable consumer** (the producer's own hosts):
drop FTS packs — `cold pack --no-fts` (the default) — and rebuild on
hydrate.
- **Constrained consumer** (slow CPU, mobile / SPV peer): ship FTS packs
— `cold pack --with-fts` — and restore them, skipping the rebuild.
`make cold-hydrate` picks automatically: it runs `cold verify` and only
rebuilds when the restored index isn't already searchable. So shipping
FTS packs makes recovery fast; dropping them makes the bucket small;
either way the result is verified before the run reports success.
**Faithful FTS5 transport (the headless-index bug, fixed 2026-05-29).**
An FTS5 index is not a plain row set — its shadow tables (`_data`,
`_idx`, `_docsize`, `_config`) are mutually-dependent internal state, and
`_data` **rowid 1 is the segment "structure" record** FTS5 reads first to
learn which segments exist. The restore copies those shadow tables
verbatim (correct in principle), but `CREATE VIRTUAL TABLE` first seeds
an *empty* structure record at id 1, and the old restore used
`INSERT OR IGNORE` — so the pack's real structure record lost the PK
conflict and was silently dropped, leaving an empty "0 segments" header
over a full body of orphaned segments (`count(*)` via `_docsize` looked
right; every `MATCH` returned 0). **Fix:** clear the seeded rows then
copy verbatim (`DELETE FROM main.<shadow>; INSERT … SELECT * FROM
fts_src.<shadow>`), so each shadow table becomes a byte-for-byte copy of
the producer's index and the real structure record survives.
## DVD-R archival workflow
```bash
# 1. Write packs to a staging dir; skip the bucket entirely.
make cold-pack-dvd LOCAL_DIR=/mnt/dvd-staging
# 2. Each pack is one disc. Burn with growisofs.
for pack in /mnt/dvd-staging/arborist-pack-*.tar.zst; do
growisofs -dvd-compat -Z /dev/sr0="$pack"
# ... eject, insert next blank, repeat ...
done
# 3. On a fresh node, copy a pack from disc and unpack:
mount /dev/sr0 /mnt/dvd
arborist --db fresh.db cold unpack \
"$(basename /mnt/dvd/arborist-pack-*.tar.zst .tar.zst | cut -d- -f3)"
```
The `pack_hash` is in the filename (`arborist-pack-<hash[:16]>.tar.zst`)
so the disc itself is self-describing — no separate index needed.
## Pack-size cap — fit on a 4.7 GB DVD-R, safely
Default cap is **4,400,000,000 bytes (4.4 GB, ~6.5 % buffer below the
4.7 GB marketing capacity)**. Targeting 4.7 GB directly is unsafe:
filesystem overhead, media manufacturing variance, growisofs
lead-in/lead-out, and older drives refusing the outer edge all eat
into nominal capacity. 4.4 GB sits between the industry-standard tool
defaults (HandBrake DVD-5 = 4,377 MiB ≈ 4.59 GB; DVDFab fit-to-DVD-5 =
4.3 GB; mkisofs default DVD = 4,377 MiB).
The cap applies to *compressed* bytes per pack. `stream_packs` uses
streaming zstd compression and peeks the compressed-buffer size after
every chunk (via `FLUSH_BLOCK`, which preserves the compressor's
dictionary so block boundaries cost almost nothing in ratio). When the
buffer reaches the cap, the pack is finalized and a new one starts. So
each disc fills to ~4.4 GB of recorded data, not 3050 % of capacity.
Overshoot bound: tar trailer (~1 KB padding) + zstd frame footer (~10 B)
get emitted after the last in-loop size check, so actual compressed
size can land at cap + ~2 KB. Trivial for a 4.4 GB cap.
For larger media:
| Media | `--max-pack-bytes` | Marketing |
|----------------------|---------------------------|-----------|
| **DVD-R (default)** | `4_400_000_000` (4.4 GB) | 4.7 GB |
| DVD+R DL | `8_000_000_000` (8.0 GB) | 8.5 GB |
| BD-R | `24_000_000_000` (24 GB) | 25 GB |
| BD-R DL | `48_000_000_000` (48 GB) | 50 GB |
## Cost model (DO Spaces, current corpus)
Pack format v2 (self-sufficient for new-peer hydration). Numbers measured
2026-05-26 against the live 4-shard corpus (14.1M total chunks; the
1.56M hot-content chunks per shard go into packs; metadata is added on
top via the v2 dump path):
| Path | Count | Storage | Cost |
|-------------------------------|-------------|----------|-------------------------------|
| v2 pack storage (per shard) | 1 pack | ~2.1 GB | — |
| v2 pack storage (all 4) | 4 packs | ~8.5 GB | $0.17/mo (@ $0.02/GB) |
| Full-corpus hydrate (CDN) | ~4 GETs | — | ~$0.00002 in requests |
| Egress (in-region) | 0 | — | $0 |
| Egress (CDN to public) | 8.5 GB/peer | — | $0.09 per fresh peer (@ $0.01/GB) |
The v1 chunks-only format produced ~1.78 GB per shard (7.1 GB total).
v2 adds ~0.30.4 GB per shard for the metadata tables (chunks-meta,
documents, audit_events, merkle_nodes, edges fan-in restructured, plus
small tables). Trade: ~20 % more storage for a self-sufficient pack
that a fresh peer can unpack into a working shard with no other inputs.
Repacking after a falsification event costs the same as the initial
pack — one full corpus serialization per event-batched run, gated by
re-pack cadence (operational policy).
## Failure modes
| Symptom | Cause | Recovery |
|------------------------------------------|------------------------------------|--------------------------------------|
| `pack chunk hash mismatch` on unpack | Pack bytes corrupted in transit or on disc | Re-download / re-burn; pack is content-addressed so a fresh fetch is verifiable. |
| `cold pack` produces no packs | No hot chunks with non-null content | `cold pack` operates on local content. Confirm shard isn't empty / fully evicted. |
| Peer's snapshot_root differs from pack's | Local corpus drifted after unpack (ingest, falsification, etc.) | Expected. Pack is a delayed snapshot; the peer has moved on. Re-pack to re-baseline. |
| Bucket missing a pack | GC'd, never uploaded, wrong bucket | Re-build pack from any shard that still has the source content. |
| `cold verify` fails: content zero-filled | Phase-2 chunk-body fill never completed (crashed/interrupted restore, or `--just-enough` mode) | Re-run `make cold-hydrate` (serial, full mode). Pre-sized zeroblob placeholders are present but never overwritten. |
| `cold verify` fails: FTS dead (`MATCH`=0) | No fts pack shipped and the rebuild didn't run (or a pre-2026-05-29 headless-restore) | `arborist cold rebuild-fts --shards-dir DIR` (rebuilds from content). |
| Search empty despite docs present | Dead / missing FTS index | `cold verify` to confirm, then `cold rebuild-fts`. |
## Future work
- **Multipart upload for packs.** Provider single-object limits (DO
Spaces = 5 GB non-multipart, AWS S3 = 5 GB; both support multipart up
to 5 TB). Today's code uses `put_object` which is single-shot. boto3
`upload_file` is the one-line drop-in.
- **Streaming pack builder.** ✅ Landed as `stream_packs`. Caps target
compressed bytes; each disc fills. `build_pack` stays for tests +
small/known-set callers.
- **Pack GC.** Stale packs (those whose `snapshot_root` is older than N
re-pack cycles) get bucket-deleted automatically.
- **Range-fetch partial pack pulls.** Manifest carries offsets;
`GET .tar.zst Range: bytes=X-Y` would let a consumer pull one chunk
from a huge pack without downloading the whole thing.
- **KMS / SSE-S3.** Server-side encryption (mesh ciphertext on a
public bucket is the v1 confidentiality path).
- **Multi-region replication.** Handled by the provider within a region;
cross-provider replication is a separate distribution-policy question.