feat: #000061 cold-pack distribution tier (boto3 S3-compat + DVD-R safe-fit)
Ship arborist corpus state to new peers and DVD-R archival via
point-in-time tar.zst packs. One artifact serves both channels —
bucket+CDN delivery and physical-media archival.
Bucket holds packs only. Pack key = hash_leaf(manifest_bytes), so same
chunk set on two writers produces the same pack_hash and upload is
idempotent. Each pack pins the corpus snapshot_root it covers in audit
+ result body — packs are delayed snapshots, not live mirrors;
falsifications between repacks produce new pack_hashes.
stream_packs runs streaming zstd over tarfile, peeking compressed-buffer
size after each chunk via FLUSH_BLOCK (preserves dictionary). Default
cap 4_400_000_000 — 4.4 GB DVD-R safe-fit, ~6.5% buffer below the
4.7 GB marketing capacity to absorb ISO9660 overhead, growisofs
lead-in/lead-out, media variance, and drive-edge refusal. Each disc
fills to ~4.4 GB recorded data, not the ~1.5 GB an uncompressed cap
produced.
One backend class (S3CompatibleBackend via boto3 + endpoint_url) covers
AWS S3, DO Spaces, R2, B2, GCS S3-interop, MinIO. Optional dep
[object-store] = boto3>=1.34; dev extras pull moto for the wire test.
Voyeur: credentials via AWS_ACCESS_KEY_ID/_SECRET_ACCESS_KEY env or
~/.aws/credentials, never printed; only endpoint URL + bucket name
surface in logs.
CLI: arborist cold {pack,unpack,stats}. Makefile: cold-pack,
cold-pack-dvd (local-dir output for growisofs), cold-unpack, cold-stats.
Sizing for current shards (14.1M chunks, ~17 GB compressed): ~4 packs
at the default cap, ~\$0.34/mo DO Spaces storage, ~\$0.0001/fresh-peer
hydrate.
Always-on raw-UTF-8 leaf store (per ticket "Hard invariants") deferred
— packs-only for now, backfill later.
2557 passed, 28 skipped, 1 xfailed.
This commit is contained in:
parent
06e6c7a918
commit
727cb1bd96
11 changed files with 1911 additions and 4 deletions
|
|
@ -111,6 +111,7 @@ Newest first. Update on every open/close.
|
|||
|
||||
| ID | Title | Status | Opened | Directive |
|
||||
|----------|------------------------------------------------|-----------------------|------------|-----------|
|
||||
| #000061 | Cold-pack distribution tier (boto3 S3-compat, DO Spaces + DVD-R targets) | **in progress** — opened 2026-05-25 (fox: "implement it now … target digital ocean first as a test"; later "I wanted a way to hydrate using tarballs (the core and important data) for bringing new machines up"). Tarball-only distribution mechanism — bucket holds `tar.zst` packs keyed by `hash_leaf(manifest)`, no individual-chunk blobs. New peers hydrate by downloading packs from the bucket's CDN edge (~4 HTTPS GETs for the current ~14.1M-chunk corpus, packs filled to 4.4 GB compressed each via streaming zstd, vs ~14M for individual blobs). Same artifact ≤4.4 GB safe-fit (~6.5 % buffer below DVD-R's 4.7 GB marketing capacity, accommodating ISO9660 overhead + media variance + drive-edge refusal) burns directly to physical media via `--local-dir` + `growisofs`. Packs are *delayed* snapshots: each pack pins the corpus `snapshot_root` it covers in audit + body, so falsifications between repacks produce new pack_hashes and stale packs stay in the bucket until explicit GC (future ticket). Packs include cores AND surfaces (full-corpus hydration). Default selection covers every hot chunk with local content in the shard. Multi-pack splitting via `stream_packs` (streaming zstd, FLUSH_BLOCK peek of compressed buffer after each chunk, cut at cap) fills each disc to ~4.4 GB compressed instead of leaving ~50% empty. One backend class (`S3CompatibleBackend` via boto3 + `endpoint_url`) covers AWS S3, DO Spaces, GCS S3-interop, R2, B2, MinIO. CDN public-read makes packs accessible to anyone; hash binding via in-tar `leaf_hash` member names makes hostile-bucket scenarios safe. Optional dep `[object-store]` = boto3>=1.34. Voyeur: credentials via standard `AWS_ACCESS_KEY_ID`/`_SECRET_ACCESS_KEY` (env or `~/.aws/credentials`), never printed; only endpoint URL + bucket name surface in logs. Initial individual-blob path (per-chunk S3 objects) was scoped+landed then **deleted same day** (fox: "what ever was blobs? I wanted a way to hydrate using tarballs"); the five-step deletion record lives inline in the doc — we'd added 14M-object storage and ~$70/hydrate request cost for a workflow that needed neither. Sizing math for current shards: ~4 packs total (17.2 GB compressed ÷ 4.4 GB compressed per pack), ~17 GB bucket storage, ~$0.34/mo DO Spaces. | 2026-05-25 | — |
|
||||
| #000060 | H-ABCDEFG same-model substrate-delta harness (+ jaggedness tensor + curvature) | open · awaiting go/no-go (2026-05-20; from Dav1dPrometheus *Protocol-Layer AGI* working report §26/§50/§82-84). The report's "decisive proof": run the SAME base model substrate-OFF vs substrate-ON over long-horizon/adversarial/non-jagged batteries, report the delta. Two new metrics: jaggedness tensor `J_norm` (§73 — variance across nearby variants, normalized by difficulty) + discrete performance curvature `κ_t` (§5.2, with the honest no-global-convexity bound, Erratum 5). A-vs-C spine (B optional, D=mesh OUT → #000012/#000016). Curvature-aware ForkScore extension folds into **#000012** (NOT a new ticket — reserved `iota`/`kappa` weight slots already exist). Budget: control arms = Hermes/Qwen, never Opus without go; heavy passes on GPU box. Held-out/mechanism-agnostic variants required so ABCDEFG doesn't self-validate. | 2026-05-20 | — |
|
||||
| #000059 | Admission discipline: claim-graveyard query + self-providence quarantine | open · awaiting go/no-go (2026-05-20; Dav1dPrometheus report §11 Priority 2 + §57 "admissible state transition" thesis). Two coupled write-path mechanisms. **(A) GraveyardCheck:** storage already exists (`falsification_state` failed/stale/quarantined records ARE the graveyard); the gap is burden-shifting — a re-asked claim family with a known `failed` history should require stronger evidence (§42.9/§62-63). **(B) Self-providence quarantine:** `make ingest-self-providence` (Makefile:769) deliberately promotes STRICT records into the corpus — the exact self-confirmation loop §70 warns of — and ships with NO guard; detect/lineage-tag self-providence-descended evidence + quarantine for high-impact claims. Both advisory-sidecar-first (run-DAG only, never `providence_cache`/`audit_events`); demote-hooks bench-gated + unwired pending net win (would fold `governance_policy_hash`). Soft signals never enter the hard proof path. **Bounded-ingestion hard constraint (fox 2026-05-20, §7):** the graveyard MUST reach a steady-state size ∝ the *recurring*-error surface, never queries-ever — earn-to-enter (recurrence-gated), fingerprints not transcripts (UTXO-set analogy), decay/compact (evicts like a surface), off the hot path. BTC's lesson is bounded self-regulating ingestion, not "store everything." Gossip-group falsifier admission inherits difficulty-adjusted stable-rate + per-window budget (#000036) → enforced in #000012/`mesh/`. If it can't be bounded, it isn't built. | 2026-05-20 | — |
|
||||
| #000058 | `cache_key_9` verifier-policy: mandatory-vs-legible decision + doc reconciliation | open · awaiting go/no-go (2026-05-20; Dav1dPrometheus report §2 Erratum 1 / §11 Priority 1 "mandatory cache_key_9"). **Five-step #1 correction:** the report's *correctness* premise is already false in arborist — verifier fields are a subset of the policy dict and so already fold into `governance_policy_hash` (`keys.py:269-275`); a verifier-rule change ALREADY changes the cache_key today. The explicit 9th `verifier_policy_hash` buys **audit legibility**, not correctness — so "mandatory" would stale every prior record for zero correctness gain. Decision: A leave-as-is (8-dim default, 9th optional) + doc reconcile [recommended] · B default-write 9-dim · C flag-staged bench-gated default-write — never a hard mandatory flip. Doc reconcile (CLAUDE.md/concepts "8-dim" → "8 + optional legible 9th") is the do-regardless. | 2026-05-20 | — |
|
||||
|
|
@ -174,4 +175,4 @@ Newest first. Update on every open/close.
|
|||
|
||||
## Next ID
|
||||
|
||||
`000061`
|
||||
`000062`
|
||||
|
|
|
|||
259
docs/cold-object-store.md
Normal file
259
docs/cold-object-store.md
Normal file
|
|
@ -0,0 +1,259 @@
|
|||
# Cold-pack distribution tier (ticket #000061)
|
||||
|
||||
A point-in-time corpus distribution mechanism. arborist serializes its
|
||||
local chunks into `tar.zst` packs, ships them to an S3-compatible bucket
|
||||
(and/or to local disk for DVD-burning), and any new peer hydrates by
|
||||
downloading those packs from the bucket's CDN edge and unpacking them
|
||||
into a fresh shard.
|
||||
|
||||
## What this is, and what it is not
|
||||
|
||||
**Is:** a backup-and-distribution unit. Pack bytes are content-addressed.
|
||||
Same chunk set on two writers → same `pack_hash`. The bucket is a
|
||||
delivery medium for a *delayed* snapshot of the corpus — repackaging
|
||||
after falsifications produces a new pack with a new hash.
|
||||
|
||||
**Is not:** a live mirror. Packs do not see falsifications that happen
|
||||
*after* the pack was built. They do not see ingests after the pack was
|
||||
built. They are frozen artifacts, identified by `snapshot_root` of the
|
||||
corpus state at pack time.
|
||||
|
||||
**Is not:** an individual-chunk fetch tier. There is no per-chunk URL in
|
||||
the bucket — corpus chunks live exclusively inside packs. New peers and
|
||||
backup consumers download packs whole.
|
||||
|
||||
## Hard invariants
|
||||
|
||||
1. **Bucket holds packs only.** Layout:
|
||||
|
||||
```
|
||||
<bucket>/packs/<pack_hash>.tar.zst # pack body
|
||||
<bucket>/packs/<pack_hash>.manifest.ndjson # pack contents sidecar
|
||||
```
|
||||
|
||||
No `blobs/` prefix, no per-chunk objects. (One pack ↔ one disc ↔ one
|
||||
bucket object.)
|
||||
|
||||
2. **`pack_hash = hash_leaf(manifest_bytes)`.** The manifest is sorted
|
||||
by `leaf_hash` and deduped before hashing, so input order and
|
||||
accidental duplicates don't move the hash. Two writers producing the
|
||||
same chunk set produce the same `pack_hash` — bucket upload is
|
||||
idempotent, DVD burns at two sites are byte-identical.
|
||||
|
||||
3. **`hash_leaf(chunk_body).hex() == leaf_hash`** is verified on every
|
||||
`open_pack` member. The pack's tar member name is `blobs/<hash[:2]>/<hash[2:]>` —
|
||||
that's a within-tar convention, not a bucket layout. Tampering with
|
||||
pack bytes is caught at unpack time, never reaches the local DB.
|
||||
|
||||
4. **Every pack pins a `snapshot_root`.** Pack creation reads the
|
||||
corpus's current snapshot root (`arborist/snapshot.py:compute_snapshot_root`)
|
||||
and records it in:
|
||||
- the audit row (`cold_pack_pushed.body.snapshot_root`)
|
||||
- the `push_pack` return body
|
||||
- the local-dir filenames implicitly (pack_hash itself encodes the
|
||||
manifest, which encodes the chunk set, which encodes that snapshot's
|
||||
content)
|
||||
|
||||
Consumers can run `arborist snapshot verify <root>` after unpack to
|
||||
detect drift between the pack and the corpus state on the consuming
|
||||
node.
|
||||
|
||||
5. **Cores never evict** (CLAUDE.md rule). Packs include cores AND
|
||||
surfaces — cores carry the distillation derivations a new peer needs
|
||||
to bootstrap the v9.8 chain.
|
||||
|
||||
6. **No credentials in audit body.** Backend identity is endpoint URL +
|
||||
bucket name only. Credentials live in env vars / `~/.aws/credentials`
|
||||
via standard boto3 discovery — Operation Voyeur.
|
||||
|
||||
## Delayed snapshots and falsifications
|
||||
|
||||
Packs are not live. Between two pack runs, three things can happen:
|
||||
|
||||
1. **New ingest.** `ingest_source` adds new documents. They aren't in
|
||||
the old pack; they show up in the next pack. The old pack stays a
|
||||
valid snapshot of *its* state.
|
||||
|
||||
2. **Falsification.** Drift detection, `arborist falsify`, or
|
||||
`rehydrate_drift` flips a `providence_cache` row to
|
||||
`falsification_state='stale'` and/or marks a document for
|
||||
re-derivation. Chunk content does NOT change (chunks are immutable;
|
||||
content-addressed). A new pack covers the same chunk *bytes* but with
|
||||
a different `providence_cache` view.
|
||||
|
||||
3. **Re-pack.** A new pack run reads the current corpus and produces a
|
||||
pack with a new `pack_hash` (because the manifest covers a different
|
||||
chunk set — newly ingested, possibly with the same hashes minus any
|
||||
superseded ones).
|
||||
|
||||
Three operational consequences:
|
||||
|
||||
- **Stale packs accumulate.** Old `pack_hash`es stay in the bucket
|
||||
until explicitly garbage-collected. They're still valid snapshots
|
||||
of past corpus states. There's no automatic cleanup; that's a future
|
||||
ticket.
|
||||
- **A peer hydrated from an old pack is honestly old.** It has the
|
||||
corpus state from the pack's `snapshot_root`. To catch up, it
|
||||
follows the same path any live peer does — ingest new sources,
|
||||
receive falsification events on the mesh, re-derive cores.
|
||||
- **The bucket is eventually consistent with intent**, not with the
|
||||
live corpus. Re-pack cadence (daily? weekly? per-event?) is an
|
||||
operational policy, not a code property.
|
||||
|
||||
## Two distribution channels — same artifact
|
||||
|
||||
The same `.tar.zst` file serves two channels:
|
||||
|
||||
| Channel | Transport | Default cap |
|
||||
|---------------------|----------------------------|---------------|
|
||||
| **Bucket + CDN** | `S3CompatibleBackend.put_pack` → public-read DO Spaces / R2 / S3, CDN edge serves consumers | 4.4 GB / pack |
|
||||
| **DVD-R archival** | `--local-dir DIR` → `growisofs -dvd-compat -Z /dev/sr0=<pack>` | 4.4 GB / pack |
|
||||
|
||||
Pack files are byte-identical between channels. A DVD burned from one
|
||||
local-dir pack and a CDN-fetched pack of the same content collide on
|
||||
`sha256sum`.
|
||||
|
||||
## DO Spaces quickstart
|
||||
|
||||
```bash
|
||||
# 1. Install the optional backend.
|
||||
make bootstrap-object-store
|
||||
|
||||
# 2. Set boto3 standard env vars (never hard-code in scripts).
|
||||
export AWS_ACCESS_KEY_ID=<your-spaces-key>
|
||||
export AWS_SECRET_ACCESS_KEY=<your-spaces-secret>
|
||||
|
||||
# 3. Set bucket config.
|
||||
export ARBORIST_COLD_ENDPOINT_URL=https://nyc3.digitaloceanspaces.com
|
||||
export ARBORIST_COLD_BUCKET=arborist-corpus
|
||||
|
||||
# 4. Build packs and push them. Default cap = 4.4 GB / pack (DVD-R safe-
|
||||
# fit). One shard typically yields 1-3 packs.
|
||||
make cold-pack
|
||||
|
||||
# 5. Confirm what's in the bucket.
|
||||
make cold-stats
|
||||
```
|
||||
|
||||
Same flow works on AWS S3 (`endpoint_url=https://s3.<region>.amazonaws.com`),
|
||||
Cloudflare R2, Backblaze B2, GCS S3-interop, MinIO.
|
||||
|
||||
## Hydrating a new peer from CDN
|
||||
|
||||
```bash
|
||||
# 1. On the fresh node, install arborist + the [object-store] extra.
|
||||
make bootstrap-object-store
|
||||
|
||||
# 2. List packs the publisher made available.
|
||||
ARBORIST_COLD_ENDPOINT_URL=... ARBORIST_COLD_BUCKET=... \
|
||||
arborist cold stats
|
||||
|
||||
# 3. For each pack, unpack into a local shard. Verifies every chunk on
|
||||
# the way in; bad bytes from a hostile CDN never reach the DB.
|
||||
for hash in <pack-hashes>; do
|
||||
arborist --db ~/.arborist/shards/000.db cold unpack $hash
|
||||
done
|
||||
|
||||
# 4. (Optional) Pin which corpus state we're at.
|
||||
arborist --db ~/.arborist/shards/000.db snapshot list | head -1
|
||||
```
|
||||
|
||||
The snapshot_root the publisher pinned at pack time is in the audit row;
|
||||
the verifier on the consumer side recomputes `snapshot_root` after
|
||||
unpack and they should match if the corpus is a clean restore.
|
||||
|
||||
## DVD-R archival workflow
|
||||
|
||||
```bash
|
||||
# 1. Write packs to a staging dir; skip the bucket entirely.
|
||||
make cold-pack-dvd LOCAL_DIR=/mnt/dvd-staging
|
||||
|
||||
# 2. Each pack is one disc. Burn with growisofs.
|
||||
for pack in /mnt/dvd-staging/arborist-pack-*.tar.zst; do
|
||||
growisofs -dvd-compat -Z /dev/sr0="$pack"
|
||||
# ... eject, insert next blank, repeat ...
|
||||
done
|
||||
|
||||
# 3. On a fresh node, copy a pack from disc and unpack:
|
||||
mount /dev/sr0 /mnt/dvd
|
||||
arborist --db fresh.db cold unpack \
|
||||
"$(basename /mnt/dvd/arborist-pack-*.tar.zst .tar.zst | cut -d- -f3)"
|
||||
```
|
||||
|
||||
The `pack_hash` is in the filename (`arborist-pack-<hash[:16]>.tar.zst`)
|
||||
so the disc itself is self-describing — no separate index needed.
|
||||
|
||||
## Pack-size cap — fit on a 4.7 GB DVD-R, safely
|
||||
|
||||
Default cap is **4,400,000,000 bytes (4.4 GB, ~6.5 % buffer below the
|
||||
4.7 GB marketing capacity)**. Targeting 4.7 GB directly is unsafe:
|
||||
filesystem overhead, media manufacturing variance, growisofs
|
||||
lead-in/lead-out, and older drives refusing the outer edge all eat
|
||||
into nominal capacity. 4.4 GB sits between the industry-standard tool
|
||||
defaults (HandBrake DVD-5 = 4,377 MiB ≈ 4.59 GB; DVDFab fit-to-DVD-5 =
|
||||
4.3 GB; mkisofs default DVD = 4,377 MiB).
|
||||
|
||||
The cap applies to *compressed* bytes per pack. `stream_packs` uses
|
||||
streaming zstd compression and peeks the compressed-buffer size after
|
||||
every chunk (via `FLUSH_BLOCK`, which preserves the compressor's
|
||||
dictionary so block boundaries cost almost nothing in ratio). When the
|
||||
buffer reaches the cap, the pack is finalized and a new one starts. So
|
||||
each disc fills to ~4.4 GB of recorded data, not 30–50 % of capacity.
|
||||
|
||||
Overshoot bound: tar trailer (~1 KB padding) + zstd frame footer (~10 B)
|
||||
get emitted after the last in-loop size check, so actual compressed
|
||||
size can land at cap + ~2 KB. Trivial for a 4.4 GB cap.
|
||||
|
||||
For larger media:
|
||||
|
||||
| Media | `--max-pack-bytes` | Marketing |
|
||||
|----------------------|---------------------------|-----------|
|
||||
| **DVD-R (default)** | `4_400_000_000` (4.4 GB) | 4.7 GB |
|
||||
| DVD+R DL | `8_000_000_000` (8.0 GB) | 8.5 GB |
|
||||
| BD-R | `24_000_000_000` (24 GB) | 25 GB |
|
||||
| BD-R DL | `48_000_000_000` (48 GB) | 50 GB |
|
||||
|
||||
## Cost model (DO Spaces, current corpus)
|
||||
|
||||
Numbers from the live shard estimator (4 shards × ~3.5M chunks each,
|
||||
14.1M chunks total, ~17 GB compressed; streaming cap fills each pack
|
||||
to ~4.4 GB compressed):
|
||||
|
||||
| Path | Count | Storage | Cost |
|
||||
|-----------------------------|-------------|----------|---------------------|
|
||||
| Bucket pack storage | ~4 packs | ~17 GB | $0.34/mo (@ $0.02/GB) |
|
||||
| Full-corpus hydrate (CDN) | ~4 GETs | — | ~$0.00002 in requests |
|
||||
| Egress (in-region) | 0 | — | $0 |
|
||||
| Egress (CDN to public) | 17 GB / peer | — | $0.17 per fresh peer (@ $0.01/GB) |
|
||||
|
||||
Repacking after a falsification event costs the same as the initial
|
||||
pack — one full corpus serialization per event-batched run, gated by
|
||||
re-pack cadence (operational policy).
|
||||
|
||||
## Failure modes
|
||||
|
||||
| Symptom | Cause | Recovery |
|
||||
|------------------------------------------|------------------------------------|--------------------------------------|
|
||||
| `pack chunk hash mismatch` on unpack | Pack bytes corrupted in transit or on disc | Re-download / re-burn; pack is content-addressed so a fresh fetch is verifiable. |
|
||||
| `cold pack` produces no packs | No hot chunks with non-null content | `cold pack` operates on local content. Confirm shard isn't empty / fully evicted. |
|
||||
| Peer's snapshot_root differs from pack's | Local corpus drifted after unpack (ingest, falsification, etc.) | Expected. Pack is a delayed snapshot; the peer has moved on. Re-pack to re-baseline. |
|
||||
| Bucket missing a pack | GC'd, never uploaded, wrong bucket | Re-build pack from any shard that still has the source content. |
|
||||
|
||||
## Future work
|
||||
|
||||
- **Multipart upload for packs.** Provider single-object limits (DO
|
||||
Spaces = 5 GB non-multipart, AWS S3 = 5 GB; both support multipart up
|
||||
to 5 TB). Today's code uses `put_object` which is single-shot. boto3
|
||||
`upload_file` is the one-line drop-in.
|
||||
- **Streaming pack builder.** ✅ Landed as `stream_packs`. Caps target
|
||||
compressed bytes; each disc fills. `build_pack` stays for tests +
|
||||
small/known-set callers.
|
||||
- **Pack GC.** Stale packs (those whose `snapshot_root` is older than N
|
||||
re-pack cycles) get bucket-deleted automatically.
|
||||
- **Range-fetch partial pack pulls.** Manifest carries offsets;
|
||||
`GET .tar.zst Range: bytes=X-Y` would let a consumer pull one chunk
|
||||
from a huge pack without downloading the whole thing.
|
||||
- **KMS / SSE-S3.** Server-side encryption (mesh ciphertext on a
|
||||
public bucket is the v1 confidentiality path).
|
||||
- **Multi-region replication.** Handled by the provider within a region;
|
||||
cross-provider replication is a separate distribution-policy question.
|
||||
194
docs/tickets/ticket-000061-cold-object-store-tier.md
Normal file
194
docs/tickets/ticket-000061-cold-object-store-tier.md
Normal file
|
|
@ -0,0 +1,194 @@
|
|||
# Ticket #000061 — Cold-pack distribution tier
|
||||
|
||||
**Status:** in progress — opened 2026-05-25
|
||||
**Opened:** 2026-05-25
|
||||
**Scope:** ship arborist corpus state to new peers (and to DVD-R archival)
|
||||
via point-in-time `tar.zst` packs hosted on an S3-compatible bucket
|
||||
and/or burned to physical media. One artifact serves both channels.
|
||||
**Audience:** dav1d (architectural inflection — new optional dep, new
|
||||
optional network egress, new public-readable surface when CDN is
|
||||
enabled, new "delayed snapshot" semantics around falsifications).
|
||||
**Hard constraint:** packs are content-addressed
|
||||
(`pack_hash = hash_leaf(manifest_bytes)` over a leaf-hash-sorted,
|
||||
deduped manifest). Same chunk set → same pack_hash. Each pack pins the
|
||||
corpus `snapshot_root` it covers in the audit chain — consumers can
|
||||
detect drift between the pack and current corpus state.
|
||||
|
||||
## Problem
|
||||
|
||||
A new peer comes up with an empty SQLite shard. How does it become a
|
||||
working arborist node?
|
||||
|
||||
Options today:
|
||||
1. Re-ingest every source from upstream (Wikipedia dumps, textbooks,
|
||||
crawls). Hours-to-days; depends on every upstream being reachable.
|
||||
2. `rsync` someone else's shard. Works but bypasses the audit chain —
|
||||
the receiving node has no proof the bytes came from a trusted
|
||||
producer with verifiable provenance.
|
||||
3. **Download a tarball.** Fast, content-addressed, audit-row-pinned,
|
||||
verifiable on unpack.
|
||||
|
||||
`snapshot.py` already covers corpus *identity* (the `snapshot_root`
|
||||
Merkle hash over sorted document_roots). What's missing is the
|
||||
*delivery* — getting the bytes to a fresh node.
|
||||
|
||||
## Design
|
||||
|
||||
### What goes in the bucket
|
||||
|
||||
Only packs. No individual chunk objects. Layout:
|
||||
|
||||
```
|
||||
<bucket>/packs/<pack_hash>.tar.zst # pack body
|
||||
<bucket>/packs/<pack_hash>.manifest.ndjson # contents sidecar
|
||||
```
|
||||
|
||||
Per-pack contents (inside the tar):
|
||||
|
||||
```
|
||||
manifest.ndjson # one line per chunk: {leaf_hash, size}
|
||||
blobs/<hash[:2]>/<hash[2:]> # one tar member per chunk, raw UTF-8 body
|
||||
```
|
||||
|
||||
The `blobs/` prefix inside the tar is an in-pack convention, not a
|
||||
bucket layout — there is no `blobs/` prefix in the bucket itself.
|
||||
|
||||
### Pack identity and idempotence
|
||||
|
||||
- `pack_hash = hash_leaf(manifest_bytes)` where the manifest is sorted
|
||||
by `leaf_hash` and deduped before hashing.
|
||||
- Same chunk set → same pack_hash. Two writers building the same pack
|
||||
collide on bucket upload — no GC after duplicate runs.
|
||||
- Pack uploads are idempotent. Pack contents are append-only by
|
||||
construction.
|
||||
|
||||
### Pack-size cap (DVD-R safe-fit)
|
||||
|
||||
Default `max_pack_bytes = 4_400_000_000` (4.4 GB) — sits ~6.5 % below
|
||||
the 4.7 GB DVD-R marketing capacity to absorb:
|
||||
|
||||
- ISO9660 / UDF filesystem overhead
|
||||
- growisofs lead-in / lead-out
|
||||
- Media manufacturing variance (~1–2 %)
|
||||
- Older drives refusing the outer edge (~1–3 %)
|
||||
|
||||
Sits between industry-standard tool defaults (HandBrake DVD-5 = 4.59 GB;
|
||||
DVDFab fit-to-DVD-5 = 4.3 GB; mkisofs default DVD = 4.59 GB). Cap
|
||||
applies to uncompressed bytes so the compressed `.tar.zst` is ≤ cap by
|
||||
construction. v1 produces ~30–50 % media fill on prose; future
|
||||
streaming-compressed cap fills discs better.
|
||||
|
||||
Multi-pack splitting is greedy first-fit by accumulated raw bytes.
|
||||
|
||||
### Two channels — one artifact
|
||||
|
||||
```
|
||||
┌── S3CompatibleBackend.put_pack ── DO Spaces / R2 / S3
|
||||
build_pack ─┬────┤
|
||||
└────└── --local-dir DIR ─── growisofs ─── /dev/sr0
|
||||
```
|
||||
|
||||
Byte-identical packs in both channels. A pack burned at one site and
|
||||
fetched from CDN at another collide on `sha256sum`.
|
||||
|
||||
### Delayed-snapshot discipline
|
||||
|
||||
Packs are *not* live mirrors. The bucket is eventually consistent with
|
||||
intent, not with the live corpus. Three operational consequences:
|
||||
|
||||
1. **Each pack pins a `snapshot_root`** (`compute_snapshot_root(conn)`
|
||||
at pack time). Recorded in the `cold_pack_pushed` audit row and in
|
||||
the `push_pack` return body.
|
||||
2. **Falsifications produce new packs.** Between repacks, drift
|
||||
detection / `arborist falsify` / `rehydrate_drift` can flip
|
||||
`providence_cache` rows to `stale`. Chunk content is immutable; what
|
||||
changes is the corpus-state envelope. A re-pack with the same chunk
|
||||
set produces a different `snapshot_doc_count` if documents were
|
||||
added; same hash if not — but the audit row's `snapshot_root`
|
||||
distinguishes the moments.
|
||||
3. **Stale packs accumulate.** No automatic GC in v1. Old `pack_hash`es
|
||||
stay in the bucket as valid snapshots of past corpus states. GC is
|
||||
a separate ticket once we have a re-pack cadence to GC against.
|
||||
|
||||
### Cores never evict; packs include them
|
||||
|
||||
Packs cover every hot chunk with local content — surfaces AND cores.
|
||||
Cores carry the distillation derivations a new peer needs to bootstrap
|
||||
the v9.8 chain. (Earlier draft restricted packs to surfaces only — a
|
||||
bug that would have shipped a fresh peer with no derivation roots.)
|
||||
|
||||
### Verification on unpack
|
||||
|
||||
Every chunk's `leaf_hash` is the tar member name. `open_pack` recomputes
|
||||
`hash_leaf(body).hex()` per member and refuses to restore on mismatch.
|
||||
Hostile-bucket bytes never reach the local DB.
|
||||
|
||||
After unpack, consumers can `arborist snapshot verify <root>` against
|
||||
the pack's recorded `snapshot_root` to confirm clean restore (or detect
|
||||
drift if the local corpus has moved on since).
|
||||
|
||||
## What was deleted from this ticket's first attempt
|
||||
|
||||
The original implementation included per-chunk individual blob storage
|
||||
(`evict_to_object`, `rehydrate_from_object`, `cold push`, `cold pull`,
|
||||
`blob_key()`, `BLOB_PREFIX`, `put_chunk`/`get_chunk`/`has_chunk`/
|
||||
`list_chunk_hashes` on the backend ABC). Fox's correction same-day
|
||||
(2026-05-25): "what ever was blobs? I wanted a way to hydrate using
|
||||
tarballs (the core and important data) for bringing new machines up".
|
||||
|
||||
Five-step deletion record:
|
||||
|
||||
- **Step 1 — Make the requirement less dumb:** the requirement was
|
||||
always "ship a corpus to a new peer," not "expose every chunk as an
|
||||
S3 object." Individual blobs solved a problem nobody asked for.
|
||||
- **Step 2 — Delete the part:** 14M-object storage, ~$70/hydrate
|
||||
request cost, 14M-entry LIST walks — all deleted. ~250 lines of
|
||||
source + 7 tests gone.
|
||||
- **Result:** packs-only design lands at ~17 GB bucket storage (vs
|
||||
38 GB for raw blobs), ~4 GETs/hydrate (vs 14M; packs filled to 4.4 GB
|
||||
compressed each via streaming zstd), DO Spaces request cost falls
|
||||
from ~$70 to ~$0.00002 per fresh peer.
|
||||
|
||||
## Implementation
|
||||
|
||||
- `arborist/cold_object.py` — `ObjectStoreBackend` ABC (raw bytes,
|
||||
pack-shaped wrappers) + `S3CompatibleBackend` (boto3) + `MemoryBackend`
|
||||
(tests) + `build_pack` / `open_pack` / `parse_manifest`.
|
||||
- `arborist/evict.py` — `push_pack` / `pull_pack`. Per-call binds the
|
||||
corpus `snapshot_root` into pack audit + return body.
|
||||
- `arborist/cli.py` — `arborist cold pack | unpack | stats`.
|
||||
- `tests/test_cold_object.py` — pack invariants, splitting, snapshot
|
||||
binding, local-dir, no-push mode. Default MemoryBackend so the
|
||||
default suite runs without boto3.
|
||||
- `tests/test_cold_object_boto3.py` — moto-mocked S3 wire test
|
||||
(boto3 + moto gated).
|
||||
- `pyproject.toml` — `[object-store] = boto3>=1.34`; dev extras pull
|
||||
`moto>=5.0`.
|
||||
- `Makefile` — `bootstrap-object-store`, `cold-pack`, `cold-pack-dvd`,
|
||||
`cold-unpack`, `cold-stats`.
|
||||
- `docs/cold-object-store.md` — full design, DO Spaces quickstart,
|
||||
CDN hydrate recipe, DVD-burn recipe, delayed-snapshot discipline,
|
||||
failure-mode table.
|
||||
|
||||
## Scope boundaries
|
||||
|
||||
In scope: pack build / push / pull / unpack, snapshot binding, hash
|
||||
verification, audit chain, DVD-R local-dir output, multi-pack
|
||||
splitting.
|
||||
|
||||
Out of scope (future tickets if needed):
|
||||
- Multipart upload (single-pack > 5 GB on DO Spaces / AWS).
|
||||
- Streaming pack builder (lift the in-memory tar ceiling, target
|
||||
compressed-bytes cap).
|
||||
- Stale-pack GC.
|
||||
- Range-fetch partial pack pulls.
|
||||
- Server-side encryption (KMS / SSE-S3); mesh ciphertext on public
|
||||
bucket is the v1 confidentiality path.
|
||||
- Cross-provider replication.
|
||||
|
||||
## Status
|
||||
|
||||
In progress. Code lands incrementally on `main`. Sizing math against
|
||||
current 4-shard / 14.1M-chunk corpus: **~4 packs total** (17.2 GB
|
||||
compressed ÷ 4.4 GB compressed-cap per pack via streaming zstd), ~17 GB
|
||||
bucket storage.
|
||||
Loading…
Add table
Add a link
Reference in a new issue