The c=4 sample-shuffled bench at 15:07Z lands the post-Sprint-1b + post-Sprint-2 + post-DRY + post-keep-alive + post-shuffle state. Headlines: quote 0.50 → 0.54 (+4pp) claim_lattice_pointer 0.23 → 0.20 (-3pp) claim_lattice (JSON) 0.44 → 0.42 (-2pp) Quote's +4pp is the cleanest lift of the sprint set: the per-mode 24KB cap (Sprint 1b) surfaces tighter retrievals that quote can ground verbatim, and the bucket data confirms quote peaks at 8-16KB (0.58 strict-rate). JSON's peak migrated to its targeted 32-64KB bucket (0.48 strict-rate, vs 0.38 at 16-32KB) — Sprint 1b's intent confirmed at the per-bucket level even though the aggregate slipped 2pp. Pointer's slight drop is consistent with Sprint 2's smoke result — the chunk-specificity Rule 9 didn't lift Hermes-3-8B's lazy-anchoring at n=3. The structural fix will need a stronger intervention than a prompt nudge. Wall-clock & throughput: 11:31Z: cell-grouped, c=4, n=2, 426 tasks, 51 min, 8.4/min 15:07Z: sample-shuffled, c=4, n=3, 639 tasks, 48 min, 13.3/min Sample-shuffled scheduling delivers +58% throughput at same concurrency. n=3 (50% more work) ran in 6% LESS wall-clock. Per-call mean latency dropped 35-42% across all modes — vLLM's continuous batcher fills better when fed a diverse request stream instead of cache_key-correlated cells. Concurrency sweep: c=3 peak, c=4 within 4% (chosen), c=5 12% slower, c=6 brutal (45% slower). vLLM saturates at c=3-4 on this endpoint. Errors: 6, all on 'tell me about the roman empire' question. Root cause traced & fixed in |
||
|---|---|---|
| aborist | ||
| bench | ||
| docs | ||
| scripts | ||
| tests | ||
| .gitignore | ||
| .gitlab-ci.yml | ||
| CLAUDE.md | ||
| LICENSE | ||
| Makefile | ||
| pyproject.toml | ||
| README.md | ||
aborist
An arborist for trees and forests of cross-linked information.
Aborist ingests documents into a content-addressed, Merkle-committed SQLite store, distills them into recursive cores, and answers questions over the resulting corpus via an OpenAI-compatible LLM. Every cached answer carries a verifiable Merkle proof tying it back to its source documents — Merkle-AGI v9.8 / Merkle Providence Reverse RAG, runnable end-to-end.
Quickstart
Two end-to-end paths. Pick whichever corpus you want first; both share the same query, verify, falsify, and inspect surfaces.
A. Wikipedia 2003 (the canonical bootstrap dataset)
make bootstrap # one-time: venv + dev extras
make fetch-cur # download the 2003-05-16 snapshot (~82 MB)
make ingest-cur-attached # ~3 min: 128k articles, 4 parallel shards
make distill-shards-parallel # surface → core (first-sentence)
make distill-shards-tfidf-parallel # core → keyword sets for retrieval
make query Q="What is anarcho-capitalism?"
A Hermes-3 inference runs against the local corpus, picks 4–8 source articles by Merkle root, and returns an answer plus a verifier label that names what the lexical verifier could confirm. The current four-rung ladder for claim-lattice modes is POINTER-LINKED → ANCHOR-WARRANTED → EVIDENCE-WARRANTED → UNGROUNDED (with -PARTIAL suffix on HYBRID). Repeat the same question and a cache hit replays in ~100 ms.
B. Crawl any live website and query it
make bootstrap-crawler # one-time: install [crawler] extras
make crawl-ingest URL=https://russell.ballestrini.net DEPTH=2 # BFS + ingest
make query Q="who is Russell Ballestrini?" # cross-shard query — picks up the new shard automatically
The crawl shard is named after the seed hostname (crawl_russell_ballestrini_net.db) under ~/.aborist/shards/. Add FAST=1 for aggressive crawling against your own sites; MAX=N to cap discovery; DEPTH=N to bound BFS. Robots Disallow is always honored. After ingest, make recrawl-check DOMAIN=russell.ballestrini.net does a conditional-HEAD freshness probe per page.
After the answer
make inspect KEY=<cache_key> # sidecar: classify each unverified span
make falsify KEY=<cache_key> REASON='…' # mark wrong, keep history
make burn KEY=<cache_key> REASON='…' # delete (kindergarten only — refuses if children exist)
make help lists every target.
Setup
Get the source
git clone ssh://git@git.unturf.com:2222/engineering/unturf/aborist.git
cd aborist
HTTPS variant if SSH isn't set up:
git clone https://git.unturf.com/engineering/unturf/aborist.git
Install prerequisites
Aborist needs Python 3.10+, GNU make, curl, and bzip2. SQLite 3.35+ ships with CPython.
macOS (Homebrew)
xcode-select --install # if you don't already have CLT
brew install python@3.12 git make
Apple's make is GNU make, no extra step needed. bzip2 and curl are bundled.
Ubuntu / Debian
sudo apt update
sudo apt install -y git python3 python3-venv python3-dev build-essential curl bzip2
22.04 ships Python 3.10; 24.04 ships 3.12 — both work.
Windows
The Makefile uses bash idioms, so the supported path is WSL2 running Ubuntu. From an admin PowerShell:
wsl --install -d Ubuntu-24.04
Then inside the WSL Ubuntu shell, follow the Ubuntu instructions above.
(Native cmd / PowerShell + Git Bash mostly works for the Python parts but several make targets call for i in $(seq…) and bash -c — easier to just use WSL2.)
OpenBSD
pkg_add git python-3.12 gmake curl
OpenBSD's default make is BSD make. Aborist's Makefile uses GNU-make features (?=, conditional functions). Substitute gmake for make in every command, e.g. gmake bootstrap, gmake query Q='…'.
Bootstrap
make bootstrap
Creates .venv/, installs the package in editable mode with the [dev,html] extras, and exposes aborist at .venv/bin/aborist. No system-wide install. Re-running make bootstrap is a no-op if the venv is up to date.
After bootstrap, every workflow lives behind a make target. Run make help to list them.
Data: Wikipedia 2003-05-16 (Phase III SQL dump)
The 2003 dataset lives at https://dumps.wikimedia.org/archive/2003/2003-05-16/en/ as three files. This is a MySQL extended-INSERT format dump; for XML-format dumps from 2006 onward see the Phase IV section below.
| file | size | what |
|---|---|---|
20030516_cur_tablesql.bz2 |
82 MB | one row per article: snapshot of every Wikipedia page on 2003-05-16 |
old_tablesqlbz2.1 |
640 MiB | first half of the revision-history table (split bzip2 stream) |
old_tablesqlbz2.2 |
252 MiB | second half — cat them together to decompress |
make fetch-cur # snapshot only (~82 MB on disk)
make fetch-old # full history (~1.4 GB on disk after concat)
make fetch # both
Files land in data/. Re-running is idempotent (curl skips if already present).
Sharded ingest (the canonical path)
Per-shard SQLite files, no write-lock contention. Each shard process owns its own DB; cross-shard reads attach all shards as UNION ALL views.
make ingest-cur-attached SHARDS=4 # cur snapshot, 4 parallel shards (~3 min)
make ingest-old-attached SHARDS=4 # full history, ~30–40 min
Shards land in ~/.aborist/shards/. Override with SHARDS_DIR=/path/to/somewhere.
Resumable
Add --resume (or just re-run the make target — --resume is the default for attached ingests). Each shard tracks its own high-water mark in meta; an interrupted ingest picks up where it left off without re-hashing already-cached docs.
Single-DB ingest (simpler, smaller)
For experiments under a few thousand docs, a single SQLite file is fine:
make ingest-cur INGEST_LIMIT=1000 # one DB at $(DB), default ~/.aborist/aborist.db
make ingest-cur-parallel SHARDS=4 # 4 processes, one shared DB (WAL serialized)
Distillation (cores feed retrieval)
Two distillers ship: first-sentence-v1 (one sentence per chunk) and tfidf-keywords-v1 (top-K distinctive terms per doc). Cores are Merkle-signed back to their source docs and serve as enriched titles for retrieval.
make distill-shards-parallel # first-sentence cores, per-shard
make distill-shards-tfidf-parallel # TF-IDF cores, per-shard
Run both — they generate independent cores per source. TF-IDF cores let neologisms (like a personal term that never appears in any title) match retrievals via keyword overlap.
Data: Wikipedia 2010-11 (and other Phase IV snapshots)
In 2006 MediaWiki swapped its dumps from MySQL INSERT INTO cur syntax to XML. Aborist reads both — the SQL path above for 2003-2005 cur dumps, and a streaming XML path for any dated snapshot in https://dumps.wikimedia.org/archive/. The largest single snapshot in that archive is enwiki 2010-11-08:
| file | size | what |
|---|---|---|
enwiki-20101011-pages-articles.xml.bz2 |
6.2 GB | latest revision of every main-namespace article on 2010-11-08 (~3.4M pages, ~1.9M after redirects) |
enwiki-20101011-abstract.xml |
2.9 GB | first-paragraph abstracts only — pre-distilled summaries at ~1/100th the chunk volume |
Other useful dated snapshots in the archive: 2006-07 (1.8 GB), 2006-12 (1.9 GB), 2010-03 (varies by language). All work via the same source class.
# defaults target enwiki 20101011 (2010-11)
make fetch-xml # ~6.2 GB compressed download
make ingest-xml-attached SHARDS=4 # ~2 hours sharded; ~95 GB on disk after
# pick any other snapshot by overriding the date variables
make fetch-xml WP_XML_YEAR=2006 WP_XML_MONTH=2006-07 WP_XML_DATE=20061104
# abstract feed (one-paragraph summaries, full coverage at ~5-10 GB total)
make fetch-abstract
make ingest-abstract
The XML source streams .xml.bz2 directly via iterparse with bounded memory (each <page> is processed and cleared). Same shard / resume / Merkle contract as the SQL source. Title-prefix namespace filtering kicks in for older export schemas that omit per-page <ns>.
To ingest historical revisions instead of just the current snapshot, point WP_XML at a pages-meta-history.xml.bz2 file and use make ingest-xml-history — the source emits one Document per revision and aborist's prior-doc detection chains them with supersedes edges.
Data: personal Grok export
If you have an xAI data-export bundle, point GROK_EXPORT at its root directory (the one containing ttl/30d/export_data/<user-id>/):
export GROK_EXPORT=$HOME/Downloads/<your-user-uuid>
make ingest-grok-attached # conversations -> $(SHARDS_DIR)/grok.db
make ingest-grok-media-attached # media prompts -> $(SHARDS_DIR)/grok.db
Both walk the export tree, find prod-grok-backend.json, and yield one Document per conversation (or media-generation post). Conversation titles, full message text, and turn ordering are preserved. Each becomes a normal queryable doc in the cluster — your prior chats become memory the corpus can consult.
--resume is the default for these targets. Re-run any time to pick up new exports.
Data: git and Mercurial repos (self-play)
Aborist can consult itself. Point a source at any local clone and every text file at HEAD becomes a queryable Document; re-ingesting after new commits chains old → new via supersedes edges, so the audit trail grows alongside the repo:
make ingest-self # this aborist tree, into ~/.aborist/shards/aborist-self.db
make ingest-git GIT_REPO=/path/to/repo # any other git clone
make ingest-hg HG_REPO=/path/to/repo # mercurial flavor
URI shape: git://<repo-name>/file/<relative-path> (no commit hash — that's what enables the supersedes chain on re-ingest). Binary files are skipped (NUL-byte heuristic + UTF-8 decode probe). extra carries the current commit hash, timestamp, and subject for informational purposes; the cryptographic identity is the content-derived document_root as for every other Document.
Data: live websites (the crawler)
Aborist can BFS-discover and ingest a website starting from a seed URL, respecting robots.txt and crawl delays. The crawler is off by default — heavy deps (aiohttp, bs4, lxml, mwparserfromhell, etc.) ship as the [crawler] extras and aren't pulled into the default test suite.
make bootstrap-crawler # one-time, install [crawler] extras
make crawl-ingest URL=https://russell.ballestrini.net DEPTH=2 # BFS + ingest the discovered pages
The shard filename is derived from the seed URL's hostname so cross-shard query picks it up automatically:
URL=https://russell.ballestrini.net → $(SHARDS_DIR)/crawl_russell_ballestrini_net.db
After ingest, the same make query searches across crawl shards alongside Wikipedia, Grok, and self-play sources:
make query Q="who is Russell Ballestrini?"
Knobs:
| variable | default | what |
|---|---|---|
URL= |
(required) | seed URL; BFS stays on its hostname (no subdomain crossover) |
DEPTH= |
2 |
max BFS depth from seed |
MAX= |
0 |
cap discovery at N URLs (0 = no cap, depth is the only bound) |
FAST=1 |
unset | flip the verbatim AsyncWebFetcher into fast_mode: 5s timeouts, CPU×3 parallel page workers, ignore robots.txt crawl-delay. Disallow is still honored. Use only against sites where aggressive fetching is acceptable. |
CRAWL_SHARD= |
derived from URL | override the destination shard path |
Feeds and sitemaps (atom, rss, sitemap.xml, wp-rss2.xml, etc.) are skipped at ingest — they're discovery infrastructure, not knowledge. ETag and Last-Modified per page are captured so a future probe can ask "does this need recrawling?" without re-fetching bodies:
make recrawl-check DOMAIN=russell.ballestrini.net
Conditional If-None-Match / If-Modified-Since HEAD requests classify each ingested doc as fresh (304), stale (200), gone (404/410), or unreachable. One tiny round trip per URL with no body transfer when content's unchanged.
The crawler is a verbatim lift from ~/git/agents.ai.unturf.com/core/ (provenance documented in aborist/sources/crawler/__init__.py); aborist-side changes drop the chat-bot fetch triggers and skip web_cache_manager.py (aborist has its own content-addressed cache). Run make test-crawler for the lift's own tests.
Data: OpenAI / ChatGPT export (planned)
Not implemented yet. The shape will be one new Source subclass at aborist/sources/openai.py plus a Makefile target. The OpenAI ChatGPT data export is a .zip containing conversations.json with the mapping/messages tree shape. Adding it follows the same pattern as aborist/sources/grok.py — see that file as the template.
# (placeholder)
make ingest-openai-attached # OPENAI_EXPORT=$HOME/Downloads/<chatgpt-export>
When this lands, conversations from both Grok and OpenAI will sit in the same shard cluster; queries fan out across all of them.
Asking the corpus
make query Q="What is the philosophy of stoicism?"
make query Q="tell me about permacomputer ?"
make query Q="…" QUERY_TOP_K=12 # widen the source set
make query Q="…" JSON=1 # raw record (cache_key, merkle_proof, timings, full sources)
make query-dry Q="…" # assemble context but skip the LLM call
Default render is human-readable: question, audit-mode summary line, answer, sources list, unverified spans, short cache_key. JSON=1 gives the full record. Same trailing-question-mark question deduplicates (question_hash strips trailing .?!,;:?!。、… after lowercasing).
The query path:
- Search — FTS5 (body) + SQL
LIKE(title) +JOINover derivations (TF-IDF core keywords) across every shard. Three accept paths to the relevance filter. - Concept overlay — per-shard
concept_relationsSQLite table (corpus-derived, not hand-curated). Synonyms widen retrieval; rivalries narrow it unless the query uses comparative phrasing ("compare X vs Y"). Built-in extractorlink_reciprocity_synonymreads the existingedgestable for reciprocal A↔B link pairs and emits synonym edges between their title-tokens — works for Wikipedia (See-also bidirectional), HTML site internal links, or any corpus with bidirectional document links. Seedocs/concept-relations-design.mdfor the storage tradeoff (1.6% tax measured on 6 GB Wikipedia). - Context assembly — top-K sources concatenated up to a 60 KB budget. Wikitext is stripped to plain prose via
aborist.wikitext.to_base()(the corpus stores raw[[wikilinks]]so the link graph is recoverable on demand; the LLM and verifier both see clean prose). - LLM — Hermes-3 with strict attribution rules in the system prompt + a user-turn grounding reminder.
- Verifier — every claim runs through a layered lexical check; the result rolls up into the v9.8 trichotomy (
audit_mode∈ STRICT / HYBRID / UNGROUNDED) at the schema layer AND a four-rung display ladder at render time (POINTER-LINKED → ANCHOR-WARRANTED → EVIDENCE-WARRANTED → UNGROUNDED). See below. - Cache — the v9.8 8-dim cache_key (
source_root | question_hash | model_profile | conversation | governance_policy | schema | canonicalization | chunking) keys the answer inqa.db. Cache hits replay in ~100 ms. Per-phase timings in every result.
LLM endpoint defaults to https://hermes.ai.unturf.com/v1 (Hermes-3 Llama-3.1-8B, 82K context, no auth). Override:
export ABORIST_LLM_ENDPOINT="https://your-vllm.example/v1"
export ABORIST_LLM_MODEL="meta-llama/Llama-3.1-70B-Instruct"
export ABORIST_LLM_API_KEY="..."
Verifying answers (audit modes & label ladder)
Two layers of labels stack on every answer:
Schema layer — v9.8 trichotomy (audit_mode column, persisted, drives cache lookups & audit chain):
| mode | meaning |
|---|---|
| STRICT | every evidence unit verifies against context |
| HYBRID | some claims source-grounded, some emerged from training |
| UNGROUNDED | no evidence, or none verifies — purely emergent |
Display layer — four-rung ladder (claim-lattice modes only; renderer-only transformation, schema unchanged):
| rung | what's actually proved |
|---|---|
| EVIDENCE-WARRANTED | pointer verified + warrant ran & passed + no soft demotes |
| ANCHOR-WARRANTED | pointer-linked + warrant passed; soft-demote violations present |
| POINTER-LINKED | pointer/source/chunk verified, but warrant either didn't apply or failed for some claim |
| UNGROUNDED | no verified pairs |
HYBRID gets a -PARTIAL suffix on whichever rung applies. The point of the display ladder: STRICT in claim_lattice mode is NOT "the answer is correct" — it's "every pointer resolved to a valid evidence object whose source_role is allowed AND citation-coverage passed." The display label spells out the actual property so users don't read STRICT as full semantic entailment.
Verifier strategies run in sequence; first to find evidence classifies. Each is lexical (substring or token-coverage), never embeddings — soft signals stay out of the proof path.
| # | strategy | what triggers it |
|---|---|---|
| 1 | quote |
model wraps claims in "..."; each verbatim-substring tested |
| 2 | span |
bullet/sentence units substring-tested as fallback |
| 3 | entity |
multi-word proper nouns tested with proximity-cluster gating |
| 4 | paraphrase |
inside the span path: ≥85% token coverage on prose-shaped spans |
| 5 | claim_lattice |
model emits claim text. [E1,E2] pointer-line OR {"claims":[{"text":..., "evidence_ids":[...]}]} JSON; verifier resolves pointers to runtime-built evidence objects & runs seven hard checks |
The claim_lattice path runs seven deterministic hard checks: parser succeeded, evidence_id resolves, source_role allowed, claim text non-empty, citation coverage threshold, pointer count cap, anchor-class warrant. Anchor-class warrant composes five lexical anchor classes — proper-noun, date, entity-list, count (with digit↔word equivalence), and why-cause — each gated on either question shape, claim content, or both. See docs/concept-relations-design.md and the whitepaper §13.9 for the architecture.
Trailing (Source: https://...) parentheticals the model appends to verbatim source sentences are stripped before substring testing, so verbatim-with-citation no longer flags HYBRID.
unverified_quotes on each record is the corpus-growth signal — model output that didn't ground anywhere. aborist emergent --aggregate ranks them by frequency (the worklist of "things to ingest more sources for"). aborist reclassify re-runs the verifier against existing live records after corpus growth without any LLM call; HYBRID promotes to STRICT, UNGROUNDED to HYBRID, and one providence_reclassify audit event per change.
To dig into a specific record's unverified spans:
make inspect KEY=<cache_key>
Read-only sidecar diagnostic: pulls source chunks, classifies each unverified span as verbatim_in_base / trailing_artifact / paraphrase / partial_paraphrase / no_overlap. Writes nothing — sidecars never enter the v9.8 hard chain (that invariant is what keeps audit_mode a binary classification rather than a soft score).
Marking and burning records
Two ways to remove a cached answer, depending on whether it has children:
make falsify KEY=<cache_key> REASON="why it was wrong"
The record stays in the database; falsification_state flips to failed and a falsifications row + falsify audit event record the act. Future lookups skip records whose falsification_state != 'live'.
make burn KEY=<cache_key> REASON="..." # providence leaves
make burn KIND=document ROOT=<hex> REASON="..." # surface documents
make burn KIND=core ROOT=<hex> REASON="..." # distilled cores
burn actually deletes the row. Refuses if the leaf has children (falsifications referencing the cache_key, derivations using the document, or other records pointing at it) — pass FORCE=1 to override. Always writes an audit event. Use during early/scratch corpus building; falsify is the audit-preserving alternative once downstream consumers exist.
make chain-check-shards # 0 chain breaks across every shard's audit_events
Mesh / federation (off by default)
Two machines that ingest the same dump compute bit-identical document_roots — that's the v9.8 admissibility property. The mesh layer is the wire-and-trust scaffolding that lets peers gossip those identities (plus derivations, falsifications, cross-witnesses) and dedup-by-content across instances.
It ships off by default. No code path touches the network unless the mesh.enabled flag is set. Initialization flow:
aborist mesh init --group myteam # mint Ed25519 + X25519 keys; create epoch 0
aborist mesh status # always-safe inspection; shows enabled/false until you flip it
aborist mesh enable # flip the gating flag on
Membership is per-epoch. Adding a member, kicking a member, or rotating the secret each bumps the epoch and writes an audit event:
aborist mesh add --member-id bob --sign-pub <hex> --dh-pub <hex>
aborist mesh kick --member-id bob --reason "..." # admin-only; bumps epoch, omits bob from new envelope
aborist mesh rotate --reason "..." # refresh secret, same roster
aborist mesh members # list current epoch's roster
The kicked member's prior signatures stay verifiable forever (their roster row at older epochs is preserved on disk). They have no entry in the new epoch's secret envelope, so any AEAD-protected gossip from epoch+1 onward is opaque to them — that is the eviction guarantee.
Once two peers have enrolled each other, the HTTP gossip wire is live:
aborist mesh serve --host 0.0.0.0 --port 8400 # blocks until SIGINT
aborist mesh sync --peer http://other.example.com:8400 # announce local document_roots
aborist mesh pull --peer http://... --root <hex> # pull a body, verify Merkle
Every envelope is Ed25519-signed; bodies can opt into AEAD encryption under the per-epoch shared secret (encrypt=True on the client API). Receivers track each sender's chain-of-claims and reject envelopes whose prev_event_hash doesn't extend the last known event for that peer — fork detection at the wire layer. Two-host setup runbook: docs/mesh-deploy.md. Protocol contract + diagrams: docs/mesh.md.
Inspecting
make stats-shards # totals across shards
make analyze-shards # compression spectrum, depth histogram, audit chain integrity
make verify-shards # round-trip Merkle proofs on a random sample
make activity # recent Q&A + ingests + derives + falsifications (agent timeline)
make activity ACTIVITY_LIMIT=20
activity is JSON; pipe through jq to drill in.
Architecture (one screenful)
For per-module reference docs see docs/modules/.
Full pipeline diagrams live in docs/diagrams/ — query,
ingest, verifier-ladder, plus the mesh sub-diagrams. Render any change
with make docs.
aborist/
├── merkle.py Python port of proxy.unturf.com/pkg/verified/merkle.go
│ (non-commutative HashCombine 0x03 prefix, explicit
│ IsLeft per sibling, self-duplicate odd elements).
├── store.py SQLite v9.8 schema (8-dim cache key, falsification
│ state, audit chain, surface/core kinds, hot/warm/
│ cold tier, document_http_meta, concept_relations,
│ mesh_*).
├── ingest.py normalize → chunk → merkle → upsert. Bulk-batched.
├── document.py Document, Edge, Chunker (tok-512-v1 default).
├── source.py Source ABC: iter_documents() -> Iterator[Document].
├── wikitext.py to_base(): wikitext → plain prose, BASE_VERSION-pinned.
├── sources/
│ ├── wikipedia.py cur + old MediaWiki SQL dumps (bz2-streamed).
│ ├── wikipedia_xml.py Phase IV XML dumps (iterparse, page + history).
│ ├── html_page.py URL list + selectolax + httpx (robots-aware).
│ ├── grok.py xAI data export (conversations + media prompts).
│ ├── vcs.py git + Mercurial repos (HEAD walk, supersedes chain).
│ └── crawler/ verbatim AsyncWebFetcher lift + ingest bridge.
├── distill/ Distiller ABC + first_sentence + tfidf + runner.
├── search/ FTS5 backend + SearchBackend ABC + AuditMode.
├── concepts/ corpus-derived synonym & rivalry layer.
│ ├── store.py append-only concept_relations CRUD + idempotent
│ │ INSERT OR IGNORE on UNIQUE re-derivation key.
│ ├── query.py cross-shard synonym_expand & rivalry_excluded;
│ │ per-process LRU keyed on shard mtime; manual /
│ │ derived index split with per-token degree caps.
│ ├── extract.py pluggable extractor registry; built-in
│ │ `link_reciprocity_synonym` reads existing edges.
│ └── seed.py one-shot migration of legacy frozensets to
│ manual_legacy rows (clique edges per group).
├── qa/
│ ├── client.py ChatClient + StubClient + OpenAICompatibleClient.
│ ├── keys.py 8-dim cache_key + question_hash (CJK-aware strip).
│ ├── verify.py quote/span/entity/paraphrase + claim_lattice
│ │ verifier (seven deterministic hard checks incl.
│ │ anchor-class warrant). Citation-strip +
│ │ wikitext-strip.
│ ├── warrant.py five anchor classes — proper-noun, date,
│ │ entity-list, count (digit↔word equivalence),
│ │ cause — each gated on question shape and/or
│ │ claim content. Lexical only; in the proof path.
│ ├── evidence.py EvidenceObject + spotlight excerpt (density
│ │ rank picks the load-bearing slice).
│ ├── parse_claims.py pointer-line parser (`claim. [E1,E2]`).
│ ├── concepts.py backwards-compat shim — delegates to aborist.concepts.
│ ├── dag.py per-run Merkle-DAG (7-stage quote / 9-stage CTI).
│ ├── inspect.py sidecar diagnostic — read-only span classifier.
│ ├── runner.py ask(): single-doc Q&A + cache + verify.
│ └── query.py query(): multi-source RAG + concept overlay
│ (FTS5 AND→OR-with-synonyms fallback).
├── mesh/ off-by-default federation: identity (Ed25519/X25519),
│ per-epoch roster, AEAD-wrapped epoch secrets, HTTP
│ gossip wire, per-peer chain-of-claims tracking.
├── evict.py hot ↔ cold tier transitions; rehydrate via source.
└── cli.py argparse entrypoint (the make targets call into here).
For diagrams of how these modules wire together, see docs/diagrams/ (rendered SVGs of the module graph, query pipeline, ingest pipeline, mesh wire, & verifier ladder). For per-module reference docs see docs/modules/ — every top-level package has its own page.
Source papers (read first if confused):
~/git/unfirehose-nextjs-logger/whitepaper/merkle-providence-reverse-rag-whitepaper.rst— canonical whitepaper (rst, builds the PDF). §13.8 covers the verifier; §13.9 covers claim-lattice / CTI; §13.4.11 covers the corpus-derived concept layer.~/Downloads/merkle-agi-dag_v7.txt— formal substrate (TLV/canonical encoding, theorems T1–T5).docs/concept-relations-design.md— synonym & rivalry layer architecture + 1.6% storage tradeoff rationale.docs/mesh.md+docs/mesh-deploy.md— protocol contract + two-host runbook for federation.
Tests
make test # 641+ tests, default suite (stdlib + pytest, no network)
make test-crawler # opt-in: tests for the verbatim crawler lift
make bench # ETL throughput across configs (serial / shared-WAL / attached)
The default suite never hits the network. The crawler suite is gated behind make bootstrap-crawler (installs the [crawler] extras).
License
License: AGPL-3.0-only · This algorithm, its implementation, & all associated code carry the GNU Affero General Public License v3.0 (only). You may use, modify, & distribute under those terms. No proprietary relicensing exists.
(Verbatim from the Merkle Providence Reverse RAG whitepaper, April 2026.)