Find a file
russell@unturf.com 546fa6d690
docs: auto-generate Makefile reference page for RTD
Adds a Sphinx extension at docs/_source/_ext/makefile_targets.py that
parses the project Makefile's '## description' annotations and writes
docs/_source/api/makefile.rst at build time. Same convention 'make help'
uses, so the reference stays in sync with the source.

Generated page is grouped by target prefix (fetch-, ingest-, distill-,
docs-, etc.) and rendered as a list-table. Shows on RTD alongside the
autodoc API modules.

Generated file is gitignored — RTD regenerates on every build.
2026-05-04 08:26:27 -04:00
aborist docs: fix inaccurate docstrings caught in contextual review 2026-05-04 08:13:13 -04:00
bench ticket(#000006): 300-cycle update — zero false-positive STRICT post-hardening 2026-05-03 23:16:45 -04:00
docs docs: auto-generate Makefile reference page for RTD 2026-05-04 08:26:27 -04:00
scripts journal: emit unfirehose/1.0 JSONL for queries + bench cycles 2026-05-02 15:19:39 -04:00
tests qa(#000011 + 4 more): SOFT_PREFLIGHT_HINT impl + 5-task fan-out 2026-05-03 23:00:56 -04:00
.gitignore docs: auto-generate Makefile reference page for RTD 2026-05-04 08:26:27 -04:00
.gitlab-ci.yml ci: disable pipeline (bench-qa runner contention) 2026-05-02 10:40:22 -04:00
.readthedocs.yaml docs: add Read the Docs configuration 2026-05-04 08:17:24 -04:00
CLAUDE.md docs: add docs/benchmarks.md — orientation doc for bench harnesses 2026-05-04 07:05:46 -04:00
LICENSE LICENSE: full AGPL-3.0 + Permacomputer Preamble 2026-04-27 16:15:16 -04:00
Makefile docs: add Sphinx API reference generation from docstrings 2026-05-04 07:55:34 -04:00
pyproject.toml speed: pytest-xdist, bench smoke, concurrency default; UTF surrogate fix 2026-05-02 09:29:40 -04:00
README.md docs: trim README — remove sections now in Sphinx/docs (31% reduction) 2026-05-04 08:24:14 -04:00

aborist

An arborist for trees and forests of cross-linked information.

Aborist ingests documents into a content-addressed, Merkle-committed SQLite store, distills them into recursive cores, and answers questions over the resulting corpus via an OpenAI-compatible LLM. Every cached answer carries a verifiable Merkle proof tying it back to its source documents — Merkle-AGI v9.8 / Merkle Providence Reverse RAG, runnable end-to-end.

Quickstart

Two end-to-end paths. Pick whichever corpus you want first; both share the same query, verify, falsify, and inspect surfaces.

A. Wikipedia 2003 (the canonical bootstrap dataset)

make bootstrap                          # one-time: venv + dev extras
make fetch-cur                          # download the 2003-05-16 snapshot (~82 MB)
make ingest-cur-attached                # ~3 min: 128k articles, 4 parallel shards
make distill-shards-parallel            # surface → core (first-sentence)
make distill-shards-tfidf-parallel      # core → keyword sets for retrieval
make query Q="What is anarcho-capitalism?"

A Hermes-3 inference runs against the local corpus, picks 48 source articles by Merkle root, and returns an answer plus a verifier label that names what the lexical verifier could confirm. The current four-rung ladder for claim-lattice modes is POINTER-LINKEDANCHOR-WARRANTEDEVIDENCE-WARRANTEDUNGROUNDED (with -PARTIAL suffix on HYBRID). Repeat the same question and a cache hit replays in ~100 ms.

B. Crawl any live website and query it

make bootstrap-crawler                                            # one-time: install [crawler] extras
make crawl-ingest URL=https://russell.ballestrini.net DEPTH=2     # BFS + ingest
make query Q="who is Russell Ballestrini?"                        # cross-shard query — picks up the new shard automatically

The crawl shard is named after the seed hostname (crawl_russell_ballestrini_net.db) under ~/.aborist/shards/. Add FAST=1 for aggressive crawling against your own sites; MAX=N to cap discovery; DEPTH=N to bound BFS. Robots Disallow is always honored. After ingest, make recrawl-check DOMAIN=russell.ballestrini.net does a conditional-HEAD freshness probe per page.

After the answer

make inspect KEY=<cache_key>           # sidecar: classify each unverified span
make falsify KEY=<cache_key> REASON='…' # mark wrong, keep history
make burn    KEY=<cache_key> REASON='…' # delete (kindergarten only — refuses if children exist)

make help lists every target.

Setup

Get the source

git clone ssh://git@git.unturf.com:2222/engineering/unturf/aborist.git
cd aborist

HTTPS variant if SSH isn't set up:

git clone https://git.unturf.com/engineering/unturf/aborist.git

Install prerequisites

Aborist needs Python 3.10+, GNU make, curl, and bzip2. SQLite 3.35+ ships with CPython.

macOS (Homebrew)

xcode-select --install        # if you don't already have CLT
brew install python@3.12 git make

Apple's make is GNU make, no extra step needed. bzip2 and curl are bundled.

Ubuntu / Debian

sudo apt update
sudo apt install -y git python3 python3-venv python3-dev build-essential curl bzip2

22.04 ships Python 3.10; 24.04 ships 3.12 — both work.

Windows

The Makefile uses bash idioms, so the supported path is WSL2 running Ubuntu. From an admin PowerShell:

wsl --install -d Ubuntu-24.04

Then inside the WSL Ubuntu shell, follow the Ubuntu instructions above.

(Native cmd / PowerShell + Git Bash mostly works for the Python parts but several make targets call for i in $(seq…) and bash -c — easier to just use WSL2.)

OpenBSD

pkg_add git python-3.12 gmake curl

OpenBSD's default make is BSD make. Aborist's Makefile uses GNU-make features (?=, conditional functions). Substitute gmake for make in every command, e.g. gmake bootstrap, gmake query Q='…'.

Bootstrap

make bootstrap

Creates .venv/, installs the package in editable mode with the [dev,html] extras, and exposes aborist at .venv/bin/aborist. No system-wide install. Re-running make bootstrap is a no-op if the venv is up to date.

After bootstrap, every workflow lives behind a make target. Run make help to list them.

Data: Wikipedia 2003-05-16 (Phase III SQL dump)

The 2003 dataset lives at https://dumps.wikimedia.org/archive/2003/2003-05-16/en/ as three files. This is a MySQL extended-INSERT format dump; for XML-format dumps from 2006 onward see the Phase IV section below.

file size what
20030516_cur_tablesql.bz2 82 MB one row per article: snapshot of every Wikipedia page on 2003-05-16
old_tablesqlbz2.1 640 MiB first half of the revision-history table (split bzip2 stream)
old_tablesqlbz2.2 252 MiB second half — cat them together to decompress
make fetch-cur                # snapshot only (~82 MB on disk)
make fetch-old                # full history (~1.4 GB on disk after concat)
make fetch                    # both

Files land in data/. Re-running is idempotent (curl skips if already present).

Sharded ingest (the canonical path)

Per-shard SQLite files, no write-lock contention. Each shard process owns its own DB; cross-shard reads attach all shards as UNION ALL views.

make ingest-cur-attached SHARDS=4    # cur snapshot, 4 parallel shards (~3 min)
make ingest-old-attached SHARDS=4    # full history, ~3040 min

Shards land in ~/.aborist/shards/. Override with SHARDS_DIR=/path/to/somewhere.

Resumable

Add --resume (or just re-run the make target — --resume is the default for attached ingests). Each shard tracks its own high-water mark in meta; an interrupted ingest picks up where it left off without re-hashing already-cached docs.

Single-DB ingest (simpler, smaller)

For experiments under a few thousand docs, a single SQLite file is fine:

make ingest-cur INGEST_LIMIT=1000       # one DB at $(DB), default ~/.aborist/aborist.db
make ingest-cur-parallel SHARDS=4       # 4 processes, one shared DB (WAL serialized)

Distillation (cores feed retrieval)

Two distillers ship: first-sentence-v1 (one sentence per chunk) and tfidf-keywords-v1 (top-K distinctive terms per doc). Cores are Merkle-signed back to their source docs and serve as enriched titles for retrieval.

make distill-shards-parallel              # first-sentence cores, per-shard
make distill-shards-tfidf-parallel        # TF-IDF cores, per-shard

Run both — they generate independent cores per source. TF-IDF cores let neologisms (like a personal term that never appears in any title) match retrievals via keyword overlap.

Data: Wikipedia 2010-11 (and other Phase IV snapshots)

In 2006 MediaWiki swapped its dumps from MySQL INSERT INTO cur syntax to XML. Aborist reads both — the SQL path above for 2003-2005 cur dumps, and a streaming XML path for any dated snapshot in https://dumps.wikimedia.org/archive/. The largest single snapshot in that archive is enwiki 2010-11-08:

file size what
enwiki-20101011-pages-articles.xml.bz2 6.2 GB latest revision of every main-namespace article on 2010-11-08 (~3.4M pages, ~1.9M after redirects)
enwiki-20101011-abstract.xml 2.9 GB first-paragraph abstracts only — pre-distilled summaries at ~1/100th the chunk volume

Other useful dated snapshots in the archive: 2006-07 (1.8 GB), 2006-12 (1.9 GB), 2010-03 (varies by language). All work via the same source class.

# defaults target enwiki 20101011 (2010-11)
make fetch-xml                # ~6.2 GB compressed download
make ingest-xml-attached SHARDS=4    # ~2 hours sharded; ~95 GB on disk after

# pick any other snapshot by overriding the date variables
make fetch-xml WP_XML_YEAR=2006 WP_XML_MONTH=2006-07 WP_XML_DATE=20061104

# abstract feed (one-paragraph summaries, full coverage at ~5-10 GB total)
make fetch-abstract
make ingest-abstract

The XML source streams .xml.bz2 directly via iterparse with bounded memory (each <page> is processed and cleared). Same shard / resume / Merkle contract as the SQL source. Title-prefix namespace filtering kicks in for older export schemas that omit per-page <ns>.

To ingest historical revisions instead of just the current snapshot, point WP_XML at a pages-meta-history.xml.bz2 file and use make ingest-xml-history — the source emits one Document per revision and aborist's prior-doc detection chains them with supersedes edges.

Data: personal Grok export

If you have an xAI data-export bundle, point GROK_EXPORT at its root directory (the one containing ttl/30d/export_data/<user-id>/):

export GROK_EXPORT=$HOME/Downloads/<your-user-uuid>
make ingest-grok-attached            # conversations  -> $(SHARDS_DIR)/grok.db
make ingest-grok-media-attached      # media prompts  -> $(SHARDS_DIR)/grok.db

Both walk the export tree, find prod-grok-backend.json, and yield one Document per conversation (or media-generation post). Conversation titles, full message text, and turn ordering are preserved. Each becomes a normal queryable doc in the cluster — your prior chats become memory the corpus can consult.

--resume is the default for these targets. Re-run any time to pick up new exports.

Data: git and Mercurial repos (self-play)

Aborist can consult itself. Point a source at any local clone and every text file at HEAD becomes a queryable Document; re-ingesting after new commits chains old → new via supersedes edges, so the audit trail grows alongside the repo:

make ingest-self                         # this aborist tree, into ~/.aborist/shards/aborist-self.db
make ingest-git GIT_REPO=/path/to/repo   # any other git clone
make ingest-hg  HG_REPO=/path/to/repo    # mercurial flavor

URI shape: git://<repo-name>/file/<relative-path> (no commit hash — that's what enables the supersedes chain on re-ingest). Binary files are skipped (NUL-byte heuristic + UTF-8 decode probe). extra carries the current commit hash, timestamp, and subject for informational purposes; the cryptographic identity is the content-derived document_root as for every other Document.

Data: live websites (the crawler)

Aborist can BFS-discover and ingest a website starting from a seed URL, respecting robots.txt and crawl delays. The crawler is off by default — heavy deps (aiohttp, bs4, lxml, mwparserfromhell, etc.) ship as the [crawler] extras and aren't pulled into the default test suite.

make bootstrap-crawler                                              # one-time, install [crawler] extras
make crawl-ingest URL=https://russell.ballestrini.net DEPTH=2       # BFS + ingest the discovered pages

The shard filename is derived from the seed URL's hostname so cross-shard query picks it up automatically:

URL=https://russell.ballestrini.net  →  $(SHARDS_DIR)/crawl_russell_ballestrini_net.db

After ingest, the same make query searches across crawl shards alongside Wikipedia, Grok, and self-play sources:

make query Q="who is Russell Ballestrini?"

Knobs:

variable default what
URL= (required) seed URL; BFS stays on its hostname (no subdomain crossover)
DEPTH= 2 max BFS depth from seed
MAX= 0 cap discovery at N URLs (0 = no cap, depth is the only bound)
FAST=1 unset flip the verbatim AsyncWebFetcher into fast_mode: 5s timeouts, CPU×3 parallel page workers, ignore robots.txt crawl-delay. Disallow is still honored. Use only against sites where aggressive fetching is acceptable.
CRAWL_SHARD= derived from URL override the destination shard path

Feeds and sitemaps (atom, rss, sitemap.xml, wp-rss2.xml, etc.) are skipped at ingest — they're discovery infrastructure, not knowledge. ETag and Last-Modified per page are captured so a future probe can ask "does this need recrawling?" without re-fetching bodies:

make recrawl-check DOMAIN=russell.ballestrini.net

Conditional If-None-Match / If-Modified-Since HEAD requests classify each ingested doc as fresh (304), stale (200), gone (404/410), or unreachable. One tiny round trip per URL with no body transfer when content's unchanged.

The crawler is a verbatim lift from ~/git/agents.ai.unturf.com/core/ (provenance documented in aborist/sources/crawler/__init__.py); aborist-side changes drop the chat-bot fetch triggers and skip web_cache_manager.py (aborist has its own content-addressed cache). Run make test-crawler for the lift's own tests.

Asking the corpus

make query Q="What is the philosophy of stoicism?"
make query Q="tell me about permacomputer ?"
make query Q="…" QUERY_TOP_K=12        # widen the source set
make query Q="…" JSON=1                # raw record (cache_key, merkle_proof, timings, full sources)
make query-dry Q="…"                    # assemble context but skip the LLM call

Default render is human-readable: question, audit-mode summary line, answer, sources list, unverified spans, short cache_key. JSON=1 gives the full record. Same trailing-question-mark question deduplicates (question_hash strips trailing .?!,;:?!。、… after lowercasing).

The query path:

  1. Search — FTS5 (body) + SQL LIKE (title) + JOIN over derivations (TF-IDF core keywords) across every shard. Three accept paths to the relevance filter.
  2. Concept overlay — per-shard concept_relations SQLite table (corpus-derived, not hand-curated). Synonyms widen retrieval; rivalries narrow it unless the query uses comparative phrasing ("compare X vs Y"). Built-in extractor link_reciprocity_synonym reads the existing edges table for reciprocal A↔B link pairs and emits synonym edges between their title-tokens. ~1.6% storage tax measured on 6 GB Wikipedia.
  3. Context assembly — top-K sources concatenated up to a 60 KB budget. Wikitext is stripped to plain prose via aborist.wikitext.to_base() (the corpus stores raw [[wikilinks]] so the link graph is recoverable on demand; the LLM and verifier both see clean prose).
  4. LLM — Hermes-3 with strict attribution rules in the system prompt + a user-turn grounding reminder.
  5. Verifier — every claim runs through a layered lexical check; the result rolls up into the v9.8 trichotomy (audit_mode ∈ STRICT / HYBRID / UNGROUNDED) at the schema layer AND a four-rung display ladder at render time (POINTER-LINKED → ANCHOR-WARRANTED → EVIDENCE-WARRANTED → UNGROUNDED). See below.
  6. Cache — the v9.8 8-dim cache_key (source_root | question_hash | model_profile | conversation | governance_policy | schema | canonicalization | chunking) keys the answer in qa.db. Cache hits replay in ~100 ms. Per-phase timings in every result.

LLM endpoint defaults to https://hermes.ai.unturf.com/v1 (Hermes-3 Llama-3.1-8B, 82K context, no auth). Override:

export ABORIST_LLM_ENDPOINT="https://your-vllm.example/v1"
export ABORIST_LLM_MODEL="meta-llama/Llama-3.1-70B-Instruct"
export ABORIST_LLM_API_KEY="..."

Verifying answers (audit modes & label ladder)

Two layers of labels stack on every answer:

Schema layer — v9.8 trichotomy (audit_mode column, persisted, drives cache lookups & audit chain):

mode meaning
STRICT every evidence unit verifies against context
HYBRID some claims source-grounded, some emerged from training
UNGROUNDED no evidence, or none verifies — purely emergent

Display layer — four-rung ladder (claim-lattice modes only; renderer-only transformation, schema unchanged):

rung what's actually proved
EVIDENCE-WARRANTED pointer verified + warrant ran & passed + no soft demotes
ANCHOR-WARRANTED pointer-linked + warrant passed; soft-demote violations present
POINTER-LINKED pointer/source/chunk verified, but warrant either didn't apply or failed for some claim
UNGROUNDED no verified pairs

HYBRID gets a -PARTIAL suffix on whichever rung applies. The point of the display ladder: STRICT in claim_lattice mode is NOT "the answer is correct" — it's "every pointer resolved to a valid evidence object whose source_role is allowed AND citation-coverage passed." The display label spells out the actual property so users don't read STRICT as full semantic entailment.

Verifier strategies run in sequence; first to find evidence classifies. Each is lexical (substring or token-coverage), never embeddings — soft signals stay out of the proof path.

# strategy what triggers it
1 quote model wraps claims in "..."; each verbatim-substring tested
2 span bullet/sentence units substring-tested as fallback
3 entity multi-word proper nouns tested with proximity-cluster gating
4 paraphrase inside the span path: ≥85% token coverage on prose-shaped spans
5 claim_lattice model emits claim text. [E1,E2] pointer-line OR {"claims":[{"text":..., "evidence_ids":[...]}]} JSON; verifier resolves pointers to runtime-built evidence objects & runs seven hard checks

The claim_lattice path runs seven deterministic hard checks: parser succeeded, evidence_id resolves, source_role allowed, claim text non-empty, citation coverage threshold, pointer count cap, anchor-class warrant. Anchor-class warrant composes five lexical anchor classes — proper-noun, date, entity-list, count (with digit↔word equivalence), and why-cause — each gated on either question shape, claim content, or both. See whitepaper §13.9 for the architecture.

Trailing (Source: https://...) parentheticals the model appends to verbatim source sentences are stripped before substring testing, so verbatim-with-citation no longer flags HYBRID.

unverified_quotes on each record is the corpus-growth signal — model output that didn't ground anywhere. aborist emergent --aggregate ranks them by frequency (the worklist of "things to ingest more sources for"). aborist reclassify re-runs the verifier against existing live records after corpus growth without any LLM call; HYBRID promotes to STRICT, UNGROUNDED to HYBRID, and one providence_reclassify audit event per change.

To dig into a specific record's unverified spans:

make inspect KEY=<cache_key>

Read-only sidecar diagnostic: pulls source chunks, classifies each unverified span as verbatim_in_base / trailing_artifact / paraphrase / partial_paraphrase / no_overlap. Writes nothing — sidecars never enter the v9.8 hard chain (that invariant is what keeps audit_mode a binary classification rather than a soft score).

Mesh / federation (off by default)

Optional gossip layer for peer-to-peer corpus sync. Default off — no code path touches the network unless mesh.enabled is set. See docs/mesh.md for protocol contract and the aborist mesh CLI subcommands.

Inspecting

make stats-shards                    # totals across shards
make analyze-shards                  # compression spectrum, depth histogram, audit chain integrity
make verify-shards                   # round-trip Merkle proofs on a random sample
make activity                        # recent Q&A + ingests + derives + falsifications (agent timeline)
make activity ACTIVITY_LIMIT=20

activity is JSON; pipe through jq to drill in.

Architecture

Aborist module graph

Generated API reference (every module, class, and function from docstrings): build with make docs-api (output at docs/_source/_build/html/) or browse the published Read the Docs site. Pipeline diagrams live in docs/diagrams/; render with make docs.

Tests & bench

make test                # 641+ tests, default suite, ~10s with pytest -n auto
make test-crawler        # opt-in: tests for the verbatim crawler lift
make test-live           # gated: live QA quality fixtures against Hermes (~1 min)
make bench               # ETL throughput across configs (serial / shared-WAL / attached)
make bench-qa            # full QA-quality sweep, sample-shuffled, 71q × 3m × 3n at c=4
make bench-qa-smoke      # 5-question smoke fixture for prompt-iteration loops (~1-3 min)

The default suite never hits the network. The crawler suite is gated behind make bootstrap-crawler (installs the [crawler] extras).

make bench-qa writes JSONL + markdown into bench/qa_results/<utc-stamp>.{jsonl,md} (gitignored); design-log entries live in docs/qa-modes-bench-<date>.md. The bench is stop/start-able via --resume <jsonl-path> (same --seed required for shuffled-task-order alignment). BENCH_QA_CONCURRENCY=N Makefile variable overrides the default 4. See docs/bench-maxing.md for the full speed playbook.

License

License: AGPL-3.0-only · This algorithm, its implementation, & all associated code carry the GNU Affero General Public License v3.0 (only). You may use, modify, & distribute under those terms. No proprietary relicensing exists.

(Verbatim from the Merkle Providence Reverse RAG whitepaper, April 2026.)