arborist/docs/crawler.md
russell@unturf.com 7f7eeefeb9
crawl central-db + query auto-include + read-seam provenance
- make crawl-ingest writes to one central crawl db (CRAWL_DB, default
  ~/.arborist/crawl/web.db) instead of per-domain shards in the
  peer-shared main dir: keeps locally-crawled content out of peer
  sharing by default and a growing domain set under SQLite's 10-attach
  cap (Makefile, docs/crawler.md).

- arborist query auto-includes the local crawl db (query() gains
  extra_shards; CLI --include-shard / --no-crawl-db, default-on when
  web.db exists). Fix latent --db single-file query AttributeError
  (cli.py). Persist used / used_pointer_ids + retrieval_purity into
  merkle_proof so read-only consumers can see which chunks fed the
  answer (qa/query.py).

- arborist.read: read-only seam for dashboards / verifiers; on a
  multi-source context root surface the real primary source instead of
  the opaque corpus://multi-source sentinel (read.py). Backs the
  arborist-viz Merkle Command Center (#000069).

- tests for extra_shards, the CLI crawl-db resolver, and the read seam.
2026-05-29 13:45:47 -04:00

259 lines
11 KiB
Markdown
Raw Permalink Blame History

This file contains ambiguous Unicode characters

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

# Web crawler — surface ingest from live sites
Arborist's crawler turns a live website into surface-layer documents: it
walks a site from a seed URL, fetches each page, and runs the HTML
through the same normalize → chunk → Merkle → upsert pipeline as any
other source. The output lands in a content-addressed store, which is
what makes the crawler's two headline diagnostics — **duplicate
detection** and **orphan finding** — nearly free.
It pairs with the textbook-ingest workflow (`docs/warrant-substrate-cookbook.md`)
and the retrieval pipeline (`arborist/qa/query.py`).
> Status: discovery + ingest + recrawl-check ship today. The webmaster
> report (`crawl-report`) described under [Diagnostics](#diagnostics-the-webmaster-tools-angle)
> is designed-not-built — this doc pins the contract.
## Default off
The crawler touches the network only when you ask it to, and its heavy
dependencies (`aiohttp`, `beautifulsoup4`, `lxml`, …) are an opt-in
extra:
```
pip install 'arborist[crawler]' # or: make bootstrap-crawler
```
The default test suite never imports the crawler module. Nothing in a
normal lookup / ingest / Q&A path reaches it.
## How discovery works
The crawl driver is `arborist/sources/crawler/bridge.py`
(`_crawl_seed_async`). It is a **breadth-first walk that stays on one
hostname**:
1. Start from `--seed-url`. Fetch it, extract its links.
2. Keep links that are **same exact hostname** (`_same_domain`: no
subdomain crossover — `x.com` does not follow into `docs.x.com`),
`http(s)`, and not already seen.
3. Enqueue the survivors at `depth + 1`. Repeat until the queue drains,
`--depth` is hit, or `--max-pages` is reached (`0` = no cap, depth is
the only bound).
Discovery fetches each page only far enough to extract its links; the
arborist-side ingest is a separate step (`--ingest`).
### What it refuses to crawl
- **`robots.txt` Disallow is always honored** — in both polite and fast
mode. A blocked URL is skipped, never fetched.
- **Feeds and sitemaps are skipped at discovery** (`_looks_like_feed_url`):
`*.atom`, `*.rss`, `*.rdf`, `/feed/`, `/feeds/`, `sitemap.xml`,
`sitemap_index.xml`, `/sitemaps/*.xml`, and the WordPress legacy feed
names. They are crawl *infrastructure*, not knowledge — `ingest_crawled`
would discard them anyway, and a multi-MB `atom.xml` is a slow
straggler that gates everything fetched alongside it. A second filter
at ingest (URL pattern **and** response `Content-Type`) catches feeds
served at unconventional paths. `xhtml` is preserved — it *is* HTML.
The crawler **never proactively probes** `/sitemap.xml` or guesses feed
locations on the path we use. (The fetcher carries an unused
sitemap-discovery `crawl()` strategy method from its upstream lift; the
arborist bridge calls only `fetch_webpage`.)
## Polite vs fast
`--fast` flips the underlying `AsyncWebFetcher` into `fast_mode`. The one
rule that matters: **fast mode drops the rate limit, never the access
rules.**
| | polite (default) | `--fast` |
|---|---|---|
| robots `Disallow` | honored | honored |
| robots `crawl-delay` | honored; **2.0s** fallback when robots is silent/absent | **ignored** (0s) |
| concurrency per wave | 1 (sequential) | `CPU × 3` |
| connections | shared keepalive session | shared keepalive session |
| per-page requests | one `GET` | one `GET` |
| page `GET` timeout | 15s (robots 5s) | 15s (robots 5s) |
Use `--fast` only where aggressive fetching is acceptable: your own
sites, or dumps you have been authorized to mirror. On someone else's
infrastructure, the polite default is the courteous neighbour.
### Why fast is fast
Each fetch wave runs up to `CPU × 3` requests concurrently through **one
shared `aiohttp` session** with a keepalive connector, so connections
are reused across the whole crawl instead of paying a fresh TCP+TLS per
page. There is **no preflight `HEAD`**`aiohttp` exposes response
headers before the body is read, so the binary content-type guard skips
images/video/audio without downloading them, and the second round trip
the old `HEAD` cost is gone. Link extraction uses `lxml` (a C parser
that releases the GIL), so parsing genuinely parallelises across a wave
rather than serialising on the interpreter lock.
Measured on a robots-less site (`russell.ballestrini.net`, 242 URIs),
these together took a full `--fast` crawl from ~26s to ~5s. The largest
single factor was making fast mode actually ignore the crawl-delay on a
site whose `robots.txt` 404s — previously it fell back to the 2s default
and every concurrent wave slept.
## Ingest and recrawl
With `--ingest`, each discovered page is fetched with `httpx`, parsed,
and run through `ingest_source`, so crawled documents land in the same
tables as Wikipedia or textbook sources. The crawler additionally
records each document's `ETag` and `Last-Modified` in
`document_http_meta`, so `arborist crawler recrawl-check` can later send
conditional `HEAD` requests and classify each URL **fresh** (304) /
**stale** (200) / **gone** (404) without re-downloading bodies.
Re-ingest is idempotent: same content → same `document_root` → no-op.
Same URI with changed content → a new document plus a `supersedes` edge
(lossless history).
## The content-addressed payoff
This is where the crawler differs from an off-the-shelf spider. Every
document's identity is its **Merkle root over canonicalized body text**
(`ingest.py:_compute_artifacts`):
```
text = canonicalize(content) # body only — norm-v1
leaves = [hash_leaf(chunk) for chunk in chunker.split(text)]
document_root = MerkleTree.build(leaves).root
```
The **URI is not part of the root** — it is a separate `document_uri`
column. Two consequences fall straight out of that:
### Duplicate detection (free)
Two different URLs whose bodies canonicalize to the same text produce
the **same `document_root`**. So exact-and-near-duplicate pages are a
`GROUP BY document_root HAVING COUNT(*) > 1` query — no re-crawl, no
heuristics. Because the body is `canonicalize()`d before hashing,
trivial whitespace/markup differences collapse too, which is the right
behaviour for a "duplicate content" signal.
Leaves are **per chunk**, so the same data also yields **partial
overlap**: pages that share chunk-leaf hashes without being full dupes
(boilerplate headers/footers, syndicated sections) cluster via a
`chunks` join. No other crawler hands you "these N URLs are byte- or
chunk-identical" for nothing — it is a property of the store's design.
### Orphan finding
Crawl gives the **link-reachable set** (what BFS discovered). A
sitemap — fetched read-only, *once*, never ingested — gives the
**declared set**. The difference is the orphans:
```
orphans = sitemap_urls bfs_reachable_urls
```
These are pages a site lists but nothing links to. Surfacing them is a
classic webmaster need and costs us a single extra fetch on top of a
crawl we already ran.
## Diagnostics: the webmaster-tools angle
The store's by-products make a small "webmaster report" cheap. Planned
shape: `arborist crawl-report --seed-url … [--sitemap …]`, off by
default, read-only, never writing `providence_cache` or `audit_events`.
| Signal | How it's derived | Extra cost |
|---|---|---|
| Orphan pages | `sitemap BFS-reached` | one sitemap fetch |
| Exact duplicate pages | group by `document_root` | free (store query) |
| Partial-overlap clusters | shared chunk-leaf hashes | free (store query) |
| Broken links / 404s | per-URL status during crawl | free |
| Redirect chains | 301/308 hops observed | free |
| Depth / crawl-budget map | BFS depth per URL | free |
| Thin / untitled pages | chunk count + missing `<title>` | free |
A likely v1 keeps it to **orphans + exact dupes + broken links**, with
partial-overlap and the rest as follow-ons.
## Pros and cons
**Pros**
- **Compliant by default.** robots `Disallow` always honored;
crawl-delay honored unless you explicitly opt into `--fast`.
- **Fast when you own the target.** Concurrent waves, keepalive
sessions, no redundant `HEAD`, GIL-releasing parse.
- **Duplicate and partial-overlap detection for free**, because identity
is content not URL.
- **Idempotent, lossless re-ingest** with a conditional-`HEAD`
recrawl-check that avoids re-downloading unchanged pages.
- **Same store as every other source** — crawled pages are queryable,
distillable, and Merkle-verifiable like any document.
**Cons / trade-offs**
- **Orphans are invisible to the crawl itself.** We follow links and do
not consume sitemaps for discovery, so a page reachable *only* via the
sitemap is never ingested. That is a deliberate default (crawl what is
linked) — orphan *finding* is a report you run on top, not a crawl
behaviour.
- **Single hostname, exact match.** No subdomain crossover and no
cross-domain following by design; multi-host sites need multiple
seeds.
- **No JavaScript execution.** Links and content rendered client-side
(SPA routes) are not seen — this is an HTML fetcher, not a headless
browser.
- **`--fast` is anti-social off your own turf.** Ignoring crawl-delay
and fanning out `CPU × 3` is fine for sites you control and rude
elsewhere. The default is polite for a reason.
- **Diverged from upstream.** Arborist's `AsyncWebFetcher` no longer
matches the `agents.ai.unturf.com/core` verbatim lift (shared session,
dropped `HEAD`, fast-mode crawl-delay fix). A candidate to upstream.
## CLI and make targets
```
arborist crawl --seed-url URL [--depth N] [--max-pages N] [--ingest] [--fast] [--author NAME]
arborist crawler recrawl-check [--domain D] [--limit N]
```
- `--depth` — max BFS depth (default 2).
- `--max-pages` — cap discovery at N URLs (`0` = no cap, depth bounds it).
- `--ingest` — ingest discovered pages; without it, the URL list is
printed only.
- `--fast` — fast mode (see table). Disallow still honored.
- `--author` — default author surname appended to titles for warrant
resolution (only with `--ingest`).
Make targets drive both crawl workflows; crawl shards land in
`~/.arborist/crawl/` (separate from the main `~/.arborist/shards` so
SQLite's 10-attached-DB limit isn't tripped, and so locally crawled
content isn't shared as a peer by default):
```
make crawl-ingest URL=https://x.com # general web crawl → ONE central
# db (CRAWL_DB, default web.db)
make crawl-textbooks # BFS-crawl every manifest entry
# with a crawl_url (warrant substrate)
make textbook ID=<id> # ingest one textbook by id
make crawl-textbooks-stats # docs-per-shard summary
```
General web crawls (`make crawl-ingest`) all flow into a **single**
central db rather than one-per-domain: content-addressing lets many
domains coexist in one file (idempotent re-ingest, `supersedes` edges
on change), and a single file always attaches under the 10-DB cap.
Query it standalone with `arborist --db ~/.arborist/crawl/web.db query
"…"`, or attach it alongside the main corpus when you want unified
results. The per-host textbook crawls stay separate — they are warrant
substrate, resolved through a different path.
## Source map
| File | Role |
|---|---|
| `arborist/sources/crawler/bridge.py` | BFS driver, feed-skip, ingest, recrawl-check |
| `arborist/sources/crawler/async_web_fetcher.py` | fetch + robots + crawl-delay + shared session |
| `arborist/ingest.py` | `document_root` derivation (content, not URI) |
| `arborist/merkle.py` | Merkle tree / proof conventions |