# Web crawler — surface ingest from live sites Arborist's crawler turns a live website into surface-layer documents: it walks a site from a seed URL, fetches each page, and runs the HTML through the same normalize → chunk → Merkle → upsert pipeline as any other source. The output lands in a content-addressed store, which is what makes the crawler's two headline diagnostics — **duplicate detection** and **orphan finding** — nearly free. It pairs with the textbook-ingest workflow (`docs/warrant-substrate-cookbook.md`) and the retrieval pipeline (`arborist/qa/query.py`). > Status: discovery + ingest + recrawl-check ship today. The webmaster > report (`crawl-report`) described under [Diagnostics](#diagnostics-the-webmaster-tools-angle) > is designed-not-built — this doc pins the contract. ## Default off The crawler touches the network only when you ask it to, and its heavy dependencies (`aiohttp`, `beautifulsoup4`, `lxml`, …) are an opt-in extra: ``` pip install 'arborist[crawler]' # or: make bootstrap-crawler ``` The default test suite never imports the crawler module. Nothing in a normal lookup / ingest / Q&A path reaches it. ## How discovery works The crawl driver is `arborist/sources/crawler/bridge.py` (`_crawl_seed_async`). It is a **breadth-first walk that stays on one hostname**: 1. Start from `--seed-url`. Fetch it, extract its links. 2. Keep links that are **same exact hostname** (`_same_domain`: no subdomain crossover — `x.com` does not follow into `docs.x.com`), `http(s)`, and not already seen. 3. Enqueue the survivors at `depth + 1`. Repeat until the queue drains, `--depth` is hit, or `--max-pages` is reached (`0` = no cap, depth is the only bound). Discovery fetches each page only far enough to extract its links; the arborist-side ingest is a separate step (`--ingest`). ### What it refuses to crawl - **`robots.txt` Disallow is always honored** — in both polite and fast mode. A blocked URL is skipped, never fetched. - **Feeds and sitemaps are skipped at discovery** (`_looks_like_feed_url`): `*.atom`, `*.rss`, `*.rdf`, `/feed/`, `/feeds/`, `sitemap.xml`, `sitemap_index.xml`, `/sitemaps/*.xml`, and the WordPress legacy feed names. They are crawl *infrastructure*, not knowledge — `ingest_crawled` would discard them anyway, and a multi-MB `atom.xml` is a slow straggler that gates everything fetched alongside it. A second filter at ingest (URL pattern **and** response `Content-Type`) catches feeds served at unconventional paths. `xhtml` is preserved — it *is* HTML. The crawler **never proactively probes** `/sitemap.xml` or guesses feed locations on the path we use. (The fetcher carries an unused sitemap-discovery `crawl()` strategy method from its upstream lift; the arborist bridge calls only `fetch_webpage`.) ## Polite vs fast `--fast` flips the underlying `AsyncWebFetcher` into `fast_mode`. The one rule that matters: **fast mode drops the rate limit, never the access rules.** | | polite (default) | `--fast` | |---|---|---| | robots `Disallow` | honored | honored | | robots `crawl-delay` | honored; **2.0s** fallback when robots is silent/absent | **ignored** (0s) | | concurrency per wave | 1 (sequential) | `CPU × 3` | | connections | shared keepalive session | shared keepalive session | | per-page requests | one `GET` | one `GET` | | page `GET` timeout | 15s (robots 5s) | 15s (robots 5s) | Use `--fast` only where aggressive fetching is acceptable: your own sites, or dumps you have been authorized to mirror. On someone else's infrastructure, the polite default is the courteous neighbour. ### Why fast is fast Each fetch wave runs up to `CPU × 3` requests concurrently through **one shared `aiohttp` session** with a keepalive connector, so connections are reused across the whole crawl instead of paying a fresh TCP+TLS per page. There is **no preflight `HEAD`** — `aiohttp` exposes response headers before the body is read, so the binary content-type guard skips images/video/audio without downloading them, and the second round trip the old `HEAD` cost is gone. Link extraction uses `lxml` (a C parser that releases the GIL), so parsing genuinely parallelises across a wave rather than serialising on the interpreter lock. Measured on a robots-less site (`russell.ballestrini.net`, 242 URIs), these together took a full `--fast` crawl from ~26s to ~5s. The largest single factor was making fast mode actually ignore the crawl-delay on a site whose `robots.txt` 404s — previously it fell back to the 2s default and every concurrent wave slept. ## Ingest and recrawl With `--ingest`, each discovered page is fetched with `httpx`, parsed, and run through `ingest_source`, so crawled documents land in the same tables as Wikipedia or textbook sources. The crawler additionally records each document's `ETag` and `Last-Modified` in `document_http_meta`, so `arborist crawler recrawl-check` can later send conditional `HEAD` requests and classify each URL **fresh** (304) / **stale** (200) / **gone** (404) without re-downloading bodies. Re-ingest is idempotent: same content → same `document_root` → no-op. Same URI with changed content → a new document plus a `supersedes` edge (lossless history). ## The content-addressed payoff This is where the crawler differs from an off-the-shelf spider. Every document's identity is its **Merkle root over canonicalized body text** (`ingest.py:_compute_artifacts`): ``` text = canonicalize(content) # body only — norm-v1 leaves = [hash_leaf(chunk) for chunk in chunker.split(text)] document_root = MerkleTree.build(leaves).root ``` The **URI is not part of the root** — it is a separate `document_uri` column. Two consequences fall straight out of that: ### Duplicate detection (free) Two different URLs whose bodies canonicalize to the same text produce the **same `document_root`**. So exact-and-near-duplicate pages are a `GROUP BY document_root HAVING COUNT(*) > 1` query — no re-crawl, no heuristics. Because the body is `canonicalize()`d before hashing, trivial whitespace/markup differences collapse too, which is the right behaviour for a "duplicate content" signal. Leaves are **per chunk**, so the same data also yields **partial overlap**: pages that share chunk-leaf hashes without being full dupes (boilerplate headers/footers, syndicated sections) cluster via a `chunks` join. No other crawler hands you "these N URLs are byte- or chunk-identical" for nothing — it is a property of the store's design. ### Orphan finding Crawl gives the **link-reachable set** (what BFS discovered). A sitemap — fetched read-only, *once*, never ingested — gives the **declared set**. The difference is the orphans: ``` orphans = sitemap_urls − bfs_reachable_urls ``` These are pages a site lists but nothing links to. Surfacing them is a classic webmaster need and costs us a single extra fetch on top of a crawl we already ran. ## Diagnostics: the webmaster-tools angle The store's by-products make a small "webmaster report" cheap. Planned shape: `arborist crawl-report --seed-url … [--sitemap …]`, off by default, read-only, never writing `providence_cache` or `audit_events`. | Signal | How it's derived | Extra cost | |---|---|---| | Orphan pages | `sitemap − BFS-reached` | one sitemap fetch | | Exact duplicate pages | group by `document_root` | free (store query) | | Partial-overlap clusters | shared chunk-leaf hashes | free (store query) | | Broken links / 404s | per-URL status during crawl | free | | Redirect chains | 301/308 hops observed | free | | Depth / crawl-budget map | BFS depth per URL | free | | Thin / untitled pages | chunk count + missing `` | free | A likely v1 keeps it to **orphans + exact dupes + broken links**, with partial-overlap and the rest as follow-ons. ## Pros and cons **Pros** - **Compliant by default.** robots `Disallow` always honored; crawl-delay honored unless you explicitly opt into `--fast`. - **Fast when you own the target.** Concurrent waves, keepalive sessions, no redundant `HEAD`, GIL-releasing parse. - **Duplicate and partial-overlap detection for free**, because identity is content not URL. - **Idempotent, lossless re-ingest** with a conditional-`HEAD` recrawl-check that avoids re-downloading unchanged pages. - **Same store as every other source** — crawled pages are queryable, distillable, and Merkle-verifiable like any document. **Cons / trade-offs** - **Orphans are invisible to the crawl itself.** We follow links and do not consume sitemaps for discovery, so a page reachable *only* via the sitemap is never ingested. That is a deliberate default (crawl what is linked) — orphan *finding* is a report you run on top, not a crawl behaviour. - **Single hostname, exact match.** No subdomain crossover and no cross-domain following by design; multi-host sites need multiple seeds. - **No JavaScript execution.** Links and content rendered client-side (SPA routes) are not seen — this is an HTML fetcher, not a headless browser. - **`--fast` is anti-social off your own turf.** Ignoring crawl-delay and fanning out `CPU × 3` is fine for sites you control and rude elsewhere. The default is polite for a reason. - **Diverged from upstream.** Arborist's `AsyncWebFetcher` no longer matches the `agents.ai.unturf.com/core` verbatim lift (shared session, dropped `HEAD`, fast-mode crawl-delay fix). A candidate to upstream. ## CLI and make targets ``` arborist crawl --seed-url URL [--depth N] [--max-pages N] [--ingest] [--fast] [--author NAME] arborist crawler recrawl-check [--domain D] [--limit N] ``` - `--depth` — max BFS depth (default 2). - `--max-pages` — cap discovery at N URLs (`0` = no cap, depth bounds it). - `--ingest` — ingest discovered pages; without it, the URL list is printed only. - `--fast` — fast mode (see table). Disallow still honored. - `--author` — default author surname appended to titles for warrant resolution (only with `--ingest`). Make targets drive both crawl workflows; crawl shards land in `~/.arborist/crawl/` (separate from the main `~/.arborist/shards` so SQLite's 10-attached-DB limit isn't tripped, and so locally crawled content isn't shared as a peer by default): ``` make crawl-ingest URL=https://x.com # general web crawl → ONE central # db (CRAWL_DB, default web.db) make crawl-textbooks # BFS-crawl every manifest entry # with a crawl_url (warrant substrate) make textbook ID=<id> # ingest one textbook by id make crawl-textbooks-stats # docs-per-shard summary ``` General web crawls (`make crawl-ingest`) all flow into a **single** central db rather than one-per-domain: content-addressing lets many domains coexist in one file (idempotent re-ingest, `supersedes` edges on change), and a single file always attaches under the 10-DB cap. Query it standalone with `arborist --db ~/.arborist/crawl/web.db query "…"`, or attach it alongside the main corpus when you want unified results. The per-host textbook crawls stay separate — they are warrant substrate, resolved through a different path. ## Source map | File | Role | |---|---| | `arborist/sources/crawler/bridge.py` | BFS driver, feed-skip, ingest, recrawl-check | | `arborist/sources/crawler/async_web_fetcher.py` | fetch + robots + crawl-delay + shared session | | `arborist/ingest.py` | `document_root` derivation (content, not URI) | | `arborist/merkle.py` | Merkle tree / proof conventions |