Full --fast crawl of russell.ballestrini.net (242 URIs): 26s -> ~5s. Three changes, biggest first: 1. fast_mode now actually ignores crawl-delay (the ~5x). The delay was only zeroed on the robots-200 path; a site with no robots.txt (404) or a robots fetch error fell back to default_crawl_delay (2s). Under --fast that made every concurrent fetch wave sleep ~2s — ~10 waves x 2s dominated the wall time. _enforce_crawl_delay now short-circuits when fast_mode, matching the documented "ignore crawl-delay" contract regardless of robots status. Disallow is still honored (separate path). 2. One shared ClientSession for the fetcher's lifetime (keepalive TCP connector sized to page-worker width) instead of a fresh session per fetch — ~3x on a 24-page wave. Lazily built in-loop via _get_session; the bridge closes it in a finally (guarded on owning the fetcher). 3. Drop the per-page preflight HEAD. aiohttp exposes response headers before the body is read, so the existing content-type binary guard skips images/video/audio without downloading them — the HEAD was a redundant round trip that doubled per-page latency. Diverges arborist's AsyncWebFetcher from the agents.ai.unturf.com/core verbatim lift (fox-approved); candidate to upstream. Regression tests pin fast=no-delay / polite=delay, shared-session lifecycle, and bridge session teardown (owned vs injected). |
||
|---|---|---|
| .. | ||
| test_async_web_fetcher.py | ||
| test_bridge.py | ||
| test_web_fetch.py | ||