# Neopig Module Refactoring Plan ## Overview This document outlines the restructuring of `serp.py` (3291 lines) and `neopig.py` (2935 lines) into proper Python packages for improved maintainability, testability, and code organization. ## Current State Analysis ### serp.py Breakdown (3291 lines) | Section | Lines | Content | |---------|-------|---------| | TRANSLATIONS dict | 1-1111 | ~1100 lines of i18n strings (26 languages) | | i18n helpers | 1113-1194 | `get_lang()`, `t()`, `inject_i18n()`, NAV_HTML, SEARCH_BOX_HTML | | App setup | 1195-1233 | FastAPI app, Jinja2, static mounts, config globals | | Job logging | 1234-1277 | `start_job_logging()`, `stop_job_logging()`, `get_job_logs()` | | Tarball mode | 1278-1333 | `_open_tarball()`, `read_from_tarball()`, `find_media_in_tarball()` | | ArchiveDB | 1334-1353 | SQLite wrapper for archive mode | | Pydantic models | 1411-1424 | `CrawlRequest` model | | Page routes | 1425-2099 | `/`, `/crawl`, `/live`, `/view`, `/page`, `/phantom`, `/about` | | API routes | 2100-2810 | `/health`, `/api/stats`, `/api/search`, `/api/media`, `/api/import/*` | | Crawl API | 2811-3082 | `/api/crawl/jobs/*` CRUD and control | | Main entry | 3199-3291 | `init_tarball_mode()`, `main()` | ### neopig.py Breakdown (2935 lines) | Section | Lines | Content | |---------|-------|---------| | Imports/globals | 1-98 | Imports, child process tracking, live queue | | Job logging | 99-115 | Duplicate of serp.py's logging helpers | | AppendOnlyStateLog | 117-229 | Fast crawl state persistence (~113 lines) | | State helpers | 231-306 | `get_state_log_path()`, `rotate_state_file()`, logging setup | | NeoPig class | 307-1751 | Core crawler class (~1444 lines) | | HTML utilities | 1752-1828 | `trim_html_wrapper()`, `extract_meta_from_html()` | | Multiprocess helpers | 1830-1954 | `_process_single_page()`, `_process_chunk()` | | backfill_markdown | 1955-2121 | Async markdown regeneration (~167 lines) | | backfill_screenshots | 2122-2384 | Async screenshot backfill (~262 lines) | | CLI main() | 2385-2935 | Argument parsing + orchestration (~550 lines) | --- ## Proposed Structure ### 1. serp/ Module ``` serp/ ├── __init__.py # Exports app, main() ├── __main__.py # Entry point: python -m serp ├── app.py # FastAPI app creation, config, events (~150 lines) ├── i18n.py # TRANSLATIONS, get_lang(), t(), inject_i18n() (~1200 lines) ├── templates.py # Jinja2 setup, NAV_HTML, SEARCH_BOX_HTML (~100 lines) ├── models.py # Pydantic request/response models (~100 lines) ├── archive.py # Tarball mode: ArchiveDB, read helpers (~150 lines) ├── logging.py # Job logging: start/stop/get_job_logs() (~50 lines) └── routes/ ├── __init__.py # Router aggregation ├── pages.py # HTML pages: /, /crawl, /live, /view, /about, /phantom (~600 lines) ├── search.py # GET /api/search, /api/search/pages, /random (~150 lines) ├── crawl.py # /api/crawl/* job management (~300 lines) ├── live.py # SSE /api/live/stream (~50 lines) ├── media.py # /media/{hash}, /api/media/{hash} (~200 lines) └── import_.py # /import, /api/import/* upload handling (~300 lines) ``` **Key changes:** - i18n becomes its own module (largest single component) - Routes split by domain (search, crawl, media, import) - Archive/tarball mode isolated for clarity - Shared templates/nav moved to templates.py ### 2. neopig/ Module ``` neopig/ ├── __init__.py # Exports NeoPig, CrawlMode, main() ├── __main__.py # Entry point: python -m neopig ├── crawler.py # NeoPig class (~1000 lines after cleanup) ├── state.py # AppendOnlyStateLog, state file helpers (~200 lines) ├── live.py # Live queue: get_live_queue(), emit_live_media() (~30 lines) ├── cli.py # main(), argparse, orchestration (~550 lines) ├── logging.py # TqdmLoggingHandler, setup_logging() (~50 lines) ├── html_utils.py # trim_html_wrapper(), extract_meta_from_html() (~80 lines) └── backfill/ ├── __init__.py ├── markdown.py # backfill_markdown(), multiprocess helpers (~250 lines) └── screenshots.py # backfill_screenshots() (~270 lines) ``` **Key changes:** - NeoPig class isolated in crawler.py - State management extracted to state.py - Backfill operations grouped in subpackage - CLI separated from business logic - HTML utilities extracted for reuse --- ## Migration Strategy ### Phase 1: Extract Without Breaking (Low Risk) 1. **Create package directories** ```bash mkdir -p serp/routes neopig/backfill ``` 2. **Extract pure utility modules first** (no circular deps): - `serp/i18n.py` - TRANSLATIONS dict + functions - `serp/models.py` - Pydantic models - `neopig/state.py` - AppendOnlyStateLog class - `neopig/html_utils.py` - HTML processing functions - `neopig/live.py` - Queue management 3. **Create compatibility shims** in original files: ```python # serp.py (temporary) from serp.i18n import TRANSLATIONS, get_lang, t, inject_i18n ``` ### Phase 2: Extract Routes (Medium Risk) 1. **Create route modules** with APIRouter: ```python # serp/routes/search.py from fastapi import APIRouter router = APIRouter(tags=["search"]) @router.get("/api/search") async def search(...): ... ``` 2. **Aggregate routers** in `serp/routes/__init__.py`: ```python from .search import router as search_router from .crawl import router as crawl_router # ... routers = [search_router, crawl_router, ...] ``` 3. **Update app.py** to include routers: ```python for router in routers: app.include_router(router) ``` ### Phase 3: Extract Core Classes (Higher Risk) 1. **Move NeoPig class** to `neopig/crawler.py` - Update all internal imports - Handle circular dependencies with TYPE_CHECKING 2. **Move CLI** to `neopig/cli.py` - Keep `main()` callable from both `neopig.py` and `python -m neopig` 3. **Create package entry points**: ```python # neopig/__init__.py from .crawler import NeoPig from .cli import main from async_web_fetcher import CrawlMode __all__ = ['NeoPig', 'CrawlMode', 'main'] ``` ### Phase 4: Deprecate Original Files 1. **Convert `serp.py` to thin wrapper**: ```python #!/usr/bin/env python3 """Legacy entry point. Use: python -m serp""" from serp import main if __name__ == "__main__": main() ``` 2. **Convert `neopig.py` to thin wrapper**: ```python #!/usr/bin/env python3 """Legacy entry point. Use: python -m neopig""" from neopig import main import asyncio if __name__ == "__main__": asyncio.run(main()) ``` --- ## Dependency Graph (Post-Refactor) ``` ┌─────────────┐ │ database │ └──────┬──────┘ │ ┌─────────────────┼─────────────────┐ │ │ │ ▼ ▼ ▼ ┌──────────┐ ┌──────────┐ ┌──────────────┐ │ neopig/ │ │ serp/ │ │ async_web_ │ │ crawler │────▶│ routes │ │ fetcher │ └────┬─────┘ └────┬─────┘ └──────────────┘ │ │ ▼ ▼ ┌──────────┐ ┌──────────┐ │ neopig/ │ │ serp/ │ │ state │ │ i18n │ └──────────┘ └──────────┘ ``` --- ## Circular Dependency Prevention ### Known Risks 1. **serp imports neopig.get_live_queue()** - Solution: Move `get_live_queue()` to shared module or `neopig/live.py` - serp imports from `neopig.live`, not `neopig.crawler` 2. **Job logging duplicated in both files** - Solution: Consolidate in `neopig/logging.py`, serp imports from there 3. **Database used by both** - Already separate module - no change needed ### TYPE_CHECKING Pattern For type hints that would cause circular imports: ```python from typing import TYPE_CHECKING if TYPE_CHECKING: from neopig.crawler import NeoPig def some_function(pig: "NeoPig") -> None: ... ``` --- ## Testing Strategy ### Unit Tests Per Module ``` tests/unit/ ├── test_serp_i18n.py # Translation coverage ├── test_serp_models.py # Pydantic validation ├── test_neopig_state.py # AppendOnlyStateLog ├── test_neopig_html.py # HTML utilities └── test_neopig_crawler.py # NeoPig class methods ``` ### Integration Tests ``` tests/integration/ ├── test_serp_routes.py # FastAPI TestClient ├── test_neopig_crawl.py # End-to-end crawl └── test_full_pipeline.py # Crawl + SERP serving ``` --- ## Backwards Compatibility ### Entry Points Preserved ```bash # These continue to work: python serp.py --port 31337 python neopig.py https://example.com # New alternatives: python -m serp --port 31337 python -m neopig https://example.com ``` ### Import Paths Preserved ```python # These continue to work: from neopig import get_live_queue from serp import app # New alternatives: from neopig.live import get_live_queue from serp.app import app ``` --- ## Estimated Line Counts (Post-Refactor) ### serp/ Package | Module | Lines | Notes | |--------|-------|-------| | i18n.py | ~1200 | Mostly TRANSLATIONS dict | | app.py | ~150 | App creation, config, events | | templates.py | ~100 | Jinja2 setup, HTML fragments | | models.py | ~100 | Pydantic models | | archive.py | ~150 | Tarball mode | | logging.py | ~50 | Job log helpers | | routes/pages.py | ~600 | HTML page routes | | routes/search.py | ~150 | Search API | | routes/crawl.py | ~300 | Crawl job API | | routes/live.py | ~50 | SSE stream | | routes/media.py | ~200 | Media serving | | routes/import_.py | ~300 | Import handling | | **Total** | **~3350** | Slight overhead from structure | ### neopig/ Package | Module | Lines | Notes | |--------|-------|-------| | crawler.py | ~1000 | NeoPig class (trimmed) | | state.py | ~200 | State management | | live.py | ~30 | Queue helpers | | cli.py | ~550 | CLI entry point | | logging.py | ~50 | tqdm-safe logging | | html_utils.py | ~80 | HTML processing | | backfill/markdown.py | ~250 | Markdown backfill | | backfill/screenshots.py | ~270 | Screenshot backfill | | **Total** | **~2430** | Cleaner than original | --- ## Open Questions 1. **Should i18n be a separate top-level package?** - Pro: Could be reused by other projects - Con: Adds complexity, tightly coupled to serp templates 2. **Should backfill be part of neopig or separate?** - Currently in neopig, but operates on database directly - Could be `backfill/` at top level if other tools need it 3. **Where should shared logging live?** - Option A: `neopig/logging.py` (crawler owns it) - Option B: New `common/` package - Option C: Keep duplicated (simplest) --- ## Implementation Checklist - [ ] Create `serp/` directory structure - [ ] Extract `serp/i18n.py` (TRANSLATIONS + functions) - [ ] Extract `serp/models.py` (Pydantic models) - [ ] Extract `serp/archive.py` (tarball mode) - [ ] Extract `serp/templates.py` (Jinja2 setup) - [ ] Extract `serp/routes/` (all route handlers) - [ ] Create `serp/app.py` (FastAPI app assembly) - [ ] Create `serp/__init__.py` and `__main__.py` - [ ] Create `neopig/` directory structure - [ ] Extract `neopig/state.py` (AppendOnlyStateLog) - [ ] Extract `neopig/live.py` (queue management) - [ ] Extract `neopig/html_utils.py` - [ ] Extract `neopig/logging.py` - [ ] Extract `neopig/backfill/` (markdown + screenshots) - [ ] Extract `neopig/crawler.py` (NeoPig class) - [ ] Extract `neopig/cli.py` (main function) - [ ] Create `neopig/__init__.py` and `__main__.py` - [ ] Convert original files to thin wrappers - [ ] Update all imports in dependent files - [ ] Run full test suite - [ ] Update CLAUDE.md with new structure