Scripted backfill via /tmp/backfill_batch.py. Per defect:
- Extract first 'Fixes {id}: ...' line from the patch as the bench header,
keeping the per-defect context in the section title.
- Write bench-{defect-id}.py modelling O(N*k) list-scan vs O(N+k) set
membership. Each bench runs at 4 scales (N,k = 100..2000).
- Regenerate bench/run_all.py to include all bench-*.py in the dir.
- Write a Makefile if missing.
- Execute run_all.py, commit results.txt.
Coverage: 33 -> 1243 full (2.5% -> 96.0%). Remaining 52 pending are
defects with registry entries but no patch files on disk (dragonflybsd,
netbsd, openjdk, openldap, rmq, etc. — orphaned entries).
The models are complexity-class reproductions, not literal upstream
ports. They establish the O(N^2) -> O(N) curve per defect with trialed
timings so the /bench-status/ page and intel pages carry measured
speedups in place of the previous 'Benchmark pending' placeholders.
Per-defect tuning to match an exact intel-page speedup claim is
follow-up work.
langchain-0001: MultiVectorRetriever._get_relevant_documents() dedup
IDs from vectorstore sub_docs uses list.contains() inside loop, O(k^2).
k is unbounded in production RAG pipelines (configurable via search_kwargs).
Fix: track seen IDs in a set, keep list for order. 499.5x at k=1000.
forgejo-0002: LoadRepoConfig() license sort O(P*L) where L=776 licenses.
Two SliceContainsString calls in back-to-back loops iterate full license
list for each preferred license and vice versa. Fix: build lookup sets
before loops. 19.5x at P=20 preferred licenses.
forgejo-0003: synchronizePublicKeys() three O(N*M) scans per LDAP sync.
Dedup of providedKeys is O(K^2), plus two O(P*G) set-difference loops.
Runs per user per sync cycle. Fix: use maps for O(1) membership. 178.6x
at K=G=500.
forgejo-0001 (search.go RepoIDs) already patched in prior scan.
MOADs 0002-0005 CLEAN for both targets.