arborist/bench/qa_questions_amp.txt
russell@unturf.com b573c592d8
feat(retrieval): accent-fold (+30pp recall@1) + fold-search factory hardening
Second MEASURED fold-search win, and the instrument correcting my own
premature call. accent-fold ON vs OFF on the mined accent fixture:
recall@1 55% -> 85% (+30pp), rank-1 22/40 -> 34/40. recall@8 was
flat (95->98) — a too-lenient k nearly got a real lever wrongly
reverted; @1/@3 is the resolution that drives primary-source
selection. _accent_fold_variants: ASCII-fold then re-tokenise so a
diacritic title ("Béla Bartók", which _TITLE_TOKEN_RE otherwise
fragments to junk) matches the ASCII form a user types. Additive+
symmetric, no-op on pure-ASCII (zero effect on non-accent
queries/titles), mirrors _hyphen_fold_variants (#000007).

Also fixes a defect I shipped in a3ac653: an orphaned duplicate
body left as dead code after `return base` in _title_query_tokens
(unreachable — numeral-fold behaviour/measurement were valid — but
cruft; removed).

Fold-search factory, fanned out across the full survey backlog
(deterministic recall, no LLM, parallel — serial-by-caution was
halting in disguise):
- recall_at_k.py: returns rank -> recall@1/@3/@k from one retrieval
  (verified offline). A coarse k hides rank-only lifts.
- mine_questions.py: numeral/accent/hyphen/honorific/amp/brit
  ground-truth classes; fixtures committed.
- Measured @1 headroom verdicts: accent SHIP (this commit);
  honorific 45% / brit 50% = real headroom (build next); hyphen
  90% = existing #000007 already delivers, NOTHING to build (the
  measure-the-unmeasured-thing check pays off); amp 82% = no fold
  needed (prevalence-overranked, instrument kills it cheaply).

CLAUDE.md bench-maxing: two measured lessons codified — report
recall@1/@3/@k (a lenient k hides rank lifts; prevalence != miss-
rate), and fan out independent measurements (serial-by-caution is
halting). Full suite 2488 passed, 0 regressions (accent-fold is
hot-path in _title_query_tokens); real-path test (FakeSource->
ingest->query()->real _Hit).
2026-05-18 19:23:22 -04:00

43 lines
1.5 KiB
Text

# AUTO-MINED (amp class) from corpus titles via bench/mine_questions.py — ground-truth-carrying.
# Graded by deterministic retrieval recall@k (bench/recall_at_k.py), NOT audit_mode. Not adversarial; complements (never replaces) qa_questions.txt.
what is Heckler and Koch?
what is Science and Environmental Policy Project?
what is Texas A and M University?
what is Pratt and Whitney?
what is Ernst and Young?
what is Waterloo and City line?
what is Duany Plater-Zyberk and Company?
what is The Sandman: Fables and Reflections?
what is Rape, Abuse and Incest National Network?
what is Hilton Hotels and Resorts?
what is North Walsham and Dilham Canal?
what is Bill and Melinda Gates Foundation?
what is Question Mark and the Mysterians?
what is A and B?
what is Chivalry and Sorcery?
what is II and III?
what is Standard and Poor's?
what is Cheech and Chong?
what is Sasha and John Digweed?
what is Murat and Jose?
what is Valleys and Cardiff Local Routes?
what is The College of William and Mary?
what is Industrial Light and Magic?
what is Barnes and Noble?
what is Law and Order: Special Victims Unit?
what is Grammy Award for Best R and B Song?
what is Ike and Tina Turner?
what is H and M?
what is Yesterday and Today?
what is Open Fire (Y and T album)?
what is Funk and Wagnalls?
what is South Park: Bigger, Longer and Uncut?
what is Lewis and Clark College?
what is Love and Pop?
what is Barnes and Barnes?
what is Sky (UK and Ireland)?
what is Sergio and The Ladies?
what is B and B?
what is Starwood Hotels and Resorts Worldwide?
what is FRANC 2D and 3D?