7 models on the 4090 against the 26-pair candidate-bench: bge-reranker-large 560MB +6.285 ← new primary (most cushion) ms-marco-electra-base 110MB +5.349 (cost alternate — 5× smaller, 87% the margin) MiniLM-L-4-v2 50MB +3.187 bge-reranker-base 280MB +3.076 MiniLM-L-6-v2 80MB +3.053 (was primary before §3.2.1) MiniLM-L-12-v2 130MB +2.663 (deeper ≠ better — same pattern as §7 #24) MiniLM-L-2-v2 30MB -1.970 ← capacity floor (NOT separable, 10/12 catch) Findings: - Biggest is best WITHIN a family (bge-large > bge-base, 2× size → 2× margin). - Across families: electra-base (110MB) beats bge-base (280MB) — 60% smaller, 75% more margin. Architecture/training corpus > parameter count. - Depth non-monotonic within MS-MARCO MiniLM: L-4 > L-6 > L-12. - Capacity floor between L-2 (30MB, fails) and L-4 (50MB, separates). - 6 of 7 clean-separate the candidate-bench. The 'reranker catches Zionist-style mis-cites' claim is a property of competent rerankers as a class — once above the floor. Manifest primary moves to bge-reranker-large for cushion. demote_below_score STAYS null — clean candidate-bench doesn't predict real-pipeline behavior (§7 #18→#27 = 6 verdict flips). Expect a walk-back at §3.2.2 step 2. |
||
|---|---|---|
| .. | ||
| batteries | ||
| fixtures | ||
| results | ||
| scripts | ||
| emergent_log.jsonl | ||
| prometheus_sigma_trigger_probe.py | ||
| qa_questions.txt | ||
| qa_questions_canonical_witness_npower.txt | ||
| qa_questions_metacog_subset.txt | ||
| qa_questions_progressive_and.txt | ||
| qa_questions_quantifier_baseline.txt | ||
| qa_questions_quantifier_subset.txt | ||
| qa_questions_smoke.txt | ||
| qa_questions_warrant_chain_aggressive.txt | ||
| qa_questions_warrant_chain_paraphrase.txt | ||
| qa_questions_warrant_chain_probe.txt | ||
| qa_sweep.py | ||
| run.sh | ||