All 4 manifest candidates clean-separate (12/12 NEG catch at 0/14 POS FP), so 'reranker discriminates Zionist-entity-style mis-cites' is a property of MS-MARCO-trained rerankers as a class — not the specific L-6 I picked first. Ranked by separation margin (per §7 #18: separation beats raw): ms-marco-electra-base +5.349 ← new primary BAAI/bge-reranker-base +3.076 ms-marco-MiniLM-L-6-v2 +3.053 (previous primary, demoted) ms-marco-MiniLM-L-12-v2 +2.663 (worst — deeper ≠ better) Manifest primary moved to electra-base for the cushion. demote_below_score STAYS null — clean candidate-bench thresholds don't predict real-pipeline behavior (the §7 #18→#27 history is 6 verdict flips on the NLI side); §3.2.2 step 2 (real-traffic shadow sweep on pooled bench-qa STRICT) is what sets it. Expect a walk-back. Reranker still doesn't catch the Kilimanjaro/Mount-Kenya recombination (aboutness ≠ truth-of-attribution; that's #000049 territory). fox's bench-maxing correction applied. |
||
|---|---|---|
| .. | ||
| batteries | ||
| fixtures | ||
| results | ||
| scripts | ||
| emergent_log.jsonl | ||
| prometheus_sigma_trigger_probe.py | ||
| qa_questions.txt | ||
| qa_questions_canonical_witness_npower.txt | ||
| qa_questions_metacog_subset.txt | ||
| qa_questions_progressive_and.txt | ||
| qa_questions_quantifier_baseline.txt | ||
| qa_questions_quantifier_subset.txt | ||
| qa_questions_smoke.txt | ||
| qa_questions_warrant_chain_aggressive.txt | ||
| qa_questions_warrant_chain_paraphrase.txt | ||
| qa_questions_warrant_chain_probe.txt | ||
| qa_sweep.py | ||
| run.sh | ||