First grid (MiniLM, n=1 bench-qa 89 STRICT cells): the §7 #22 'fails the gate' verdict was config-specific (defaults k=6/θc=0.5/θe=0.9) — the grid finds k=1/θc=0.93/θe=0.5 → catch 8/12 falsification-hard (incl. BOTH recombination fixtures hard-003 Mercury + hard-005 Einstein), 0/89 STRICT FP, 0/26 synthetic-legit FP; 18/28 synthetic recombination. So the lexical-candidate approach is NOT a dead end. (k=1 = no haystack; θe=0.5 tighter than the clean-set 0.9.) Full multi-model + n=3 characterization next. |
||
|---|---|---|
| .. | ||
| batteries | ||
| fixtures | ||
| results | ||
| scripts | ||
| emergent_log.jsonl | ||
| prometheus_sigma_trigger_probe.py | ||
| qa_questions.txt | ||
| qa_questions_canonical_witness_npower.txt | ||
| qa_questions_metacog_subset.txt | ||
| qa_questions_progressive_and.txt | ||
| qa_questions_quantifier_baseline.txt | ||
| qa_questions_quantifier_subset.txt | ||
| qa_questions_smoke.txt | ||
| qa_questions_warrant_chain_aggressive.txt | ||
| qa_questions_warrant_chain_paraphrase.txt | ||
| qa_questions_warrant_chain_probe.txt | ||
| qa_sweep.py | ||
| run.sh | ||