candidate_clauses() — NLI now runs only on the top-N source clauses by content-token overlap with the answer claim (max_candidate_clauses=6), not the whole context; records n_candidate_clauses / best_clause_overlap / recombination_risk. Synthetic sweep unchanged (28/28 recombination, 0/26 legit FP, mean 1.45 candidate clauses/record). Real-traffic smoke re-run: STRICT would-demote 30% → 20%, overall 47% → 33% — better, not fixed; recombination-risk split doesn't separate either. Residual STRICT false-contras at ~0.83-0.92 → θc would need ≈ 0.90 (vs the clean-set 0.5); at θc=0.90 the data in hand gives 27/28 synthetic recall, 0/26 legit FP, 0/10 smoke STRICT FP — but n=10 is too small to set on. Next: a fuller ARBORIST_NLI_SHADOW=1 bench-qa run → sweep θc on hundreds of STRICT cells → confirm → set it. θc stays 0.5; runtime NLI demotion stays off. Production verifier unchanged; falsification-hard stays 10/12. |
||
|---|---|---|
| .. | ||
| batteries | ||
| fixtures | ||
| results | ||
| scripts | ||
| emergent_log.jsonl | ||
| prometheus_sigma_trigger_probe.py | ||
| qa_questions.txt | ||
| qa_questions_canonical_witness_npower.txt | ||
| qa_questions_metacog_subset.txt | ||
| qa_questions_progressive_and.txt | ||
| qa_questions_quantifier_baseline.txt | ||
| qa_questions_quantifier_subset.txt | ||
| qa_questions_smoke.txt | ||
| qa_questions_warrant_chain_aggressive.txt | ||
| qa_questions_warrant_chain_paraphrase.txt | ||
| qa_questions_warrant_chain_probe.txt | ||
| qa_sweep.py | ||
| run.sh | ||