ARBORIST_NLI_SHADOW=1 make bench-qa BENCH_QA_N=5 → 1125 cells, 444 real STRICT (5x n=1, 1.6x n=3). Re-ran the 7-model mega-grid: the lexical- candidate NLI veto robustly clears the §7 #12 gate with microsoft/deberta-large-mnli / k=2 / agg=max / guard=max_entail / θc≈0.96 / θe=0.9 → 28/28 synthetic recombinations (incl. both fixtures) · 0/444 real STRICT FP · 0/26 synthetic legit FP, ~4 pts θc headroom; roberta-large-mnli equally good (k=2/max/θc=0.95). Resolved: agg=max + max_entail guard is the robust score-shape (§7 #24's margin win was a sample tie); the large checkpoints (~350-400M deberta-large-mnli / roberta-large-mnli) hit 1.0/0.0, bart-large (similar size) only ~0.71, deberta-base at the cliff (27/28 here, 11/28 at n=3), small models (MiniLM-82M, deberta-v3-small) cap at ~0.82 — so §7 #18's 'MiniLM is the cost-pick' is OVERTURNED by the proper-n evidence; k=2 consistent winner; int8-ONNX costs ~1 catch. recommended_operating_point updated. Remaining: a bench-qa-derived recombination set (the one check not done); n=9 if dav1d wants more; still SHADOW; runtime promotion is fox+dav1d-decides. Production verifier unchanged; falsification-hard stays 10/12. This run is the worked example behind CLAUDE.md's new bench-maxing line. |
||
|---|---|---|
| .. | ||
| batteries | ||
| fixtures | ||
| results | ||
| scripts | ||
| emergent_log.jsonl | ||
| prometheus_sigma_trigger_probe.py | ||
| qa_questions.txt | ||
| qa_questions_canonical_witness_npower.txt | ||
| qa_questions_metacog_subset.txt | ||
| qa_questions_progressive_and.txt | ||
| qa_questions_quantifier_baseline.txt | ||
| qa_questions_quantifier_subset.txt | ||
| qa_questions_smoke.txt | ||
| qa_questions_warrant_chain_aggressive.txt | ||
| qa_questions_warrant_chain_paraphrase.txt | ||
| qa_questions_warrant_chain_probe.txt | ||
| qa_sweep.py | ||
| run.sh | ||