arborist/bench
russell@unturf.com d9760bc5bb
#000049 §7 #25: n=5 confirmation (444 STRICT cells) — verdict settles; the large models are the robust ones, MiniLM-cost-pick overturned
ARBORIST_NLI_SHADOW=1 make bench-qa BENCH_QA_N=5 → 1125 cells, 444 real
STRICT (5x n=1, 1.6x n=3). Re-ran the 7-model mega-grid: the lexical-
candidate NLI veto robustly clears the §7 #12 gate with
microsoft/deberta-large-mnli / k=2 / agg=max / guard=max_entail /
θc≈0.96 / θe=0.9 → 28/28 synthetic recombinations (incl. both fixtures)
· 0/444 real STRICT FP · 0/26 synthetic legit FP, ~4 pts θc headroom;
roberta-large-mnli equally good (k=2/max/θc=0.95). Resolved: agg=max +
max_entail guard is the robust score-shape (§7 #24's margin win was a
sample tie); the large checkpoints (~350-400M deberta-large-mnli /
roberta-large-mnli) hit 1.0/0.0, bart-large (similar size) only ~0.71,
deberta-base at the cliff (27/28 here, 11/28 at n=3), small models
(MiniLM-82M, deberta-v3-small) cap at ~0.82 — so §7 #18's 'MiniLM is
the cost-pick' is OVERTURNED by the proper-n evidence; k=2 consistent
winner; int8-ONNX costs ~1 catch. recommended_operating_point updated.
Remaining: a bench-qa-derived recombination set (the one check not
done); n=9 if dav1d wants more; still SHADOW; runtime promotion is
fox+dav1d-decides. Production verifier unchanged; falsification-hard
stays 10/12. This run is the worked example behind CLAUDE.md's new
bench-maxing line.
2026-05-12 19:42:55 -04:00
..
batteries #000025 §10.11 + §10.13 + §10.14 — close the 5F battery 2026-05-11 07:41:37 -04:00
fixtures #000048 step 2.4 — parse_pointer_claims clause segmentation 2026-05-11 17:09:06 -04:00
results #000049 §7 #25: n=5 confirmation (444 STRICT cells) — verdict settles; the large models are the robust ones, MiniLM-cost-pick overturned 2026-05-12 19:42:55 -04:00
scripts nli_shadow_grid: expand to no-stone-unturned sweep — 7 aggregations (incl. margin + paired-entail guard), finer θc grid (20 pts up to 0.999), --extra-models flag, global-best-fp=0 + Pareto-frontier output per model 2026-05-12 18:30:40 -04:00
emergent_log.jsonl #000006 — +30 emergent cycles (2026-05-12); verifier-ladder health re-confirmed 2026-05-12 11:28:57 -04:00
prometheus_sigma_trigger_probe.py #000012 Phase 1c follow-through: wire #000037 §12 Trigger 1 probe to fork_score_branches 2026-05-11 06:56:09 -04:00
qa_questions.txt aborist/arborist 2026-05-07 09:31:49 -04:00
qa_questions_canonical_witness_npower.txt three-thread session output: stale TODOs, N-power probe, ForkScore Phase 1c 2026-05-10 07:46:35 -04:00
qa_questions_metacog_subset.txt qa(#000011 + 4 more): SOFT_PREFLIGHT_HINT impl + 5-task fan-out 2026-05-03 23:00:56 -04:00
qa_questions_progressive_and.txt bench: progressive-AND fixture + 2026-05-09 A/B baseline report 2026-05-10 06:35:18 -04:00
qa_questions_quantifier_baseline.txt bench(#000008): harness extension — FC rate, violation kinds, raw brackets 2026-05-02 18:35:08 -04:00
qa_questions_quantifier_subset.txt ticket(#000008): §12 dry-run bench findings + --policy harness flag 2026-05-03 08:39:20 -04:00
qa_questions_smoke.txt speed: pytest-xdist, bench smoke, concurrency default; UTF surrogate fix 2026-05-02 09:29:40 -04:00
qa_questions_warrant_chain_aggressive.txt bench: aggressive warrant fixture confirms Phase 3 is rescue-only, not default-path 2026-05-10 09:53:21 -04:00
qa_questions_warrant_chain_paraphrase.txt bench: Phase 3 paraphrase fixture investigation — empirically dormant on current corpus 2026-05-10 10:06:23 -04:00
qa_questions_warrant_chain_probe.txt bench: #000031 Phase 3 A/B finds mechanism dormant on warrant-targeted fixture 2026-05-10 09:37:27 -04:00
qa_sweep.py bench/qa_sweep: scrub lone surrogates from the NLI-shadow answer_text/context fields before json.dumps 2026-05-12 17:04:33 -04:00
run.sh aborist/arborist 2026-05-07 09:31:49 -04:00