Speedup (§3 plan): ShadowNLI._nli_batch batches forwards (ARBORIST_NLI_BATCH=64); device auto-detect (ARBORIST_NLI_DEVICE, else cuda-if-available); auto-prefer an ONNX export — bench/scripts/export_nli_onnx.py / make export-nli-onnx exports + int8-dynamic-quantizes the pinned checkpoint into ~/.arborist/models/nli/<ver>/onnx/ (operator state, NOT committed), _ensure_loaded loads model_quantized.onnx via optimum.onnxruntime (backend onnx-int8), falls back to torch silently. torch-cpu-batch1 ~120ms/pair → onnx-int8-cpu-batched ~32ms/pair (~4x); seconds on a 4090. optimum[onnxruntime] added to the [nli] extra; 24 tests. Gate-item-4 verdict at proper n: ARBORIST_NLI_SHADOW=1 make bench-qa BENCH_QA_N=1 → 223 cells (89 STRICT / 90 HYBRID / 44 UNGROUNDED; also surfaced + fixed a lone-surrogate bug). Shadow sweep over those: NLI-as- runtime-veto on STRICT has ~26% FP at θc 0.5, ~8% at θc 0.90, ~0% only at θc 0.99 — and θc 0.99 gives up most recombination recall (hard synthetic recombinations bottom out ~0.76). FAILS the §7 #12 gate on this design. Only untried path that might pass: a Phase-3 runtime hook running NLI on the verifier's actual matched clauses (1-3), not top-6-by-overlap. Until then: runtime NLI demotion stays off; the 2 fixtures stay permanent boundary markers; θc stays 0.5. Production verifier unchanged; falsification-hard stays 10/12. |
||
|---|---|---|
| .. | ||
| batteries | ||
| fixtures | ||
| results | ||
| scripts | ||
| emergent_log.jsonl | ||
| prometheus_sigma_trigger_probe.py | ||
| qa_questions.txt | ||
| qa_questions_canonical_witness_npower.txt | ||
| qa_questions_metacog_subset.txt | ||
| qa_questions_progressive_and.txt | ||
| qa_questions_quantifier_baseline.txt | ||
| qa_questions_quantifier_subset.txt | ||
| qa_questions_smoke.txt | ||
| qa_questions_warrant_chain_aggressive.txt | ||
| qa_questions_warrant_chain_paraphrase.txt | ||
| qa_questions_warrant_chain_probe.txt | ||
| qa_sweep.py | ||
| run.sh | ||