arborist/qa/nli/ — SHADOW ONLY (never an audit_mode input; manifest not yet in governance_policy_hash per §7 #2). manifest.json pins cross-encoder/nli-MiniLM2-L6-H768 @ a fixed HF revision + the bench-validated θc 0.5/θe 0.9 + 2 alternates + the Phase-3 TODO; shadow.py = ShadowNLI/shadow_check (lazy transformers+torch behind a new [nli] extra, clauses() segmenter, the §7 #5 clause-level Demote() decision, degrades to available=False when [nli] absent); bench/scripts/nli_shadow_sweep.py + make bootstrap-nli / bench-nli-shadow (the gate-item-4 instrument); 16 tests. First sweep (116 records — 5f-falsification packs + the arborist-nli-bench eval sets): 28/28 synth recombination demoted, 0/26 FP on legit summaries, 0/9 fires on already-STRICT_SPAN records, 25/50 on UNGROUNDED (the contradiction half; quiet on non-sequiturs). Gate items 1/2/3/5/6 clear on available data; item 4 — shadow FP rate on a real live-bench-qa sample — remains the open measurement. Production verifier unchanged; falsification-hard stays 10/12.
This commit is contained in:
parent
87c92162a1
commit
70ecda3d6c
10 changed files with 2543 additions and 8 deletions
|
|
@ -103,7 +103,7 @@ Newest first. Update on every open/close.
|
|||
|----------|------------------------------------------------|-----------------------|------------|-----------|
|
||||
| #000051 | Federated vecpack distribution (gossip the embedding backfill) | open · awaiting go/no-go · doc-only scaffold. Makes `chunk_vecs` a distributable artifact: backfill once on any CPU box (cloud / Prometheus-Σ sweep — #000037 §3.1), publish a **vecpack** `(shard_root, vec_backend_version, [(leaf_hash, embedding_blob)…])` over the mesh wire layer, every peer pulls + bulk-loads (sub-ms/chunk on the receiver — the laptop never runs the transformer). Keyed on `leaf_hash` (portable) not `chunk_id` (shard-local). Vecpacks are **soft data** — embeddings are `UNGROUNDED`, never proof path — so a cheap structural sanity gate (chunk exists locally w/ matching leaf_hash, right blob length for (dim,quant), finite norm, backend_version matches) suffices, no Merkle-proof-grade verification needed. Supplies #000050's prereq #1 ("a vecpack exists & is imported on the bench box", not "fox embedded the corpus locally"). GPU producer (the fast path): bge-small-en-v1.5 batched on a CUDA box (4090) ≈ 10³–10⁴ chunks/s → full 6.24M-chunk corpus in *minutes*, not days — drop a CUDA `Embedder` into `default_embedder()`; CUDA stack lives only on the producer box, never in arborist's `python+sqlite3` core. The mechanism behind whitepaper §1's "the embedding pass runs off the device". #000039 / #000050 sibling | 2026-05-12 | — |
|
||||
| #000050 | Vec RRF hybrid fusion (#000039 Phase 2) | open · awaiting go/no-go · doc-only scaffold; design in #000039 §4.2 (RRF) + §8 (the gate). Wire `VecBackend` as a 5th retrieval route in `query.py`, RRF-merged (route provenance carried) with the 4 FTS5 routes; UNGROUNDED hits, additive not replacement. Phase-2 sub-items now explicit: **accept-path-5** in `_filter_by_title_relevance` (low-title-overlap vec hits survive only via a stronger span-level warrant, never similarity-score alone — else the title gate drops exactly the semantic candidates vec exists for & the bench shows no lift); **six** vec config fields fold into `governance_policy_hash` (recipe-named quant `int8sym`) **+ a cache-write guard** blocking `providence_cache` persistence for vec/hybrid runs until that's wired; **run-DAG records the vec stage** (backend version, six fields, top_k, query-embedding hash, candidate chunk_ids+distances). **Gated** on (a) a corpus backfill **distributed via #000051** AND (b) a **four-condition** recall bench (A FTS5-only / B vec-only / C RRF hybrid / D candidate-union-no-RRF) clearing the 5pp floor incl. C-beats-D, on semantic-allusion + curated + **adversarial-semantic-neighbor** fixtures (else park, vec stays opt-in `--backend vec`; if C≈D ship the union, drop RRF). #000039 follow-up | 2026-05-12 | — |
|
||||
| #000049 | Attribution-aware grounding check (the recombination boundary) | open · boundary accepted · production no-go · shadow-path approved (de novo review 2026-05-13 — ticket §7) · doc-only; the home for #000048's deferred §2.3 — closing the 2 recombination over-grounds in `falsification-hard` (hard-003 Mercury / hard-005 Einstein) needs an attribution / dependency-parse or mini-NLI check, which is *not lexical* (#000048 §5). Discipline question answered: a small fixed purpose-built NLI/entailment *model* may influence `audit_mode` only as an opt-in, hash-pinned, governance-hashed, **demotion-only contradiction veto** after shadow-mode evidence (never promotes — `MODEL_ASSISTED_DEMOTION`, never `MODEL_ASSISTED_PROMOTION`). Production verifier unchanged; `falsification-hard` stays 10/12 as an honest boundary marker. Roadmap: Phase 0 (this amendment) → Phase 1 (shadow design: NLI manifest, fetch/verify, `nli_pair@v1` canonicalization, recombination-risk trigger) → Phase 2 (bench-only shadow impl, `[nli]` extra, `make fetch-nli`) → Phase 3 (demotion-only runtime, gated) → Phase 4 (mesh blob sync); §7 #12 six-condition bench gate required before Phases 2–4; if NLI ever affects `audit_mode`, `nli_policy_hash` folds into `governance_policy_hash`. **Phase-2 candidate bench done 2026-05-12** (`~/git/arborist-nli-bench/`, commits `829f9a4` + `a1cb28d`; ticket §7 #18): checkpoint-agnostic harness runs the §7 #5 clause-level algorithm over 28 synth recombination cases (incl. the 2 fixtures + harder shapes) + 26 legit cases (true summaries + near-miss decoys). 4 working candidates; `nli-MiniLM2-L6-H768` (82M, 45ms p50 CPU), `deberta-v3-base-mnli-fever-anli` (184M, 223ms), `bart-large-mnli` (407M, 259ms) all 28/28 catch · 0/26 FP with the standard θe=0.9 entailment guard; `cross-encoder/nli-deberta-v3-base` 27/28; deberta-large repo-id TODO. **Key finding: the §7 #5 two-threshold rule is load-bearing** — 3 of 4 candidates argmax-contradict 1/26 legit cases on the *wrong* source clause (competing-superlative confusion, e.g. "largest hot desert" vs "largest desert overall"); the entailment guard filters every one because another clause restates the claim → 0% guarded FP vs ~4% single-threshold. Picture: recombination is *easy* for any modern NLI checkpoint — differentiator is cost/robustness, MiniLM is the cost-pick, bart-large the threshold-robust pick. Open risk = real-traffic FP rate, measurable only by a shadow run on actual `bench-qa` (gate item 4, TODO). #000048 follow-up | 2026-05-12 | — |
|
||||
| #000049 | Attribution-aware grounding check (the recombination boundary) | open · boundary accepted · production no-go · shadow-path approved (de novo review 2026-05-13 — ticket §7) · doc-only; the home for #000048's deferred §2.3 — closing the 2 recombination over-grounds in `falsification-hard` (hard-003 Mercury / hard-005 Einstein) needs an attribution / dependency-parse or mini-NLI check, which is *not lexical* (#000048 §5). Discipline question answered: a small fixed purpose-built NLI/entailment *model* may influence `audit_mode` only as an opt-in, hash-pinned, governance-hashed, **demotion-only contradiction veto** after shadow-mode evidence (never promotes — `MODEL_ASSISTED_DEMOTION`, never `MODEL_ASSISTED_PROMOTION`). Production verifier unchanged; `falsification-hard` stays 10/12 as an honest boundary marker. Roadmap: Phase 0 (this amendment) → Phase 1 (shadow design: NLI manifest, fetch/verify, `nli_pair@v1` canonicalization, recombination-risk trigger) → Phase 2 (bench-only shadow impl, `[nli]` extra, `make fetch-nli`) → Phase 3 (demotion-only runtime, gated) → Phase 4 (mesh blob sync); §7 #12 six-condition bench gate required before Phases 2–4; if NLI ever affects `audit_mode`, `nli_policy_hash` folds into `governance_policy_hash`. **Phase-2 candidate bench done 2026-05-12** (`~/git/arborist-nli-bench/`, commits `829f9a4` + `a1cb28d`; ticket §7 #18): checkpoint-agnostic harness runs the §7 #5 clause-level algorithm over 28 synth recombination cases (incl. the 2 fixtures + harder shapes) + 26 legit cases (true summaries + near-miss decoys). 4 working candidates; `nli-MiniLM2-L6-H768` (82M, 45ms p50 CPU), `deberta-v3-base-mnli-fever-anli` (184M, 223ms), `bart-large-mnli` (407M, 259ms) all 28/28 catch · 0/26 FP with the standard θe=0.9 entailment guard; `cross-encoder/nli-deberta-v3-base` 27/28; deberta-large repo-id TODO. **Key finding: the §7 #5 two-threshold rule is load-bearing** — 3 of 4 candidates argmax-contradict 1/26 legit cases on the *wrong* source clause (competing-superlative confusion, e.g. "largest hot desert" vs "largest desert overall"); the entailment guard filters every one because another clause restates the claim → 0% guarded FP vs ~4% single-threshold. Picture: recombination is *easy* for any modern NLI checkpoint — differentiator is cost/robustness, MiniLM is the cost-pick, bart-large the threshold-robust pick. **Phase-2 shadow scaffold landed in arborist 2026-05-12** (ticket §7 #19): `arborist/qa/nli/` (manifest pins MiniLM @ a fixed HF revision + θc 0.5/θe 0.9 + 2 alternates; `ShadowNLI`/`shadow_check` lazy-imports `transformers`+`torch` behind a new `[nli]` extra, degrades to `available=False` when absent — SHADOW ONLY, never an `audit_mode` input, manifest not yet in `governance_policy_hash` per §7 #2) + `bench/scripts/nli_shadow_sweep.py` + `make bootstrap-nli` / `make bench-nli-shadow` + 16 tests. First sweep (116 records: 5f-falsification packs + the sibling-repo eval sets): 28/28 synth recombination demoted, 0/26 FP on legit summaries, 0/9 fires on already-`STRICT_SPAN` records, 25/50 on `UNGROUNDED` (the contradiction half; quiet on non-sequiturs — correct). Gate items 1/2/3/5/6 look clear on available data; **item 4 — shadow FP rate on a real live-`bench-qa` sample — remains the one open measurement** (instrument in place; the run is slow/live, fox-decides). Production verifier unchanged; `falsification-hard` stays 10/12. #000048 follow-up | 2026-05-12 | — |
|
||||
| #000048 | Verifier upgrade — recombination-aware grounding + clause segmentation | **closed · 2026-05-12** — steps 2.1 + 2.4 landed 2026-05-11 (12 of 16 residual items: 4 HYBRID_ENTITY over-grounds + 8 Formulate mis-segments → `formulate-hard` 12/12, `falsification-hard` 10/12; each bench-gated, no STRICT-rate regression — 2.1's gate fired on 0 QA answers, 2.4's segmenter touched 7 of 450 lattice cells both verdict changes correct). Step 2.2 (single-clause-containment paraphrase check) attempted + reverted — catches the 2 recombination fixtures but also rejects legit cross-sentence summaries with no threshold separating the two; recombination-vs-summary isn't lexical (§5 "What we learned"). The attribution-aware path moved to **#000049** (fox 2026-05-12). 2 live-pack `expected_reason` updated HYBRID_ENTITY→UNGROUNDED; 12+ tests; `make bench-5f-falsification-hard` / `bench-5f-formulate-hard` / `bench-fork-baseline-hard`. #000046 follow-up; #000047 closed | 2026-05-11 | — |
|
||||
| #000047 | ForkScore `_delta_*` aggregator (mean vs max vs sum) | **closed · 2026-05-11** — Option D: `WeightSet.delta_aggregator` ∈ {`mean`,`max`,`sum`} (default `mean` unchanged → no `ESTIMATOR_VERSION` bump), `fork_score._delta_5{s,t,f}` dispatch via `_aggregate`, recorded in `ScoredFork.weights`, per-sub `HARD_REGRESSION_FLOOR` flags aggregator-independent; bench data behind keeping `mean` in `5f-threshold-calibration-2026-05-11.md` §5; 8+1 tests. #000012-revision / #000025 §10.14 follow-up | 2026-05-11 | — |
|
||||
| #000046 | Harder 5S/5T/5F fixture tier (below-ceiling baselines) | **closed · 2026-05-11** — Phase 1 `falsification-hard-v1.jsonl` (12 near-misses) + Phase 2 `formulate-hard-v1.jsonl` (12 mis-segments, rate 4/12) + Phase 3 `verify_quotes` paraphrase numeric-agreement gate (`_numeric_signature`; demotes a token-covering span asserting a digit-number the source lacks modulo thousands-comma) → falsification-hard rate 4/12 → 6/12 on a real change; bench-gated (`make bench-qa` n=3×75×3 before/after — no STRICT-rate regression on legit answers; only gate-caused QA shift was correctly demoting a fictional-year claim STRICT→HYBRID); `fork_score` γ·Δ5f went positive on it. Headroom now down to 2 falsification-hard over-grounds (#000048 step 2.1 closed the 4 entity over-grounds; step 2.4 closed the 8 Formulate mis-segments → that pack 12/12; step 2.2 attempted + reverted — the last 2 recombination fixtures need an attribution-aware verifier, now tracked as **#000049**, and stand as documented residue). `make bench-5f-falsification-hard` / `bench-5f-formulate-hard` / `bench-fork-baseline-hard`; 7+ tests. #000025 §10.14 follow-up; #000047 closed; #000048 closed | 2026-05-11 | — |
|
||||
|
|
|
|||
|
|
@ -1,12 +1,16 @@
|
|||
# Ticket #000049 — Attribution-aware grounding check (the recombination boundary)
|
||||
|
||||
**Status:** open · boundary accepted · production no-go · shadow-path
|
||||
approved (de novo review 2026-05-13 — see §7) · Phase-2 candidate bench
|
||||
done (§7 #18 — `~/git/arborist-nli-bench/`; 4 working candidates, 3 hit
|
||||
28/28 catch · 0/26 FP on a 54-case synth set incl. harder shapes;
|
||||
confirmed the §7 #5 entailment guard is load-bearing — filters spurious
|
||||
competing-superlative contradictions; MiniLM-82M is the cost-pick;
|
||||
real-QA shadow FP rate still TODO = gate item 4)
|
||||
approved (de novo review 2026-05-13 — see §7) · Phase-1 + Phase-2-scaffold
|
||||
landed 2026-05-12 (§7 #18/#19 — candidate bench in `~/git/arborist-nli-bench/`
|
||||
[MiniLM-82M cost-pick, §7 #5 entailment guard confirmed load-bearing];
|
||||
`arborist/qa/nli/` shadow module + `[nli]` extra + `make bench-nli-shadow`
|
||||
in-repo, SHADOW ONLY — never an `audit_mode` input; first sweep:
|
||||
28/28 recombination demoted, 0/26 FP on legit summaries, never fires on
|
||||
already-STRICT records). Remaining: §7 #12 gate item 4 — shadow FP rate
|
||||
on a real live-`bench-qa` sample (instrument in place; the run is the
|
||||
next step, slow/live). Production verifier unchanged; `falsification-hard`
|
||||
stays 10/12.
|
||||
**Opened:** 2026-05-12
|
||||
**Scope:** Decide whether — and if so how — to add a verifier check
|
||||
that catches a *recombination*: a claim whose content tokens are all
|
||||
|
|
@ -641,3 +645,51 @@ two-threshold algorithm is the right shape (the entailment guard
|
|||
earns its place); a small cross-encoder NLI suffices; the one open
|
||||
risk is real-traffic false positives, measurable only by a shadow
|
||||
run on actual QA (Phase 2 → gate item 4).
|
||||
|
||||
**19. Phase-2 shadow scaffold landed in arborist (2026-05-12).**
|
||||
`arborist/qa/nli/` — SHADOW ONLY (writes nothing to `providence_cache`
|
||||
/ `audit_events`, never an `audit_mode` input; per §7 #2 the manifest
|
||||
does *not* yet fold into `governance_policy_hash` because shadow
|
||||
output can't touch `audit_mode`). Pieces: `manifest.json` (pins
|
||||
`cross-encoder/nli-MiniLM2-L6-H768` @ a fixed HF revision + the
|
||||
bench-validated θc 0.5 / θe 0.9 + two alternates; carries the Phase-3
|
||||
TODO — add single-blob `checkpoint_sha256` + `tokenizer_sha256` +
|
||||
`nli_policy_hash` and ONNX-export to drop torch); `shadow.py`
|
||||
(`ShadowNLI` / `shadow_check` — lazy `transformers`+`torch` import
|
||||
behind the new `[nli]` extra, `clauses()` segmenter, the §7 #5
|
||||
clause-level `Demote()` decision; construction never raises, degrades
|
||||
to `available=False` when `[nli]` absent); `[nli]` extra in
|
||||
`pyproject.toml` (hard out of core/dev — fresh checkout stays
|
||||
`python3.12 + venv + sqlite3`; `make bootstrap-nli` to opt in);
|
||||
`bench/scripts/nli_shadow_sweep.py` + `make bench-nli-shadow` (the
|
||||
gate-item-4 instrument — sweeps `(answer, context)` records, reports
|
||||
the *would-demote* rate bucketed by verifier label; renders even
|
||||
without `[nli]`, marked `available:false`); <!--AUTOCOUNT:tests:tests/test_nli_shadow.py-->16<!--/AUTOCOUNT--> tests in
|
||||
`tests/test_nli_shadow.py` (pure-Python parts + graceful degradation
|
||||
+ the bench-sweep parser — run in the default suite).
|
||||
|
||||
**First shadow sweep (2026-05-12, `nli-shadow-v1-minilm2-l6-h768`,
|
||||
θc 0.5 / θe 0.9, 116 records, ~14 s CPU; `bench/results/nli-shadow-sweep.json`)** —
|
||||
inputs: `bench/fixtures/5f/falsification-{hard,live}-v1.jsonl` (all
|
||||
*false* claims the lexical verifier already rejects) + the
|
||||
`arborist-nli-bench` 28 recombination / 26 legit-summary eval sets:
|
||||
|
||||
| verifier-label bucket | n | would_demote | rate | reading |
|
||||
|---|---|---|---|---|
|
||||
| `contradiction` (synth recombination, incl. the 2 fixtures) | 28 | 28 | **1.00** | every recombination falsehood demoted — closes the fixtures (gate item 1) |
|
||||
| `not_contradiction` (synth legit summaries + decoys) | 26 | 0 | **0.00** | zero false positives (gate item 2) |
|
||||
| `STRICT_SPAN` | 9 | 0 | 0.00 | never fires on an already-STRICT-verified record (gate item 3, in this sample) |
|
||||
| `STRICT_PARAPHRASE` | 1 | 0 | 0.00 | " |
|
||||
| `UNGROUNDED` (5f-fal-live, all false) | 50 | 25 | 0.50 | NLI demotes the *attribution-contradiction* half, stays quiet on the *non-sequitur* half ("random claim" vs "unrelated context" is *neutral*, not contradicted — correct: NLI is a contradiction veto, not a grounding check) |
|
||||
| `HYBRID_ENTITY` | 2 | 1 | 0.50 | n too small to read |
|
||||
|
||||
So gate items 1, 2, 3, 5 (latency ≈ 60 ms/pair CPU, and it'd run only
|
||||
on the unresolved subset), 6 (deterministic — `ShadowNLI` is) all look
|
||||
clear on available data. **Item 4 — the shadow FP rate on a *real*
|
||||
`bench-qa` sample (live-LLM answers that actually took the
|
||||
paraphrase/entity path) — remains the one open measurement**; this
|
||||
sweep used synthetic legit + the 5f packs, not real Wikipedia QA
|
||||
output. The instrument is in place (`make bench-nli-shadow
|
||||
INPUT="…"`); feeding it a live `bench-qa` run is the next concrete
|
||||
step, and a slow/live one (fox-decides). Nothing here changes the
|
||||
production verifier; `falsification-hard` stays 10/12.
|
||||
|
|
|
|||
Loading…
Add table
Add a link
Reference in a new issue