Three workstreams, full suite 2482 passed, experimental paths default-OFF.
#000055 — Windows quickstart without make
tasks.py (pure-stdlib runner) + make.bat shim + .gitattributes;
README Windows section rewritten. Quickstart needs only Python
3.10+ (no make/bzip2/curl/bash). Mirrors the Makefile quickstart
subset; drift-pinned by tests/test_tasks_runner.py.
#000001 §7 Phase 0 — deterministic cross-language guard
arborist/qa/crosslang.py: non-English signal (¿/¡/non-ASCII) + an
es function-word stoppack. Fail-closed to UNGROUNDED before
retrieval/LLM (mirrors the quantifier reject-DAG) when no content
token survives, else strips es stopwords from the retrieval query
only. English path byte-identical by construction. Default OFF
(crosslang_guard_enabled). Measured: the anarcocapitalismo field
case 10.4s -> 1.6s.
#000056 — Operation Sandwich (cross-language grounding)
arborist/qa/mt/: opus-mt es/fr/ru<->en, lazy per-pair memoised
singleton (fixes the 88%-engine-error concurrency defect),
manifest-pinned, [mt] extra; entity_mask wrapper. Sandwich =
translate query in (retrieval + LLM prompt) -> English answer ->
UNTOUCHED verifier grounds English-vs-English -> translate the
verified answer out as display-only (banner-labelled, zero
grounding). question_hash + verifier_policy_hash invariant; MT
engine identity binds into RetrievalPlan, not governance. CLI
--crosslang-translate / make XLANG_MT=1. Default OFF; entity_mask
default OFF (measured net-negative at bench scale). Fan-out bench
(bench/*.py): Spanish ~0% -> 71% grounded vs the real no-support
baseline; the round-trip predictor was tried and refuted; the
entity-mask lever failed at scale (corpus-title anchoring untried).
CLAUDE.md: cross-language bright-line convention + module map.
Pre-existing modified diagram files are intentionally excluded.
Previously linked to the bare domain, which serves a marketing page. The actual OpenAI-compatible endpoint is /v1; /v1/models is the clickable verification (returns the served model card on the live deployment).
Shard capacity convention is documented in arborist/search/fts5.py:113 ("~50ms cold per token on a 10GB shard"). The earlier "~2 GB for a Wikipedia-sized corpus" line in both pagers was a fabricated figure off by 20×. Replaced with the real numbers: per-shard ~10 GB design target, live deployment of four shards totalling ~38 GB and holding 3.5 M documents / 6.2 M chunks.
gitlab-runner rotates between build dirs (builds/RUNNER_ID/0, /1, /3, ...). The cached .venv/ embeds the absolute build-dir path into __editable__.arborist-0.0.1.pth + the dist-info RECORD via pip's editable install. When the cache is restored into a different build dir, pip's implicit uninstall-then-reinstall step fails with `OSError: No such file or directory` looking for files at the old build dir.
Fix: shared .warm-venv-setup hidden job referenced from every real job's before_script. It wipes arborist-*.dist-info, the __editable__*.pth marker, the finder, and the bin/arborist entry script before pip install — so pip does a fresh install at the current build dir. Cache stays warm (deps don't re-download), only the editable-install bookkeeping is rebuilt (~1s per job).
DRY'd via `!reference [.warm-venv-setup, before_script]` so test / bench-suite / substrate-score / test-crawler all share the same setup; each job adds only its own install line ('.[dev]' vs '.[dev,crawler]').
Old drafts read as internal substrate notes. Rewrites lead with what the system does for a consumer or evaluator and what it costs to run, with no references to internal tickets, table names, schema-version strings, governance hash dimensions, or per-record audit-mode tokens. Appendix diagrams updated in lockstep: friendly labels ("grounded / partly grounded / not grounded") replace the schema-column trichotomy, layer names paraphrased away from SURFACE/CORE/PROVIDENCE.
1-pager (docs/_source/arborist-one-pager.rst, 1 page) for AI-literate readers: the trichotomy, the 8-dim cache key, CTI synthetic-elision-impossible, soft-channel separation, real-traffic bench numbers (mis-cite 100% @ 0% FP, warrant 92/92, quote 0.54 STRICT-rate).
2-pager (docs/_source/arborist-two-pager.rst, 3 pages = 2 body + 1 appendix) for technical reviewers: letterhead, Permacomputer Preamble license box, six numbered sections, plus appendix figures (pager-arch-stack 3-layer architecture, pager-verifier-flow question→pointer→verifier→trichotomy).
Both pages live under docs/_source/ so the same RST renders into the Sphinx readthedocs site (toctree caption "Summary pages" added to docs/_source/index.rst) AND into standalone PDFs via rst2pdf (docs/pager.style, lazy install into .venv).
Makefile targets: docs-one-pager, docs-two-pager, docs-pagers, docs-pagers-clean. Diagrams render through the existing DOT pipeline.
Synthesizes the bench-maxing work across:
#000049 NLI recombination veto (bart-large-mnli/k=12/margin/θ=0.999
→ 48% real-haystack recall at 0/808 STRICT FP — partial closure,
Phase-3 semantic candidate selector for full closure)
#000052 §3.1 diagnose_coherence (lexical sidecar, advisory-only,
1.1% real-STRICT FP after round-2 patch)
#000052 §3.2 relevance reranker (bge-reranker-large + cleaned +
θ ≤ -2.42 → 100% mis-cite / 55% deflection / 0% STRICT FP —
motivating Zionist-shape failure fully covered)
The three are architecturally orthogonal (§3.2.2 step 3C verified:
combining lexical sidecars with the relevance reranker gives no
lift; each owns its own failure-shape slice). Three structurally
distinct demote-only signals layered on the binary verifier.
Three runtime-promotion decisions for fox+dav1d:
- §3.1: keep advisory or wire policy hook? (probably advisory)
- §3.2: promote at the fp=0 operating point? (sign-off folds
relevance_policy_hash into governance_policy_hash)
- #000049: promote at 48% partial closure, or wait for Phase 3
semantic candidate selector?
Bench-maxing methodology codified in CLAUDE.md is the transferable
artifact: 'clean candidate-bench can mis-predict in BOTH directions
— real-data fixtures on both precision AND recall axes are the
only load-bearing measurement'. Eight instances across the two
arcs; the discipline applies to any future model-based addition.
Indexed in docs/TICKETS.md 'Distinction from other docs' section as
a non-ticket reference doc. Production verifier unchanged; nothing
in audit_mode; all work SHADOW pending sign-off.
Ran diagnose_deflection and diagnose_coherence on the 20+20 NEG
fixtures from step 3 parts A+B, plus all 808 STRICT.
Result: 0/40 NEG fire on diagnose_deflection (because the fixtures
are token-overlap-correct by construction — the question's subject
appears in the answer; that's exactly the failure mode #000052 was
built to catch beyond lexical). 0/40 on diagnose_coherence (the
fixtures are well-formed sentences).
So:
UNION (relevance OR deflection) = relevance alone (no lift)
INTERSECT (relevance AND deflection) = 0/40
Multi-signal combination doesn't help on these failure shapes.
But the architectural finding is positive in a different way: the
sidecars and the relevance reranker cover NON-OVERLAPPING failure
shapes cleanly:
- diagnose_coherence: structural breakage (word-salad, vacuous,
phrase-component-reuse). Owns the 'incoherent answer' slice.
- diagnose_deflection: token-overlap mismatch (subject anchor not
in answer). Owns the 'wholesale topic drift' slice.
- bge-large relevance: semantic aboutness mismatch despite shared
tokens. Owns the 'topic-collision / mis-cite / on-topic-but-
not-answering' slice — what §3.2 was built for.
Each signal owns its own slice; combining is redundant on these
cases. That's the architectural validation of the §3.1 + §3.2 +
existing-lexical-sidecars split as orthogonal, not overlapping.
Manifest runtime_viability.as_multi_signal_factor updated from
'viable' to 'TESTED — does not lift; sidecars are complementary
not combinatorial on these shapes'. Cleaned up stale
nli-shadow-grid-n1-minilm.json.
Also cleaned up a stale nli-shadow-grid JSON.
The §3.2 arc is now complete:
step 1 candidate-bench (overclaim, contrived data)
step 2 real STRICT FP (over-pessimistic 'NOT VIABLE')
step 3A real-context deflection (positive reversal, 55% at fp=0)
step 3B real-context mis-cite (100% at fp=0, motivating failure covered)
step 3C multi-signal combination (no lift; clean architectural split)
The recommended operating point holds: bge-reranker-large + cleaned
+ θ ≤ -2.42 → 100% mis-cite, 55% deflect, 0% STRICT FP.
Built bench/fixtures/5f/relevance-miscite-realcontext-v1.jsonl —
20 hand-crafted (claim, source) mis-cite pairs: claim about X, source
about Y, X≠Y but shared tokens. Each pair survives the lexical
title-relevance + verifier sidecars by construction. This is the
Zionist-entity failure mode (claim about a different entity than the
cited source, both lexically related).
Examples:
- 'Mercury is the smallest planet' / Roman-god Mercury source
- 'Java is a programming language' / Java-the-island source
- 'Apple Inc. was co-founded by Steve Jobs' / apple-the-fruit source
- 'The Eiffel Tower is in Paris' / Gustave-Eiffel-person source
- 'Mozart composed The Magic Flute' / Mozart-effect-theory source
Headline (bge-reranker-large + cleaned + θ ≤ -2.42):
Mis-cite catch: 20/20 = 100% ← FULL COVERAGE of motivating shape
Deflection catch: 11/20 = 55%
Combined NEG: 31/40 = 78%
Real-STRICT FP: 0/808 = 0% ← strictly safe
bge-large mis-cite scores: -9.37 to -3.81 (max). STRICT min: -2.42.
Mis-cite is STRICTLY SEPARABLE from real STRICT — there's a 1.4-pt
gap with no overlap. (Deflection harder; some overlap with weak STRICT.)
MiniLM-L-6 cost-pick (5× smaller, cleaned, θ=+3):
100% mis-cite + 65% deflect + 0.4% STRICT FP
Mis-cite is structurally MUCH easier than deflection — both models
hit 95-100% mis-cite catch at modest θ; deflection is harder because
'answer doesn't quite address question' can look like a weak STRICT.
That's appropriate: mis-cite is 'wrong topic entirely'; deflection
is 'right topic, not answering'.
Manifest:
- runtime_viability flipped (step 2 → step 3): NOT VIABLE → VIABLE
at the bge-large fp=0 operating point.
- demote_below_score still null pending fox+dav1d sign-off (setting
it folds relevance_policy_hash into governance_policy_hash per
#000049 §7 #2).
- PRIMARY RECOMMENDATION: bge-reranker-large + cleaned +
θ ≤ -2.42 → 100% mis-cite, 55% deflect, 0% STRICT FP.
Production verifier unchanged; still SHADOW. The §3.2 arc went:
candidate-bench (overclaim) → real STRICT step 2 (over-pessimistic
NOT VIABLE) → real-context deflection step 3 (positive reversal,
55% at fp=0) → real-context mis-cite step 3 part B (full closure
at fp=0 for the motivating failure mode). Eight meta-lesson
instances over the §000049 + §000052 arc, with the sharpest one
yet: clean candidate-bench can mis-predict in BOTH directions —
real-data fixtures on BOTH precision and recall axes are the
only load-bearing measurement.
Built bench/fixtures/5f/relevance-deflection-realcontext-v1.jsonl —
20 hand-crafted deflection answers against real bench-qa questions
(each: coherent well-formed answer using context-tokens, but NOT
addressing the question — the kind of failure §3.2 was built for).
Real-context deflection scores (cleaned, MiniLM-L-6):
min -6.71, median +1.56, max +6.34
Real-context deflection scores (cleaned, bge-reranker-large):
min -5.57, median -3.02, max +5.01
vs STRICT (cleaned):
MiniLM-L-6: min -5.29, p10 +6.90, median +9.32
bge-large: min -2.42, p10 +3.56, median +6.48
Distributions are CLEAN-separable on real data — 0% of MiniLM
deflections score above p10(STRICT); bge-large median is -3.02 vs
STRICT median +6.48.
Pareto frontier (MiniLM-L-6, cleaned):
θ=-5.29: catch 10% / fp 0% (fp=0 floor, low signal but safe)
θ=+3: catch 65% / fp 0.4% ← strong runtime soft-veto
θ=+5: catch 90% / fp 1.9% ← aggressive soft-veto
θ=+7: catch 100% / fp 10.4% (too FP for runtime)
bge-reranker-large:
θ=-2.42: catch 55% / fp 0% ← VIABLE runtime soft-veto at strict fp=0
The step-2 'NOT VIABLE' verdict was an artifact of using the
candidate-bench NEG distribution (tight contrived band, overlapped
real STRICT) as the recall denominator. Real-context deflections sit
in a much lower band (median around -3 for bge-large) than real
STRICT, so absolute thresholds DO separate them cleanly. The
relevance reranker IS a viable runtime soft-veto on this design —
the candidate-bench-only step 2 measurement misled us.
Manifest's runtime_viability.as_runtime_demotion_veto flipped from
'NOT VIABLE' to 'VIABLE at low-to-moderate FP', with both per-model
Pareto frontiers recorded. demote_below_score stays null pending
fox+dav1d sign-off; setting it folds relevance_policy_hash into
governance_policy_hash per #000049 §7 #2 discipline.
Still SHADOW; production verifier unchanged; no audit_mode effect.
Eighth instance of the meta-lesson, with the lesson sharpening AGAIN:
clean candidate-bench can MIS-PREDICT in BOTH directions — over-
optimistic on threshold (step 2) AND over-pessimistic on viability
(step 3 reversal). Real-data fixtures are the only load-bearing
denominators on either axis.
Tested Q→A score vs A→context_lead score on 808 pooled STRICT pairs
with cleaned MiniLM-L-6. The hypothesis was: deflected answers
should have HIGH A→ctx_lead but LOW Q→A (still grounded but
off-topic). Reality:
Q→A quantiles: p10 +6.90 p50 +9.32 p90 +10.57
A→ctx quantiles: p10 -6.73 p50 +1.51 p90 +8.68
Δ=Q→A-Actx: p10 +0.90 p50 +6.95 p90 +14.62
Of the 73 cleaned STRICT-fires at cb-θ=6.754:
- 71 (97%) have BOTH axes low (co-varying — not deflection)
- 2 (3%) match the deflection signature; both are the same
'who wrote GNU linux?' replica
Why it doesn't separate: the A→ctx axis isn't measuring what the
hypothesis assumed. A focused-claim against a 30KB topic-broad
Wikipedia haystack scores LOW by default — that IS the normal
STRICT shape (the answer is one clause in a sprawling document).
The cross-encoder expects the document to be ABOUT the query
(MS-MARCO retrieval shape); a STRICT (answer, full-haystack) pair
violates that. So both axes co-vary and the delta is a poor
discriminator.
This closes the third precision-side rescue path I'd left open
after step 2's verdict:
✗ absolute threshold — universal walk-back across 6 models
✗ cleaning preprocessing — helps but doesn't separate
✗ contrastive Q→A vs A→ctx — doesn't separate either
The relevance reranker conclusively cannot be a runtime
demotion-only veto on this design. The soft-signal uses
(render-tail, multi-signal advisory, etc.) remain viable.
runtime_viability.as_contrastive_signal updated from 'worth
measuring' to 'tested, doesn't separate'. Production verifier
unchanged; no audit_mode effect.
Hand-inspection of the bottom-15 STRICT-fires from the raw §3.2.2 step 2
sweep showed claim-lattice overlay markup ([E\d+ | title | hash: '…'])
depressing scores on correct concise answers (the 6× Henry-VIII case),
while true-positive deflections (broad-question / narrow-answer like
'winners of all major sports?' → just-one-sport) remained correctly
low-scored. So the noise FP class is the bracket metadata; cleaning it
should reduce FP without losing true-positive signal.
Built clean_for_relevance() in arborist/qa/relevance/shadow.py — strips
[E\d+ | ... ] blocks + trailing '...']' tails. Baked into
ShadowRelevance.check_question_answer / check_claim_source by default
(opt out with clean_input=False). relevance_shadow_sweep.py applies it
to inputs before _score_batch (opt out with --no-clean).
Re-ran the full 6-model sweep on the 808-cell pooled STRICT with
cleaning:
bge-reranker-large 21.8% → 18.6% (-3.2)
MiniLM-L-4-v2 15.0% → 9.4% (-5.6 pts, -37% rel)
MiniLM-L-6-v2 11.6% → 9.0% (-2.6)
MiniLM-L-12-v2 11.0% → 8.0% (-3.0)
bge-reranker-base 9.5% → 9.0% (-0.5)
MiniLM-L-2-v2 4.5% → 1.5% (-3.0 pts, -67% rel)
Universal improvement, every model better. Big surprise: MiniLM-L-2-v2
— the model that FAILED the candidate-bench separability (margin
-1.97, declared 'capacity floor') — has the LOWEST real-traffic FP
rate at its own cb θ (1.5%). Because L-2's compressed score range
gives it a low cb θ which few real STRICT pairs score below.
SEVENTH instance of 'candidate-bench doesn't predict real-traffic'.
Runtime-veto verdict UNCHANGED — still not viable; smallest fp=0 θ on
real STRICT is below the cb NEG max for every model, so at any
runtime-safe θ the catch on cb NEG is 0/12. But cleaning is now FREE
improvement for any soft-signal / advisory / contrastive use of the
relevance score. Hand-inspected bottom-10 post-cleaning confirms true-
positive deflection signal preserved.
All 6 swept rerankers (bge-large/base, MiniLM-L-2/L-4/L-6/L-12)
false-fire on 4.5–21.8% of real bench-qa STRICT at their
candidate-bench fp=0 θ. The smallest θ that yields fp=0 on real
STRICT is BELOW the candidate-bench NEG max for every model —
meaning at the runtime-safe θ, catch on the 12 candidate-bench NEG
= 0/12 across the board.
Structural reason: real bench-qa STRICT answers have a much wider
score distribution (bge-large STRICT: min -2.20, p10 +2.99, p50
+5.80, p90 +7.54) than the tight contrived candidate-bench POS
band. The bottom 10% of legitimate STRICT score below where the
candidate-bench NEG cases sat. Distributions overlap heavily; no
threshold separates them.
This is the §3.2 mirror of #000049 §7 #27's recall-side walk-back —
clean candidate-bench → fails real-pipeline gate. Same diagnosis:
lexical-candidate selection + cross-encoder scoring + hard threshold
doesn't survive real-pipeline heterogeneity.
Verdict: relevance reranker CANNOT be promoted to a runtime
demotion-only veto on this design. demote_below_score stays null;
manifest gains runtime_viability block documenting the negative
result + the still-viable advisory soft-signal uses (render-tail,
multi-signal advisory, contrastive Q→A vs claim→source delta).
bench/scripts/relevance_shadow_sweep.py + bench/results/
relevance-shadow-sweep-pooled808.json committed. Production
verifier unchanged; advisory only; no audit_mode effect.
Five rule tightenings, each targeting a specific bench-qa STRICT
false-positive shape (the xfail regressions from the previous commit):
1. claim-lattice bracket-artifact skip — sentences matching
`\[E\d+\s*\|` (pointer markup) or `..."\]` (truncation tail)
are no longer parsed as natural-language assertions; `Such a
thesis was..."]` no longer fires vacuous.
2. Circular-rule differentia cap — circular now requires the
predicate to be (a) entirely vacuous OR (b) leads with a subject
token AND has ≤ 2 non-subject non-filler differentia tokens.
"Michael Jordan's Restaurant was a restaurant in Chicago,
Illinois, named after the basketball player Michael Jordan"
has 6 differentia → no longer fires. Pre-existing positive test
"The entity is the entity referring to the State of Israel"
has 2 differentia (state, israel) → still fires (under threshold).
3. phrase_component_reuse translation-chain exception — when the
predicate ALSO contains a quoted phrase (translation /
definition / etymology context), token reuse with the subject's
quoted phrase is legitimate, not circular. "The name 'Rosebud
River' is a translation … 'the river of the roses'" no longer
fires.
4. Vacuous-rule short-acronym escape — `_coherence_predicate_has_short_acronym_content`
recognizes title-cased element-symbols (Au, Fe, Pb…) and all-caps
2-5-char acronyms (DNA, FBI, USB, NASA…) as content even though
they're below the ≥3-char content-token filter. "The chemical
symbol for gold is Au." no longer fires; "Iron has the chemical
symbol Fe." also clean; tautology "DNA stands for DNA." still
correctly flagged circular.
5. (the 'term <X>' idiom xfail stays xfail — borderline, no clean
lexical fix.)
5 xfail → passing (the regression coverage is now executable proof
of fix); 1 xfail remains. Pooled bench-qa STRICT FP rate test ceiling
tightened from 7% to 2%. All 7 pre-existing positive coherence tests
still fire correctly. 110 total tests pass; 1 xfailed; no regressions
on doc-counts / nli / relevance.
§3.1 diagnose_coherence (now 19 tests, +10 from the parallel session's 9):
- 3 more positive shapes (multi-sentence vacuous, named-entity circular,
grammar-term phrase_component_reuse).
- 5 xfail regression tests for SHAPES THAT FALSE-FIRE on real bench-qa
STRICT data (44/808 = 5.4% FP rate measured on the pooled n=1+3+5
STRICT answers). Each xfail names the exact shape + why it should
ideally be 'ok' + which rule needs tightening:
* 'The chemical symbol for gold is Au.' → vacuous (short predicate)
* 'Michael Jordan's Restaurant was a restaurant ... named after
Michael Jordan.' → circular (named-after re-use)
* 'The Western X was the western half of the X' → circular
* 'The name <Phrase> is a translation ... of the <derivative>' →
phrase_component_reuse (translation/etymology)
* claim-lattice [E1 | … …"] tails → vacuous (truncated bracket
fragment)
* 'The term <X>' → phrase_component_reuse (idiomatic English)
- 1 load-bearing real-traffic test: FP rate on 808-cell pooled STRICT
must stay ≤ 7% (current 5.4%) — fires loud if a future change
regresses it. Skips on fresh-checkout (bench/qa_results/ gitignored).
§3.2 ShadowRelevance (now 20 tests, +7 from the round-1 scaffold):
- Manifest tests for round-2 primary (bge-reranker-large), the size
spectrum coverage (50-560MB), the candidate-bench findings block
(biggest-within-family / not-across-families / deeper-not-better /
capacity-floor).
- Pair-kind distinction (question_answer vs claim_source recorded
separately for downstream telemetry / governance hashing).
- Batch-order preservation (_score_batch must return scores in input
order — load-bearing for downstream zip-back).
- Empty-input handling (Q empty, D empty, whitespace-only).
- Zionist-entity discriminator sanity (on-topic > off-topic logit).
- demote_below_score-stays-null invariant (the §7 #18→#27 discipline:
no hardcoded threshold; must come from a real-traffic shadow sweep).
Total: 101 passed + 6 xfailed (5 §3.1 regressions documented + 1 from
parallel session). The 5 xfails are the bench-maxing receipts — they
document EXACTLY which shapes §3.1 false-fires on, with the rule that
needs tightening named in each reason.
`make bench-qa BENCH_QA_N=3 BENCH_QA_LIMIT=5` (2026-05-13T14:24Z) on
the bench/qa_questions.txt set: 30/45 STRICT (67%), zero regressions
on mona lisa / capital of france / new london bridge (each 9/9 STRICT
across quote / pointer / lattice modes). The bench-qa-smoke n=1
flickers I saw ("dinosaurs" S→U, "soviet union" S→H) were LLM
stochasticity at n=1, not retrieval bugs. GNU-linux + python failures
look like genuine hard questions, not Phase-2 side effects.
End-to-end gap-close from the Phase 1 extractor. Five interlocking
fixes; live-verified that `what is a CPU?` → "Central processing unit"
at #1, `what is a GPU?` → "Graphics processing unit" at #1
EVIDENCE-WARRANTED 1/1; Mount Kilimanjaro / Soviet Union queries
unchanged (no regression).
(a) `synonym_expand` over-cap path is now rank-and-truncate by
source-frequency (descending) instead of hard-skip. CPU has 13
legitimate homonym expansions across the corpus; the prior
MAX_NEIGHBORS_PER_TOKEN=8 cap contributed *zero* expansion → no
canonical-article surfacing. Now: keep the 8 dominant by per-(token,
target) source-root count via _load_neighbor_source_freq.
(b) `_search_titles` orders by FTS5 bm25 ASC instead of LENGTH(title)
ASC on the FTS5-MATCH path. The length-asc tie-break was correct for
the 2026-05-02 "Back to the Future" LIKE-substring case but
counter-productive on FTS5 (tokenized; no substring junk; length-asc
preferred "Unit" / "Unite" / "B unit" over "Graphics processing
unit"). LIKE fallback keeps length-asc since the substring issue
persists there.
(c) `accept_tokens` (feeds title-search, core-keyword, title-rerank)
uses the synonym-expanded set instead of qtokens-only. The Phase 1
expansion existed but was only used in the FTS5 OR-fallback; satellite
articles saturated the budget before the canonical article entered.
(d) ARCHITECTURAL: `synonym_expand_strict()` (new — high-trust
evidence-kind subset: manual + manual_legacy + acronym_parens,
**excludes** link_reciprocity) for use in the multiplicative
`_rerank_by_title_purity` and as the source for `accept_tokens`.
Reciprocal-wikilink edges express *topical adjacency*, not synonymy
(a `Dinosaurs` page reciprocally links to `Curious George Brigade` →
edge that should not amplify retrieval); a multiplicative ranker over
them blows up. Strict view preserves the acronym-parens surfacing
(those edges ARE the phrase=expansion identity) while keeping
link-reciprocity to additive retrieval-route boosts via the broad
synonym_expand (still wired to `or_synonym_pool` for the FTS5 OR
fallback).
(e) Extractor regex tightened `[A-Z]{2,6}` → `[A-Z]{3,6}` and purged
~21K 2-letter acronym edges from shards. 2-letter acronyms (AI/ML/OS/
US/UK/IT/PC/TV) homonym-collide too often with common 2-letter QUERY
tokens like `go`/`is`/`am` — without this, "why did the dinosaurs go
extinct?" pulled Curious George Brigade via GO-acronym edges. The
high-value acronyms (CPU/GPU/RAM/DNA/FBI/WHO/…) all clear 3 chars.
Also: CLAUDE.md gains a "prefer existing ticket; only split for
Dav1d-review audience" discipline note (saved as feedback memory) —
this work is itself an example: would have been #000055 + #000056 +
#000057 under the prior pattern; instead extends #000054.
Suite: 2531 passed (no regression). bench-qa in flight separately.
All 4 manifest candidates clean-separate (12/12 NEG catch at 0/14 POS FP),
so 'reranker discriminates Zionist-entity-style mis-cites' is a property
of MS-MARCO-trained rerankers as a class — not the specific L-6 I picked
first. Ranked by separation margin (per §7 #18: separation beats raw):
ms-marco-electra-base +5.349 ← new primary
BAAI/bge-reranker-base +3.076
ms-marco-MiniLM-L-6-v2 +3.053 (previous primary, demoted)
ms-marco-MiniLM-L-12-v2 +2.663 (worst — deeper ≠ better)
Manifest primary moved to electra-base for the cushion. demote_below_score
STAYS null — clean candidate-bench thresholds don't predict real-pipeline
behavior (the §7 #18→#27 history is 6 verdict flips on the NLI side);
§3.2.2 step 2 (real-traffic shadow sweep on pooled bench-qa STRICT) is
what sets it. Expect a walk-back. Reranker still doesn't catch the
Kilimanjaro/Mount-Kenya recombination (aboutness ≠ truth-of-attribution;
that's #000049 territory). fox's bench-maxing correction applied.
arborist/qa/relevance/ — manifest pins cross-encoder/ms-marco-MiniLM-L-6-v2
(~80MB, Apache) as primary; alternates: L-12, BAAI/bge-reranker-base,
ms-marco-electra-base. demote_below_score=null on purpose — the
#000049 §7 #18→#27 discipline (proved 6× that clean-eval thresholds
don't transfer to bench-qa data) requires the threshold to be set by a
shadow sweep against pooled real STRICT, not by a literature number.
ShadowRelevance class mirrors ShadowNLI (lazy [nli]-extra import, cuda
auto-detect via ARBORIST_RELEVANCE_DEVICE or ARBORIST_NLI_DEVICE,
batched _score_batch, graceful degrade-to-available=False). Two surface
methods: check_question_answer (deflection / Q-A drift) and
check_claim_source (topic-collision mis-cite). 13 tests.
Sanity on the motivating field case (Zionist entity): ON-topic +9.96
vs OFF-topic -9.04 → 18-pt margin. Mona Lisa Q→A deflection: on +10.45
vs deflect +3.56 → ~7-pt margin. The model CLEANLY discriminates the
failure modes #000052 §1 named. It does NOT catch the
recombination-where-the-different-entity-clause-also-mentions-the-target
case (Kilimanjaro/Mount Kenya) — and that's the right architectural
split: aboutness (#000052 §3.2) and entailment (#000049 NLI) are
orthogonal axes; the Kilimanjaro recombination case needs the semantic
candidate selector (#000050/#000051 vec hybrid).
Remaining: build candidate-bench eval (~20-30 deflection + mis-cite
fixtures), shadow-sweep θ over pooled bench-qa STRICT (expect another
walk-back per the #000049 lesson), recall-side realism check, then
fox+dav1d sign-off. Still SHADOW; production verifier unchanged.
test_concepts_extract.py grew 20 -> 28 with the acronym_parens
extractor's 8 tests; the AUTOCOUNT discipline (CLAUDE.md) fires the
doc-counts regression on stale numeric claims.
`arborist/concepts/extract.py:acronym_parens_synonym` — new
corpus-agnostic extractor. Scans each doc's lead chunk (first 4000
chars) for `<Multi-Word Phrase> (ACRO)` where the all-caps acronym's
letters strictly match the content-word initials of the phrase, in
order, after function-word filtering. Emits bidirectional synonym
edges between the lowercased acronym and each ≥3-char content token
of the phrase, evidence_kind="acronym_parens", anchored to that doc's
document_root. Idempotent like link_reciprocity_synonym.
Why this complements link_reciprocity: Wikipedia represents
abbreviation→expansion as a one-way *redirect* (CPU →
Central processing unit), which the ingest does not record as an
edge — so the existing reciprocal-link extractor never learned the
relation. The relation IS in body text by near-universal convention
("Central processing unit (CPU) is..."), which this extractor reads.
Corpus-agnostic: HTML, blogs, textbooks benefit equally.
Conservative: strict 1:1 acronym-to-atom match (rejects HTTP-shape,
where letters land mid-word), function words filtered, repeated
definitions deduped per doc, ≥3-char target floor. 8 new tests
covering CPU bidirectional emit, RAM idempotency, FBI function-word
filter, HTTP length-mismatch reject, XYZ initial-mismatch reject,
ROM hyphenated-word handling, per-doc dedupe, registry presence.
Retrieval-side only — synonym edges reshape FTS5 candidate selection
via synonym_expand at query time, never enter audit_mode / cache_key
/ audit_event_hash. No governance hash bump, no cache invalidation.
Closes#000050 §2a's CPU/GPU abbreviation rows *upstream* of vec;
the Orwell-shape conceptual-allusion row remains the genuine #000050
justification. Operational follow-up (not code): run on each shard
via `arborist concepts derive --extractor acronym_parens` (CLI
surface itself is aspirational in docstrings; extractors are called
programmatically today). Next ID 000054 -> 000055.
Built bench/fixtures/5f/recombination-realcontext-v1.jsonl: 25 hand-built
recombination claims (deliberate re-attributions within real bench-qa
30 KB Wikipedia haystacks the pipeline produced — Mount Kenya into a
Kilimanjaro answer, Galileo onto all 63 Jupiter moons, Joshua Reynolds
onto the Mona Lisa, etc.). Ran the full 7-model × full-grid sweep vs
the 808 pooled STRICT. Result:
- bart-large-mnli / k=12 / max / θc=0.999 → 12/25 = 0.48 catch · 0/808 FP
(the real-haystack fp=0 leader)
- deberta-large-mnli / k=3 / margin / θc=0.995 → 6/25 = 0.24 (§7 #26's
'settled' config — 28/28 synthetic, 0.24 real-haystack: 4× over-estimate)
- roberta-large 0.12, MiniLM 0.08, deberta-base 0.04
So §7 #26's 'boundary closed' walks back to 'boundary PARTIALLY closed'
on real haystacks. The bottleneck is architectural: top-k by token
overlap misses the contradicting clause when it shares few subject-area
tokens with the answer (e.g. the Mount Kenya clause only shares 'Kenya'
with a Kilimanjaro claim — ranked low, NLI never sees it). Threshold
tuning doesn't lift the ceiling; a SEMANTIC candidate selector
(vec-driven, sibling of #000050/#000051's hybrid retrieval) does.
bart's pareto above fp=0: fp=0.011 catch=0.52, fp=0.057 catch=0.84,
fp=0.068 catch=0.92 — permissive operating points are on the menu if
fox+dav1d sign off. recommended_operating_point updated to
bart-large-mnli/k=12/max/θc=0.999; deberta-large/margin kept as the
synthetic-eval reference. Sixth meta-lesson instance: clean synthetic
eval doesn't predict bench-qa precision OR recall — neither contrived
dataset axis is load-bearing, only the real pipeline shape is.
Production verifier unchanged; falsification-hard stays 10/12. Still
SHADOW; runtime promotion fox+dav1d-decides.
(a) Mined the pooled n=1+3+5 bench-qa runs (808 distinct STRICT answers)
for natural recombinations at lowered θc≥0.7 → 37 would-fires, ALL
token-collision FPs on inspection (Mount Kenya pulled into a Kilimanjaro
answer, Dalí into da Vinci, Donovan into Superman). ZERO genuine
recombination errors — the boundary is theoretical-in-practice; the
failure mode is the candidate selector (top-k by token overlap) pulling
different-entity same-subject-area clauses.
(b) Re-ran the grid against the 808-cell pooled STRICT set: the §7 #25
'max@0.96' was itself a small-sample artifact — the n=5 444-cell set
lacked the high-confidence spurious hits the pooled set has. On 808
cells θc goes back to ~0.995, and at θc=0.995 only agg=margin still
catches 28/28 (max gets 27/28). microsoft/deberta-large-mnli / k=3 /
agg=margin / θc=0.995 → 28/28 synthetic recombinations · 0/808 pooled
real STRICT FP · 0/26 synthetic legit — the ONLY config in the
7-model×full-grid sweep that hits 1.0/0.0 on 808 cells, held at n=3
too. recommended_operating_point reverted to margin@0.995.
Realistic next check: ~20-30 hand-built synthetic-recombination-vs-
real-bench-qa-context fixtures (real haystack, deliberate re-attribution).
Still SHADOW; runtime promotion fox+dav1d-decides. Production verifier
unchanged; falsification-hard stays 10/12.
Meta-lesson instance five: a bigger sample can vindicate a config a
smaller one made look unnecessary — re-confirm the config choice (not
just the threshold) each time the denominator grows.
Enumerates the concrete query-words-share-zero-tokens-with-target-title
cases the §2 gate's "semantic-allusion fixtures" must include, as a
running list: Orwell→Eastasia (genuine conceptual allusion — the case
that justifies the vec layer), "what is a CPU?"→Central processing
unit, "what is a GPU?"→Graphics processing unit (abbreviation→expansion
subclass — also fixable upstream by a concepts/ synonym edge; bench
records which fix closes each row). Field cases 2026-05-13, fox.
#000053 fixed the verifier's separate acronym blind spot but not this
retrieval gap.
ARBORIST_NLI_SHADOW=1 make bench-qa BENCH_QA_N=5 → 1125 cells, 444 real
STRICT (5x n=1, 1.6x n=3). Re-ran the 7-model mega-grid: the lexical-
candidate NLI veto robustly clears the §7 #12 gate with
microsoft/deberta-large-mnli / k=2 / agg=max / guard=max_entail /
θc≈0.96 / θe=0.9 → 28/28 synthetic recombinations (incl. both fixtures)
· 0/444 real STRICT FP · 0/26 synthetic legit FP, ~4 pts θc headroom;
roberta-large-mnli equally good (k=2/max/θc=0.95). Resolved: agg=max +
max_entail guard is the robust score-shape (§7 #24's margin win was a
sample tie); the large checkpoints (~350-400M deberta-large-mnli /
roberta-large-mnli) hit 1.0/0.0, bart-large (similar size) only ~0.71,
deberta-base at the cliff (27/28 here, 11/28 at n=3), small models
(MiniLM-82M, deberta-v3-small) cap at ~0.82 — so §7 #18's 'MiniLM is
the cost-pick' is OVERTURNED by the proper-n evidence; k=2 consistent
winner; int8-ONNX costs ~1 catch. recommended_operating_point updated.
Remaining: a bench-qa-derived recombination set (the one check not
done); n=9 if dav1d wants more; still SHADOW; runtime promotion is
fox+dav1d-decides. Production verifier unchanged; falsification-hard
stays 10/12. This run is the worked example behind CLAUDE.md's new
bench-maxing line.
`arborist.qa.evidence._content_tokens` dropped every token under 4
chars, so a short all-caps acronym (CPU, GPU, DNA, FBI, USB…) never
registered as a content token — which defeated Rule 8
(_claim_title_overlap / TITLE_MISMATCH), the subject-tokens-absent
check (Rule 9), the bare-name-claim guard, and spotlight-excerpt token
selection whenever a question/claim's topic IS an acronym. The field
case: `what is a CPU?` cited to the "CPU design" article tripped
TITLE_MISMATCH even though claim and title both contain "CPU".
Fix: keep a token if it's an all-caps 2-3-char alpha run in the source
text; everything else unchanged. The change only ever ADDS tokens, so
TITLE_MISMATCH / SUBJECT_TOKENS_ABSENT / BARE_NAME_CLAIM can only stop
firing, never start — monotone toward fewer spurious demotes; no
STRICT→non-STRICT transition is possible from it.
Versioned: `content_token_rules: "v2-acronym-aware"` added to
runner.DEFAULT_POLICY + query.DEFAULT_QUERY_POLICY +
keys._VERIFIER_POLICY_FIELDS → folds into verifier_policy_hash, prior
cache records orphan on lookup (by design; same discipline as
base_version / hyphen_fold_v1). Does NOT touch the retrieval
abbreviation→expansion gap (CPU→Central processing unit — #000050 vec
hybrid / concepts/ synonym edges; the root cause of the satellite
retrieval). 8 new tests; full suite green (2502); bench-qa-smoke clean.
Next ID 000053 -> 000054.
ARBORIST_NLI_SHADOW=1 make bench-qa BENCH_QA_N=3 → 275 real STRICT cells
(3x the n=1 sample). Re-ran the expanded grid (7 aggregations incl.
margin = max_clause(p_contra - p_entail), paired-entail guard variant,
θc to 0.999, --extra-models) over all manifest models + 4 extra
xsmall→large (microsoft/deberta-large-mnli, roberta-large-mnli,
nli-deberta-v3-small, deberta-v3-xsmall), synth-28 recombination vs the
275 STRICT cells. Result: the §7 #23 deberta-base/k=2/θc=0.99 config
does NOT survive — it catches only 11/28 at θc=0.995 (which the larger
STRICT sample forces). BUT the broader sweep found the config that does:
microsoft/deberta-large-mnli / k=3 / agg=margin / θc=0.995 → 28/28
synthetic recombinations (incl. both fixtures) + 8/12 falsification-hard,
0/275 real STRICT FP, 0/26 synthetic legit FP — a passing config at
proper n. Findings: margin is the right score-shape (single threshold,
folds the guard in); the specific checkpoint matters more than param
count (deberta-large-mnli wins clean, deberta-base collapses,
roberta/bart ~0.71-0.75 — no 'bigger is better' law). recommended_operating_point
updated. Still SHADOW; runtime promotion needs a bigger STRICT sample +
a bigger recombination set + fox/dav1d sign-off. Production verifier
unchanged; falsification-hard stays 10/12.
Meta-lesson sharpened twice: clean eval ≠ bench-qa precision (§7 #18→#20),
default config ≠ best config (§7 #22→#23), small FP sample ≠ large FP
rate (§7 #23→#24).
The §7 #22 'fails the gate' was the verdict for the DEFAULT config
(k=6/θc=0.5/θe=0.9, tuned on the clean synthetic set), not the
approach. A {model × candidate-cap k × aggregation × θc × θe} grid
sweep (bench/scripts/nli_shadow_grid.py — NLI runs once per
(model,record) over the top-12 candidate clauses, the k/agg/θ grid is
then arithmetic on cached scores; ~10s on the 4090 for 4 models) finds
clean passing configs: on the 89 real STRICT cells (n=1 bench-qa),
deberta-base-184M / k=2 / agg=max / θc=0.99 / θe=0.9 → 27/28 synthetic
recombinations caught (incl. both 5f-fal-hard fixtures), 0/89 STRICT
FP. MiniLM-82M passes too (24/28 · 0/89). Model science: 184M > 82M >
407M for fp=0 recombination recall; int8-ONNX costs ~1 catch vs fp32.
Caveats: FP side is n=1 (BENCH_QA_N=3 run in flight); recall is on the
synthetic set; flipping to a runtime demotion-only veto is fox-decides
(then nli_policy_hash folds into governance_policy_hash per §7 #2).
Manifest active defaults stay k=6/θc=0.5; recommended_operating_point
(deberta-base, k=2, agg=max, θc=0.99, θe=0.9) documented in the
manifest. 3 grid result JSONs committed. Standing lesson: neither the
clean synthetic eval NOR the default config predicts bench-qa precision
— you have to sweep. Production verifier unchanged; falsification-hard
stays 10/12.
First grid (MiniLM, n=1 bench-qa 89 STRICT cells): the §7 #22 'fails the
gate' verdict was config-specific (defaults k=6/θc=0.5/θe=0.9) — the grid
finds k=1/θc=0.93/θe=0.5 → catch 8/12 falsification-hard (incl. BOTH
recombination fixtures hard-003 Mercury + hard-005 Einstein), 0/89 STRICT
FP, 0/26 synthetic-legit FP; 18/28 synthetic recombination. So the
lexical-candidate approach is NOT a dead end. (k=1 = no haystack;
θe=0.5 tighter than the clean-set 0.9.) Full multi-model + n=3
characterization next.
bootstrap-nli-only: a venv + [nli] only — no [dev] extras (no crawler/
hessian/vec/sympy). On a CUDA host PyPI's torch wheel is the CUDA build,
so ShadowNLI auto-detects cuda and bench-nli-shadow / export-nli-onnx /
the candidate bench all run on the GPU with no further wiring. Verified
on the ai box (RTX 4090): torch 2.11+cu130, cuda True; 82M cross-encoder
batched ≈ 0.09 ms/pair (512 pairs in 0.047s — ~350x onnx-int8-cpu,
~1300x torch-cpu-batch1); 62-record synthetic shadow sweep in 2.4s wall;
24/24 tests pass. The GPU only accelerates the NLI half — make bench-qa
(Hermes LLM) still runs wherever the shard corpus is.
Speedup (§3 plan): ShadowNLI._nli_batch batches forwards
(ARBORIST_NLI_BATCH=64); device auto-detect (ARBORIST_NLI_DEVICE, else
cuda-if-available); auto-prefer an ONNX export — bench/scripts/export_nli_onnx.py
/ make export-nli-onnx exports + int8-dynamic-quantizes the pinned
checkpoint into ~/.arborist/models/nli/<ver>/onnx/ (operator state, NOT
committed), _ensure_loaded loads model_quantized.onnx via
optimum.onnxruntime (backend onnx-int8), falls back to torch silently.
torch-cpu-batch1 ~120ms/pair → onnx-int8-cpu-batched ~32ms/pair (~4x);
seconds on a 4090. optimum[onnxruntime] added to the [nli] extra; 24
tests.
Gate-item-4 verdict at proper n: ARBORIST_NLI_SHADOW=1 make bench-qa
BENCH_QA_N=1 → 223 cells (89 STRICT / 90 HYBRID / 44 UNGROUNDED; also
surfaced + fixed a lone-surrogate bug). Shadow sweep over those: NLI-as-
runtime-veto on STRICT has ~26% FP at θc 0.5, ~8% at θc 0.90, ~0% only
at θc 0.99 — and θc 0.99 gives up most recombination recall (hard
synthetic recombinations bottom out ~0.76). FAILS the §7 #12 gate on
this design. Only untried path that might pass: a Phase-3 runtime hook
running NLI on the verifier's actual matched clauses (1-3), not
top-6-by-overlap. Until then: runtime NLI demotion stays off; the 2
fixtures stay permanent boundary markers; θc stays 0.5. Production
verifier unchanged; falsification-hard stays 10/12.
Per-sentence shape check (no model) emitting kind ∈
{phrase_component_reuse, circular, vacuous, ok, empty}:
- circular: subject content-tokens ⊆ predicate's and the predicate
leads with a subject token ("Water is water").
- phrase_component_reuse: subject quotes a phrase, predicate reuses
one of that phrase's own tokens as a bare "the/a/an <token>"
referent — the 2026-05-12 field case ("the phrase 'Zionist entity'
is used as the entity"), a token collision the verifier +
deflection + title-relevance all pass and NLI returns neutral on.
Copulas inside a quoted span are skipped so 'war is peace' doesn't
break the subject/predicate split.
- vacuous: predicate is only placeholder hypernyms + filler ("X is
a thing").
Conservative — no full token-salad parsing; legit definitions pass ok.
Surfaced in inspect_cache_key + the `arborist inspect` human view
(· incoherent: <kind>). Advisory only — never writes providence_cache
/ audit_events / run_dag_root; demote-only verifier hook deliberately
not wired. 9 tests; full suite green (2500 passed).
ARBORIST_NLI_SHADOW=1 carries the raw verifier-input text into bench
rows; real Wikipedia context occasionally has U+D800–U+DFFF code points
(mangled source encoding) that json.dumps(..., ensure_ascii=False) then
refuses to UTF-8-encode → the run died at row 224/225. Scrub via
encode('utf-8','replace').decode() — U+FFFD is fine for a measurement
field. Only the two new shadow fields are touched.
Doc-only scaffold. Two more read-only/demote-only/never-in-proof-path
sidecars joining the diagnose_deflection family: (1) diagnose_coherence
— word-salad/circular/vacuous answers; lexical, no model; the near-term
win. (2) diagnose_relevance — semantic 'aboutness' (does the answer
address the question / is each claim about its cited source?); today's
checks are lexical and a token collision defeats them; a small
aboutness/reranker model (NOT NLI) under #000049 §7's discipline cage
verbatim; gated on evidence, travels with #000049's model question.
Motivating field case: the 'Zionist entity' claim_lattice query —
incoherent token-collision recombination NLI can't catch (returns
neutral) and both lexical relevance checks waved through. Flags an
upstream retrieval (polysemy/title-soup) root-cause ticket, not scoped
here. Next ID 000052 → 000053. #000049 sibling
candidate_clauses() — NLI now runs only on the top-N source clauses by
content-token overlap with the answer claim (max_candidate_clauses=6),
not the whole context; records n_candidate_clauses / best_clause_overlap
/ recombination_risk. Synthetic sweep unchanged (28/28 recombination,
0/26 legit FP, mean 1.45 candidate clauses/record). Real-traffic smoke
re-run: STRICT would-demote 30% → 20%, overall 47% → 33% — better, not
fixed; recombination-risk split doesn't separate either. Residual STRICT
false-contras at ~0.83-0.92 → θc would need ≈ 0.90 (vs the clean-set
0.5); at θc=0.90 the data in hand gives 27/28 synthetic recall, 0/26
legit FP, 0/10 smoke STRICT FP — but n=10 is too small to set on.
Next: a fuller ARBORIST_NLI_SHADOW=1 bench-qa run → sweep θc on hundreds
of STRICT cells → confirm → set it. θc stays 0.5; runtime NLI demotion
stays off. Production verifier unchanged; falsification-hard stays 10/12.
Live hook: ARBORIST_NLI_SHADOW=1 makes query() surface the verifier-input
text (gated off-by-default, never a cache_key/governance/audit_mode
input); qa_sweep.py carries it + the answer into bench rows; the shadow
sweep reads them and buckets by audit_mode. ARBORIST_NLI_SHADOW=1 make
bench-qa-smoke (15 cells) → the naive 'NLI on every context clause'
scaffold has a ~30% would-demote rate on STRICT answers — a haystack /
multiple-comparisons artifact (real contexts segment into 100-336
clauses; max-over-all almost always finds a tangential clause the model
reads as contradicting; a paraphrased STRICT answer often isn't verbatim-
entailed by any single clause so the entailment guard doesn't rescue it).
Lesson: the §7 #5 'candidate source clauses' + recombination-risk gating
is load-bearing, not optional. Do NOT enable runtime NLI demotion on the
current scaffold; next step is the candidate-clause restriction, then
re-run, then gate item 4 is meaningful. Production verifier unchanged;
falsification-hard stays 10/12.
arborist/qa/nli/ — SHADOW ONLY (never an audit_mode input; manifest not
yet in governance_policy_hash per §7 #2). manifest.json pins
cross-encoder/nli-MiniLM2-L6-H768 @ a fixed HF revision + the
bench-validated θc 0.5/θe 0.9 + 2 alternates + the Phase-3 TODO;
shadow.py = ShadowNLI/shadow_check (lazy transformers+torch behind a new
[nli] extra, clauses() segmenter, the §7 #5 clause-level Demote()
decision, degrades to available=False when [nli] absent);
bench/scripts/nli_shadow_sweep.py + make bootstrap-nli / bench-nli-shadow
(the gate-item-4 instrument); 16 tests.
First sweep (116 records — 5f-falsification packs + the arborist-nli-bench
eval sets): 28/28 synth recombination demoted, 0/26 FP on legit summaries,
0/9 fires on already-STRICT_SPAN records, 25/50 on UNGROUNDED (the
contradiction half; quiet on non-sequiturs). Gate items 1/2/3/5/6 clear
on available data; item 4 — shadow FP rate on a real live-bench-qa
sample — remains the open measurement. Production verifier unchanged;
falsification-hard stays 10/12.