#000052 §3.1 round-2 patch: tighten coherence rules — FP rate on pooled bench-qa STRICT drops 5.4% → 1.1% (80% relative reduction)

Five rule tightenings, each targeting a specific bench-qa STRICT
false-positive shape (the xfail regressions from the previous commit):

1. claim-lattice bracket-artifact skip — sentences matching
   `\[E\d+\s*\|` (pointer markup) or `..."\]` (truncation tail)
   are no longer parsed as natural-language assertions; `Such a
   thesis was..."]` no longer fires vacuous.

2. Circular-rule differentia cap — circular now requires the
   predicate to be (a) entirely vacuous OR (b) leads with a subject
   token AND has ≤ 2 non-subject non-filler differentia tokens.
   "Michael Jordan's Restaurant was a restaurant in Chicago,
   Illinois, named after the basketball player Michael Jordan"
   has 6 differentia → no longer fires. Pre-existing positive test
   "The entity is the entity referring to the State of Israel"
   has 2 differentia (state, israel) → still fires (under threshold).

3. phrase_component_reuse translation-chain exception — when the
   predicate ALSO contains a quoted phrase (translation /
   definition / etymology context), token reuse with the subject's
   quoted phrase is legitimate, not circular. "The name 'Rosebud
   River' is a translation … 'the river of the roses'" no longer
   fires.

4. Vacuous-rule short-acronym escape — `_coherence_predicate_has_short_acronym_content`
   recognizes title-cased element-symbols (Au, Fe, Pb…) and all-caps
   2-5-char acronyms (DNA, FBI, USB, NASA…) as content even though
   they're below the ≥3-char content-token filter. "The chemical
   symbol for gold is Au." no longer fires; "Iron has the chemical
   symbol Fe." also clean; tautology "DNA stands for DNA." still
   correctly flagged circular.

5. (the 'term <X>' idiom xfail stays xfail — borderline, no clean
   lexical fix.)

5 xfail → passing (the regression coverage is now executable proof
of fix); 1 xfail remains. Pooled bench-qa STRICT FP rate test ceiling
tightened from 7% to 2%. All 7 pre-existing positive coherence tests
still fire correctly. 110 total tests pass; 1 xfailed; no regressions
on doc-counts / nli / relevance.
This commit is contained in:
russell@unturf.com 2026-05-13 13:09:33 -04:00
parent 5f4f4ceb1c
commit 4444478153
No known key found for this signature in database
2 changed files with 124 additions and 20 deletions

View file

@ -16,6 +16,7 @@ from __future__ import annotations
import json
import os
import re
import sqlite3
import unicodedata
from pathlib import Path
@ -965,11 +966,64 @@ def _coherence_bare_referents(predicate_text: str) -> set[str]:
return out
_COHERENCE_BRACKET_ARTIFACT_RE = re.compile(
# claim-lattice-mode pointer markup ([E1 | ...]) or a truncated
# trailing quote-bracket fragment (`..."]`, `..."`, etc.) — these
# are verifier-overlay artifacts in the answer, not assertions
r"""\[E\d+\s*\||\.\.\.["'’”]?\s*[\]>]?\s*$""",
re.VERBOSE,
)
_COHERENCE_SHORT_ACRONYM_RE = re.compile(r"\b([A-Z][A-Za-z0-9]{1,4})\b")
# 2-char tokens that aren't real content — copulas, prepositions, etc.
_COHERENCE_SHORT_STOPWORDS = frozenset({
"is", "of", "to", "in", "on", "at", "by", "an", "or", "as", "be",
"if", "it", "no", "so", "up", "us", "we", "do", "he", "my", "go",
"am", "pm", "ok",
})
# Max non-subject differentia tokens in the predicate that still
# qualifies as "circular". The original rule fires whenever the
# subject's tokens are ⊆ the predicate's AND the predicate leads with
# a subject token — too lax for encyclopedic "X's Y was a Y in PLACE,
# named after X" shapes (real content). Cap at 2: the existing
# positive test "The entity is the entity referring to the State of
# Israel" has 2 non-subject differentia tokens {state, israel} and
# stays circular; bench-qa real-traffic FPs ("Michael Jordan's
# Restaurant was a restaurant in Chicago, Illinois, named after the
# basketball player Michael Jordan." → 6 differentia tokens) no
# longer fire.
_COHERENCE_CIRCULAR_MAX_DIFFERENTIA = 2
def _coherence_predicate_has_short_acronym_content(pred_text: str) -> bool:
"""True iff the predicate text contains a short acronym/symbol-like
token that the >=3-char content filter misses but a reader would
treat as the actual answer e.g. 'Au' in 'The chemical symbol for
gold is Au.', 'DNA', 'RNA', 'FBI', 'USB'. Used to prevent the
vacuous rule from firing on concrete-fact short-predicate
answers."""
for m in _COHERENCE_SHORT_ACRONYM_RE.finditer(pred_text):
t = m.group(1)
low = t.lower()
if len(low) == 2 and low in _COHERENCE_SHORT_STOPWORDS:
continue
# accept 2-letter title-cased element-symbols (Au, Fe, Pb…) and
# any all-caps 2-5-char acronym (DNA, FBI, USB, NASA, …)
if t.isupper() or (len(t) == 2 and t[0].isupper() and t[1].islower()):
return True
return False
def _coherence_classify_sentence(sentence: str) -> dict[str, Any] | None:
"""Return a finding dict for the most severe shape problem in this
sentence, or None if it looks well-formed.
Severity: phrase_component_reuse > circular > vacuous."""
# Skip claim-lattice-mode pointer-markup artifacts ([E\d | …]) and
# truncated trailing quote-bracket fragments — they're verifier
# overlay, not natural-language assertions (#000052 §3.1 round-2
# bench-qa regression).
if _COHERENCE_BRACKET_ARTIFACT_RE.search(sentence):
return None
split = _coherence_split_on_copula(sentence)
if split is None:
return None
@ -980,8 +1034,13 @@ def _coherence_classify_sentence(sentence: str) -> dict[str, Any] | None:
# 1. phrase_component_reuse — subject quotes a phrase; predicate
# re-uses one of that phrase's own tokens as a bare referent.
# Suppressed when the predicate ALSO contains a quoted phrase —
# that signals a translation / definition / etymology chain
# where token reuse is legitimate ("The name 'Rosebud River' is
# a translation of … 'the river of the roses'" — §3.1 round-2
# bench-qa regression).
quoted_sets = _coherence_quoted_token_sets(subj_text)
if quoted_sets:
if quoted_sets and not _coherence_quoted_token_sets(pred_text):
bare = _coherence_bare_referents(pred_text)
for qs in quoted_sets:
reused = sorted(qs & bare)
@ -994,8 +1053,12 @@ def _coherence_classify_sentence(sentence: str) -> dict[str, Any] | None:
}
# 2. circular — subject content tokens are a non-empty subset of the
# predicate's, and the predicate leads with a subject token (or
# the predicate is itself entirely vacuous).
# predicate's, AND (a) the predicate is itself entirely vacuous,
# OR (b) the predicate leads with a subject token AND carries
# at most _COHERENCE_CIRCULAR_MAX_DIFFERENTIA non-subject content
# tokens (the original rule fired on (b) alone, which over-
# flagged encyclopedic "X's Y was a Y in PLACE, named after X"
# real definitions — §3.1 round-2 regression).
pred_non_vacuous = [
t for t in pred_in_order
if t not in _COHERENCE_FILLER and t not in _COHERENCE_VACUOUS_HEADS
@ -1003,7 +1066,16 @@ def _coherence_classify_sentence(sentence: str) -> dict[str, Any] | None:
pred_all_vacuous = len(pred_toks) > 0 and not pred_non_vacuous
if subj_toks and subj_toks <= pred_toks:
leads_with_subject = bool(pred_in_order) and pred_in_order[0] in subj_toks
if leads_with_subject or pred_all_vacuous:
differentia = [
t for t in pred_in_order
if t not in subj_toks
and t not in _COHERENCE_FILLER
and t not in _COHERENCE_VACUOUS_HEADS
]
if pred_all_vacuous or (
leads_with_subject
and len(set(differentia)) <= _COHERENCE_CIRCULAR_MAX_DIFFERENTIA
):
return {
"kind": "circular",
"sentence": sentence,
@ -1012,8 +1084,14 @@ def _coherence_classify_sentence(sentence: str) -> dict[str, Any] | None:
}
# 3. vacuous — predicate has no content beyond placeholder hypernyms
# + filler (and the subject did carry a real topic).
# + filler (and the subject did carry a real topic). Don't fire
# if the predicate contains a short acronym/symbol token like
# "Au"/"DNA"/"FBI" — those are content the >=3-char filter misses
# but a reader would treat as the actual answer (§3.1 round-2
# bench-qa regression).
if subj_toks and (not pred_toks or pred_all_vacuous):
if _coherence_predicate_has_short_acronym_content(pred_text):
return None
return {
"kind": "vacuous",
"sentence": sentence,

View file

@ -975,14 +975,29 @@ def test_coherence_more_vacuous_pure_hypernym_chain():
import pytest as _pytest
@_pytest.mark.xfail(reason="§3.1 false-positive: short-predicate concrete facts like 'The chemical symbol for gold is Au.' flagged as vacuous. The trailing 'Au' is a 2-char content-free token to the rule, but it IS the answer. Real bench-qa STRICT regression: this exact sentence appears 4× in the 808-cell pool, flagged each time. Tighten the vacuous rule to recognize a single-token-named entity as a sufficient predicate when it's not on the hypernym placeholder list.")
# The 5 tests below were xfail regressions documenting bench-qa STRICT
# false-positives; the rules have now been tightened in
# arborist/qa/inspect.py to fix them (§3.1 round-2 patch 2026-05-13).
def test_coherence_ok_on_short_concrete_fact_with_acronym_value():
"""Was xfail: 'The chemical symbol for gold is Au.' → vacuous.
Fix: vacuous rule now treats short acronym/symbol tokens (Au, Fe,
DNA, FBI, etc.) as content via
``_coherence_predicate_has_short_acronym_content``."""
d = diagnose_coherence("The chemical symbol for gold is Au.")
assert d["kind"] == "ok"
# other short-acronym variants:
assert diagnose_coherence("DNA stands for DNA.")["kind"] == "circular" # still tautology
assert diagnose_coherence("Iron has the chemical symbol Fe.")["kind"] == "ok"
@_pytest.mark.xfail(reason="§3.1 false-positive: 'X's Y was a Y ... named after X' shape — subject's proper-noun tokens reappear at the predicate's end via 'named after / called'. Subject's tokens ⊆ predicate's tokens, but the predicate has real differentia. Tighten circular to require the predicate to be NEARLY-EMPTY-modulo-determiners, not just superset on head tokens. Real bench-qa STRICT regression: 'Michael Jordan's Restaurant ... named after the basketball player Michael Jordan' (the trailing 'Michael Jordan' is what trips it; the shorter variant without that trailing reference does NOT trip).")
def test_coherence_ok_on_named_subject_with_repeated_head_noun():
"""Was xfail: 'X's Y was a Y ... named after X' shape.
Fix: circular now requires (predicate-all-vacuous) OR
(leads-with-subject AND non-subject differentia 2). The MJ
Restaurant predicate has 6 non-subject differentia
(restaurant, chicago, illinois, basketball, player, named) so
no longer fires."""
d = diagnose_coherence(
"Michael Jordan's Restaurant was a restaurant in Chicago, Illinois, "
"named after the basketball player Michael Jordan."
@ -990,8 +1005,9 @@ def test_coherence_ok_on_named_subject_with_repeated_head_noun():
assert d["kind"] == "ok"
@_pytest.mark.xfail(reason="§3.1 false-positive: 'the Western X was the western half of the X' encyclopedic compound-noun shape — tracked from real bench-qa STRICT regression. The 'half/part of X' construction is a legitimate definitional relation, not a circular tautology.")
def test_coherence_ok_on_compound_noun_definition():
"""Was xfail: 'the Western X was the western half of the X'.
Same circular-differentia-cap fix as MJ."""
d = diagnose_coherence(
"The Western Roman Empire was the western half of the Roman Empire, "
"from its division by Diocletian in 285 AD until its fall."
@ -999,8 +1015,11 @@ def test_coherence_ok_on_compound_noun_definition():
assert d["kind"] == "ok"
@_pytest.mark.xfail(reason="§3.1 false-positive: a quoted phrase containing a noun whose plural/derivative appears legitimately in the predicate (translation/etymology context) — 'Rosebud River' phrase, 'roses' in the predicate. Tracked from real bench-qa STRICT regression. Tighten phrase_component_reuse to require the bare-referent token to be USED AS THE SAME REFERENT as the phrase's head, not just lexically related.")
def test_coherence_ok_on_translation_etymology_with_derivative_token():
"""Was xfail: 'The name <Phrase> is a translation … <derivative>'.
Fix: phrase_component_reuse now suppresses when the predicate
contains another quoted phrase (signals a translation /
definition / etymology chain, not circularity)."""
d = diagnose_coherence(
"The name 'Rosebud River' is a translation of the Blackfoot word "
"Akokiniskway, meaning 'the river of the roses'."
@ -1008,11 +1027,11 @@ def test_coherence_ok_on_translation_etymology_with_derivative_token():
assert d["kind"] == "ok"
@_pytest.mark.xfail(reason="§3.1 false-positive: a bracketed claim-lattice-mode answer fragment (closing `..\"]`) tail-flagged as vacuous on a truncated/list-tail sentence. Tracked from real bench-qa STRICT regression. Tighten the sentence splitter to skip bracket-fragment tails.")
def test_coherence_ok_on_claim_lattice_truncated_bracket_tail():
# Real bench-qa STRICT shape — list-mode answer with a quoted
# excerpt that ends in `..."]`. The trailing fragment is parsed as
# a sentence and flagged as `vacuous` (predicate is empty).
"""Was xfail: claim-lattice [E\\d | …\"] artifacts parsed as
vacuous sentences. Fix: sentences matching the bracket-artifact
regex (``\\[E\\d+\\s*\\|`` opener or ``...\"]`` truncation
tail) are skipped entirely in classify."""
d = diagnose_coherence(
"Linux and BSD are both Unix-like operating systems. "
"[E1 | Linux | abcd1234: \"Linux is a Unix-like operating system kernel "
@ -1032,12 +1051,18 @@ def test_coherence_ok_on_term_idiom():
def test_coherence_pooled_bench_qa_strict_fp_rate_documented():
"""Document the load-bearing real-traffic FP rate so future work
can be measured against it. As of 2026-05-13, §3.1 false-fires on
44/808 = 5.4% of pooled bench-qa STRICT answers (the data set
pooled from n=1 + n=3 + n=5 ARBORIST_NLI_SHADOW=1 bench-qa runs).
can be measured against it. The pooled bench-qa STRICT sample
(n=1 + n=3 + n=5 ARBORIST_NLI_SHADOW=1 bench-qa runs = 808 cells)
has been the §3.1 regression target:
2026-05-13 (initial measurement, pre-patch): 44/808 = 5.4% FP
2026-05-13 (round-2 rule tightening): 9/808 = 1.1% FP
bracket-artifact skip + circular-differentia cap +
translation-phrase-pair exception + short-acronym vacuous
escape (5 of 6 §3.1 round-2 xfail regressions fixed; the
'term <X>' idiomatic-encyclopedic-English case stays xfail).
This test asserts the *upper bound* tightening the rules should
keep it at or below this. If a future change pushes it UP, this
test will fail and the regression is loud."""
keep it at or below the post-patch number. If a future change
pushes it UP past 2%, this test fails loud."""
import json as _json
from pathlib import Path as _Path
# this is a slow-ish test (~1s for 808 rows of lexical regex);
@ -1058,6 +1083,7 @@ def test_coherence_pooled_bench_qa_strict_fp_rate_documented():
flagged = sum(1 for r in rows
if diagnose_coherence(r["answer_text"])["kind"] not in ("ok", "empty"))
assert len(rows) >= 800, f"pooled STRICT sample shrunk unexpectedly ({len(rows)} rows; expected ~808)"
# current rate: 44/808 = 0.0545; ceiling at 0.07 leaves a tiny bit of headroom for fixture churn
assert flagged / len(rows) <= 0.07, \
# post round-2 rate: 9/808 = 0.011; ceiling at 0.02 leaves a tiny
# bit of headroom for fixture churn but fires loud on regressions.
assert flagged / len(rows) <= 0.02, \
f"§3.1 FP rate on real STRICT regressed: {flagged}/{len(rows)} = {flagged/len(rows):.3f}"