#000052 §3.2.1: candidate-bench landed (4 models × 26 pairs on the 4090 — proper bench, not the 6-pair smoke)
All 4 manifest candidates clean-separate (12/12 NEG catch at 0/14 POS FP), so 'reranker discriminates Zionist-entity-style mis-cites' is a property of MS-MARCO-trained rerankers as a class — not the specific L-6 I picked first. Ranked by separation margin (per §7 #18: separation beats raw): ms-marco-electra-base +5.349 ← new primary BAAI/bge-reranker-base +3.076 ms-marco-MiniLM-L-6-v2 +3.053 (previous primary, demoted) ms-marco-MiniLM-L-12-v2 +2.663 (worst — deeper ≠ better) Manifest primary moved to electra-base for the cushion. demote_below_score STAYS null — clean candidate-bench thresholds don't predict real-pipeline behavior (the §7 #18→#27 history is 6 verdict flips on the NLI side); §3.2.2 step 2 (real-traffic shadow sweep on pooled bench-qa STRICT) is what sets it. Expect a walk-back. Reranker still doesn't catch the Kilimanjaro/Mount-Kenya recombination (aboutness ≠ truth-of-attribution; that's #000049 territory). fox's bench-maxing correction applied.
This commit is contained in:
parent
ba1b6b904b
commit
aeeef603a1
3 changed files with 278 additions and 27 deletions
|
|
@ -1,10 +1,10 @@
|
|||
{
|
||||
"_comment": "Pinned aboutness/relevance reranker for the #000052 §3.2 SHADOW path. SHADOW MODE ONLY: this manifest does NOT yet fold into governance_policy_hash because shadow output cannot touch audit_mode (#000049 §7 #2 discipline inherited verbatim). If/when promoted to a runtime demotion-only veto, add a single-blob checkpoint_sha256 + tokenizer_sha256 + an explicit relevance_policy_hash that folds into governance_policy_hash before any audit_mode effect. Cross-encoder MS-MARCO reranker — single relevance score per (query, document) pair, NOT entailment (that's NLI / arborist.qa.nli; the boundary failure mode this catches is *topic-collision recombination* — answer/claim is on a different topic than its cited source despite shared tokens, the 'Zionist entity' case — which NLI returns neutral on and the lexical title-relevance sidecar passes because the cited title shares a polysemous token). Reuses the [nli] extra (transformers+torch+optimum); the checkpoint is loaded as AutoModelForSequenceClassification (single regression head emitting a relevance logit), no sentence-transformers dependency.",
|
||||
"relevance_model_version": "relevance-shadow-v1-ms-marco-MiniLM-L-6-v2",
|
||||
"hf_repo": "cross-encoder/ms-marco-MiniLM-L-6-v2",
|
||||
"_comment": "Pinned aboutness/relevance reranker for the #000052 §3.2 SHADOW path. SHADOW MODE ONLY: this manifest does NOT yet fold into governance_policy_hash because shadow output cannot touch audit_mode (#000049 §7 #2 discipline inherited verbatim). If/when promoted to a runtime demotion-only veto, add a single-blob checkpoint_sha256 + tokenizer_sha256 + an explicit relevance_policy_hash that folds into governance_policy_hash before any audit_mode effect. Cross-encoder MS-MARCO reranker — single relevance score per (query, document) pair, NOT entailment (that's NLI / arborist.qa.nli; the boundary failure mode this catches is *topic-collision recombination* — answer/claim is on a different topic than its cited source despite shared tokens, the 'Zionist entity' case — which NLI returns neutral on and the lexical title-relevance sidecar passes because the cited title shares a polysemous token). Reuses the [nli] extra (transformers+torch+optimum); the checkpoint is loaded as AutoModelForSequenceClassification (single regression head emitting a relevance logit), no sentence-transformers dependency. PRIMARY = ms-marco-electra-base picked from the candidate-bench (#000052 §3.2.1 — sep margin +5.35 vs +3.05/+3.08/+2.66 for the others). NOTE: the candidate-bench is contrived; the §7 #18→#27 history says clean-eval doesn't predict real-traffic precision OR recall — this primary is provisional pending §3.2.2 step 2 (real-traffic shadow sweep on pooled bench-qa STRICT) + step 3 (recall-side realism on real-context fixtures). Expect a walk-back.",
|
||||
"relevance_model_version": "relevance-shadow-v1-ms-marco-electra-base",
|
||||
"hf_repo": "cross-encoder/ms-marco-electra-base",
|
||||
"license": "Apache-2.0",
|
||||
"source_url": "https://huggingface.co/cross-encoder/ms-marco-MiniLM-L-6-v2",
|
||||
"approx_mb": 90,
|
||||
"source_url": "https://huggingface.co/cross-encoder/ms-marco-electra-base",
|
||||
"approx_mb": 110,
|
||||
"score_shape": "single logit per (query, document) pair, NOT a softmax — higher = more relevant; trained on MS MARCO passage-ranking, so the score is well-calibrated for 'is this doc relevant to this query?' and uncalibrated for any other notion of similarity",
|
||||
"max_length": 256,
|
||||
"truncation_policy": "query_then_document@v1",
|
||||
|
|
@ -12,9 +12,20 @@
|
|||
"demote_below_score": null,
|
||||
"demote_below_score_comment": "the demote threshold MUST be set by a shadow sweep against bench-qa traffic — DO NOT default to a literature number. The #000049 §7 #18→#27 arc proved repeatedly that clean-eval thresholds don't transfer to real-pipeline data; recommended_operating_point is left null until a candidate-bench + a bench-qa shadow sweep produces a θ with the FP discipline §3.2 inherits from #000049 (zero false-demote on the pooled-808 STRICT denominator at the chosen θ).",
|
||||
"alternates": [
|
||||
{"relevance_model_version": "relevance-shadow-v1-ms-marco-MiniLM-L-12-v2", "hf_repo": "cross-encoder/ms-marco-MiniLM-L-12-v2", "license": "Apache-2.0", "approx_mb": 130, "note": "deeper L-12 variant of the primary; bench against the L-6 on the off-topic eval set"},
|
||||
{"relevance_model_version": "relevance-shadow-v1-bge-reranker-base", "hf_repo": "BAAI/bge-reranker-base", "license": "MIT", "approx_mb": 280, "note": "BAAI's general-purpose reranker; different training corpus, useful as a non-MS-MARCO comparison point"},
|
||||
{"relevance_model_version": "relevance-shadow-v1-ms-marco-electra-base", "hf_repo": "cross-encoder/ms-marco-electra-base", "license": "Apache-2.0", "approx_mb": 110, "note": "ELECTRA architecture for variety; same MS MARCO supervision"}
|
||||
{"relevance_model_version": "relevance-shadow-v1-bge-reranker-base", "hf_repo": "BAAI/bge-reranker-base", "license": "MIT", "approx_mb": 280, "candidate_bench_margin": 3.076, "note": "BAAI general-purpose reranker; second-best candidate-bench margin after electra-base"},
|
||||
{"relevance_model_version": "relevance-shadow-v1-ms-marco-MiniLM-L-6-v2", "hf_repo": "cross-encoder/ms-marco-MiniLM-L-6-v2", "license": "Apache-2.0", "approx_mb": 80, "candidate_bench_margin": 3.053, "note": "the original primary before §3.2.1 candidate-bench; smallest/cheapest, comparable margin to BGE on the 26-pair candidate-bench"},
|
||||
{"relevance_model_version": "relevance-shadow-v1-ms-marco-MiniLM-L-12-v2", "hf_repo": "cross-encoder/ms-marco-MiniLM-L-12-v2", "license": "Apache-2.0", "approx_mb": 130, "candidate_bench_margin": 2.663, "note": "deeper L-12 MiniLM; worst margin of the 4 on the candidate-bench — deeper ≠ better here, same pattern as #000049 §7 #24's 'bart-large worse than deberta-large'"}
|
||||
],
|
||||
"candidate_bench_results_2026_05_13": {
|
||||
"fixtures": "bench/fixtures/5f/relevance-aboutness-v1.jsonl (14 POS + 12 NEG)",
|
||||
"grid_output": "bench/results/relevance-aboutness-candidatebench.json",
|
||||
"rank_by_sep_margin": [
|
||||
{"model": "ms-marco-electra-base", "margin": 5.349, "min_pos": 0.050, "max_neg": -5.300, "fp0_theta": 0.050, "catch_neg": "12/12"},
|
||||
{"model": "bge-reranker-base", "margin": 3.076, "min_pos": 2.395, "max_neg": -0.681, "fp0_theta": 2.395, "catch_neg": "12/12"},
|
||||
{"model": "ms-marco-MiniLM-L-6-v2", "margin": 3.053, "min_pos": 6.754, "max_neg": 3.701, "fp0_theta": 6.754, "catch_neg": "12/12"},
|
||||
{"model": "ms-marco-MiniLM-L-12-v2","margin": 2.663, "min_pos": 7.138, "max_neg": 4.475, "fp0_theta": 7.138, "catch_neg": "12/12"}
|
||||
],
|
||||
"note": "All 4 candidates clean-separate on this candidate-bench. Picked electra-base as primary by margin (per §7 #18: rank by separation, not raw score). The thresholds above are CANDIDATE-BENCH ONLY — they will move when the denominator grows (§7 #25→#26 happened three times on the NLI side; assume the same here). demote_below_score stays null until §3.2.2 step 2 (real-traffic shadow sweep) sets it."
|
||||
},
|
||||
"rationale_for_cross_encoder_not_bi_encoder": "A bi-encoder (sentence-transformers, vec-style) gives an unsupervised cosine similarity that's good for retrieval *ranking* but not calibrated for the 'is this on-topic?' question we need here. A cross-encoder reranker is trained on (query, doc) pairs with relevance labels — its single logit IS the 'is this doc relevant to this query?' score we want. The architectural cross-link to #000050/#000051's vec hybrid: that work uses *bi-encoders* (fastembed, bge-small) for retrieval recall; this work uses a *cross-encoder* for downstream aboutness scoring. They are complementary — vec finds candidates, the cross-encoder scores them."
|
||||
}
|
||||
|
|
|
|||
216
bench/results/relevance-aboutness-candidatebench.json
Normal file
216
bench/results/relevance-aboutness-candidatebench.json
Normal file
|
|
@ -0,0 +1,216 @@
|
|||
{
|
||||
"generated_at": "2026-05-13T13:53:25Z",
|
||||
"fixtures_file": "/home/fox/nli-data/relevance-aboutness-v1.jsonl",
|
||||
"n_pos": 14,
|
||||
"n_neg": 12,
|
||||
"models": {
|
||||
"relevance-shadow-v1-ms-marco-MiniLM-L-6-v2": {
|
||||
"hf_repo": "cross-encoder/ms-marco-MiniLM-L-6-v2",
|
||||
"backend": "torch",
|
||||
"device": "cuda",
|
||||
"n_pairs": 26,
|
||||
"infer_seconds": 0.18,
|
||||
"pos_scores": {
|
||||
"pos-qa-001": 10.3561,
|
||||
"pos-qa-002": 7.6135,
|
||||
"pos-qa-003": 10.2615,
|
||||
"pos-qa-004": 9.6113,
|
||||
"pos-qa-005": 10.761,
|
||||
"pos-qa-006": 10.6969,
|
||||
"pos-qa-007": 10.189,
|
||||
"pos-cs-001": 8.6417,
|
||||
"pos-cs-002": 6.754,
|
||||
"pos-cs-003": 10.2053,
|
||||
"pos-cs-004": 9.1802,
|
||||
"pos-cs-005": 10.58,
|
||||
"pos-cs-006": 9.8569,
|
||||
"pos-cs-007": 9.986
|
||||
},
|
||||
"neg_scores": {
|
||||
"neg-qa-001": 3.7006,
|
||||
"neg-qa-002": -0.0128,
|
||||
"neg-qa-003": 3.1369,
|
||||
"neg-qa-004": 0.2058,
|
||||
"neg-qa-005": -5.2781,
|
||||
"neg-qa-006": -9.0515,
|
||||
"neg-cs-001": -8.1389,
|
||||
"neg-cs-002": -2.9362,
|
||||
"neg-cs-003": -5.963,
|
||||
"neg-cs-004": -6.7567,
|
||||
"neg-cs-005": -5.9802,
|
||||
"neg-cs-006": -9.8485
|
||||
},
|
||||
"frontier": {
|
||||
"separation_margin": 3.0533,
|
||||
"fp0_best": {
|
||||
"theta": 6.754,
|
||||
"catch_neg": 1.0,
|
||||
"fp_on_pos": 0.0
|
||||
},
|
||||
"pareto": [
|
||||
{
|
||||
"theta": 6.754,
|
||||
"catch_neg": 1.0,
|
||||
"fp_on_pos": 0.0
|
||||
}
|
||||
]
|
||||
}
|
||||
},
|
||||
"relevance-shadow-v1-ms-marco-MiniLM-L-12-v2": {
|
||||
"hf_repo": "cross-encoder/ms-marco-MiniLM-L-12-v2",
|
||||
"backend": "torch",
|
||||
"device": "cuda",
|
||||
"n_pairs": 26,
|
||||
"infer_seconds": 0.05,
|
||||
"pos_scores": {
|
||||
"pos-qa-001": 10.1692,
|
||||
"pos-qa-002": 8.139,
|
||||
"pos-qa-003": 10.1767,
|
||||
"pos-qa-004": 9.3511,
|
||||
"pos-qa-005": 10.5464,
|
||||
"pos-qa-006": 10.0394,
|
||||
"pos-qa-007": 9.4536,
|
||||
"pos-cs-001": 8.6255,
|
||||
"pos-cs-002": 7.1379,
|
||||
"pos-cs-003": 9.4853,
|
||||
"pos-cs-004": 9.6927,
|
||||
"pos-cs-005": 10.1209,
|
||||
"pos-cs-006": 9.9942,
|
||||
"pos-cs-007": 10.2065
|
||||
},
|
||||
"neg_scores": {
|
||||
"neg-qa-001": 4.4752,
|
||||
"neg-qa-002": -0.2735,
|
||||
"neg-qa-003": 3.1724,
|
||||
"neg-qa-004": -0.3941,
|
||||
"neg-qa-005": -5.2772,
|
||||
"neg-qa-006": -10.3666,
|
||||
"neg-cs-001": -6.9534,
|
||||
"neg-cs-002": -2.7428,
|
||||
"neg-cs-003": -6.2688,
|
||||
"neg-cs-004": -6.3582,
|
||||
"neg-cs-005": -7.7344,
|
||||
"neg-cs-006": -10.1564
|
||||
},
|
||||
"frontier": {
|
||||
"separation_margin": 2.6627,
|
||||
"fp0_best": {
|
||||
"theta": 7.1379,
|
||||
"catch_neg": 1.0,
|
||||
"fp_on_pos": 0.0
|
||||
},
|
||||
"pareto": [
|
||||
{
|
||||
"theta": 7.1379,
|
||||
"catch_neg": 1.0,
|
||||
"fp_on_pos": 0.0
|
||||
}
|
||||
]
|
||||
}
|
||||
},
|
||||
"relevance-shadow-v1-bge-reranker-base": {
|
||||
"hf_repo": "BAAI/bge-reranker-base",
|
||||
"backend": "torch",
|
||||
"device": "cuda",
|
||||
"n_pairs": 26,
|
||||
"infer_seconds": 0.08,
|
||||
"pos_scores": {
|
||||
"pos-qa-001": 8.0599,
|
||||
"pos-qa-002": 7.8094,
|
||||
"pos-qa-003": 6.7217,
|
||||
"pos-qa-004": 7.1199,
|
||||
"pos-qa-005": 8.3108,
|
||||
"pos-qa-006": 7.723,
|
||||
"pos-qa-007": 6.2936,
|
||||
"pos-cs-001": 2.395,
|
||||
"pos-cs-002": 2.7944,
|
||||
"pos-cs-003": 4.9885,
|
||||
"pos-cs-004": 4.7798,
|
||||
"pos-cs-005": 6.3818,
|
||||
"pos-cs-006": 8.417,
|
||||
"pos-cs-007": 5.2747
|
||||
},
|
||||
"neg_scores": {
|
||||
"neg-qa-001": -0.7671,
|
||||
"neg-qa-002": -3.387,
|
||||
"neg-qa-003": -1.9937,
|
||||
"neg-qa-004": -0.6807,
|
||||
"neg-qa-005": -6.0444,
|
||||
"neg-qa-006": -7.8323,
|
||||
"neg-cs-001": -6.4084,
|
||||
"neg-cs-002": -3.9995,
|
||||
"neg-cs-003": -5.247,
|
||||
"neg-cs-004": -4.4321,
|
||||
"neg-cs-005": -7.8685,
|
||||
"neg-cs-006": -6.9767
|
||||
},
|
||||
"frontier": {
|
||||
"separation_margin": 3.0757,
|
||||
"fp0_best": {
|
||||
"theta": 2.395,
|
||||
"catch_neg": 1.0,
|
||||
"fp_on_pos": 0.0
|
||||
},
|
||||
"pareto": [
|
||||
{
|
||||
"theta": 2.395,
|
||||
"catch_neg": 1.0,
|
||||
"fp_on_pos": 0.0
|
||||
}
|
||||
]
|
||||
}
|
||||
},
|
||||
"relevance-shadow-v1-ms-marco-electra-base": {
|
||||
"hf_repo": "cross-encoder/ms-marco-electra-base",
|
||||
"backend": "torch",
|
||||
"device": "cuda",
|
||||
"n_pairs": 26,
|
||||
"infer_seconds": 0.1,
|
||||
"pos_scores": {
|
||||
"pos-qa-001": 3.448,
|
||||
"pos-qa-002": 2.3775,
|
||||
"pos-qa-003": 3.8591,
|
||||
"pos-qa-004": 1.7753,
|
||||
"pos-qa-005": 3.4742,
|
||||
"pos-qa-006": 1.5494,
|
||||
"pos-qa-007": 3.3069,
|
||||
"pos-cs-001": 0.0497,
|
||||
"pos-cs-002": 2.3681,
|
||||
"pos-cs-003": 4.3109,
|
||||
"pos-cs-004": 2.4378,
|
||||
"pos-cs-005": 2.0715,
|
||||
"pos-cs-006": 5.0473,
|
||||
"pos-cs-007": 4.9582
|
||||
},
|
||||
"neg_scores": {
|
||||
"neg-qa-001": -5.7487,
|
||||
"neg-qa-002": -8.6513,
|
||||
"neg-qa-003": -6.7965,
|
||||
"neg-qa-004": -10.2494,
|
||||
"neg-qa-005": -10.6582,
|
||||
"neg-qa-006": -10.3879,
|
||||
"neg-cs-001": -10.4413,
|
||||
"neg-cs-002": -5.9059,
|
||||
"neg-cs-003": -5.2999,
|
||||
"neg-cs-004": -10.5911,
|
||||
"neg-cs-005": -9.6781,
|
||||
"neg-cs-006": -10.8435
|
||||
},
|
||||
"frontier": {
|
||||
"separation_margin": 5.3495,
|
||||
"fp0_best": {
|
||||
"theta": 0.0497,
|
||||
"catch_neg": 1.0,
|
||||
"fp_on_pos": 0.0
|
||||
},
|
||||
"pareto": [
|
||||
{
|
||||
"theta": 0.0497,
|
||||
"catch_neg": 1.0,
|
||||
"fp_on_pos": 0.0
|
||||
}
|
||||
]
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
|
|
@ -254,27 +254,51 @@ structure verbatim. Pieces:
|
|||
- `tests/test_relevance_shadow.py` — <!--AUTOCOUNT:tests:tests/test_relevance_shadow.py-->13<!--/AUTOCOUNT--> tests (pure-Python + graceful
|
||||
degradation + the Zionist sanity check); all green.
|
||||
|
||||
**Sanity check on real (claim, source) pairs** (run on the loaded
|
||||
`ms-marco-MiniLM-L-6-v2`, CPU, batch=1; scores are logits, sign matters):
|
||||
**Candidate-bench (2026-05-13)** — 26-pair hand-built fixture set
|
||||
(`bench/fixtures/5f/relevance-aboutness-v1.jsonl`: 14 on-topic POS +
|
||||
12 off-topic / topic-collision NEG, incl. the Zionist field case),
|
||||
all 4 manifest candidates swept on the 4090 (104 GPU NLI pairs total,
|
||||
~1 s per model; `bench/results/relevance-aboutness-candidatebench.json`).
|
||||
Ranked by separation margin (= `min(POS_score) − max(NEG_score)`;
|
||||
>0 = clean-separable):
|
||||
|
||||
| pair | claim | source | score |
|
||||
|---|---|---|---|
|
||||
| Zionist ON-topic | "...used as a pejorative for the State of Israel" | the real `Zionist entity` sentence | **+9.96** |
|
||||
| Zionist OFF-topic (the field case) | "...used as the entity, referring to the State of Israel" | a grammar `Phrase` definition | **−9.04** |
|
||||
| Mona Lisa Q→A on | "who painted the mona lisa?" | "Leonardo da Vinci painted the Mona Lisa." | +10.45 |
|
||||
| Mona Lisa Q→A deflect | "who painted the mona lisa?" | "The Mona Lisa is a half-length portrait housed in the Louvre." | +3.56 |
|
||||
| **Kilimanjaro recombination** | "Kilimanjaro is in central Kenya south of the equator" | the Mount Kenya clause | **+7.59** (false positive — see below) |
|
||||
| Kilimanjaro on | "Kilimanjaro is in north-eastern Tanzania" | the Kilimanjaro lead | +6.06 |
|
||||
| model (size) | min(POS) | max(NEG) | **sep margin** | fp=0 θ | catch on NEG |
|
||||
|---|---|---|---|---|---|
|
||||
| **`cross-encoder/ms-marco-electra-base`** (~110 MB) | +0.050 | −5.300 | **+5.349** | 0.050 | **12/12 = 1.000** |
|
||||
| `BAAI/bge-reranker-base` (~280 MB) | +2.395 | −0.681 | +3.076 | 2.395 | 12/12 = 1.000 |
|
||||
| `cross-encoder/ms-marco-MiniLM-L-6-v2` (~80 MB) | +6.754 | +3.701 | +3.053 | 6.754 | 12/12 = 1.000 |
|
||||
| `cross-encoder/ms-marco-MiniLM-L-12-v2` (~130 MB) | +7.138 | +4.475 | +2.663 | 7.138 | 12/12 = 1.000 |
|
||||
|
||||
**Reading.** The reranker **cleanly discriminates the Zionist-entity
|
||||
field case (18-point margin)** and the Q→A deflection case (~7-point
|
||||
margin) — i.e. it catches exactly the failure mode #000052 §1 named
|
||||
("the lexical title-relevance sidecar passes a token-collision
|
||||
mis-cite"). It does **NOT** catch the *recombination* case (Mount
|
||||
Kenya clause scores *higher* than the Kilimanjaro lead for a
|
||||
"Kilimanjaro is in Kenya" claim) — because the Mount Kenya clause is
|
||||
genuinely about Kilimanjaro-as-comparison and the reranker is scoring
|
||||
*aboutness*, not *truth*. That is the right architectural split:
|
||||
**Read carefully:**
|
||||
- **All 4 candidates clean-separate** the candidate-bench — every model
|
||||
has a θ that catches all 12 off-topic NEG (incl. the Zionist field
|
||||
case) with 0 FP on the 14 POS. So "the reranker discriminates the
|
||||
Zionist-style failure mode" is *not a property of any specific
|
||||
checkpoint*; it's a property of MS-MARCO-trained rerankers (and
|
||||
BGE) as a class on contrived data.
|
||||
- **The deepest MiniLM (L-12) is the *worst* by margin** (+2.66 vs
|
||||
+5.35 for electra-base). Same pattern as #000049 §7 #24's "bart-large
|
||||
is worse than deberta-large" — *deeper ≠ better*; specific
|
||||
fine-tune and architecture dominate. The §3.2 manifest primary
|
||||
moves to `ms-marco-electra-base` for the most cushion (~75% more
|
||||
margin than the next-best, ~100% more than the smallest L-6).
|
||||
- **None of this predicts real-traffic behavior.** The
|
||||
#000049 §7 #18→#27 history is the receipts: clean candidate-bench
|
||||
→ real-bench-qa precision flipped six times on the NLI side; assume
|
||||
the same here. The thresholds above (electra-base at θ=0.05, L-6 at
|
||||
θ=6.75) are *candidate-bench numbers* and will move when the
|
||||
denominator grows. The manifest's `demote_below_score` stays `null`
|
||||
until the real-traffic shadow sweep (§3.2.2 step 2) sets it.
|
||||
|
||||
**What the candidate-bench does NOT establish** (preserved from the
|
||||
earlier 6-pair sanity check — the architectural split remains the
|
||||
same point, just with proper data behind the catch-rate claim):
|
||||
the reranker does **NOT** catch the *recombination*-where-the-
|
||||
different-entity-clause-also-mentions-the-target-subject case (the
|
||||
Mount Kenya clause scores higher than the Kilimanjaro lead for a
|
||||
"Kilimanjaro is in Kenya" claim) — because the Mount Kenya clause
|
||||
is genuinely *about* Kilimanjaro-as-comparison and the reranker is
|
||||
scoring *aboutness*, not *truth*. The right architectural split:
|
||||
|
||||
- **#000052 §3.2 relevance reranker** = catches *topic-collision /
|
||||
off-topic* failures (Zionist entity, Q→A deflection). ✓ shown.
|
||||
|
|
|
|||
Loading…
Add table
Add a link
Reference in a new issue