arborist/docs/qa-modes-bench.md
russell@unturf.com 42f614a501
docs(#000057): Addendum 8 — control sweep retrieval × model × framing × reasoning
Durable record of the #000057 sweep in the bench journal. Captures:

- The question: is Hermes-8B's confident present-day-officeholder
  fabrication an 8B weakness, a framing artefact, or does retrieval
  fix it? Crosses {hermes, qwen-nothink, qwen-think} × {plain,
  source_relative, as_of_corpus} × {solo, arborist} on a 386-item
  office-holder fixture with corpus-vintage gold.

- The judge methodology: Opus headless judge burned quota (79.5%
  JUDGE_ERROR), replaced with the deterministic code judge
  (bench/judge_code.py), calibrated against Opus's gradeable records
  (CG agreement 13->47%, WRONG 56->89%, ABSTAINED 80->95%).

- Consolidated CG% scorecard, all arms on the identical final judge.

- Three findings:
  1. Retrieval dominates — arb/qwen-nothink/plain 82% vs 7% solo;
     no solo config approaches the retrieval arms.
  2. Reasoning does NOT improve raw correctness — qwen-think/as_of
     44% vs nothink 50%.
  3. Reasoning's real cost is broken honest-abstention —
     qwen-nothink/source_relative abstains 97% (clean); qwen-think
     only 61%, reasoning itself into wrong parametric answers.

- Production recommendation: arborist + qwen-nothink, plain framing,
  reasoning OFF (82% CG, ~0% abstain, 11% wrong-assert).

- Held cell noted: arborist+qwen-think running at write time, result
  to be appended.

Bench %s are point-in-time measurements (not repo-derived counts),
so no AUTOCOUNT tags — consistent with addenda 1-7. test_doc_counts
3/3.
2026-05-20 06:52:42 -04:00

35 KiB
Raw Blame History

QA-modes bench — 2026-05-02

Date: 2026-05-02 Endpoint: https://hermes.ai.unturf.com/v1 (Hermes-3-Llama-3.1-8B-FP8-Dynamic, vLLM, 82K ctx) Corpus: Wikipedia 2003-05-16 cur snapshot, sharded under ~/.arborist/shards

Two sweeps landed today, each 71 questions × 3 modes:

stamp samples concurrency scheduling runs wall-clock
11:31Z n=2 c=4 cell-grouped 426 51 min
15:07Z n=3 c=4 sample-shuffled, --seed 0 639 48 min

The 15:07Z sweep is the authoritative state of the substrate at end-of-day; the 11:31Z sweep is the intermediate witnessed before the second wave of hardening landed. Headlines below show both.

Hardening since the prior bench (2026-04-30 post-retry)

Substrate-side (between 2026-04-30T18:46Z and 2026-05-02T11:31Z):

  • Rule 8 — title-relevance promoted to hard verifier check (TITLE_MISMATCH violation demotes STRICT → HYBRID when no cited evidence's source title shares a stemmed token with the claim).
  • Anchor-class warrant generalized to entity-list / count / why-cause shapes (Ticket #000003).
  • Retrieval-plan hash bound into the run-DAG retrieval stage (Ticket #000001).
  • Reference-frame polarity contract (Ticket #000002).
  • Four-rung ladder display (Ticket #000005) — POINTER-LINKED → ANCHOR-WARRANTED → EVIDENCE-WARRANTED → (ENTAILMENT-VERIFIED reserved) mapped from the v9.8 audit_mode trichotomy at render time.
  • Bench harness extended with directive coverage + log-scale buckets to 1M (Ticket #000004).

Substrate-side (between 2026-05-02T11:31Z and 2026-05-02T15:07Z):

  • Sprint 1b — per-mode max_context_chars. Bench's recommended-context-budget table flows back into DEFAULT_QUERY_POLICY["max_context_chars_by_mode"]: quote 24 KB, pointer 24 KB, JSON 48 KB. Folds into governance_policy_hash.
  • Sprint 2 — pointer Rule 9: chunk-specificity instruction added to the lattice-pointer system prompt.
  • DRY collapse — the four lattice prompts (system + grounding × pointer + JSON) lifted to arborist/qa/prompts.py as a single source of truth, imported by both runner.DEFAULT_POLICY and query.DEFAULT_QUERY_POLICY.
  • httpx persistent client — the chat-completion path used to construct a fresh httpx.Client per call, paying a TLS handshake every request. Move to __init__; HTTP/1.1 keep-alive across calls. Save 1-2 min on a 426-call bench.
  • Sample-level shuffled bench scheduling — every (question, mode, sample_idx) is a task, shuffled with --seed, dispatched concurrently. Per-cell Lock dict serializes burn-then-write on the shared cache_key. True i.i.d. n=3 variance; vLLM batcher fed a diverse request stream.
  • Bench --resume — read existing JSONL, skip done tasks, append fresh rows. Stop/start-able.
  • Concurrency sweep on the smoke fixture. vLLM peaks at c=3-4; saturates badly past c=4 (see Wall-clock & throughput below).
  • UTF-16 surrogate fix v2 (41d1d9b) — corpus chunks with lone surrogates broke httpx's outbound JSON encode. Scrub at the client boundary by WTF-8 → UTF-8-with-replace roundtrip. (The earlier v1 fix 3b91223 hardened the OUTPUT side — SHA-256 hashers — but missed the INPUT side. Both paths now safe.)

Aggregate (15:07Z, authoritative)

mode runs STRICT HYBRID UNGROUNDED err strict-rate mean ratio mean latency
quote 213 116 51 46 0 0.54 0.699 16.9 s
claim_lattice_pointer 213 43 146 21 3 0.20 0.642 18.5 s
claim_lattice (JSON) 213 89 82 39 3 0.42 0.698 18.3 s

Δ across the day

mode 2026-04-30 (post-retry) 2026-05-02T11:31Z 2026-05-02T15:07Z net Δ
quote 0.47 0.50 (+3pp) 0.54 (+4pp) +7pp
claim_lattice_pointer 0.24 0.23 (1pp) 0.20 (3pp) 4pp
claim_lattice (JSON) 0.50 0.44 (6pp) 0.42 (2pp) 8pp

Quote climbed the most across the day, ending at 0.54. The Sprint 1b 24 KB cap surfaces tighter retrievals that quote can ground verbatim (the bucket data confirms — quote's peak migrated to 8-16 KB at 0.58 strict-rate).

Pointer & JSON each took an honesty cost. Rule 8 promotion + warrant-class generalization demote more cases that would have classified STRICT under the looser pre-2026-05-02 verifier. The architectural win — 99% directive coverage on lattice modes — is the price for those drops; false-positive STRICT was corruption, and we converted it to honest HYBRID.

Per-bucket strict-rate (15:07Z)

mode bucket runs strict-rate note
quote 8-16 KB 105 0.58 peak
quote 16-32 KB 105 0.52
claim_lattice_pointer 16-32 KB 210 0.20 peak (only bucket)
claim_lattice (JSON) 32-64 KB 99 0.48 peak
claim_lattice (JSON) 16-32 KB 108 0.38

Sprint 1b's intent confirmed. JSON's peak at 32-64 KB justifies the 48 KB default. Quote's peak migrated to 8-16 KB after the 24 KB cap; the cap may be tighter than optimal — quote could plausibly be cut to 16 KB (mid of 8-16 KB bucket) for another small lift in a follow-up sprint.

claim_lattice (JSON) peaks at LARGER context than the other modes. The structured per-claim evidence linkage that JSON enforces benefits from more evidence per claim. Pointer & quote degrade past 16-32 KB.

5-run minimum to reduce noise; from the 15:07Z bucket data:

mode peak bucket strict-rate n
quote 8-16 KB 0.58 105
claim_lattice_pointer 16-32 KB 0.20 210
claim_lattice (JSON) 32-64 KB 0.48 99

These flow into DEFAULT_QUERY_POLICY["max_context_chars_by_mode"] (currently 24 / 24 / 48 KB — mid of each peak bucket). When the bench shifts those peaks, retune the policy and let governance_policy_hash partition the cache.

Directive coverage (seven-point program)

mode D2 pointer D3 cti-ready D4 ev-map bound D6 warrant D7 honest label
quote 0/213 (0%) 0/213 (0%) 213/213 (100%) 0/213 (0%) 213/213 (100%)
claim_lattice_pointer 210/213 (99%) 210/213 (99%) 210/213 (99%) 210/213 (99%) 210/213 (99%)
claim_lattice 210/213 (99%) 210/213 (99%) 210/213 (99%) 210/213 (99%) 210/213 (99%)

Lattice modes hit 99% on every observable directive; the 1% gap is the 6 surrogate errors (already fixed in 41d1d9b, next bench will hit 100%). Quote shows 0% on D2 / D3 / D6 by construction (those are lattice-only directives) and 100% on D4 / D7.

Pointer-mode signal — where the gap lives

Pointer mode's strict-rate did not lift on Sprint 2's chunk-specificity prompt nudge. The load-bearing failure pattern is lazy-anchoring: the model cites a topic-overview chunk for every claim instead of the specific chunk that supports each individual claim. Lazy-anchor ratio histogram (11:31Z bench, 142 pointer rows):

lazy_anchor_ratio bucket rows
0.00 25
<0.25 9
<0.5 11
<0.75 46
>=0.75 49 (35%)

35% of pointer rows anchor ≥75% of claims to a topic chunk. This is the failure pattern Rule 8 was promoted to catch — when the lazy-anchor target's source title shares zero stemmed content tokens with the claim, the run demotes STRICT → HYBRID via TITLE_MISMATCH. The honesty surfaces; the structural lift requires a stronger intervention (evidence-map ranking change, model fine-tune, or a different prompt frame).

19 of 71 questions show JSON ≥ 50pp above pointer, including:

pointer=HH  json=SS   did napoleon really die on saint helena?
pointer=HH  json=SS   name simpsons family members including pets?
pointer=HH  json=SS   what are the planets of our solar system?
pointer=HH  json=SS   what country is the city of prague in?
pointer=HH  json=SS   what is the boltzmann constant?
pointer=UU  json=SS   what is the difference between http and ftp?
pointer=HH  json=SS   when did the soviet union dissolve?
pointer=HH  json=SS   where does the nile river begin?
pointer=HH  json=SS   where is mount kilimanjaro located?
pointer=HH  json=SS   who is supermans girlfriend?
pointer=HH  json=SS   who painted the mona lisa?
pointer=HH  json=SS   who said may the force be with you?
pointer=HH  json=SS   who were the original seven mercury astronauts?
pointer=HH  json=SS   why did the dinosaurs go extinct?

(plus 5 more.) These are factual single-fact questions where JSON's structured per-claim evidence linkage forces the model to pick the specific chunk; pointer's prose-with-tags lets the model lazy-anchor to the topic article and Rule 8 catches it.

subject_in_answer is high across all rungs (STRICT 100%, HYBRID 92%, UNGROUNDED 84%) — the subject is in the answer; the answer just doesn't anchor cleanly to a single chunk. This is structural-honesty signal, not deflection.

Wall-clock & throughput evolution

stamp concurrency scheduling n tasks wall-clock tasks/min mean latency (mode)
11:31Z 4 cell-grouped 2 426 51 min 8.4 26-29 s
15:07Z 4 sample-shuffled 3 639 48 min 13.3 17-19 s

Sample-level shuffled scheduling at the same concurrency delivers +58% throughput. n=3 (50% more total work) ran in 6% LESS wall-clock than n=2 cell-grouped. Per-call mean latency dropped 35-42% across all modes (quote 29.1 s → 16.9 s, pointer 26.7 s → 18.5 s, JSON 28.2 s → 18.3 s) because vLLM's continuous batcher fills better when fed a diverse, uncorrelated request stream.

Concurrency sweep (smoke fixture, 15-task)

c wall-clock tasks/min normalized
3 102 s 8.8 1.00 (peak)
4 106 s 8.5 0.96 (chosen — within 4% of peak, more forgiving on single-call hiccups)
5 119 s 7.6 0.86
6 185 s 4.9 0.55 (vLLM batching ceiling, brutal)

vLLM saturates at c=3-4 on this endpoint. More concurrent requests fill the batch better, but only up to the point where per-call latency growth outpaces parallelism gain. Re-sweep when the endpoint is upgraded, when corpus shape changes the average context size, or when other tenants change the queue.

Errors

Across both benches: 2 errors at 11:31Z, 6 errors at 15:07Z — all on the same question (tell me about the roman empire). Lone UTF-16 surrogates in Wikipedia chunk content, two distinct paths:

  • v1 (3b91223) — hardened SHA-256 hashers on the OUTPUT side (arborist/qa/dag.py, keys.py, evidence.py, store.py audit chain) with errors='surrogatepass' so the run-DAG roots survive surrogate-bearing model output.
  • v2 (41d1d9b) — scrubs message content INSIDE OpenAICompatibleClient.chat_completion before httpx's outbound JSON encode. The corpus chunk text was the path; httpx's .encode('utf-8') on the request body raised before the call left the client.

Verified: tell me about the roman empire under claim_lattice now classifies HYBRID 5/7 instead of erroring. Next bench will land 0 errors.

Verdict & per-mode recommendation

Quote leads on raw lexical grounding (0.54 strict-rate). Best when the four-rung ladder isn't needed and the answer is short-and-cite-able. Doesn't surface anchor-class warrants.

JSON leads among lattice modes (0.42). Recommended default for cache-grade provenance. Structured per-claim evidence linkage; peaks at larger context (32-64 KB) so it scales with evidence growth.

Pointer (0.20) stays useful for low-context-budget scenarios (16-32 KB peak) and prose-distribution models that struggle with JSON grammar; the lazy-anchor honesty cost is the trade.

The directive-coverage table (99% across every observable D2-D7 on lattice modes) is the architectural win. False-positive STRICT was corruption; converting it to honest HYBRID was the price of calibration. Strict-rate slipped on lattice modes; the substrate is more honest.

Outputs

  • bench/qa_results/2026-05-02T11-31-55Z.{jsonl,md} — 426 rows
  • bench/qa_results/2026-05-02T15-07-24Z.{jsonl,md} — 639 rows

(Both gitignored under bench/qa_results/. Headlines & analysis live in this journal.)


Addendum — 2026-05-03 broad-quantifier A/B (Ticket #000008)

This journal froze the 2026-05-02 substrate baseline. On 2026-05-03 Ticket #000008 (broad-quantifier preflight guard) landed Phases 04 and ran a four-cell A/B on a 9-question broad subset (bench/qa_questions_quantifier_subset.txt, 81 runs per cell).

Findings relevant to this journal's per-mode recommendations:

  • JSON-mode strict-rate moves under cap-only: 0.19 baseline → 0.33 with quantifier_guard_apply_caps=true (+14pp on the broad subset). The recommendation here that "JSON leads among lattice modes" still holds, but the headroom above 0.42 globally is partly dependent on broad-vs-narrow question mix.
  • Pointer-mode stays at 0/27 STRICT on broad questions across all four A/B cells. The 0.20 strict-rate above is across the full 71-question set; the broad subset alone is structurally unfavorable to pointer mode regardless of cap or reminder.
  • Reminder eliminates FORMAT_COLLAPSED: 2/27 → 0/27 on pointer mode under quantifier_reminder_enabled=true.
  • Cap and reminder help DIFFERENT failure modes: reminder rescues UNGROUNDED → HYBRID (restates citation rule); cap rescues HYBRID → STRICT (forces fewer-but-better claims). Compound effect on pointer mean-ratio is best at 0.684 (vs 0.473 baseline).

See docs/tickets/ticket-000008-broad-quantifier-preflight-guard.md §12 for the four-cell data and the §10.8 decision-tree verdict. Bench artifacts: bench/qa_results/2026-05-03T12-{29-53,38-53,47-23,54-11}Z.{jsonl,md}.

Addendum 2 — preflight ON vs OFF validation (2026-05-03T23-06-21Z)

After ticket #000010 flipped quantifier_reminder_enabled=True and shipped metacognition_enabled=True defaults, ran a preflight-OFF cell to validate the flip didn't regress baseline behavior. Same 9-question broad subset (bench/qa_questions_quantifier_subset.txt), n=3, policy override --policy metacognition_enabled=false --policy quantifier_reminder_enabled=false.

Compared against §12.6 reminder-only baseline (cleanest single-knob on-cell). On this 9-question subset, none of the metacognition detectors fire (no temporal / contradiction / false-premise / out- of-corpus shapes), so the comparison effectively isolates the reminder contribution.

Metric OFF ON Δ
quote strict-rate 0.63 0.52 11pp (noise band on 27 samples)
pointer strict-rate 0.00 0.00 0
JSON strict-rate 0.26 0.22 4pp
pointer mean ratio 0.483 0.643 +16pp
JSON mean ratio 0.570 0.735 +17pp
JSON UNGROUNDED rate 7/27 (26%) 1/27 (4%) 22pp
pointer FORMAT_COLLAPSED 2/27 0/27 100%
pointer NO_EVIDENCE_POINTER 8/27 6/27 7pp

Interpretation:

  • Mean-ratio improvement is robust (+16-17pp on lattice modes). Grounded rows are MORE thoroughly grounded under preflight ON.
  • JSON UNGROUNDED collapse is dramatic (22pp). The "didn't ground" pool reclassifies into "partially grounded" — operator- visible win.
  • FORMAT_COLLAPSED elimination on pointer mode (2 → 0). The reminder restates the [E\d+] citation rule and Hermes follows it.
  • STRICT-rate moves are within noise. Quote-mode dipped 11pp, but quote is mode-gated off the guard so this is pure Hermes nondeterminism — the §10.8 5pp floor exists exactly to filter this. JSON SR moved 4pp (within floor).

Verdict: the #000010 default-on flip is doing what was claimed. Mean-ratio + UNGROUNDED + FORMAT_COLLAPSED metrics all improve by ≥5pp on lattice modes; STRICT-rate is within noise. Defaults stay on.

Bench artifact: bench/qa_results/2026-05-03T23-06-21Z.{jsonl,md}.

Addendum 3 — full-bench regression check (2026-05-03T23-30-12Z)

Validates that the #000010 default flip (reminder default-on for lattice modes) doesn't regress narrow-question performance. The prior validations in Addendum 1 + Addendum 2 used the 9-question broad subset only; this run sweeps the full 75-question bench/qa_questions.txt (~10% broad, ~89% narrow factoid / descriptive / list shapes).

Same harness, n=3 × 3 modes × 75 questions = 225 runs per mode. Compared against the frozen 15:07Z baseline (the authoritative pre-#000008/9/10 state of the substrate, 71 questions × n=3 = 213 runs per mode):

Mode Pre-flip (15:07Z) Post-flip (23:30Z) Δ SR Δ mean ratio
quote 116S/51H/46U · SR 0.54 · ratio 0.699 116S/60H/49U · SR 0.52 · ratio 0.692 2pp 1pp
claim_lattice_pointer 43S/146H/21U · SR 0.20 · ratio 0.642 48S/145H/32U · SR 0.21 · ratio 0.663 +1pp +2pp
claim_lattice (JSON) 89S/82H/39U · SR 0.42 · ratio 0.698 99S/80H/46U · SR 0.44 · ratio 0.730 +2pp +3pp

All STRICT-rate deltas within the 5pp signal floor. Same question set ±2 (3 added: winners of all major sports?, name all members of the beatles, list all planets in the solar system).

Substrate-level findings:

  • Pointer-mode FORMAT_COLLAPSED: 0/225 across the full sweep. The default-on reminder eliminates the collapse mode globally, not just on broad questions where it was originally measured.
  • Pointer-mode NO_EVIDENCE_POINTER: 29/225 (13%) — lower than the 33% we saw on broad-only (9/27) when reminder was off. Reminder discipline propagates to non-broad questions even though the reminder text only fires on broad shapes (the broader signal is that the reminder text reinforces the citation rule across the model's session attention).
  • Quote-mode is essentially unchanged (-2pp SR, -1pp ratio). Quote opts out of the guard (quantifier_guard_modes default excludes quote), so this is pure Hermes nondeterminism on a 225-sample cell.
  • JSON-mode mean ratio +3pp — small but positive, consistent with the broad-subset finding that reminder boosts per-claim grounding density.

Latency note. Wall-clock per call jumped from ~17-19s (pre-flip baseline) to ~33-35s on this run. This is endpoint contention on Hermes-3 at the time of measurement, not a feature regression (the new substrate adds zero LLM calls — preflight is pure regex). The throughput numbers in §6 of the journal pre-date this evening's network conditions; future bench cycles should normalize.

Verdict — no regression. The #000010 default flip ships clean across the full corpus, not just the broad-quantifier subset where it was originally validated. Defaults stay on. The substrate is strictly more honest under preflight (FORMAT_COLLAPSED → 0, NO_EVIDENCE_POINTER ↓) without sacrificing throughput on non-broad questions.

Bench artifact: bench/qa_results/2026-05-03T23-30-12Z.{jsonl,md}.

Addendum 4 — metacog-trigger detector validation (2026-05-04T02-18-42Z)

The validations above (Addendum 1-3) covered the broad-quantifier subset and the full corpus, neither of which contains questions that fire the #000010 metacog detectors (temporal sensitivity, contradiction, false-premise-lite, out-of-corpus). This addendum closes that empirical gap.

Fixture: bench/qa_questions_metacog_subset.txt — 6 questions, one per detector kind plus a control:

who is the current CEO of OpenAI?               # stale_risk
which character is alive and dead?              # contradictory
when did George Washington stop being president  # false_premise +
  of France?                                    #   stale_risk (multi)
what does my uploaded contract say about        # out_of_corpus
  clause 9?
when did the current CEO of Twitter stop        # stale_risk +
  being CEO of Tesla?                           #   false_premise
who painted the mona lisa?                      # control (well_formed)

Run: n=3 × 6 questions × 3 modes = 54 rows.

Per-question verdict matrix:

question (detector) quote pointer JSON
Current CEO (stale_risk) 0/0/3 U 0/0/3 U 0/0/3 U
Alive+dead (contradictory) 0/0/3 U 3/0/0 S 3/0/0 S
Wash. stop France (FP) 0/0/3 U 0/3/0 H 3/0/0 S
Uploaded contract (OOC) 0/3/0 H 0/1/2 U 3/0/0 S
Twitter→Tesla (multi) 0/0/3 U 3/0/0 S 0/0/3 U
Mona Lisa (control) 3/0/0 S 1/2/0 S 3/0/0 S

(S/H/U = STRICT/HYBRID/UNGROUNDED; n=3 each cell.)

Findings:

  1. Detector accuracy is 6/6. All trigger questions fire the expected logical_statuses value during preflight (verified programmatically before the bench: stale_risk, contradictory_question, false_premise_suspected, out_of_corpus_risk — matched 1:1 with fixture intent). The detectors are doing what their unit tests claim.

  2. Quote mode is the most honest fallback. 4 of 5 trigger questions land all-UNGROUNDED on quote mode. The paraphrase verifier won't substring-match across the corpus when the question's premise has no anchor. Quote mode's mode-gated-off guard works in our favor here.

  3. Lattice modes accidentally ground 2 trigger questions to STRICT. Schrödinger's cat (alive+dead) JSON STRICT is defensible — the corpus contains quantum-mechanics articles that legitimately discuss the state. But:

    • JSON STRICT on "when did George Washington stop being president of France?" is NOT defensible. False premise; the model invented an answer that lexically grounded against some chunk. The metacog detector correctly flagged false_premise_suspected; the audit-line tail surfaced · false premise; but the verdict still lands STRICT.
    • JSON STRICT on "what does my uploaded contract say about clause 9?" is the same shape: out-of-corpus reference, model fabricates a grounding.
  4. Audit-line tails are doing operator-warning duty correctly. The bench rows persist preflight_logical_statuses and the render layer tails (· stale risk, · false premise, etc.) even on STRICT verdicts — operator sees the warning. But the substrate doesn't refuse execution by default for these shapes.

Implication for #000011 (SOFT_PREFLIGHT_HINT). This bench validates the design rationale for the soft-sidecar ticket: the deterministic metacog detectors flag these shapes correctly, but the corpus accidentally grounds 2/5 of them to STRICT. A model- assisted soft preflight could add independent semantic skepticism ("does George Washington being president of France match historical reality?") that the lexical detectors can't supply. Soft sidecar output would surface as · soft: false_premise_* on the audit-line, distinct from the hard · false premise tail, giving operators a stronger warning when both signals fire.

Implication for default policy. Keep metacognition_block_on_contradiction=False as the default. Schrödinger's cat (alive+dead) would have been rejected unnecessarily under a hard-block, and that's a real-world question with a legitimate answer. The label-only default is correct; operators wanting strictness opt in via --block-on-contradiction.

Bench artifact: bench/qa_results/2026-05-04T02-18-42Z.{jsonl,md}.

Addendum 5 — #000046 Phase 3 verifier numeric-gate regression check (2026-05-11)

arborist/qa/verify.py gained a paraphrase numeric-agreement gate (_numeric_signature + a check in _check_each_with_paraphrase): a span that token-covers the source ≥ paraphrase_coverage but asserts a digit-number the source lacks (modulo thousands-comma) is no longer paraphrase-grounded — it goes to unverified. Closes #000046 by lifting the falsification-hard-v1.jsonl rate 4/12 → 6/12 (5f-fal-hard-004 50-vs-100 and -007 300-vs-300,000 → UNGROUNDED).

Before/after make bench-qa (n=3 × 75 questions × 3 modes = 675 cells):

mode STRICT-rate before → after Δ
quote 0.53 → 0.50 3pp
claim_lattice_pointer 0.23 → 0.25 +2pp
claim_lattice 0.44 → 0.45 +1pp

All within the 5-pp noise floor. Per-row diff (675 common cells, 74 changed audit_mode): the only clearly gate-attributable QA shift was the fictional "our cold fusion breakthrough" year-claim demoting STRICT → HYBRID across all 3 samples — a correct demotion (the year isn't grounded). Every other transition was quote→quote / claim_lattice→claim_lattice LLM re-answer variance — the answer text changed on the re-run, not the verifier (the gate touches only the paraphrase fallback, never the verbatim/span/entity/claim-lattice paths). No regression on legitimate answers; the gate ships.

Bench artifacts: bench/qa_results/2026-05-11T13-42-38Z.{jsonl,md} (before) · bench/qa_results/2026-05-11T14-19-51Z.{jsonl,md} (after). Full per-ticket detail: docs/tickets/ticket-000046-harder-5sf-fixture-tier.md §5 Phase 3.

Addendum 6 — #000048 step 2.1 verifier entity-gate regression check (2026-05-11)

arborist/qa/verify.py gained an entity salient-token-disagreement gate (_entity_salient_disagrees + _is_single_sentence): in verify_quotes' entity branch (proximity policy), the weakest grounding — not cluster AND len(verified) <= 1 AND single sentence — is declined (→ UNGROUNDED) when the answer asserts a > 4-char capitalized content token (stopword-filtered) or a digit-number the source lacks. Catches "Insulin was discovered by Alexander Fleming" against "Penicillin was discovered by Alexander Fleming" (shared "Alexander Fleming" matched; swapped subject "Insulin" trips the gate). Lifts the falsification-hard-v1.jsonl rate 6/12 → 10/12 (the 4 HYBRID_ENTITY over-grounds: Insulin / Berlin / 1889 / Pacific).

Before/after make bench-qa (n=3 × 75 questions × 3 modes = 675 cells):

mode STRICT-rate before → after Δ
quote 0.50 → 0.54 +4pp
claim_lattice_pointer 0.25 → 0.22 3pp
claim_lattice 0.45 → 0.43 2pp

All within the 5-pp noise floor. Per-row diff (675 common cells, 30 quote-mode rows changed audit_mode): 0 quote-mode rows demoted to UNGROUNDED from the entity path — the gate didn't fire on a single legitimate QA answer in the whole bench. Every quote-mode transition was LLM re-answer variance (verifier quote→quote with the verdict flipping = a different answer text); the pointer/lattice deltas are noise too (the gate is in verify_quotes / quote mode, not the claim-lattice verifier). The gate is provably narrow on real traffic; it ships.

Bench artifacts: bench/qa_results/2026-05-11T14-19-51Z.{jsonl,md} (before — HEAD's verify.py) · bench/qa_results/2026-05-11T17-12-41Z.{jsonl,md} (after). Full per-ticket detail: docs/tickets/ticket-000048-verifier-upgrade-recombination-segmentation.md §5 step 2.1.

Addendum 7 — #000048 step 2.4 parse_pointer_claims clause-segmentation regression check (2026-05-11)

arborist/qa/parse_claims.py gained a clause segmenter: a line that crams several well-pointered claims onto one row is split into one claim per clause (split on ;, sentence boundaries, spaced dashes, and/or/because/although/since/while, inline (N) enumeration markers, commas — with (?![^\[]*\]) so a comma inside a [E1, E2] bracket never splits it). _segment_line keeps the split only if every resulting segment is well-pointered (a legit single claim like "The cat is black and white [E1]." or "The cast: A, B, C [E1]." is never broken — splitting would create pointer-less fragments → guard rejects; a leading colon-terminated header — "Two facts:", "Key points:" — is dropped). Plus a wrapped-bullet join: a continuation line (leading whitespace then a lowercase letter, no bullet) folds into the previous claim. Closes the 8 mis-segments in formulate-hard-v1.jsonl → rate 4/12 → 12/12 (that pack now at ceiling).

Before/after make bench-qa (n=3 × 75 questions × 3 modes = 675 cells; parse_pointer_claims feeds the 450 claim_lattice_pointer + claim_lattice cells):

mode STRICT-rate before → after Δ
quote 0.54 → 0.55 +1pp
claim_lattice_pointer 0.22 → 0.22 0pp
claim_lattice 0.43 → 0.45 +2pp

All within the 5-pp noise floor. Per-row diff (675 common cells): the segmenter changed the parsed-claim count on the same answer text for 7 of the 450 lattice cells (0 in claim_lattice, 7 in claim_lattice_pointer); of those, 2 caused an audit_mode change — both correct: (a) a wrap-join recovered an answer's intended structure (4 claims, 2 of which were pointer-less wrap-fragments → HYBRID) into 2 well-pointered claims → STRICT; (b) a crammed-one-line blob (1 monolithic claim, all pointers → STRICT) split into 8 claims, some of which don't individually verify → HYBRID — the honest verdict (false-positive STRICT was the corruption). Every other lattice/quote delta is LLM re-answer variance (answer_chars changed, often drastically). No regression — the segmenter's only visible effects on real traffic are honest improvements.

Bench artifacts: bench/qa_results/2026-05-11T17-12-41Z.{jsonl,md} (before — HEAD's parse_claims.py) · bench/qa_results/2026-05-11T20-26-37Z.{jsonl,md} (after). Full per-ticket detail: docs/tickets/ticket-000048-verifier-upgrade-recombination-segmentation.md §5 step 2.4.

Addendum 8 — #000057 control sweep: retrieval × model × framing × reasoning (2026-05-19/20)

The question. Hermes-3-8B confidently states the present-day office-holder ("the president of France is Emmanuel Macron") against a ~2010-vintage corpus (Sarkozy). Is that an 8B-model weakness, an unfair question framing, or does retrieval fix it regardless? The sweep crosses {model} × {framing} × {retrieval on/off} over a 386-item office-holder fixture (bench/qa_questions_stale_map.json), each item with a fixed corpus-vintage gold article.

  • models: hermes (Hermes-3-8B), qwen-nothink (Qwen3.6-27B, reasoning off), qwen-think (same weights, reasoning on). Qwen runs on llama.cpp; Hermes on vLLM.
  • framings: plain ("who is the president of France?"), source_relative ("According to the reference knowledge base, …"), as_of_corpus ("As of 2010, who was …").
  • arms: solo (model alone — measures parametric prior) vs arborist (model + retrieval over the 2010 corpus, answer_mode= claim_lattice).

Judge methodology — the code judge. The original sweep used the Opus headless judge (bench/judge.py); it burned our Anthropic quota and 79.5 % of its verdicts came back JUDGE_ERROR (rate-limited). Replaced with a deterministic code judge (bench/judge_code.py): abstention regex → short-answer entity-grounding fast path → NLI contradiction (θ=0.85) → lexical verifier (quote/span/entity/ paraphrase) → WRONG-vs-FABRICATED tie-break on subject-in-gold. No LLM, no quota, fully replayable. Calibrated against the records Opus did grade: agreement rose CG 13 %→47 %, WRONG 56 %→89 %, ABSTAINED 80 %→95 % across four targeted fixes (raise NLI contradiction threshold off the manifest's 0.5; expand abstention patterns for the "I do not have … information/access" family; demote FABRICATED→WRONG when gold mentions the question subject; short-answer entity-grounding for terse-name answers the verifier's prose-shape extractor whiffs on). A claim-lattice JSON-unwrap lets the same judge grade the Arborist arm's {"claims":[…]} envelopes as plain prose. Verdicts are a deterministic proxy for grounding (lexical + NLI), not Opus-grade reading — JUDGE_ERROR residue (HYBRID-via-entity with no NLI corroboration) is the natural input to a later LLM-batch pass.

Consolidated scorecard — CG% (correct-grounded rate), all arms on the identical calibrated judge. arborist/hermes is n=40 (phase 1's smaller arborist arm); every other cell is n=386.

arm / model plain source_relative as_of_corpus
solo / hermes 9 % 5 % 18 %
solo / qwen-nothink 7 % 0 % 50 %
solo / qwen-think 6 % 5 % 44 %
arborist / hermes (n=40) 60 % 62 % 25 %
arborist / qwen-nothink 82 % 65 % 50 %

Grounding-fidelity detail (abstain% / wrong-assert% = W+F over n) for the cells the findings turn on:

cell CG% abstain% wrong-assert%
solo qwen-nothink / source_relative 0 % 97 % 1 %
solo qwen-think / source_relative 5 % 61 % 32 %
solo qwen-nothink / plain 7 % 9 % 79 %
solo qwen-think / plain 6 % 43 % 48 %
arborist qwen-nothink / plain 82 % 0 % 11 %

Findings.

  1. Retrieval dominates every other lever. arborist/qwen-nothink/ plain = 82 % CG vs 7 % solo; arborist/hermes/plain = 60 % vs 9 % solo. No solo configuration — no model size, no framing — approaches the retrieval arms. The production answer is retrieval.

  2. Reasoning does not improve raw correctness. solo qwen-think vs qwen-nothink: as_of_corpus 44 % vs 50 % (nothink wins), plain 6 %/7 %, source_relative 5 %/0 %. Chain-of-thought trades correctness for caution; it does not recall corpus-vintage facts better.

  3. Reasoning's real cost is broken honest-abstention. The cleanest grounding signal in the sweep is solo qwen-nothink/source_relative: 97 % abstain — told to use a reference it lacks, it declines. Reasoning breaks this: qwen-think/source_relative abstains only 61 % and reasons itself into wrong parametric answers (wrong-assert 1 %→32 %). On plain framing reasoning helps caution (wrong-assert 79 %→48 %); on the source-grounding framing it hurts.

Production recommendation: arborist + qwen-nothink, plain framing — 82 % CG, ~0 % abstain, 11 % wrong-assert. Reasoning OFF: it degrades the honest-abstention property and buys nothing once retrieval supplies the source.

Held / in flight: arborist + qwen-think (the last cell) is running at write time — tests whether reasoning hurts the retrieval arm the way it hurt solo/source_relative. Will append the result.

Artifacts. Solo + arborist-hermes: control_sweep_2026-05-19T21-52-56Z* (phase 1). qwen-think solo: …23-22-21Z* (phase 2). arborist+qwen-nothink: …2026-05-20T00-07-38Z* (phase 3). All re-graded under the final judge (*_code_judge_final.{md,jsonl}). Judge + sweep tooling: bench/judge_code.py, bench/score_with_code_judge.py, bench/analyze_judge_disagreement.py, bench/control_sweep.py (--judge {code,opus}, --skip-solo). Engine-agnostic JSON-schema enforcement for the Arborist arm: arborist.qa.verify.claim_lattice_structured_output_extras.