# QA-modes bench — 2026-05-02 **Date:** 2026-05-02 **Endpoint:** `https://hermes.ai.unturf.com/v1` (Hermes-3-Llama-3.1-8B-FP8-Dynamic, vLLM, 82K ctx) **Corpus:** Wikipedia 2003-05-16 cur snapshot, sharded under `~/.arborist/shards` Two sweeps landed today, each 71 questions × 3 modes: | stamp | samples | concurrency | scheduling | runs | wall-clock | |-------|---------|-------------|------------|------|------------| | `11:31Z` | n=2 | c=4 | cell-grouped | 426 | 51 min | | `15:07Z` | n=3 | c=4 | sample-shuffled, `--seed 0` | 639 | **48 min** | The 15:07Z sweep is the **authoritative state of the substrate** at end-of-day; the 11:31Z sweep is the intermediate witnessed before the second wave of hardening landed. Headlines below show both. ## Hardening since the prior bench (2026-04-30 post-retry) Substrate-side (between `2026-04-30T18:46Z` and `2026-05-02T11:31Z`): - Rule 8 — title-relevance promoted to hard verifier check (`TITLE_MISMATCH` violation demotes `STRICT → HYBRID` when no cited evidence's source title shares a stemmed token with the claim). - Anchor-class warrant generalized to entity-list / count / why-cause shapes (Ticket #000003). - Retrieval-plan hash bound into the run-DAG retrieval stage (Ticket #000001). - Reference-frame polarity contract (Ticket #000002). - Four-rung ladder display (Ticket #000005) — `POINTER-LINKED → ANCHOR-WARRANTED → EVIDENCE-WARRANTED → (ENTAILMENT-VERIFIED reserved)` mapped from the v9.8 audit_mode trichotomy at render time. - Bench harness extended with directive coverage + log-scale buckets to 1M (Ticket #000004). Substrate-side (between `2026-05-02T11:31Z` and `2026-05-02T15:07Z`): - **Sprint 1b** — per-mode `max_context_chars`. Bench's recommended-context-budget table flows back into `DEFAULT_QUERY_POLICY["max_context_chars_by_mode"]`: quote 24 KB, pointer 24 KB, JSON 48 KB. Folds into `governance_policy_hash`. - **Sprint 2** — pointer Rule 9: chunk-specificity instruction added to the lattice-pointer system prompt. - **DRY collapse** — the four lattice prompts (system + grounding × pointer + JSON) lifted to `arborist/qa/prompts.py` as a single source of truth, imported by both `runner.DEFAULT_POLICY` and `query.DEFAULT_QUERY_POLICY`. - **httpx persistent client** — the chat-completion path used to construct a fresh `httpx.Client` per call, paying a TLS handshake every request. Move to `__init__`; HTTP/1.1 keep-alive across calls. Save 1-2 min on a 426-call bench. - **Sample-level shuffled bench scheduling** — every `(question, mode, sample_idx)` is a task, shuffled with `--seed`, dispatched concurrently. Per-cell `Lock` dict serializes burn-then-write on the shared cache_key. True i.i.d. n=3 variance; vLLM batcher fed a diverse request stream. - **Bench `--resume`** — read existing JSONL, skip done tasks, append fresh rows. Stop/start-able. - **Concurrency sweep on the smoke fixture.** vLLM peaks at c=3-4; saturates badly past c=4 (see Wall-clock & throughput below). - **UTF-16 surrogate fix v2** (`41d1d9b`) — corpus chunks with lone surrogates broke httpx's outbound JSON encode. Scrub at the client boundary by WTF-8 → UTF-8-with-replace roundtrip. (The earlier v1 fix `3b91223` hardened the OUTPUT side — SHA-256 hashers — but missed the INPUT side. Both paths now safe.) ## Aggregate (15:07Z, authoritative) | mode | runs | STRICT | HYBRID | UNGROUNDED | err | strict-rate | mean ratio | mean latency | |------|------|--------|--------|------------|-----|-------------|------------|--------------| | `quote` | 213 | 116 | 51 | 46 | 0 | **0.54** | 0.699 | **16.9 s** | | `claim_lattice_pointer` | 213 | 43 | 146 | 21 | 3 | **0.20** | 0.642 | 18.5 s | | `claim_lattice` (JSON) | 213 | 89 | 82 | 39 | 3 | **0.42** | 0.698 | 18.3 s | ## Δ across the day | mode | 2026-04-30 (post-retry) | 2026-05-02T11:31Z | 2026-05-02T15:07Z | net Δ | |------|------------------------|-------------------|-------------------|-------| | `quote` | 0.47 | 0.50 (+3pp) | **0.54** (+4pp) | **+7pp** | | `claim_lattice_pointer` | 0.24 | 0.23 (−1pp) | 0.20 (−3pp) | −4pp | | `claim_lattice` (JSON) | 0.50 | 0.44 (−6pp) | 0.42 (−2pp) | −8pp | **Quote climbed the most across the day**, ending at 0.54. The Sprint 1b 24 KB cap surfaces tighter retrievals that quote can ground verbatim (the bucket data confirms — quote's peak migrated to 8-16 KB at 0.58 strict-rate). **Pointer & JSON each took an honesty cost.** Rule 8 promotion + warrant-class generalization demote more cases that *would* have classified STRICT under the looser pre-2026-05-02 verifier. The architectural win — 99% directive coverage on lattice modes — is the price for those drops; false-positive STRICT was corruption, and we converted it to honest HYBRID. ## Per-bucket strict-rate (15:07Z) | mode | bucket | runs | strict-rate | note | |------|--------|------|-------------|------| | `quote` | 8-16 KB | 105 | **0.58** | peak | | `quote` | 16-32 KB | 105 | 0.52 | | | `claim_lattice_pointer` | 16-32 KB | 210 | **0.20** | peak (only bucket) | | `claim_lattice` (JSON) | 32-64 KB | 99 | **0.48** | peak | | `claim_lattice` (JSON) | 16-32 KB | 108 | 0.38 | | **Sprint 1b's intent confirmed.** JSON's peak at 32-64 KB justifies the 48 KB default. Quote's peak migrated to 8-16 KB after the 24 KB cap; the cap may be tighter than optimal — quote could plausibly be cut to 16 KB (mid of 8-16 KB bucket) for another small lift in a follow-up sprint. `claim_lattice` (JSON) **peaks at LARGER context** than the other modes. The structured per-claim evidence linkage that JSON enforces benefits from more evidence per claim. Pointer & quote degrade past 16-32 KB. ## Recommended context budget (per-mode) 5-run minimum to reduce noise; from the 15:07Z bucket data: | mode | peak bucket | strict-rate | n | |------|-------------|-------------|---| | `quote` | 8-16 KB | 0.58 | 105 | | `claim_lattice_pointer` | 16-32 KB | 0.20 | 210 | | `claim_lattice` (JSON) | **32-64 KB** | **0.48** | 99 | These flow into `DEFAULT_QUERY_POLICY["max_context_chars_by_mode"]` (currently 24 / 24 / 48 KB — mid of each peak bucket). When the bench shifts those peaks, retune the policy and let `governance_policy_hash` partition the cache. ## Directive coverage (seven-point program) | mode | D2 pointer | D3 cti-ready | D4 ev-map bound | D6 warrant | D7 honest label | |------|-----------|--------------|-----------------|-----------|----------------| | `quote` | 0/213 (0%) | 0/213 (0%) | 213/213 (100%) | 0/213 (0%) | 213/213 (100%) | | `claim_lattice_pointer` | 210/213 (99%) | 210/213 (99%) | 210/213 (99%) | 210/213 (99%) | 210/213 (99%) | | `claim_lattice` | 210/213 (99%) | 210/213 (99%) | 210/213 (99%) | 210/213 (99%) | 210/213 (99%) | Lattice modes hit **99% on every observable directive**; the 1% gap is the 6 surrogate errors (already fixed in `41d1d9b`, next bench will hit 100%). Quote shows 0% on D2 / D3 / D6 by construction (those are lattice-only directives) and 100% on D4 / D7. ## Pointer-mode signal — where the gap lives Pointer mode's strict-rate did not lift on Sprint 2's chunk-specificity prompt nudge. The load-bearing failure pattern is **lazy-anchoring**: the model cites a topic-overview chunk for every claim instead of the specific chunk that supports each individual claim. Lazy-anchor ratio histogram (11:31Z bench, 142 pointer rows): | `lazy_anchor_ratio` bucket | rows | |----------------------------|------| | `0.00` | 25 | | `<0.25` | 9 | | `<0.5` | 11 | | `<0.75` | 46 | | `>=0.75` | **49** (35%) | 35% of pointer rows anchor ≥75% of claims to a topic chunk. **This is the failure pattern Rule 8 was promoted to catch** — when the lazy-anchor target's source title shares zero stemmed content tokens with the claim, the run demotes `STRICT → HYBRID` via `TITLE_MISMATCH`. The honesty surfaces; the structural lift requires a stronger intervention (evidence-map ranking change, model fine-tune, or a different prompt frame). 19 of 71 questions show JSON ≥ 50pp above pointer, including: ``` pointer=HH json=SS did napoleon really die on saint helena? pointer=HH json=SS name simpsons family members including pets? pointer=HH json=SS what are the planets of our solar system? pointer=HH json=SS what country is the city of prague in? pointer=HH json=SS what is the boltzmann constant? pointer=UU json=SS what is the difference between http and ftp? pointer=HH json=SS when did the soviet union dissolve? pointer=HH json=SS where does the nile river begin? pointer=HH json=SS where is mount kilimanjaro located? pointer=HH json=SS who is supermans girlfriend? pointer=HH json=SS who painted the mona lisa? pointer=HH json=SS who said may the force be with you? pointer=HH json=SS who were the original seven mercury astronauts? pointer=HH json=SS why did the dinosaurs go extinct? ``` (plus 5 more.) These are factual single-fact questions where JSON's structured per-claim evidence linkage forces the model to pick the specific chunk; pointer's prose-with-tags lets the model lazy-anchor to the topic article and Rule 8 catches it. `subject_in_answer` is high across all rungs (STRICT 100%, HYBRID 92%, UNGROUNDED 84%) — the subject is in the answer; the answer just doesn't anchor cleanly to a single chunk. This is **structural-honesty signal**, not deflection. ## Wall-clock & throughput evolution | stamp | concurrency | scheduling | n | tasks | wall-clock | tasks/min | mean latency (mode) | |-------|-------------|-----------|---|-------|------------|-----------|---------------------| | 11:31Z | 4 | cell-grouped | 2 | 426 | 51 min | 8.4 | 26-29 s | | 15:07Z | 4 | sample-shuffled | 3 | 639 | **48 min** | **13.3** | **17-19 s** | Sample-level shuffled scheduling at the same concurrency delivers **+58% throughput**. n=3 (50% more total work) ran in 6% LESS wall-clock than n=2 cell-grouped. Per-call mean latency dropped 35-42% across all modes (quote 29.1 s → 16.9 s, pointer 26.7 s → 18.5 s, JSON 28.2 s → 18.3 s) because vLLM's continuous batcher fills better when fed a diverse, uncorrelated request stream. ### Concurrency sweep (smoke fixture, 15-task) | c | wall-clock | tasks/min | normalized | |---|------------|-----------|------------| | 3 | 102 s | 8.8 | **1.00** (peak) | | 4 | 106 s | 8.5 | 0.96 (chosen — within 4% of peak, more forgiving on single-call hiccups) | | 5 | 119 s | 7.6 | 0.86 | | 6 | 185 s | 4.9 | 0.55 (vLLM batching ceiling, brutal) | vLLM saturates at c=3-4 on this endpoint. More concurrent requests fill the batch better, but only up to the point where per-call latency growth outpaces parallelism gain. Re-sweep when the endpoint is upgraded, when corpus shape changes the average context size, or when other tenants change the queue. ## Errors Across both benches: **2 errors at 11:31Z, 6 errors at 15:07Z** — all on the same question (`tell me about the roman empire`). Lone UTF-16 surrogates in Wikipedia chunk content, two distinct paths: - v1 (`3b91223`) — hardened SHA-256 hashers on the OUTPUT side (`arborist/qa/dag.py`, `keys.py`, `evidence.py`, `store.py` audit chain) with `errors='surrogatepass'` so the run-DAG roots survive surrogate-bearing model output. - v2 (`41d1d9b`) — scrubs message content INSIDE `OpenAICompatibleClient.chat_completion` before httpx's outbound JSON encode. The corpus chunk text was the path; httpx's `.encode('utf-8')` on the request body raised before the call left the client. Verified: `tell me about the roman empire` under `claim_lattice` now classifies HYBRID 5/7 instead of erroring. Next bench will land 0 errors. ## Verdict & per-mode recommendation **Quote leads on raw lexical grounding (0.54 strict-rate).** Best when the four-rung ladder isn't needed and the answer is short-and-cite-able. Doesn't surface anchor-class warrants. **JSON leads among lattice modes (0.42).** Recommended default for cache-grade provenance. Structured per-claim evidence linkage; peaks at larger context (32-64 KB) so it scales with evidence growth. **Pointer (0.20) stays useful for low-context-budget scenarios** (16-32 KB peak) and prose-distribution models that struggle with JSON grammar; the lazy-anchor honesty cost is the trade. **The directive-coverage table (99% across every observable D2-D7 on lattice modes) is the architectural win.** False-positive STRICT was corruption; converting it to honest HYBRID was the price of calibration. Strict-rate slipped on lattice modes; the substrate is more honest. ## Outputs - `bench/qa_results/2026-05-02T11-31-55Z.{jsonl,md}` — 426 rows - `bench/qa_results/2026-05-02T15-07-24Z.{jsonl,md}` — 639 rows (Both gitignored under `bench/qa_results/`. Headlines & analysis live in this journal.) --- ## Addendum — 2026-05-03 broad-quantifier A/B (Ticket #000008) This journal froze the 2026-05-02 substrate baseline. On 2026-05-03 Ticket #000008 (broad-quantifier preflight guard) landed Phases 0–4 and ran a four-cell A/B on a 9-question broad subset (`bench/qa_questions_quantifier_subset.txt`, 81 runs per cell). Findings relevant to this journal's per-mode recommendations: - **JSON-mode strict-rate moves under cap-only**: 0.19 baseline → 0.33 with `quantifier_guard_apply_caps=true` (+14pp on the broad subset). The recommendation here that "JSON leads among lattice modes" still holds, but the *headroom above 0.42 globally* is partly dependent on broad-vs-narrow question mix. - **Pointer-mode stays at 0/27 STRICT** on broad questions across all four A/B cells. The 0.20 strict-rate above is *across the full 71-question set*; the broad subset alone is structurally unfavorable to pointer mode regardless of cap or reminder. - **Reminder eliminates FORMAT_COLLAPSED**: 2/27 → 0/27 on pointer mode under `quantifier_reminder_enabled=true`. - **Cap and reminder help DIFFERENT failure modes**: reminder rescues UNGROUNDED → HYBRID (restates citation rule); cap rescues HYBRID → STRICT (forces fewer-but-better claims). Compound effect on pointer mean-ratio is best at 0.684 (vs 0.473 baseline). See `docs/tickets/ticket-000008-broad-quantifier-preflight-guard.md` §12 for the four-cell data and the §10.8 decision-tree verdict. Bench artifacts: `bench/qa_results/2026-05-03T12-{29-53,38-53,47-23,54-11}Z.{jsonl,md}`. ### Addendum 2 — preflight ON vs OFF validation (2026-05-03T23-06-21Z) After ticket #000010 flipped `quantifier_reminder_enabled=True` and shipped `metacognition_enabled=True` defaults, ran a preflight-OFF cell to validate the flip didn't regress baseline behavior. Same 9-question broad subset (`bench/qa_questions_quantifier_subset.txt`), n=3, policy override `--policy metacognition_enabled=false --policy quantifier_reminder_enabled=false`. Compared against §12.6 reminder-only baseline (cleanest single-knob on-cell). On this 9-question subset, none of the metacognition detectors fire (no temporal / contradiction / false-premise / out- of-corpus shapes), so the comparison effectively isolates the reminder contribution. | Metric | OFF | ON | Δ | |------------------------------|-----------|-----------|----------| | `quote` strict-rate | 0.63 | 0.52 | −11pp (noise band on 27 samples) | | `pointer` strict-rate | 0.00 | 0.00 | 0 | | `JSON` strict-rate | 0.26 | 0.22 | −4pp | | **`pointer` mean ratio** | 0.483 | **0.643** | **+16pp** | | **`JSON` mean ratio** | 0.570 | **0.735** | **+17pp** | | **`JSON` UNGROUNDED rate** | 7/27 (26%) | **1/27 (4%)** | **−22pp** | | **`pointer` FORMAT_COLLAPSED** | 2/27 | **0/27** | **−100%** | | `pointer` NO_EVIDENCE_POINTER | 8/27 | 6/27 | −7pp | **Interpretation:** - **Mean-ratio improvement is robust** (+16-17pp on lattice modes). Grounded rows are MORE thoroughly grounded under preflight ON. - **JSON UNGROUNDED collapse is dramatic** (−22pp). The "didn't ground" pool reclassifies into "partially grounded" — operator- visible win. - **FORMAT_COLLAPSED elimination on pointer mode** (2 → 0). The reminder restates the [E\d+] citation rule and Hermes follows it. - **STRICT-rate moves are within noise.** Quote-mode dipped 11pp, but quote is mode-gated off the guard so this is pure Hermes nondeterminism — the §10.8 5pp floor exists exactly to filter this. JSON SR moved −4pp (within floor). **Verdict:** the #000010 default-on flip is doing what was claimed. Mean-ratio + UNGROUNDED + FORMAT_COLLAPSED metrics all improve by ≥5pp on lattice modes; STRICT-rate is within noise. Defaults stay on. Bench artifact: `bench/qa_results/2026-05-03T23-06-21Z.{jsonl,md}`. ### Addendum 3 — full-bench regression check (2026-05-03T23-30-12Z) Validates that the #000010 default flip (reminder default-on for lattice modes) doesn't regress narrow-question performance. The prior validations in Addendum 1 + Addendum 2 used the 9-question broad subset only; this run sweeps the full 75-question `bench/qa_questions.txt` (~10% broad, ~89% narrow factoid / descriptive / list shapes). Same harness, n=3 × 3 modes × 75 questions = 225 runs per mode. Compared against the frozen `15:07Z` baseline (the authoritative pre-#000008/9/10 state of the substrate, 71 questions × n=3 = 213 runs per mode): | Mode | Pre-flip (15:07Z) | **Post-flip (23:30Z)** | Δ SR | Δ mean ratio | |-------------------------|-------------------|------------------------|------|--------------| | `quote` | 116S/51H/46U · SR 0.54 · ratio 0.699 | 116S/60H/49U · **SR 0.52** · **ratio 0.692** | −2pp | −1pp | | `claim_lattice_pointer` | 43S/146H/21U · SR 0.20 · ratio 0.642 | 48S/145H/32U · **SR 0.21** · **ratio 0.663** | +1pp | +2pp | | `claim_lattice` (JSON) | 89S/82H/39U · SR 0.42 · ratio 0.698 | 99S/80H/46U · **SR 0.44** · **ratio 0.730** | +2pp | +3pp | All STRICT-rate deltas within the 5pp signal floor. Same question set ±2 (3 added: `winners of all major sports?`, `name all members of the beatles`, `list all planets in the solar system`). **Substrate-level findings:** - **Pointer-mode `FORMAT_COLLAPSED`: 0/225** across the full sweep. The default-on reminder eliminates the collapse mode globally, not just on broad questions where it was originally measured. - **Pointer-mode `NO_EVIDENCE_POINTER`: 29/225** (13%) — lower than the 33% we saw on broad-only (9/27) when reminder was off. Reminder discipline propagates to non-broad questions even though the reminder text only fires on broad shapes (the broader signal is that the reminder text reinforces the citation rule across the model's session attention). - **Quote-mode is essentially unchanged** (-2pp SR, -1pp ratio). Quote opts out of the guard (`quantifier_guard_modes` default excludes quote), so this is pure Hermes nondeterminism on a 225-sample cell. - **JSON-mode mean ratio +3pp** — small but positive, consistent with the broad-subset finding that reminder boosts per-claim grounding density. **Latency note.** Wall-clock per call jumped from ~17-19s (pre-flip baseline) to ~33-35s on this run. This is **endpoint contention** on Hermes-3 at the time of measurement, not a feature regression (the new substrate adds zero LLM calls — preflight is pure regex). The throughput numbers in §6 of the journal pre-date this evening's network conditions; future bench cycles should normalize. **Verdict — no regression.** The #000010 default flip ships clean across the full corpus, not just the broad-quantifier subset where it was originally validated. Defaults stay on. The substrate is strictly more honest under preflight (FORMAT_COLLAPSED → 0, NO_EVIDENCE_POINTER ↓) without sacrificing throughput on non-broad questions. Bench artifact: `bench/qa_results/2026-05-03T23-30-12Z.{jsonl,md}`. ### Addendum 4 — metacog-trigger detector validation (2026-05-04T02-18-42Z) The validations above (Addendum 1-3) covered the broad-quantifier subset and the full corpus, neither of which contains questions that fire the #000010 metacog detectors (temporal sensitivity, contradiction, false-premise-lite, out-of-corpus). This addendum closes that empirical gap. Fixture: `bench/qa_questions_metacog_subset.txt` — 6 questions, one per detector kind plus a control: ``` who is the current CEO of OpenAI? # stale_risk which character is alive and dead? # contradictory when did George Washington stop being president # false_premise + of France? # stale_risk (multi) what does my uploaded contract say about # out_of_corpus clause 9? when did the current CEO of Twitter stop # stale_risk + being CEO of Tesla? # false_premise who painted the mona lisa? # control (well_formed) ``` Run: n=3 × 6 questions × 3 modes = 54 rows. **Per-question verdict matrix:** | question (detector) | quote | pointer | JSON | |---------------------------|---------|----------|---------| | Current CEO (stale_risk) | 0/0/3 U | 0/0/3 U | 0/0/3 U | | Alive+dead (contradictory)| 0/0/3 U | 3/0/0 S | 3/0/0 S | | Wash. stop France (FP) | 0/0/3 U | 0/3/0 H | 3/0/0 S | | Uploaded contract (OOC) | 0/3/0 H | 0/1/2 U | 3/0/0 S | | Twitter→Tesla (multi) | 0/0/3 U | 3/0/0 S | 0/0/3 U | | Mona Lisa (control) | 3/0/0 S | 1/2/0 S | 3/0/0 S | (S/H/U = STRICT/HYBRID/UNGROUNDED; n=3 each cell.) **Findings:** 1. **Detector accuracy is 6/6.** All trigger questions fire the expected `logical_statuses` value during preflight (verified programmatically before the bench: `stale_risk`, `contradictory_question`, `false_premise_suspected`, `out_of_corpus_risk` — matched 1:1 with fixture intent). The detectors are doing what their unit tests claim. 2. **Quote mode is the most honest fallback.** 4 of 5 trigger questions land all-UNGROUNDED on quote mode. The paraphrase verifier won't substring-match across the corpus when the question's premise has no anchor. Quote mode's mode-gated-off guard works in our favor here. 3. **Lattice modes accidentally ground 2 trigger questions to STRICT.** Schrödinger's cat (alive+dead) JSON STRICT is *defensible* — the corpus contains quantum-mechanics articles that legitimately discuss the state. But: - JSON STRICT on **"when did George Washington stop being president of France?"** is **NOT defensible**. False premise; the model invented an answer that lexically grounded against some chunk. The metacog detector correctly flagged `false_premise_suspected`; the audit-line tail surfaced `· false premise`; but the verdict still lands STRICT. - JSON STRICT on **"what does my uploaded contract say about clause 9?"** is the same shape: out-of-corpus reference, model fabricates a grounding. 4. **Audit-line tails are doing operator-warning duty correctly.** The bench rows persist `preflight_logical_statuses` and the render layer tails (`· stale risk`, `· false premise`, etc.) even on STRICT verdicts — operator sees the warning. But the substrate doesn't refuse execution by default for these shapes. **Implication for #000011 (SOFT_PREFLIGHT_HINT).** This bench validates the design rationale for the soft-sidecar ticket: the deterministic metacog detectors flag these shapes correctly, but the corpus accidentally grounds 2/5 of them to STRICT. A model- assisted soft preflight could add independent semantic skepticism ("does George Washington being president of France match historical reality?") that the lexical detectors can't supply. Soft sidecar output would surface as `· soft: false_premise_*` on the audit-line, distinct from the hard `· false premise` tail, giving operators a stronger warning when both signals fire. **Implication for default policy.** Keep `metacognition_block_on_contradiction=False` as the default. Schrödinger's cat (alive+dead) would have been rejected unnecessarily under a hard-block, and that's a real-world question with a legitimate answer. The label-only default is correct; operators wanting strictness opt in via `--block-on-contradiction`. Bench artifact: `bench/qa_results/2026-05-04T02-18-42Z.{jsonl,md}`. ### Addendum 5 — #000046 Phase 3 verifier numeric-gate regression check (2026-05-11) `arborist/qa/verify.py` gained a paraphrase numeric-agreement gate (`_numeric_signature` + a check in `_check_each_with_paraphrase`): a span that token-covers the source ≥ `paraphrase_coverage` but asserts a digit-number the source lacks (modulo thousands-comma) is no longer paraphrase-grounded — it goes to `unverified`. Closes #000046 by lifting the `falsification-hard-v1.jsonl` rate 4/12 → 6/12 (`5f-fal-hard-004` 50-vs-100 and `-007` 300-vs-300,000 → UNGROUNDED). Before/after `make bench-qa` (n=3 × 75 questions × 3 modes = 675 cells): | mode | STRICT-rate before → after | Δ | |------|----------------------------|---| | quote | 0.53 → 0.50 | −3pp | | claim_lattice_pointer | 0.23 → 0.25 | +2pp | | claim_lattice | 0.44 → 0.45 | +1pp | All within the 5-pp noise floor. Per-row diff (675 common cells, 74 changed audit_mode): the *only* clearly gate-attributable QA shift was the fictional "our cold fusion breakthrough" year-claim demoting STRICT → HYBRID across all 3 samples — a **correct** demotion (the year isn't grounded). Every other transition was `quote→quote` / `claim_lattice→claim_lattice` LLM re-answer variance — the answer text changed on the re-run, not the verifier (the gate touches only the paraphrase fallback, never the verbatim/span/entity/claim-lattice paths). No regression on legitimate answers; the gate ships. Bench artifacts: `bench/qa_results/2026-05-11T13-42-38Z.{jsonl,md}` (before) · `bench/qa_results/2026-05-11T14-19-51Z.{jsonl,md}` (after). Full per-ticket detail: `docs/tickets/ticket-000046-harder-5sf-fixture-tier.md` §5 Phase 3. ### Addendum 6 — #000048 step 2.1 verifier entity-gate regression check (2026-05-11) `arborist/qa/verify.py` gained an entity salient-token-disagreement gate (`_entity_salient_disagrees` + `_is_single_sentence`): in `verify_quotes`' entity branch (proximity policy), the *weakest* grounding — `not cluster AND len(verified) <= 1 AND single sentence` — is declined (→ UNGROUNDED) when the answer asserts a > 4-char capitalized content token (stopword-filtered) or a digit-number the source lacks. Catches "Insulin was discovered by Alexander Fleming" against "Penicillin was discovered by Alexander Fleming" (shared "Alexander Fleming" matched; swapped subject "Insulin" trips the gate). Lifts the `falsification-hard-v1.jsonl` rate 6/12 → 10/12 (the 4 HYBRID_ENTITY over-grounds: Insulin / Berlin / 1889 / Pacific). Before/after `make bench-qa` (n=3 × 75 questions × 3 modes = 675 cells): | mode | STRICT-rate before → after | Δ | |------|----------------------------|---| | quote | 0.50 → 0.54 | +4pp | | claim_lattice_pointer | 0.25 → 0.22 | −3pp | | claim_lattice | 0.45 → 0.43 | −2pp | All within the 5-pp noise floor. Per-row diff (675 common cells, 30 quote-mode rows changed audit_mode): **0 quote-mode rows demoted to UNGROUNDED from the entity path** — the gate didn't fire on a single legitimate QA answer in the whole bench. Every quote-mode transition was LLM re-answer variance (`verifier quote→quote` with the verdict flipping = a different answer text); the pointer/lattice deltas are noise too (the gate is in `verify_quotes` / quote mode, not the claim-lattice verifier). The gate is provably narrow on real traffic; it ships. Bench artifacts: `bench/qa_results/2026-05-11T14-19-51Z.{jsonl,md}` (before — HEAD's verify.py) · `bench/qa_results/2026-05-11T17-12-41Z.{jsonl,md}` (after). Full per-ticket detail: `docs/tickets/ticket-000048-verifier-upgrade-recombination-segmentation.md` §5 step 2.1. ### Addendum 7 — #000048 step 2.4 parse_pointer_claims clause-segmentation regression check (2026-05-11) `arborist/qa/parse_claims.py` gained a clause segmenter: a line that crams several well-pointered claims onto one row is split into one claim per clause (split on `;`, sentence boundaries, spaced dashes, ` and `/` or `/` because `/` although `/` since `/` while `, inline `(N)` enumeration markers, commas — with `(?![^\[]*\])` so a comma inside a `[E1, E2]` bracket never splits it). `_segment_line` keeps the split *only if every resulting segment is well-pointered* (a legit single claim like "The cat is black and white [E1]." or "The cast: A, B, C [E1]." is never broken — splitting would create pointer-less fragments → guard rejects; a leading colon-terminated header — "Two facts:", "Key points:" — is dropped). Plus a wrapped-bullet join: a continuation line (leading whitespace then a lowercase letter, no bullet) folds into the previous claim. Closes the 8 mis-segments in `formulate-hard-v1.jsonl` → rate 4/12 → 12/12 (that pack now at ceiling). Before/after `make bench-qa` (n=3 × 75 questions × 3 modes = 675 cells; `parse_pointer_claims` feeds the 450 `claim_lattice_pointer` + `claim_lattice` cells): | mode | STRICT-rate before → after | Δ | |------|----------------------------|---| | quote | 0.54 → 0.55 | +1pp | | claim_lattice_pointer | 0.22 → 0.22 | 0pp | | claim_lattice | 0.43 → 0.45 | +2pp | All within the 5-pp noise floor. Per-row diff (675 common cells): the segmenter changed the parsed-claim count *on the same answer text* for **7 of the 450 lattice cells** (0 in `claim_lattice`, 7 in `claim_lattice_pointer`); of those, 2 caused an `audit_mode` change — both **correct**: (a) a wrap-join recovered an answer's intended structure (4 claims, 2 of which were pointer-less wrap-fragments → HYBRID) into 2 well-pointered claims → STRICT; (b) a crammed-one-line blob (1 monolithic claim, all pointers → STRICT) split into 8 claims, some of which don't individually verify → HYBRID — the honest verdict (false-positive STRICT was the corruption). Every other lattice/quote delta is LLM re-answer variance (`answer_chars` changed, often drastically). No regression — the segmenter's only visible effects on real traffic are honest improvements. Bench artifacts: `bench/qa_results/2026-05-11T17-12-41Z.{jsonl,md}` (before — HEAD's parse_claims.py) · `bench/qa_results/2026-05-11T20-26-37Z.{jsonl,md}` (after). Full per-ticket detail: `docs/tickets/ticket-000048-verifier-upgrade-recombination-segmentation.md` §5 step 2.4. ### Addendum 8 — #000057 control sweep: retrieval × model × framing × reasoning (2026-05-19/20) **The question.** Hermes-3-8B confidently states the *present-day* office-holder ("the president of France is Emmanuel Macron") against a ~2010-vintage corpus (Sarkozy). Is that an 8B-model weakness, an unfair question framing, or does retrieval fix it regardless? The sweep crosses **{model} × {framing} × {retrieval on/off}** over a 386-item office-holder fixture (`bench/qa_questions_stale_map.json`), each item with a fixed corpus-vintage gold article. - **models:** `hermes` (Hermes-3-8B), `qwen-nothink` (Qwen3.6-27B, reasoning off), `qwen-think` (same weights, reasoning on). Qwen runs on llama.cpp; Hermes on vLLM. - **framings:** `plain` ("who is the president of France?"), `source_relative` ("According to the reference knowledge base, …"), `as_of_corpus` ("As of 2010, who was …"). - **arms:** `solo` (model alone — measures parametric prior) vs `arborist` (model + retrieval over the 2010 corpus, `answer_mode= claim_lattice`). **Judge methodology — the code judge.** The original sweep used the Opus headless judge (`bench/judge.py`); it burned our Anthropic quota and **79.5 % of its verdicts came back JUDGE_ERROR** (rate-limited). Replaced with a deterministic code judge (`bench/judge_code.py`): abstention regex → short-answer entity-grounding fast path → NLI contradiction (θ=0.85) → lexical verifier (quote/span/entity/ paraphrase) → WRONG-vs-FABRICATED tie-break on subject-in-gold. **No LLM, no quota, fully replayable.** Calibrated against the records Opus *did* grade: agreement rose CG 13 %→47 %, WRONG 56 %→89 %, ABSTAINED 80 %→95 % across four targeted fixes (raise NLI contradiction threshold off the manifest's 0.5; expand abstention patterns for the "I do not have … information/access" family; demote FABRICATED→WRONG when gold mentions the question subject; short-answer entity-grounding for terse-name answers the verifier's prose-shape extractor whiffs on). A claim-lattice JSON-unwrap lets the same judge grade the Arborist arm's `{"claims":[…]}` envelopes as plain prose. Verdicts are a deterministic *proxy* for grounding (lexical + NLI), not Opus-grade reading — JUDGE_ERROR residue (HYBRID-via-entity with no NLI corroboration) is the natural input to a later LLM-batch pass. **Consolidated scorecard — CG% (correct-grounded rate), all arms on the identical calibrated judge.** `arborist/hermes` is n=40 (phase 1's smaller arborist arm); every other cell is n=386. | arm / model | plain | source_relative | as_of_corpus | |-------------|-------|-----------------|--------------| | solo / hermes | 9 % | 5 % | 18 % | | solo / qwen-nothink | 7 % | 0 % | **50 %** | | solo / qwen-think | 6 % | 5 % | 44 % | | arborist / hermes (n=40) | 60 % | 62 % | 25 % | | **arborist / qwen-nothink** | **82 %** | **65 %** | 50 % | Grounding-fidelity detail (abstain% / wrong-assert% = W+F over n) for the cells the findings turn on: | cell | CG% | abstain% | wrong-assert% | |------|-----|----------|---------------| | solo qwen-nothink / source_relative | 0 % | **97 %** | 1 % | | solo qwen-think / source_relative | 5 % | 61 % | 32 % | | solo qwen-nothink / plain | 7 % | 9 % | 79 % | | solo qwen-think / plain | 6 % | 43 % | 48 % | | arborist qwen-nothink / plain | 82 % | 0 % | 11 % | **Findings.** 1. **Retrieval dominates every other lever.** arborist/qwen-nothink/ plain = 82 % CG vs 7 % solo; arborist/hermes/plain = 60 % vs 9 % solo. No solo configuration — no model size, no framing — approaches the retrieval arms. The production answer is retrieval. 2. **Reasoning does not improve raw correctness.** solo qwen-think vs qwen-nothink: as_of_corpus **44 % vs 50 %** (nothink wins), plain 6 %/7 %, source_relative 5 %/0 %. Chain-of-thought trades correctness for caution; it does not recall corpus-vintage facts better. 3. **Reasoning's real cost is broken honest-abstention.** The cleanest grounding signal in the sweep is solo qwen-nothink/source_relative: **97 % abstain** — told to use a reference it lacks, it declines. Reasoning *breaks* this: qwen-think/source_relative abstains only 61 % and reasons itself into wrong parametric answers (wrong-assert 1 %→32 %). On plain framing reasoning helps caution (wrong-assert 79 %→48 %); on the source-grounding framing it hurts. **Production recommendation:** arborist + qwen-nothink, plain framing — 82 % CG, ~0 % abstain, 11 % wrong-assert. Reasoning OFF: it degrades the honest-abstention property and buys nothing once retrieval supplies the source. **Held / in flight:** arborist + qwen-think (the last cell) is running at write time — tests whether reasoning hurts the retrieval arm the way it hurt solo/source_relative. Will append the result. **Artifacts.** Solo + arborist-hermes: `control_sweep_2026-05-19T21-52-56Z*` (phase 1). qwen-think solo: `…23-22-21Z*` (phase 2). arborist+qwen-nothink: `…2026-05-20T00-07-38Z*` (phase 3). All re-graded under the final judge (`*_code_judge_final.{md,jsonl}`). Judge + sweep tooling: `bench/judge_code.py`, `bench/score_with_code_judge.py`, `bench/analyze_judge_disagreement.py`, `bench/control_sweep.py` (`--judge {code,opus}`, `--skip-solo`). Engine-agnostic JSON-schema enforcement for the Arborist arm: `arborist.qa.verify.claim_lattice_structured_output_extras`.