Full QA-quality sweep, 75 questions × 3 modes × n=3 = 675 cells
against Hermes-3 8B at concurrency=4. No mode regressed past
the 5pp signal floor:
quote 53.8% STRICT vs 54% baseline = -0pp (stable)
claim_lattice_pointer 23.1% STRICT vs 20% baseline = +3pp (within floor)
claim_lattice (JSON) 46.7% STRICT vs 42% baseline = +5pp (at floor — marginal positive)
JSON-mode +5pp is right at the noise threshold per
docs/bench-maxing.md — could be the 2026-05-10 substrate work
(92/92 warrant chains + 18 textbook ingests + cascade tuning +
Phase 3 wiring) translating to retrieval-quality lift, OR
sample variance. Follow-up bench in 1-2 days disambiguates.
Format-collapse 0/225 in every mode. Latency stable
(10.6-11.1s median). Lattice-mode directive coverage 100%.
Phase 3 per-claim warrant-chain tail still dormant on this
fixture (rescue-only by design; no warrant-shape question
retrieves a chain-backed chunk while failing lexical anchor —
parallel shift's same finding). Source-level warrants tail
fires on claim-pack-targeted retrievals as documented in
phase3-live-validation-2026-05-10.md.
Substrate + Phase 3 sprint shipped clean.