Bench-emergent stress test ran another 100 cycles under the
post-#000008/9/10/11 substrate. Total accumulated: 300 cycles.
Verdict distribution shift on last 100 vs 134-cycle baseline:
STRICT 5% (7/134) → 0% (0/100) -5pp
HYBRID 22% (29/134) → 16% (16/100) -6pp
UNGROUNDED 73% (98/134) → 84% (84/100) +11pp
Zero false-positive STRICTs across 100 random-word triplets.
The 5pp drop in STRICT-rate isn't a regression — it's the
verifier ladder + new preflight contracts doing their job.
Random-word triplets are genuinely ungrounded for the most
part; the prior 5% STRICT rate included false-positives that
the post-hardening verifier now catches.
Violation profile (last 100 cycles, claim_lattice JSON):
CITATION_MISMATCH: 86 dominant gate
TOO_MANY_EVIDENCE_IDS: 24
SUBJECT_TOKENS_ABSENT: 12 Rule 9 firing on parroting
DEFLECTION_DETECTED: 12
TITLE_MISMATCH: 10
...
metaphor_deflection fires 6/100 — still rare. Item 3
(calibration) is now closer to sample-size threshold (~30
signals across 300 cycles; needs ~50-100 to calibrate).
No new tuning candidates surface. Original three remain at
their resolution states.