Bench-emergent stress test ran another 100 cycles under the post-#000008/9/10/11 substrate. Total accumulated: 300 cycles. Verdict distribution shift on last 100 vs 134-cycle baseline: STRICT 5% (7/134) → 0% (0/100) -5pp HYBRID 22% (29/134) → 16% (16/100) -6pp UNGROUNDED 73% (98/134) → 84% (84/100) +11pp Zero false-positive STRICTs across 100 random-word triplets. The 5pp drop in STRICT-rate isn't a regression — it's the verifier ladder + new preflight contracts doing their job. Random-word triplets are genuinely ungrounded for the most part; the prior 5% STRICT rate included false-positives that the post-hardening verifier now catches. Violation profile (last 100 cycles, claim_lattice JSON): CITATION_MISMATCH: 86 dominant gate TOO_MANY_EVIDENCE_IDS: 24 SUBJECT_TOKENS_ABSENT: 12 Rule 9 firing on parroting DEFLECTION_DETECTED: 12 TITLE_MISMATCH: 10 ... metaphor_deflection fires 6/100 — still rare. Item 3 (calibration) is now closer to sample-size threshold (~30 signals across 300 cycles; needs ~50-100 to calibrate). No new tuning candidates surface. Original three remain at their resolution states. |
||
|---|---|---|
| .. | ||
| emergent_log.jsonl | ||
| qa_questions.txt | ||
| qa_questions_metacog_subset.txt | ||
| qa_questions_quantifier_baseline.txt | ||
| qa_questions_quantifier_subset.txt | ||
| qa_questions_smoke.txt | ||
| qa_sweep.py | ||
| run.sh | ||