Single canonical entry point that ties together the four existing
bench-related docs (qa-modes-bench.md / bench-maxing.md /
bench-emergent-design.md / qa-modes-bench-2026-04-30.md) plus the
make targets, fixtures, and bench-row schema.
Sections:
1. Two harnesses, two purposes
bench/qa_sweep.py — curated regression bench
scripts/bench_emergent.py — random-word stress test
2. Four question fixtures (75 / 28 / 9 / 1 / smoke) with cell
sizes + use cases per fixture
3. Signal floor (5pp / n=3 × 9 = 27 / vLLM c=3-4 saturation)
4. Make targets cheat sheet (bench-qa, bench-emergent,
--policy KEY=VALUE A/B pattern, --resume)
5. Bench-row schema — every field a row carries
(identity / verdict / diagnostics / preflight projection /
quantifier classifier / capacity / directive compliance / time)
6. Where headlines live (qa-modes-bench.md addenda, per-ticket
§12/§13 bench sections)
7. bench-emergent log shape + #000006 rolling-amend pattern
8. How to run a focused A/B (the four-cell pattern from #000008)
9. How to interpret results (STRICT-rate, mean-ratio,
UNGROUNDED-rate, FORMAT_COLLAPSED rate, violation kind
distribution, audit-line tails)
10. Operator commands cheat sheet (--show-preflight,
--apply-quantifier-caps, --reject-broad, --soft-preflight,
Makefile shortcuts)
CLAUDE.md docs index updated to point to benchmarks.md as the
"read first" entry for bench work, and to add bench-emergent-design.md
to the index (was missing).