scripts/bench_emergent.py + make bench-emergent + design doc.
Random word triangulation surfaces failure modes the curated
bench/qa_questions.txt doesn't reach.
Loop:
/usr/share/dict/words → random.sample(3) →
Hermes @ temp=0.8 weaves a creative question →
aborist student answers via query() →
append journey to bench/emergent_log.jsonl
(teacher review = separate manual step, fox brings entries to
Opus & gets judgment to append)
Word filter: ^[a-z]{5,12}$ after lowercasing. Skips short words
(too vague) + very long words (Hermes can't weave them).
Cadence: NOT every commit. ~20s per cycle (Hermes generator +
aborist student); N=10 ≈ 4 min, N=50 ≈ 17 min. Most cycles land
UNGROUNDED-by-corpus-design (random triplets rarely overlap with
2010-11 Wikipedia coverage); the interesting cases are STRICT/
HYBRID surprises and the verifier-disagreement cases the teacher
catches.
Teacher review is intentionally out of the bench script:
- separation of concerns: generation is automated, judgment is
contextual & needs the corpus-knowledge frame ("is this a
2010 Wikipedia gap or a substrate failure?")
- future flexibility: today the teacher is Claude Opus 4.7
in this conversation; tomorrow GPT-5 or a review committee.
Swapping teachers is a workflow change, not a code change.
Teacher output schema (appended to the same JSONL line):
teacher.match bool
teacher.audit_agreement agree|disagree|unsure
teacher.novelty_class known_truth_grounding | emergent_synthesis
| novel_claim | no_signal
teacher.score_0_5 0..5
teacher.bench_max_signal retrieval | warrant | prompt | nil
teacher.reasoning one sentence
teacher.reviewed_by model id
teacher.reviewed_ts unix ts
Smoke verified (N=2, seed=42): 41s wall-clock, both UNGROUNDED
(expected — random triplets rarely overlap 2010 Wikipedia).
Append-only log seeded with the smoke entries.
Future flag (not yet wired): --generator-endpoint &
--student-endpoint to swap LLM upstreams per role.
Full design + teacher protocol: docs/bench-emergent-design.md.
7.8 KiB
Bench-emergent: random-word triangulation stress test
Status: landed — scripts/bench_emergent.py, 2026-05-02
Cadence: blue-moon — runs on demand, NOT every commit
Audience: anyone hunting for failure modes the curated
bench/qa_questions.txt doesn't surface
Why
bench/qa_questions.txt is 71 hand-curated questions that pin
known-good answers + known-failure shapes. It tells us whether the
substrate handles the failure shapes we already named. It can't
tell us what we haven't named.
The curated bench is a scoreboard. This is a sonar — random pings into the combinatoric space of "what happens when somebody asks something we never thought to ask." The 2010-11 Wikipedia corpus is fixed; the question space is infinite; emergent surfacing is how we find what we don't know.
Loop shape
/usr/share/dict/words
│ random.sample(3)
▼
╔════════════════════╗
║ generator: Hermes ║ temp=0.8
║ "weave 3 words → ║ creative paragraph
║ question paragraph║ with question framing
╚════════════════════╝
│
▼
╔════════════════════╗
║ student: aborist ║ full retrieval pipeline,
║ query() against ║ claim_lattice mode,
║ live shards ║ burn=True (always fresh)
╚════════════════════╝
│
▼
┌────────────────────┐
│ append to │ one JSONL line per cycle
│ bench/emergent_ │ fields: words, question,
│ log.jsonl │ answer, audit_mode, sources,
│ │ timings, teacher: null
└────────────────────┘
│
│ (teacher review — separate, manual)
▼
╔════════════════════╗
║ teacher: Opus ║ read pending entries,
║ (Claude 4.7, this ║ judge match / novelty /
║ conversation) ║ bench-max signal,
║ ║ append `teacher: {...}`
╚════════════════════╝
Why "blue moon"
- Wall-clock cost: ~20s per cycle (Hermes generator @ 5-10s + aborist student @ 5-15s). N=10 ≈ 4 min, N=50 ≈ 17 min.
- Most cycles are UNGROUNDED-by-corpus-design (random word triplets rarely overlap with 2010 Wikipedia coverage). The interesting cases are the rare HYBRID/STRICT verdicts on triplets that surprise us, plus the verifier-disagreement cases the teacher catches.
- Curated bench (
make bench-qa{,-quick,-smoke}) is the every- iteration signal. This is the once-a-month sonar.
Word filter
scripts/bench_emergent.py uses ^[a-z]{5,12}$ after lowercasing
to admit a word. Skips:
- length ≤ 4 (short words like "the", "and" produce trivially vague questions)
- length > 12 (rare scientific terms / loanwords; Hermes struggles to weave them)
- non-alpha (apostrophes, hyphens — the unix words file mixes these in)
- ALLCAPS (filtered after the lowercasing step admits "Goldman" → "goldman" as a generic noun, which is fine; Hermes treats it as a name regardless)
Override the filter or word source via:
python scripts/bench_emergent.py --words-path /custom/words.txt
Teacher review protocol
The bench script does NOT auto-invoke a teacher. Two reasons:
-
Separation of concerns: generation is automated, judgment is contextual. The teacher (Opus, currently) needs the loop's raw output PLUS access to the corpus-knowledge frame ("is this a 2010 Wikipedia gap or a substrate failure?"). That frame lives in this repo's docs, not in a per-call API spec.
-
Future flexibility: today the teacher is Claude Opus 4.7 in this conversation. Tomorrow it might be GPT-5, Claude 5, or a Mixture-of-Experts review committee. Keeping teacher review out of the bench script means swapping teachers is a workflow change, not a code change.
How fox brings entries to the teacher
make bench-emergent-pending # print every entry with teacher==null
# … pipe into a Claude session, paste, ask for judgment
The teacher's output is a JSON dict to append to the same line:
{
"teacher": {
"match": true|false,
"audit_agreement": "agree"|"disagree"|"unsure",
"novelty_class": "known_truth_grounding"
| "emergent_synthesis"
| "novel_claim"
| "no_signal",
"score_0_5": 0..5,
"reasoning": "<one sentence>",
"bench_max_signal": "<retrieval | warrant | prompt | nil>",
"reviewed_by": "claude-opus-4-7[1m]",
"reviewed_ts": <unix>
}
}
The bench_max_signal field is the actionable bit: which subsystem
should be tuned to address this kind of failure? retrieval /
warrant / prompt / nil (no action — corpus genuinely lacks
the answer).
Fields the teacher considers
- match: did the answer address the question? Not "is the answer correct" — the corpus may legitimately not have the answer. The check is "is this answer about the same topic as the question?"
- audit_agreement: does the substrate's audit_mode (STRICT/ HYBRID/UNGROUNDED) or display rung (POINTER-LINKED → ANCHOR-WARRANTED → EVIDENCE-WARRANTED) match what the teacher thinks the answer's grounding deserves?
- novelty_class: how does the answer relate to 2010 Wikipedia
knowledge?
known_truth_grounding: cites well-known facts present verbatim in the corpus. The expected case for narrow factoids ("capital of France", "founder of Microsoft").emergent_synthesis: the answer connects facts in a way no single source contains; the substrate did real work composing across sources. The interesting case.novel_claim: the answer goes beyond what the 2010 corpus can ground (post-2010 science, personal opinion, made-up detail). The dangerous case if STRICT/EVIDENCE-WARRANTED.no_signal: UNGROUNDED, or the question was incoherent. Most random-word triplets land here — that's fine.
Future: multiple upstreams
When the substrate supports multiple inference endpoints, the generator and student can use different upstreams to surface upstream-specific failure modes. The CLI flags for this aren't implemented yet but the shape is clear:
python scripts/bench_emergent.py \
--generator-endpoint https://hermes.ai.unturf.com/v1 \
--student-endpoint https://other-llm.example/v1
Today both default to Hermes. The student is always aborist's
query() against the configured shard set; only the LLM behind
that pipeline is plugged.
How to run
make bench-emergent # 10 cycles, default seed (random)
make bench-emergent EMERGENT_N=50 EMERGENT_SEED=42 # bigger sample, reproducible
make bench-emergent-pending # print pending teacher review
# Direct:
python scripts/bench_emergent.py --n 50 --seed 42
python scripts/bench_emergent.py --print-pending
Append-only log
bench/emergent_log.jsonl accumulates every cycle ever run,
across branches, across days. Never rewritten. Each line is one
self-describing JSON record (timestamp, words, question, answer,
audit, sources, timings, teacher). The log is the substrate's
own version of /var/log/syslog — every interaction with the
emergent harness leaves a trace, and a future you can grep for
"every entry where the answer mentioned X" or "every UNGROUNDED
result on triplets containing Y."
A line whose teacher field is null is awaiting review. A line
with a populated teacher field is closed. There is no DELETE
path; corrections add new lines that supersede old ones.