arborist/bench/qa_questions_smoke.txt
russell@unturf.com 3b9122395c
speed: pytest-xdist, bench smoke, concurrency default; UTF surrogate fix
Bench-max sprint 1a + sprint 3 + speed audit. Five wins, none of
them traded calibration.

UTF-16 surrogate fix (sprint 1a)
================================
Hermes occasionally emits text with lone UTF-16 surrogates. Bare
.encode('utf-8') raises UnicodeEncodeError on those, which aborted
the run with no Merkle root. Two errors per lattice mode in the
2026-05-02 bench were this exact path on the 'tell me about the
roman empire' question.

Fix: errors='surrogatepass' on the four sha256 helpers that hash
model-derived text, plus the two audit-chain encode sites in
store.py for defense-in-depth (audit body could carry user text in
some flows). The hash stays deterministic because WTF-8 bytes are
reversible & unique per input.

Touched:
  aborist/qa/dag.py:_sha256_hex            (the loud one)
  aborist/qa/keys.py:_sha256
  aborist/qa/evidence.py:_sha256_hex
  aborist/store.py: chain_audit_events + append_audit

Predicted Δ on next bench: +1pp on lattice modes (the 2 errors
become valid runs).

Smoke fixture (sprint 3)
========================
bench/qa_questions_smoke.txt — 5 questions, all anchor classes,
each currently failing pointer mode 100% while JSON aces 100% per
the 2026-05-02 bench. Wired as 'make bench-qa-smoke', --n 1
--concurrency 4, ~30-90s wall-clock depending on vLLM warmth. The
inner loop for prompt iteration; the full 71-question sweep stays
the scoreboard.

Smoke verified: pointer=0/5 STRICT, JSON=5/5, quote=2/5. Confirms
the gap pattern from the journal.

Concurrency default
===================
Makefile bench-qa now defaults to BENCH_QA_CONCURRENCY=4 (was
sequential). Override via BENCH_QA_CONCURRENCY=N. Combined with
the --concurrency landing in 0870af6, full sweep drops from ~107
min projected to ~51 min actual.

pytest-xdist (test-speed)
=========================
Added pytest-xdist>=3.5 to dev extras. 'make test' now uses
-n auto (= one worker per logical CPU). Measured: 36s → 10s on
the 641-test suite. 3.6× speedup, no test changes required.

Bench-max scoreboard (predicted lift from this commit alone):
  +1pp lattice modes (UTF fix)
  +cycle-time enabler (smoke fixture, xdist)
  no calibration cost — none of the verifier checks moved.
2026-05-02 09:29:40 -04:00

12 lines
653 B
Text

# Smoke fixture — 5 questions covering all anchor classes (date, relation,
# entity-list, count/place, why-cause), each chosen because the 2026-05-02
# bench shows pointer mode failing 100% (HH) while JSON mode aces 100% (SS).
# These are the load-bearing test cases for prompt-iteration loops aimed at
# closing the pointer-vs-JSON gap. Run with --concurrency 4 --n 1 → ~30s
# wall-clock. The full bench (71 questions) stays the scoreboard; this is
# the inner loop.
when did the soviet union dissolve?
where is mount kilimanjaro located?
who painted the mona lisa?
who were the original seven mercury astronauts?
why did the dinosaurs go extinct?