Durable record of the #000057 sweep in the bench journal. Captures:
- The question: is Hermes-8B's confident present-day-officeholder
fabrication an 8B weakness, a framing artefact, or does retrieval
fix it? Crosses {hermes, qwen-nothink, qwen-think} × {plain,
source_relative, as_of_corpus} × {solo, arborist} on a 386-item
office-holder fixture with corpus-vintage gold.
- The judge methodology: Opus headless judge burned quota (79.5%
JUDGE_ERROR), replaced with the deterministic code judge
(bench/judge_code.py), calibrated against Opus's gradeable records
(CG agreement 13->47%, WRONG 56->89%, ABSTAINED 80->95%).
- Consolidated CG% scorecard, all arms on the identical final judge.
- Three findings:
1. Retrieval dominates — arb/qwen-nothink/plain 82% vs 7% solo;
no solo config approaches the retrieval arms.
2. Reasoning does NOT improve raw correctness — qwen-think/as_of
44% vs nothink 50%.
3. Reasoning's real cost is broken honest-abstention —
qwen-nothink/source_relative abstains 97% (clean); qwen-think
only 61%, reasoning itself into wrong parametric answers.
- Production recommendation: arborist + qwen-nothink, plain framing,
reasoning OFF (82% CG, ~0% abstain, 11% wrong-assert).
- Held cell noted: arborist+qwen-think running at write time, result
to be appended.
Bench %s are point-in-time measurements (not repo-derived counts),
so no AUTOCOUNT tags — consistent with addenda 1-7. test_doc_counts
3/3.