arborist/docs
russell@unturf.com 42f614a501
docs(#000057): Addendum 8 — control sweep retrieval × model × framing × reasoning
Durable record of the #000057 sweep in the bench journal. Captures:

- The question: is Hermes-8B's confident present-day-officeholder
  fabrication an 8B weakness, a framing artefact, or does retrieval
  fix it? Crosses {hermes, qwen-nothink, qwen-think} × {plain,
  source_relative, as_of_corpus} × {solo, arborist} on a 386-item
  office-holder fixture with corpus-vintage gold.

- The judge methodology: Opus headless judge burned quota (79.5%
  JUDGE_ERROR), replaced with the deterministic code judge
  (bench/judge_code.py), calibrated against Opus's gradeable records
  (CG agreement 13->47%, WRONG 56->89%, ABSTAINED 80->95%).

- Consolidated CG% scorecard, all arms on the identical final judge.

- Three findings:
  1. Retrieval dominates — arb/qwen-nothink/plain 82% vs 7% solo;
     no solo config approaches the retrieval arms.
  2. Reasoning does NOT improve raw correctness — qwen-think/as_of
     44% vs nothink 50%.
  3. Reasoning's real cost is broken honest-abstention —
     qwen-nothink/source_relative abstains 97% (clean); qwen-think
     only 61%, reasoning itself into wrong parametric answers.

- Production recommendation: arborist + qwen-nothink, plain framing,
  reasoning OFF (82% CG, ~0% abstain, 11% wrong-assert).

- Held cell noted: arborist+qwen-think running at write time, result
  to be appended.

Bench %s are point-in-time measurements (not repo-derived counts),
so no AUTOCOUNT tags — consistent with addenda 1-7. test_doc_counts
3/3.
2026-05-20 06:52:42 -04:00
..
_source docs: pagers — rename 'not grounded' → 'ungrounded' (agreed label set: grounded / partly grounded / ungrounded) 2026-05-14 12:56:13 -04:00
diagrams modified: docs/diagrams/arborist-modules.png 2026-05-19 17:05:16 -04:00
tickets feat(#000049 §7 #28): tinygrad NLI backend + deterministic engine-agreement A/B; ONNX-immunity rationale 2026-05-19 12:34:04 -04:00
bench-maxing.md aborist/arborist 2026-05-07 09:31:49 -04:00
benchmarks.md aborist/arborist 2026-05-07 09:31:49 -04:00
calculator-test-patterns.md ticket #000036 Tier-2: dav1d Option B (conservative B1 envelope) applied in v1 2026-05-11 07:06:50 -04:00
cti-architecture.md aborist/arborist 2026-05-07 09:31:49 -04:00
lexical-first-rationale.md docs: lexical-first-rationale.md — why the cheap retrieval path is the default 2026-05-12 09:23:50 -04:00
mesh.md aborist/arborist 2026-05-07 09:31:49 -04:00
onnx-vendor-capture-immunity.md feat(#000049 §7 #28): tinygrad NLI backend + deterministic engine-agreement A/B; ONNX-immunity rationale 2026-05-19 12:34:04 -04:00
pager.style docs: arborist-one-pager + arborist-two-pager — Dav1d/fox-signoff summaries with letterhead, license, and 2 strategic appendix diagrams 2026-05-14 09:48:50 -04:00
pi-star-composition.md pi_star: land ticket #000015 (π* domain library + composition algebra) 2026-05-07 16:51:33 -04:00
qa-modes-bench.md docs(#000057): Addendum 8 — control sweep retrieval × model × framing × reasoning 2026-05-20 06:52:42 -04:00
relevance-and-veto-synthesis-for-dav1d.md docs: relevance-and-veto-synthesis-for-dav1d.md — single decision brief synthesizing #000049 + #000052 §3.1 + §3.2 for forward review 2026-05-13 15:34:21 -04:00
seven-point-program.md tests/doc_counts: regression test for numeric claims in docs/ (4x drift fix) 2026-05-10 16:15:52 -04:00
soft-hash-channel-analysis.md docs/#000018 §9.2: mark resolved — φ_PRG = HMAC-SHA-512 (#000035 closed) 2026-05-11 17:20:48 -04:00
soft-hash-channel-t3-bound.md ticket #000036: add KAT-regen tooling + close 2026-05-11 08:02:25 -04:00
spec-methodology.md docs: land ticket #000019 (spec methodology for π*, V, policy fields) 2026-05-07 16:53:28 -04:00
TICKETS.md feat(#000049 §7 #28): tinygrad NLI backend + deterministic engine-agreement A/B; ONNX-immunity rationale 2026-05-19 12:34:04 -04:00
tool-action-dag-design.md docs: add tool-action-dag-design.md research path (pre-ticket) 2026-05-07 19:47:50 -04:00
v7w-frontier-catalog.md #000013 closed: v7-W spatial-temporal substrate paper + namespace 2026-05-09 15:00:05 -04:00
v8-fork-score.md CLI: arborist v8 score → arborist substrate score 2026-05-10 09:12:34 -04:00
warrant-substrate-cookbook.md docs: bump warrant-substrate-cookbook AUTOCOUNT 20 -> 28 for #000054 tests 2026-05-13 07:01:02 -04:00
zk-frontier-bench.md #000016 parked: ZK frontier-proof bench plan + wire protocol 2026-05-09 15:05:08 -04:00
zk-wire-protocol.md #000016 parked: ZK frontier-proof bench plan + wire protocol 2026-05-09 15:05:08 -04:00