arborist/bench/qa_questions_quantifier_baseline.txt
russell@unturf.com 2ffed001a4
bench(#000008): harness extension — FC rate, violation kinds, raw brackets
Closes the bench-side gap surfaced in §5.2: JSONL was carrying summary
numbers only, blinding the harness to FORMAT_COLLAPSED rate and per-
violation-kind distributions. Without these, A/B/D bench measurements
on the broad-quantifier preflight guard would be guesses.

- query() result dict surfaces format_collapsed + raw_answer (lattice
  modes only) so the bench can read them directly instead of re-deriving
  from cache rows that --burn overwrites.
- Each bench row gains format_collapsed, violation_kinds (sorted unique
  list — full payloads stay off the row to keep size bounded), and
  answer_brackets (count of [E\d+] in raw_answer for lattice modes).
- _summarize aggregates per-mode FC count (only explicit True; None
  means check didn't apply), per-kind tallies (each kind once per row),
  and lattice-only bracket sum/n.
- Markdown renderer adds a `## format-collapse + violation kinds`
  section with per-mode FC rate, mean raw brackets, and one column per
  observed violation kind. Degrades gracefully when the sweep produces
  no violations.
- 5 new bench-harness tests pin the aggregation rules.

Re-baseline (2026-05-02T20-58-57Z) sharpens §5.1 analysis dramatically:
NO_EVIDENCE_POINTER fires 3/3 in pointer mode and is the dominant gate,
not TITLE_MISMATCH (1/3) as §5.1 inferred from JSONL alone. FORMAT_
COLLAPSED actually fires 1/3 — not the rare corner the first baseline
called it. Implies Option B (prompt reminder) is the load-bearing fix
for the verdict; Option A (cap reduction) only moves secondary kinds.

§5.3 sub-investigation closed on first read — SCHEMA_INVALID:1 in
pointer mode is a legitimate kind emitted by verify_claim_lattice for
empty-claim-text (verify.py:1242) and bare-name-claim (verify.py:1270),
not a JSON-mode leak.
2026-05-02 18:35:08 -04:00

6 lines
353 B
Text

# One-question bench file — Ticket #000008 baseline.
# Isolates the under-specified "all" failure shape so we can measure
# FORMAT_COLLAPSED rate, audit_mode distribution, and claim count
# under the *current* policy (cap 12) before proposing changes.
# Once Option A or D lands, re-run against this same file to compare.
winners of all major sports?