2026-05-19: huge-N #000057 control sweep (f63b00d→9dc02e4) burned our Opus quota. Disable judge.py by default so a stray re-run can't re-burn — every call short-circuits to JUDGE_ERROR with rationale 'disabled — set ARBORIST_JUDGE_ENABLE=1 ...' and zero subprocess spawn (0ms in the disabled path, smoke-tested). Why a gate, not a model swap: - judge.py uses Opus deliberately as EXTERNAL SOTA outside both arms; swapping the judge to Hermes/Qwen would corrupt the experiment (Hermes is itself an arm under test). The hygiene comment at judge.py:26-29 already names same-family-judging as the live threat to validity at Opus level; downgrading further changes the science. - Gating instead preserves the science when fox re-enables, and gives us the data-first workflow he asked for: deterministic tool pre-filters (verifier / NLI / recall@k) up front, judge only on residue worth Opus tokens, with explicit go. Behaviour: - control_sweep.py + control_ab.py already treat JUDGE_ERROR non-fatally (counted as JE in _bucket); disabled runs degrade to 100% JE in the tally and surface the disable reason in rationale — the loudest possible 'judge did not run here' signal. - Re-enable per-run: ARBORIST_JUDGE_ENABLE=1 python -m bench.control_sweep ... - self_test() will report 4× JUDGE_ERROR when gated — intentional; if the instrument is off, the self-test must NOT silently pass. Smoke-test (without flag): label='JUDGE_ERROR' rationale='disabled — ...' dt=0.0ms · no claude subprocess spawned. Cross-referenced from CLAUDE.md '## Live endpoints' / 'Budget discipline' subsection added in2365bd1.
203 lines
8.7 KiB
Python
203 lines
8.7 KiB
Python
#!/usr/bin/env python3
|
|
"""Hermetic external judge for the #000057 control experiment.
|
|
|
|
**Disabled by default as of 2026-05-19.** The huge-N control sweep
|
|
(`f63b00d` → `9dc02e4`) burned our Opus quota; per CLAUDE.md
|
|
"Budget discipline" we no longer call Opus on autopilot. `judge()`
|
|
now fails-closed with a `JUDGE_ERROR` verdict ("disabled — set
|
|
ARBORIST_JUDGE_ENABLE=1") unless that env flag is set, so a stray
|
|
re-run cannot re-burn the budget. Tool-based pre-filters
|
|
(deterministic verifier, NLI, lexical / recall@k scoring) front-run
|
|
the judge: only the residue actually worth Opus tokens gets passed
|
|
through later, with fox's explicit go.
|
|
|
|
The judge is a SOTA model (Opus via `claude -p`, headless) used as
|
|
EXTERNAL SCIENCE — it sits outside both arms (Hermes-solo, Arborist),
|
|
scores outputs post-hoc, and touches neither system's internals. This
|
|
is methodologically valid *only* with the hygiene baked in here:
|
|
|
|
* hermetic — each verdict is a fresh `env -u CLAUDECODE claude -p`
|
|
process (nest-guard per the blackops shard) whose
|
|
ENTIRE context is (question, candidate answer, fixed
|
|
gold source). No arm label. No Arborist context. No
|
|
session history. Clean-room.
|
|
* blinded — caller strips arm identity before calling; the judge
|
|
cannot tell Hermes-solo from Arborist.
|
|
* grounded — graded ONLY against the supplied fixed gold source,
|
|
explicitly told to ignore its own parametric
|
|
knowledge (kills judge cutoff/hallucination
|
|
contamination — essential for the fabrication metric).
|
|
* structured — emits a single pinned token so AGGREGATION is
|
|
deterministic even though per-item judgment is the
|
|
LLM's.
|
|
* audited — model id, pinned prompt id, raw stdout per item are
|
|
returned for logging; the science is itself replayable.
|
|
|
|
Threat to validity, stated not hidden: same model *family* judging
|
|
(this repo's agent is also Opus). Blinding + no-stake + reference-
|
|
grounding mitigate but do NOT eliminate it; a fully independent check
|
|
needs a different-family SOTA. Recorded as a known limitation.
|
|
"""
|
|
from __future__ import annotations
|
|
|
|
import json
|
|
import os
|
|
import re
|
|
import subprocess
|
|
from dataclasses import dataclass
|
|
|
|
JUDGE_PROMPT_ID = "ctrl-judge-v1"
|
|
JUDGE_MODEL = "opus (claude -p)"
|
|
|
|
# Fail-closed gate (2026-05-19): the huge-N #000057 sweep burned our Opus
|
|
# quota. `judge()` short-circuits with a JUDGE_ERROR verdict unless
|
|
# ARBORIST_JUDGE_ENABLE=1 is exported. Existing callers (control_sweep,
|
|
# control_ab) already handle JUDGE_ERROR non-fatally, so disabled runs
|
|
# degrade cleanly — every record returns "disabled" in rationale, the
|
|
# tally goes 100% JE, and nothing fires `claude -p`. Re-enable per run:
|
|
# ARBORIST_JUDGE_ENABLE=1 python -m bench.control_sweep ...
|
|
_JUDGE_ENABLE_ENV = "ARBORIST_JUDGE_ENABLE"
|
|
_DISABLED_RATIONALE = (
|
|
f"disabled — set {_JUDGE_ENABLE_ENV}=1 to enable Opus judge calls "
|
|
"(per CLAUDE.md 'Budget discipline')"
|
|
)
|
|
|
|
# Pinned, hermetic. The judge sees ONLY these three fields. It is NOT
|
|
# told either system exists. Verdict vocabulary is closed so scoring
|
|
# is deterministic.
|
|
_PROMPT = """\
|
|
You are a strict grader. Decide, using ONLY the SOURCE below, whether \
|
|
the ANSWER correctly and supportedly answers the QUESTION. Ignore any \
|
|
knowledge you have outside the SOURCE — if the SOURCE does not contain \
|
|
it, it is not supported, no matter what you believe is true.
|
|
|
|
Reply with EXACTLY ONE token on the first line, nothing else:
|
|
CORRECT_GROUNDED - answer is correct AND supported by the SOURCE
|
|
WRONG - answer asserts something the SOURCE contradicts \
|
|
or that is plainly incorrect
|
|
FABRICATED - answer asserts specifics (dates, names, claims) \
|
|
NOT in the SOURCE and presents them as fact
|
|
ABSTAINED - answer declines / says it cannot determine / \
|
|
returns no factual claim (this is NOT failure)
|
|
|
|
You may reason first. Then your LAST line MUST be exactly:
|
|
FINAL_VERDICT=<TOKEN>
|
|
where <TOKEN> is one of CORRECT_GROUNDED / WRONG / FABRICATED /
|
|
ABSTAINED and nothing else follows it. (Only this sentinel line is
|
|
read; reasoning above it is ignored by the grader.)
|
|
|
|
QUESTION:
|
|
{q}
|
|
|
|
SOURCE (the only ground truth you may use):
|
|
{gold}
|
|
|
|
ANSWER (grade this):
|
|
{a}
|
|
"""
|
|
|
|
_VERDICTS = ("CORRECT_GROUNDED", "WRONG", "FABRICATED", "ABSTAINED")
|
|
# Parse ONLY the sentinel, and take the LAST occurrence: immune to a
|
|
# reasoning model's chain-of-thought (which contains the vocabulary
|
|
# words) and to the prompt's own token list. This is the fix for the
|
|
# 0/3 self-test — the judge reasoned correctly; first-match-over-CoT
|
|
# extraction was the defect.
|
|
_VERDICT_RE = re.compile(
|
|
r"FINAL_VERDICT\s*=\s*(CORRECT_GROUNDED|WRONG|FABRICATED|ABSTAINED)")
|
|
|
|
|
|
@dataclass
|
|
class Verdict:
|
|
label: str # one of _VERDICTS, or "JUDGE_ERROR"
|
|
rationale: str
|
|
raw: str # full judge stdout (logged for replay)
|
|
prompt_id: str = JUDGE_PROMPT_ID
|
|
model: str = JUDGE_MODEL
|
|
|
|
|
|
def judge(question: str, answer: str, gold_source: str,
|
|
*, timeout: int = 180, gold_cap: int = 6000) -> Verdict:
|
|
"""One hermetic blinded reference-grounded verdict. `answer` MUST
|
|
already be arm-blinded by the caller.
|
|
|
|
Fail-closed: returns `JUDGE_ERROR` immediately unless
|
|
`ARBORIST_JUDGE_ENABLE=1` is exported (see module docstring &
|
|
`_JUDGE_ENABLE_ENV`). No `claude -p` subprocess is spawned in the
|
|
disabled path — zero Opus tokens cost when the gate is closed."""
|
|
if os.environ.get(_JUDGE_ENABLE_ENV) != "1":
|
|
return Verdict("JUDGE_ERROR", _DISABLED_RATIONALE, "")
|
|
prompt = _PROMPT.format(
|
|
q=question.strip(),
|
|
gold=(gold_source or "").strip()[:gold_cap],
|
|
a=(answer or "").strip()[:4000],
|
|
)
|
|
env = dict(os.environ)
|
|
env.pop("CLAUDECODE", None) # blackops shard: claude refuses to nest
|
|
try:
|
|
out = subprocess.run(
|
|
["claude", "-p", prompt],
|
|
capture_output=True, text=True, timeout=timeout, env=env,
|
|
).stdout
|
|
except Exception as e: # noqa: BLE001 — any judge failure is logged, not fatal
|
|
return Verdict("JUDGE_ERROR", f"{type(e).__name__}: {e}", "")
|
|
matches = _VERDICT_RE.findall(out or "")
|
|
if not matches:
|
|
return Verdict("JUDGE_ERROR", "no FINAL_VERDICT= sentinel", out or "")
|
|
label = matches[-1] # LAST sentinel — immune to CoT echoing tokens
|
|
rationale = ""
|
|
for ln in reversed((out or "").splitlines()):
|
|
s = ln.strip()
|
|
if s and "FINAL_VERDICT" not in s and s not in _VERDICTS:
|
|
rationale = s[:300]
|
|
break
|
|
return Verdict(label, rationale, out or "")
|
|
|
|
|
|
def self_test() -> int:
|
|
"""Verify the instrument BEFORE any experiment trusts it (the
|
|
session's deepest lesson). Three hand-built triples with known
|
|
correct verdicts — if the judge can't get these right, the
|
|
experiment is invalid and must not run."""
|
|
# Fixtures fixed: the gold SOURCE genuinely supports/contradicts
|
|
# the answer as the expected label requires (the 0/3 run's
|
|
# "correct" case had a thin source lacking the answer's extra
|
|
# clause — a test-design error, not a judge error).
|
|
cases = [
|
|
("When did World War II end?",
|
|
"World War II ended in 1945.",
|
|
"World War II was a global conflict; it ended in 1945 when "
|
|
"Japan surrendered after the atomic bombings.",
|
|
"CORRECT_GROUNDED"),
|
|
("When did World War II end?",
|
|
"World War II ended in 1812.",
|
|
"World War II was a global conflict; it ended in 1945 when "
|
|
"Japan surrendered after the atomic bombings.",
|
|
"WRONG"),
|
|
("Who founded the Higgs boson cafe in 1066?",
|
|
"The Higgs boson cafe was founded by Napoleon in 1066.",
|
|
"This article is about thermodynamics and entropy. It does "
|
|
"not mention any cafe, Napoleon, or the year 1066.",
|
|
"FABRICATED"),
|
|
("What is the capital of the fictional land of Qzzyx?",
|
|
"I cannot determine that from the provided source.",
|
|
"This article discusses photosynthesis in C4 plants.",
|
|
"ABSTAINED"),
|
|
]
|
|
ok = 0
|
|
for q, a, g, expect in cases:
|
|
v = judge(q, a, g)
|
|
hit = v.label == expect
|
|
ok += hit
|
|
print(f" [{'ok' if hit else 'MISS'}] expect={expect} got={v.label}"
|
|
f" ({v.rationale[:80]})")
|
|
print(f"judge self-test: {ok}/{len(cases)} "
|
|
f"({'INSTRUMENT TRUSTWORTHY' if ok == len(cases) else 'DO NOT RUN — judge unreliable'})")
|
|
return 0 if ok == len(cases) else 1
|
|
|
|
|
|
if __name__ == "__main__":
|
|
import sys
|
|
if len(sys.argv) > 1 and sys.argv[1] == "--self-test":
|
|
raise SystemExit(self_test())
|
|
print(json.dumps(judge(sys.argv[1], sys.argv[2], sys.argv[3]).__dict__,
|
|
indent=2))
|