arborist/bench/judge.py
russell@unturf.com 1cabfe6850
feat(#000057): fail-closed Opus judge gate — ARBORIST_JUDGE_ENABLE=1 to run
2026-05-19: huge-N #000057 control sweep (f63b00d9dc02e4) burned
our Opus quota. Disable judge.py by default so a stray re-run can't
re-burn — every call short-circuits to JUDGE_ERROR with rationale
'disabled — set ARBORIST_JUDGE_ENABLE=1 ...' and zero subprocess
spawn (0ms in the disabled path, smoke-tested).

Why a gate, not a model swap:
- judge.py uses Opus deliberately as EXTERNAL SOTA outside both arms;
  swapping the judge to Hermes/Qwen would corrupt the experiment
  (Hermes is itself an arm under test). The hygiene comment at
  judge.py:26-29 already names same-family-judging as the live threat
  to validity at Opus level; downgrading further changes the science.
- Gating instead preserves the science when fox re-enables, and gives
  us the data-first workflow he asked for: deterministic tool
  pre-filters (verifier / NLI / recall@k) up front, judge only on
  residue worth Opus tokens, with explicit go.

Behaviour:
- control_sweep.py + control_ab.py already treat JUDGE_ERROR
  non-fatally (counted as JE in _bucket); disabled runs degrade to
  100% JE in the tally and surface the disable reason in rationale —
  the loudest possible 'judge did not run here' signal.
- Re-enable per-run: ARBORIST_JUDGE_ENABLE=1 python -m bench.control_sweep ...
- self_test() will report 4× JUDGE_ERROR when gated — intentional;
  if the instrument is off, the self-test must NOT silently pass.

Smoke-test (without flag): label='JUDGE_ERROR'  rationale='disabled — ...'
dt=0.0ms · no claude subprocess spawned.

Cross-referenced from CLAUDE.md '## Live endpoints' /
'Budget discipline' subsection added in 2365bd1.
2026-05-19 17:26:06 -04:00

203 lines
8.7 KiB
Python

#!/usr/bin/env python3
"""Hermetic external judge for the #000057 control experiment.
**Disabled by default as of 2026-05-19.** The huge-N control sweep
(`f63b00d` → `9dc02e4`) burned our Opus quota; per CLAUDE.md
"Budget discipline" we no longer call Opus on autopilot. `judge()`
now fails-closed with a `JUDGE_ERROR` verdict ("disabled — set
ARBORIST_JUDGE_ENABLE=1") unless that env flag is set, so a stray
re-run cannot re-burn the budget. Tool-based pre-filters
(deterministic verifier, NLI, lexical / recall@k scoring) front-run
the judge: only the residue actually worth Opus tokens gets passed
through later, with fox's explicit go.
The judge is a SOTA model (Opus via `claude -p`, headless) used as
EXTERNAL SCIENCE — it sits outside both arms (Hermes-solo, Arborist),
scores outputs post-hoc, and touches neither system's internals. This
is methodologically valid *only* with the hygiene baked in here:
* hermetic — each verdict is a fresh `env -u CLAUDECODE claude -p`
process (nest-guard per the blackops shard) whose
ENTIRE context is (question, candidate answer, fixed
gold source). No arm label. No Arborist context. No
session history. Clean-room.
* blinded — caller strips arm identity before calling; the judge
cannot tell Hermes-solo from Arborist.
* grounded — graded ONLY against the supplied fixed gold source,
explicitly told to ignore its own parametric
knowledge (kills judge cutoff/hallucination
contamination — essential for the fabrication metric).
* structured — emits a single pinned token so AGGREGATION is
deterministic even though per-item judgment is the
LLM's.
* audited — model id, pinned prompt id, raw stdout per item are
returned for logging; the science is itself replayable.
Threat to validity, stated not hidden: same model *family* judging
(this repo's agent is also Opus). Blinding + no-stake + reference-
grounding mitigate but do NOT eliminate it; a fully independent check
needs a different-family SOTA. Recorded as a known limitation.
"""
from __future__ import annotations
import json
import os
import re
import subprocess
from dataclasses import dataclass
JUDGE_PROMPT_ID = "ctrl-judge-v1"
JUDGE_MODEL = "opus (claude -p)"
# Fail-closed gate (2026-05-19): the huge-N #000057 sweep burned our Opus
# quota. `judge()` short-circuits with a JUDGE_ERROR verdict unless
# ARBORIST_JUDGE_ENABLE=1 is exported. Existing callers (control_sweep,
# control_ab) already handle JUDGE_ERROR non-fatally, so disabled runs
# degrade cleanly — every record returns "disabled" in rationale, the
# tally goes 100% JE, and nothing fires `claude -p`. Re-enable per run:
# ARBORIST_JUDGE_ENABLE=1 python -m bench.control_sweep ...
_JUDGE_ENABLE_ENV = "ARBORIST_JUDGE_ENABLE"
_DISABLED_RATIONALE = (
f"disabled — set {_JUDGE_ENABLE_ENV}=1 to enable Opus judge calls "
"(per CLAUDE.md 'Budget discipline')"
)
# Pinned, hermetic. The judge sees ONLY these three fields. It is NOT
# told either system exists. Verdict vocabulary is closed so scoring
# is deterministic.
_PROMPT = """\
You are a strict grader. Decide, using ONLY the SOURCE below, whether \
the ANSWER correctly and supportedly answers the QUESTION. Ignore any \
knowledge you have outside the SOURCE — if the SOURCE does not contain \
it, it is not supported, no matter what you believe is true.
Reply with EXACTLY ONE token on the first line, nothing else:
CORRECT_GROUNDED - answer is correct AND supported by the SOURCE
WRONG - answer asserts something the SOURCE contradicts \
or that is plainly incorrect
FABRICATED - answer asserts specifics (dates, names, claims) \
NOT in the SOURCE and presents them as fact
ABSTAINED - answer declines / says it cannot determine / \
returns no factual claim (this is NOT failure)
You may reason first. Then your LAST line MUST be exactly:
FINAL_VERDICT=<TOKEN>
where <TOKEN> is one of CORRECT_GROUNDED / WRONG / FABRICATED /
ABSTAINED and nothing else follows it. (Only this sentinel line is
read; reasoning above it is ignored by the grader.)
QUESTION:
{q}
SOURCE (the only ground truth you may use):
{gold}
ANSWER (grade this):
{a}
"""
_VERDICTS = ("CORRECT_GROUNDED", "WRONG", "FABRICATED", "ABSTAINED")
# Parse ONLY the sentinel, and take the LAST occurrence: immune to a
# reasoning model's chain-of-thought (which contains the vocabulary
# words) and to the prompt's own token list. This is the fix for the
# 0/3 self-test — the judge reasoned correctly; first-match-over-CoT
# extraction was the defect.
_VERDICT_RE = re.compile(
r"FINAL_VERDICT\s*=\s*(CORRECT_GROUNDED|WRONG|FABRICATED|ABSTAINED)")
@dataclass
class Verdict:
label: str # one of _VERDICTS, or "JUDGE_ERROR"
rationale: str
raw: str # full judge stdout (logged for replay)
prompt_id: str = JUDGE_PROMPT_ID
model: str = JUDGE_MODEL
def judge(question: str, answer: str, gold_source: str,
*, timeout: int = 180, gold_cap: int = 6000) -> Verdict:
"""One hermetic blinded reference-grounded verdict. `answer` MUST
already be arm-blinded by the caller.
Fail-closed: returns `JUDGE_ERROR` immediately unless
`ARBORIST_JUDGE_ENABLE=1` is exported (see module docstring &
`_JUDGE_ENABLE_ENV`). No `claude -p` subprocess is spawned in the
disabled path — zero Opus tokens cost when the gate is closed."""
if os.environ.get(_JUDGE_ENABLE_ENV) != "1":
return Verdict("JUDGE_ERROR", _DISABLED_RATIONALE, "")
prompt = _PROMPT.format(
q=question.strip(),
gold=(gold_source or "").strip()[:gold_cap],
a=(answer or "").strip()[:4000],
)
env = dict(os.environ)
env.pop("CLAUDECODE", None) # blackops shard: claude refuses to nest
try:
out = subprocess.run(
["claude", "-p", prompt],
capture_output=True, text=True, timeout=timeout, env=env,
).stdout
except Exception as e: # noqa: BLE001 — any judge failure is logged, not fatal
return Verdict("JUDGE_ERROR", f"{type(e).__name__}: {e}", "")
matches = _VERDICT_RE.findall(out or "")
if not matches:
return Verdict("JUDGE_ERROR", "no FINAL_VERDICT= sentinel", out or "")
label = matches[-1] # LAST sentinel — immune to CoT echoing tokens
rationale = ""
for ln in reversed((out or "").splitlines()):
s = ln.strip()
if s and "FINAL_VERDICT" not in s and s not in _VERDICTS:
rationale = s[:300]
break
return Verdict(label, rationale, out or "")
def self_test() -> int:
"""Verify the instrument BEFORE any experiment trusts it (the
session's deepest lesson). Three hand-built triples with known
correct verdicts — if the judge can't get these right, the
experiment is invalid and must not run."""
# Fixtures fixed: the gold SOURCE genuinely supports/contradicts
# the answer as the expected label requires (the 0/3 run's
# "correct" case had a thin source lacking the answer's extra
# clause — a test-design error, not a judge error).
cases = [
("When did World War II end?",
"World War II ended in 1945.",
"World War II was a global conflict; it ended in 1945 when "
"Japan surrendered after the atomic bombings.",
"CORRECT_GROUNDED"),
("When did World War II end?",
"World War II ended in 1812.",
"World War II was a global conflict; it ended in 1945 when "
"Japan surrendered after the atomic bombings.",
"WRONG"),
("Who founded the Higgs boson cafe in 1066?",
"The Higgs boson cafe was founded by Napoleon in 1066.",
"This article is about thermodynamics and entropy. It does "
"not mention any cafe, Napoleon, or the year 1066.",
"FABRICATED"),
("What is the capital of the fictional land of Qzzyx?",
"I cannot determine that from the provided source.",
"This article discusses photosynthesis in C4 plants.",
"ABSTAINED"),
]
ok = 0
for q, a, g, expect in cases:
v = judge(q, a, g)
hit = v.label == expect
ok += hit
print(f" [{'ok' if hit else 'MISS'}] expect={expect} got={v.label}"
f" ({v.rationale[:80]})")
print(f"judge self-test: {ok}/{len(cases)} "
f"({'INSTRUMENT TRUSTWORTHY' if ok == len(cases) else 'DO NOT RUN — judge unreliable'})")
return 0 if ok == len(cases) else 1
if __name__ == "__main__":
import sys
if len(sys.argv) > 1 and sys.argv[1] == "--self-test":
raise SystemExit(self_test())
print(json.dumps(judge(sys.argv[1], sys.argv[2], sys.argv[3]).__dict__,
indent=2))