qa(#000009): preflight stage binds into run_dag_root

Closes ticket #000009 zero-shot. Scope expanded to cover BOTH
ticket #000008 (broad-quantifier) AND ticket #000010
(meta-cognition) preflight contracts in a single combined node —
both share the same audit-replay gap and inserting two separate
nodes between question and retrieval was operationally awkward.

aborist/qa/dag.py:
  + preflight_node_hash() — combines QuestionState +
    quantifier classifier output + behavioral policy_state into
    one canonical SHA-256 hex.
  + build_run_dag() gains optional preflight_hash parameter.
    When supplied, inserts {"stage": "preflight", "hash": ...}
    at position 1 (between question and retrieval).
    Backward-compat: None → original 7/9-stage shapes preserved
    for legacy run_dag_root re-validation.

  Quote-mode: 7 → 8 stages with preflight.
  Pointer-mode CTI: 9 → 10 stages with preflight.

aborist/qa/query.py + runner.py:
  Both build the preflight payload from question_state +
  quantifier dict + 10-field policy_state (guard_enabled,
  guard_apply_caps, guard_apply_caps_mode_gated,
  claim_cap_resolved, claim_cap_actually_applied,
  reminder_enabled, reminder_eligible, reject_broad_active,
  metacognition_enabled, block_on_contradiction).

  This means two cache rows that share the same question + same
  model output + same verifier verdict but DIFFERENT preflight
  policy state now produce different run_dag_root values. Audit
  replay can pin the policy decision per row.

9 new tests in tests/test_dag.py:
  - preflight_node_hash determinism
  - hash bumps on question_state change
  - hash bumps on policy_state change (the audit-replay payoff)
  - all-None defensive shape
  - 7→8 stage transition (quote mode)
  - 9→10 stage transition (pointer mode), preflight at index 1
  - run_dag_root bumps when preflight_hash bumps
  - verify_run_dag round-trips through preflight stage

Live verification: latest providence_cache row carries
['question', 'preflight', 'retrieval', 'evidence_map', 'prompt',
'raw_answer', 'parsed_claim_lattice', 'verify', 'render',
'final_label'] — preflight stage living in the persisted DAG.

987 tests passing (9 new); 36 skipped.

Tickets:
  #000009 status: closed · landed 2026-05-03 (zero-shot)
  #000010 cross-ref updated: "DAG binding shipped via #000009"

What's NOT in this ticket (logged in §7.2):
  - CLI flag for inspecting preflight node from cache_key
  - Bench harness preflight_hash field for cross-row comparison
  - SOFT_PREFLIGHT_HINT (model-assisted preflight sidecar)
This commit is contained in:
russell@unturf.com 2026-05-03 18:34:16 -04:00
parent f2bbe512db
commit c36e85c86c
No known key found for this signature in database
6 changed files with 367 additions and 16 deletions

View file

@ -99,6 +99,47 @@ def localize_failure(
return "answer"
def preflight_node_hash(
*,
question_state: dict | None = None,
quantifier: dict | None = None,
policy_state: dict | None = None,
) -> str:
"""Hash the preflight decision into a stable hex string.
Bundles three sources of preflight provenance into one
canonical hash that contributes to ``run_dag_root`` via the
new ``preflight`` stage in :func:`build_run_dag`:
- ``question_state`` `QuestionState.to_dict()` from
``aborist.qa.metacognition`` (#000010): logical_statuses,
question_shape, false_premise_hints, contradiction_pairs,
temporal_sensitivity, etc. Carries
``preflight_policy_hash`` so policy flips invalidate the
stage hash automatically.
- ``quantifier`` classifier output from
``aborist.qa.quantifier.classify_question_quantifier``
(#000008): intensity, matched_token, explicit_count,
scope_bound_hint, classifier_version.
- ``policy_state`` the *behavioral* decisions taken on this
run that aren't already in the above (e.g. whether the cap
was actually applied vs just looked up, whether the reminder
was injected, whether reject-broad path was taken). These
are what differentiate two cache rows with the same
classifier output but different downstream effects.
Any of the three may be None (legacy / opt-out); the hash is
still stable. Returns SHA-256 hex.
"""
payload = {
"stage": "preflight",
"question_state": question_state or {},
"quantifier": quantifier or {},
"policy_state": policy_state or {},
}
return _sha256_hex(_canonical_json(payload))
def build_run_dag(
*,
question_hash: str,
@ -119,6 +160,7 @@ def build_run_dag(
parsed_lattice: list | None = None,
rendered_text: str | None = None,
retrieval_plan_hash: str | None = None,
preflight_hash: str | None = None,
) -> dict:
"""Return ``{"root": <hex>, "nodes": [<stage>, <hash>], ...}``.
@ -127,13 +169,15 @@ def build_run_dag(
byte across machines (as long as the Merkle conventions stay pinned;
they do, via ``aborist.merkle``).
Two DAG shapes:
Two base DAG shapes; both gain an optional ``preflight`` stage
when ``preflight_hash`` is supplied (Ticket #000009):
- **Quote mode (default).** 7 stages
``question / retrieval / context / prompt / answer / verify /
final_label``. Triggered when ``evidence_map_root`` is None.
Backward-compatible with all run_dag_root values written by code
that pre-dates G0.
that pre-dates G0. With ``preflight_hash``, becomes 8 stages
``question / preflight / retrieval / ...``.
- **Claim-lattice-pointer mode (G0 / CTI).** 9 stages
``question / retrieval / evidence_map / prompt / raw_answer /
@ -145,10 +189,15 @@ def build_run_dag(
map). All three of ``raw_answer_text`` / ``parsed_lattice`` /
``rendered_text`` should be supplied; missing args fall back to
``answer_text`` for the raw_answer & render hashes and ``[]`` for
the parsed_lattice hash.
the parsed_lattice hash. With ``preflight_hash``, becomes 10
stages.
``answer_mode`` & ``violations`` fold into the verify & final_label
payloads when provided.
payloads when provided. ``preflight_hash`` (Ticket #000009) is
optional; when None, the DAG shape remains 7/9 stages exactly so
pre-#000009 records can be re-validated. When supplied, the
preflight stage inserts at position 1 (between ``question`` and
``retrieval``) per ticket #000009 §3.1.
"""
sources_summary = [
{
@ -239,6 +288,18 @@ def build_run_dag(
{"stage": "render", "hash": rendered_hash},
{"stage": "final_label", "hash": final_label_hash},
]
# Ticket #000009 — preflight stage binding. When supplied,
# insert ``preflight`` between ``question`` and ``retrieval``.
# Optional so legacy run_dag_root values from pre-#000009 code
# remain reproducible (None → original 7/9-stage shape). The
# preflight_hash bundles #000008 quantifier output, #000010
# QuestionState, AND the policy decisions taken on this run
# — see preflight_node_hash() for the canonical payload.
if preflight_hash is not None:
nodes.insert(
1,
{"stage": "preflight", "hash": preflight_hash},
)
leaves = [bytes.fromhex(n["hash"]) for n in nodes]
root_hex = MerkleTree.build(leaves).root.hex()
return {"root": root_hex, "nodes": nodes}

View file

@ -2663,6 +2663,47 @@ def query(
),
)
plan_hash = retrieval_plan_hash(plan)
# Ticket #000009 — preflight node binding. Combines #000008
# quantifier classifier output + #000010 metacognition
# QuestionState + the behavioral policy decisions (cap
# actually applied, reminder actually injected, reject path
# taken) into a single hash that contributes to run_dag_root.
# Audit replay can now distinguish two cache rows that have
# the same question + same model output but different
# preflight policy state.
from aborist.qa.dag import preflight_node_hash
preflight_hash = preflight_node_hash(
question_state=question_state.to_dict(),
quantifier=quantifier,
policy_state={
"guard_enabled": quantifier_guard_on,
"guard_apply_caps": quantifier_apply_caps,
"guard_apply_caps_mode_gated": quantifier_caps_mode_gated,
"claim_cap_resolved": claim_cap_lookup,
"claim_cap_actually_applied": (
quantifier_apply_caps
and quantifier_caps_mode_gated
and claim_cap_lookup is not None
),
"reminder_enabled": bool(
policy.get("quantifier_reminder_enabled", False)
),
"reminder_eligible": (
quantifier_guard_on
and quantifier_mode_gated
and quantifier.get("is_broad", False)
),
"reject_broad_active": bool(
policy.get("quantifier_reject_broad", False)
),
"metacognition_enabled": bool(
policy.get("metacognition_enabled", True)
),
"block_on_contradiction": bool(
policy.get("metacognition_block_on_contradiction", False)
),
},
)
run_dag = build_run_dag(
question_hash=qhash,
sources=proof_obj["sources"],
@ -2682,6 +2723,7 @@ def query(
parsed_lattice=parsed_lattice,
rendered_text=answer_text if is_lattice_mode else None,
retrieval_plan_hash=plan_hash,
preflight_hash=preflight_hash,
)
run_dag_blob = json.dumps(run_dag, separators=(",", ":"))

View file

@ -852,6 +852,40 @@ def ask(
}
for i, cs in enumerate(verdict.get("claim_statuses") or [])
]
# Ticket #000009 — preflight node binding (mirror of query()).
from aborist.qa.dag import preflight_node_hash
preflight_hash = preflight_node_hash(
question_state=question_state.to_dict(),
quantifier=quantifier,
policy_state={
"guard_enabled": quantifier_guard_on,
"guard_apply_caps": quantifier_apply_caps,
"guard_apply_caps_mode_gated": quantifier_caps_mode_gated,
"claim_cap_resolved": claim_cap_lookup,
"claim_cap_actually_applied": (
quantifier_apply_caps
and quantifier_caps_mode_gated
and claim_cap_lookup is not None
),
"reminder_enabled": bool(
policy.get("quantifier_reminder_enabled", False)
),
"reminder_eligible": (
quantifier_guard_on
and quantifier_mode_gated
and quantifier.get("is_broad", False)
),
"reject_broad_active": bool(
policy.get("quantifier_reject_broad", False)
),
"metacognition_enabled": bool(
policy.get("metacognition_enabled", True)
),
"block_on_contradiction": bool(
policy.get("metacognition_block_on_contradiction", False)
),
},
)
run_dag = build_run_dag(
question_hash=qhash,
sources=[{
@ -875,6 +909,7 @@ def ask(
raw_answer_text=raw_answer if is_lattice_mode else None,
parsed_lattice=parsed_lattice,
rendered_text=answer_text if is_lattice_mode else None,
preflight_hash=preflight_hash,
)
run_dag_blob = json.dumps(run_dag, separators=(",", ":"))

View file

@ -57,8 +57,8 @@ Newest first. Update on every open/close.
| ID | Title | Status | Opened | Directive |
|----------|------------------------------------------------|-----------------------|------------|-----------|
| #000010 | Meta-Cognition Preflight Guard (M0 / MCTL) | closed · landed 2026-05-03 (Phases 14); Phase 5 DAG join #000009 | 2026-05-03 | D1, D3 |
| #000009 | Quantifier preflight run-DAG node binding | open · awaiting go/no-go | 2026-05-03 | D3, D4 |
| #000010 | Meta-Cognition Preflight Guard (M0 / MCTL) | closed · landed 2026-05-03 (Phases 14); DAG binding shipped via #000009 | 2026-05-03 | D1, D3 |
| #000009 | Preflight run-DAG node binding (#000008+#000010) | closed · landed 2026-05-03 (zero-shot) | 2026-05-03 | D3, D4 |
| #000008 | Broad-quantifier preflight guard | closed · landed in `4f2b5a6`; Phase 5 DAG binding split into #000009 | 2026-05-02 | — |
| #000007 | Query-layer hyphen folding | closed · 2026-05-02 | 2026-05-02 | — |
| #000006 | Bench-emergent findings (rolling research log) | open · rolling | 2026-05-02 | — |

View file

@ -1,11 +1,13 @@
# Ticket #000009Quantifier preflight run-DAG node binding
# Ticket #000009Preflight run-DAG node binding
**Status:** open · awaiting go/no-go
**Status:** closed · landed 2026-05-03 (zero-shot)
**Opened:** 2026-05-03
**Scope:** Bind the broad-quantifier preflight contract from
ticket #000008 into the per-run Merkle-DAG so the classifier output
+ cap-application decision are hash-bound rather than just
surface-level result-dict fields.
**Scope (expanded 2026-05-03):** Bind BOTH ticket #000008's broad-
quantifier preflight contract AND ticket #000010's meta-cognition
QuestionState into the per-run Merkle-DAG. Both share the same
audit-replay gap (per #000010 §12.6) and inserting two separate
nodes between `question` and `retrieval` is operationally awkward;
single combined `preflight` node carries both payloads.
**Audience:** fox + future blackops shifts.
**Hard constraint:** No `schema_version`, `canonicalization_version`,
or `chunking_version` bumps. Same constraint #000008 §1 carried.
@ -210,6 +212,71 @@ This ticket closes the loop.
## 7. Status
Open · awaiting go/no-go. Mechanism is well-understood from
#000008 §9.5; this ticket is the audit-binding cleanup, not new
behavior.
**Closed · landed 2026-05-03 (zero-shot).** Mechanism shipped end-
to-end:
- `aborist/qa/dag.py:preflight_node_hash()` — combines
QuestionState + quantifier classifier output + behavioral
policy_state into one canonical SHA-256 hex.
- `aborist/qa/dag.py:build_run_dag()` — new optional
`preflight_hash` parameter inserts a `preflight` stage at
position 1 (between `question` and `retrieval`). Quote-mode
shape becomes 8 stages; pointer-mode CTI shape becomes 10
stages. Backward-compatible: when `preflight_hash` is None,
shapes stay 7/9 exactly so legacy `run_dag_root` values
re-validate.
- `aborist/qa/query.py` + `aborist/qa/runner.py` — both build
the preflight payload from `question_state`, the quantifier
dict, and a 10-field `policy_state` capturing the *behavioral*
decisions taken on this run (guard_enabled, apply_caps_active,
claim_cap_resolved, claim_cap_actually_applied, reminder_enabled,
reminder_eligible, reject_broad_active, metacognition_enabled,
block_on_contradiction, mode-gating bits).
9 new tests in `tests/test_dag.py` pin: hash determinism,
question_state/policy_state independence, all-None defensive
shape, 7→8 / 9→10 stage transitions, root change on policy flip,
verify_run_dag round-trip with preflight stage.
987 tests passing (9 new); 36 skipped.
### 7.1 What this enables (audit replay)
Two cache rows that have:
- The same question
- The same model output
- The same verifier verdict
But different behind-the-scenes preflight policy state (e.g. cap
applied vs not, reminder injected vs not, reject-broad path
taken vs not) now produce **different `run_dag_root`** values.
Audit replay can:
```
hash(preflight node) = h_pre
→ pin: classifier output, cap decision, reminder decision,
reject decision, metacog gates
If h_pre changes between two cache_keys for the same question,
that row reflects a different preflight policy.
Allows: regression bisection ("which day did the cap default
flip break our STRICT-rate?"), policy A/B reconstruction
("show me all rows where cap was applied vs not"), cross-model
diff ("hermes vs qwen on the same question with the same
preflight contract").
```
### 7.2 What's NOT in this ticket
- **CLI flag for inspecting preflight node**: a future
`aborist providence --show-preflight <cache_key>` would render
the preflight payload from `run_dag_blob`. Out of scope here.
- **Bench harness preflight-hash field**: bench rows could
surface `preflight_hash` (12-char prefix like `cache_key`)
for cross-row comparison. Out of scope; can add later if
bench analysis needs it.
- **`SOFT_PREFLIGHT_HINT` (model-assisted preflight)**:
reserved per #000010 §18 / source doc. Hard rule preserved.

View file

@ -13,7 +13,12 @@ exactly as recorded.
from __future__ import annotations
from aborist.qa.dag import build_run_dag, localize_failure, verify_run_dag
from aborist.qa.dag import (
build_run_dag,
localize_failure,
preflight_node_hash,
verify_run_dag,
)
def _kw(**overrides):
@ -169,3 +174,144 @@ def test_verify_dag_accepts_json_string():
out = build_run_dag(**_kw())
blob = json.dumps(out, separators=(",", ":"))
assert verify_run_dag(blob) is True
# ---------------------------------------------------------------- Ticket #000009: preflight stage
def test_preflight_node_hash_is_deterministic():
"""Same inputs → same hex string, byte-for-byte."""
qs = {"logical_statuses": ["well_formed"], "preflight_result": "PREFLIGHT_OK"}
quant = {"intensity": "SINGULAR", "is_broad": False}
pol = {"guard_enabled": True, "guard_apply_caps": False}
a = preflight_node_hash(question_state=qs, quantifier=quant, policy_state=pol)
b = preflight_node_hash(question_state=qs, quantifier=quant, policy_state=pol)
assert a == b
assert len(a) == 64 # SHA-256 hex
def test_preflight_node_hash_changes_with_question_state():
base_quant = {"intensity": "SINGULAR", "is_broad": False}
pol = {"guard_enabled": True}
h_a = preflight_node_hash(
question_state={"preflight_result": "PREFLIGHT_OK"},
quantifier=base_quant, policy_state=pol,
)
h_b = preflight_node_hash(
question_state={"preflight_result": "PREFLIGHT_PARTIAL"},
quantifier=base_quant, policy_state=pol,
)
assert h_a != h_b
def test_preflight_node_hash_changes_with_policy_state():
"""Apply-caps flip MUST bump the preflight node hash so audit
replay can distinguish guard-on vs guard-off rows that
otherwise share the same classifier output."""
qs = {"logical_statuses": ["broad_quantifier_unbounded"]}
quant = {"intensity": "ALL", "is_broad": True}
h_off = preflight_node_hash(
question_state=qs, quantifier=quant,
policy_state={"guard_apply_caps": False},
)
h_on = preflight_node_hash(
question_state=qs, quantifier=quant,
policy_state={"guard_apply_caps": True},
)
assert h_off != h_on
def test_preflight_node_hash_handles_all_none():
"""Defensive — all three components may be None during gradual
rollout. Hash stays stable."""
a = preflight_node_hash(
question_state=None, quantifier=None, policy_state=None,
)
b = preflight_node_hash(
question_state=None, quantifier=None, policy_state=None,
)
assert a == b
assert len(a) == 64
def test_dag_without_preflight_keeps_seven_stage_shape():
"""Backward-compat: omitting preflight_hash preserves the
pre-#000009 7-stage shape so legacy run_dag_root values
re-validate."""
out = build_run_dag(**_kw())
stages = [n["stage"] for n in out["nodes"]]
assert stages == [
"question", "retrieval", "context", "prompt",
"answer", "verify", "final_label",
]
assert len(out["nodes"]) == 7
def test_dag_with_preflight_inserts_eight_stage_shape():
"""Quote-mode + preflight_hash → 8 stages, preflight at
position 1 (between question and retrieval)."""
pre_hash = preflight_node_hash(
question_state={"preflight_result": "PREFLIGHT_OK"},
quantifier={"intensity": "SINGULAR", "is_broad": False},
policy_state={"guard_enabled": True},
)
out = build_run_dag(**_kw(preflight_hash=pre_hash))
stages = [n["stage"] for n in out["nodes"]]
assert stages == [
"question", "preflight", "retrieval", "context",
"prompt", "answer", "verify", "final_label",
]
assert len(out["nodes"]) == 8
def test_dag_with_preflight_lattice_mode_ten_stages():
"""Pointer-mode + preflight_hash → 10 stages."""
pre_hash = preflight_node_hash(
question_state={"preflight_result": "PREFLIGHT_OK"},
quantifier={"intensity": "ALL", "is_broad": True},
policy_state={"guard_enabled": True},
)
out = build_run_dag(**_kw(
preflight_hash=pre_hash,
evidence_map_root="d" * 64,
verifier_method="claim_lattice_pointer",
raw_answer_text="Some claim. [E1]",
parsed_lattice=[{"claim_text": "Some claim", "evidence_ids": ["e1"]}],
rendered_text="Some claim. [E1 | source]",
))
stages = [n["stage"] for n in out["nodes"]]
assert "preflight" in stages
assert stages.index("preflight") == 1 # right after question
assert len(stages) == 10
def test_dag_root_changes_when_preflight_hash_changes():
"""Different preflight inputs → different run_dag_root. This is
the audit-replay payoff: same model output + same verifier
verdict + DIFFERENT preflight policy = different cache row."""
pre_a = preflight_node_hash(
question_state={"preflight_result": "PREFLIGHT_OK"},
quantifier={"intensity": "SINGULAR"},
policy_state={"guard_apply_caps": False},
)
pre_b = preflight_node_hash(
question_state={"preflight_result": "PREFLIGHT_OK"},
quantifier={"intensity": "SINGULAR"},
policy_state={"guard_apply_caps": True},
)
a = build_run_dag(**_kw(preflight_hash=pre_a))
b = build_run_dag(**_kw(preflight_hash=pre_b))
assert a["root"] != b["root"]
def test_dag_with_preflight_round_trips_through_verify():
"""The preflight stage's hash is part of the leaf list so
verify_run_dag must reconstruct the same root."""
import json
pre_hash = preflight_node_hash(
question_state={"preflight_result": "PREFLIGHT_PARTIAL"},
quantifier={"intensity": "ALL", "is_broad": True},
policy_state={"guard_apply_caps": True, "claim_cap_resolved": 8},
)
out = build_run_dag(**_kw(preflight_hash=pre_hash))
blob = json.dumps(out, separators=(",", ":"))
assert verify_run_dag(blob) is True