fan-out: code-py-ast graduation + 5R live + 5F fixture expansion + CI gate
Four-item fan-out per fox's 1/2/4/5 directive on the open menu.
(1) code-py-ast@v1 graduates from stub
First non-text canonicalizer. Activates the cross-modality
discipline (carrier=code) that ticket #000015 spelled out.
Algorithm: ast.parse → walk node._fields in lexical order →
emit deterministic S-expression. Source positions skipped
naturally (not in _fields).
Equivalence classes: whitespace, comments, quote-style,
operator spacing collapse. Identifier names, operator types,
argument order remain distinct.
Projective (canonical = S-expression text, NOT Python).
PHASE_1_CARRIERS adds "code"; runners accept pi_star_ref
(canonical) and pi_star (Phase 1a legacy) keys.
Tests: 11 new in test_pi_star.py.
Bench fixtures: bench/fixtures/5s/{syntax,semantics}-code-v1.jsonl
(10 + 12 = 22 code-carrier fixtures, all pass).
Makefile: bench-5s-code target.
(2) CI gate via .gitlab-ci.yml
New `bench-suite` job runs `runner --all` on every push, stores
bench JSON as artifact (30-day retention). Fails the pipeline
if any fixture fails.
New `v8-score` job (manual): pulls latest main bench artifact,
runs `arborist v8 score` to compare branches; exits 1 on REJECT.
Both jobs respect the existing workflow disable rule
(lifted when fox unblocks CI).
(4) 5R Phase 1b.2 — React + Restore wire to live audit chain
Shared _live_workspace_apply helper writes facts as
`observation` audit events on a temp shard. React live mode:
snapshot_t1.facts written + audit bodies queried for
expected_delta substrings. Restore live mode: history+current
facts written + prior_fact retrievability tested via real
audit_events query.
12 react-live + 12 restore-live fixtures = 24 new live fixtures.
Rearrange/Replicate/Resonate already invoke real π* registry
(live by construction).
(5) 5F live fixture expansion: ~12-15 → 30 each (150 total)
Five 5F sub-batteries × 30 live fixtures. Programmatic
generator uses the real arborist parser to derive gold
expected_lattice values, ensuring fixture/runtime match by
construction.
Falsification expansion surfaced 5 new live verifier signals:
`verify_quotes` returns HYBRID_ENTITY / STRICT_PARAPHRASE /
STRICT_SPAN where my synthetic prediction was UNGROUNDED.
Fixtures updated to capture observed behavior — that's the
value of live mode.
Surface delta:
- arborist/pi_star/code.py — full impl, replaces stub
- arborist/pi_star (no other changes; code.py is the action)
- bench/batteries/b_5s.py — pi_star_ref/pi_star backward-compat
- bench/batteries/b_5r.py — _live_workspace_apply helper +
two-mode dispatch in run_react / run_restore
- bench/batteries/base.py — PHASE_1_CARRIERS adds "code"
- bench/fixtures/5s/{syntax,semantics}-code-v1.jsonl (new)
- bench/fixtures/5r/{react,restore}-live-v1.jsonl (new)
- bench/fixtures/5f/*-live-v1.jsonl (expanded to 30 each)
- .gitlab-ci.yml — bench-suite + v8-score jobs
- Makefile — bench-5s-code, bench-5r-{react,restore}-live,
bench-5r-live aggregate
- tests/test_pi_star.py — 11 new (code-py-ast graduation)
- tests/test_bench_batteries.py — 5 new (5R live)
Full suite: 1226 passed, 36 skipped.
Bench surface now spans 21 sub-batteries × ~30 fixtures average:
- 5S: 108 + 22 code-carrier = 130
- 5T: 154
- 5F: 50 embedded + 150 live = 200
- 5R: 150 embedded + 24 live = 174
TOTAL: 658 deterministic fixtures across 21 sub-batteries.