arborist/docs/TICKETS.md
russell@unturf.com 5428b3b1d6
#000048 — fold the step-2.2 finding through the ticket + the cross-refs
What we learned from the (reverted) step-2.2 attempt, stated as a
general principle in #000048 §5 "What we learned":

  A recombination ("Mercury is the largest planet …" reusing the
  source's "largest planet …" with its "Mercury") and a legitimate
  cross-sentence summary ("Batman, who is the alias of Bruce Wayne,
  lives in Gotham City." reusing two adjacent source sentences) are
  the same SHAPE to any lexical signal — both scatter the answer's
  content tokens across source clauses, and the recombination's
  best-single-clause coverage (4/5 = 0.8) sits ABOVE the legit
  summary's (4/6 = 0.67), so no token-coverage / clause-containment /
  bigram threshold separates them in the safe direction. The
  discriminating thing is *attribution* — in the source, are these
  tokens attached to the same subject/predicate the answer attaches
  them to? — which is a dependency / NLI question, not a string
  metric. That's the boundary of the deterministic, no-LLM-judge
  lexical verifier: absence signals (#000046 numeric gate, #000048
  step 2.1 entity gate) and structure-of-the-model's-own-output
  signals (step 2.4 segmenter) are lexical and work; "the source
  contradicts this pairing" is not, and proxying it with a coverage
  cut trades a small contrived-fixture win for honest demotions of
  real summaries — a net loss against bench-maxing's 5-pp floor.

Updated: #000048 §2.2 (the attempted idea kept as design log + the
no-threshold-separates finding), §2.3 (now framed as the only path to
the last 2 — attribution / mini-NLI, its own ticket if ever), §3
(original plan annotated with the LANDED / ATTEMPTED+REVERTED / NOT
DONE outcome), §5 (the step-2.2 receipt + the "What we learned"
subsection + the closure recommendation). Stale cross-refs fixed:
TICKETS.md #000048 + #000046 rows, #000046 ticket Headroom section,
#000012 §8 #4 — all of which said "#000048 step 2.2 closes the last
2", now corrected to "step 2.2 attempted + reverted; the 2 recombination
fixtures stand as documented residue; #000048 §2.3 is the path if
ever wanted".

Recommendation unchanged: close #000048 at 2.1+2.4 (12 of 16 residual
items closed — formulate-hard 12/12, falsification-hard 10/12). Doc-
only — no code change.
2026-05-11 17:34:18 -04:00

18 KiB
Raw Blame History

Arborist tickets — index

Arborist uses a flat ticket-NNNNNN-<slug>.md convention for actionable proposals: defects with a known root cause, design proposals awaiting go/no-go, or scoped enhancements queued for future implementation.

Convention

  • Filename: docs/tickets/ticket-NNNNNN-<slug>.md. The ID is a six-digit zero-padded integer; the slug is hyphen-lowercase and short enough to stay readable on a 100-column terminal.
  • Numbering: monotonic, never reused. Pick the next free integer greater than every existing ticket. The current next ID is on the Next ID line below; bump it when you open a ticket and commit the bump in the same commit as the new ticket file.
  • Header: every ticket starts with
    # Ticket #NNNNNN — <short title>
    
    **Status:** open · awaiting go/no-go      (or "in progress", "closed")
    **Opened:** YYYY-MM-DD
    **Scope:** <one-liner — what this ticket is for>
    **Audience:** <who reads this>
    **Hard constraint:** <invariants this ticket commits to NOT break>
    
  • Body: problem statement, design choices with trade-offs, recommendation, implementation sketch, scope boundaries, status. See ticket-000001-retrieval-keywords-audit-gap.md for a worked template.
  • Closure: when a ticket lands, flip Status to closed · landed in commit <sha> and leave the file in place. Do not delete tickets; they are the design log.

Distinction from other docs

docs/ also contains architecture references and bench journey notes that are NOT tickets — they describe state of the world rather than proposing change. Examples that stay un-numbered:

  • seven-point-program.md (north-star architectural directive distilled 2026-05-01; every new ticket / feature / prompt edit walks past this; bench harness reports per-mode directive coverage)
  • cti-architecture.md (architecture reference for the Clause Tree Intelligence claim-lattice-pointer mode)
  • mesh.md (mesh wire format + group-key state machine)
  • benchmarks.md (bench harness orientation — fixtures, signal floor, make targets, bench-row schema, addenda index; read first when running a bench)
  • bench-maxing.md (bench discipline — 5pp signal floor, n=3 default, avoid negation in prompts, honest verdicts over optimistic ones)
  • qa-modes-bench.md (rolling bench-journal addenda with cross-references to per-ticket bench data)
  • soft-hash-channel-analysis.md (analysis doc closing #000018 open questions §9.1 / §9.2 / §9.3; companion to soft-hash-channel-t3-bound.md)
  • soft-hash-channel-t3-bound.md (#000036 formal derivation of the per-window covert-channel budget; companion to the analysis doc, both feed the t3_bound_calculator)
  • tool-action-dag-design.md (research path for an action- provenance layer downstream of final_label. Captured pre-ticket per the 2026-05-07 arborist-vs-donto comparison; promotes to a ticket when an actual agent use case surfaces)
  • warrant-substrate-cookbook.md (architecture reference for the 18 open-licensed textbook substrates backing every claim-pack record, written 2026-05-10 after the 18/92 → 92/92 push under #000031. Per-pillar substrate map, ingest patterns, alias discipline, cascade tuning, honest tier breakdown.)
  • calculator-test-patterns.md (architecture reference for testing calculator / probe-style code. Twelve patterns split into two tiers: the original 9-item core (KAT regression, hand-computed formula assertions, monotonicity, closure / sum-of-parts invariants, parametrized invalid-input cones, CLI subprocess, doc parity, module-export shape) plus three domain-specific contract pins (projective-contract for one-way canonicalizers, dispatch-order pin for Python type-hierarchy gotchas, tie-breaking-rule pin for banker's rounding). Codified 2026-05-10 from the original three-module pattern bench across t3_bound_calculator / phi_alignment_probe / anchor_prg; extended same day with the four π* kernel exemplars (pi_star_arithmetic / _logic / _code / _time_series). Explicit "doesn't apply to verifier-style code" caveat.)

If a doc proposes change AND awaits a decision AND has scoped implementation cost, it's a ticket. Otherwise it's reference.

Index

Newest first. Update on every open/close.

ID Title Status Opened Directive
#000048 Verifier upgrade — recombination-aware grounding + clause segmentation in progress · steps 2.1 + 2.4 landed 2026-05-11; step 2.2 attempted + reverted. 2.1: verify_quotes entity salient-token-disagreement gate (_entity_salient_disagrees + _is_single_sentence) → falsification-hard 6/12 → 10/12. 2.4: parse_pointer_claims clause segmentation (_SEGMENT_SEP_RE split kept only when every segment is well-pointered; leading colon-header dropped; wrapped-bullet join) → formulate-hard 4/12 → 12/12 (at ceiling). Each bench-gated (make bench-qa n=3×75×3 — 0 / 7-lattice-cells-touched-both-correct; all deltas within the 5-pp floor). 2 live-pack expected_reason updated HYBRID_ENTITY→UNGROUNDED. 2.2 (single-clause-containment paraphrase check): catches the 2 recombination fixtures but also rejects legit cross-sentence summaries (the Batman case) with no threshold separating the two — recombination-vs-summary isn't lexical; reverted. Recommendation: close at 2.1+2.4 (12 of 16 residual items closed; the 2 falsification-hard recombination fixtures a documented marker; an attribution-aware verifier = §2.3, its own ticket if ever). #000046 follow-up; awaiting fox's close-or-keep-open call 2026-05-11
#000047 ForkScore _delta_* aggregator (mean vs max vs sum) closed · 2026-05-11 — Option D: WeightSet.delta_aggregator ∈ {mean,max,sum} (default mean unchanged → no ESTIMATOR_VERSION bump), fork_score._delta_5{s,t,f} dispatch via _aggregate, recorded in ScoredFork.weights, per-sub HARD_REGRESSION_FLOOR flags aggregator-independent; bench data behind keeping mean in 5f-threshold-calibration-2026-05-11.md §5; 8+1 tests. #000012-revision / #000025 §10.14 follow-up 2026-05-11
#000046 Harder 5S/5T/5F fixture tier (below-ceiling baselines) closed · 2026-05-11 — Phase 1 falsification-hard-v1.jsonl (12 near-misses) + Phase 2 formulate-hard-v1.jsonl (12 mis-segments, rate 4/12) + Phase 3 verify_quotes paraphrase numeric-agreement gate (_numeric_signature; demotes a token-covering span asserting a digit-number the source lacks modulo thousands-comma) → falsification-hard rate 4/12 → 6/12 on a real change; bench-gated (make bench-qa n=3×75×3 before/after — no STRICT-rate regression on legit answers; only gate-caused QA shift was correctly demoting a fictional-year claim STRICT→HYBRID); fork_score γ·Δ5f went positive on it. Headroom now down to 2 falsification-hard over-grounds (#000048 step 2.1 closed the 4 entity over-grounds; step 2.4 closed the 8 Formulate mis-segments → that pack 12/12; step 2.2 attempted + reverted — the last 2 recombination fixtures need an attribution-aware verifier, #000048 §2.3, and stand as documented residue). make bench-5f-falsification-hard / bench-5f-formulate-hard / bench-fork-baseline-hard; 7+ tests. #000025 §10.14 follow-up; #000047 closed; #000048 follow-up 2026-05-11
#000045 Prometheus-Σ Phase 3 sleep-sweep scheduler (gating ticket) open · doc-only scaffold 2026-05-10; pins 8 governance parameters + 4 retrigger gates; opens implementation only after one retrigger fires 2026-05-10
#000044 AUTOCOUNT doc-drift discipline closed · landed across fc5ba50 / 03c0f6a / 6c6defb / f5dbfab / 3b30126 2026-05-10 (mechanism + 4 metrics + 54 tags across 7 doc files; harness catches drift at test time, refresh is 60-second turnaround) 2026-05-10
#000043 Empirical tightening of T3 bound constants (#000036 Phase 2) open · awaiting go/no-go (parks until v7 deployment surfaces measurement data) 2026-05-10
#000042 Term-aliases table (vocabulary-mismatch bridge) closed · 13 rows live across geometry + classical-physics + arithmetic domains by 2026-05-10 2026-05-09
#000041 Citation-aliases table (PD substitutes for proprietary cites) closed · 74 rows live as of 2026-05-10 (count grew 40 → 54 → 74; Goldstein/Newton, Mendelson/Enderton/Jech/Landau/Gödel→{Russell IMP, Russell PoM, De Morgan, Boole, Cantor, Peano, Dedekind, SF-LF}, Stanley/Brualdi/Knuth → Bogart+Levin+Keller-Trotter, Dummit-Foote/Barendregt/Böhm-Jacopini → Judson/PLFA/SF, Kolmogorov → Grinstead-Snell+Laplace) 2026-05-09
#000040 Phase 5 resolver fix — phrase + content-token cascade (Hilbert terminology mismatch surfaced) closed · cascade landed 2026-05-09; lift blocked by 1902-vs-modern vocab; follow-up #000042 2026-05-09
#000039 Optional sqlite-vec retrieval backend (A/B vs FTS5, hybrid not replacement) in progress · Phase 0 doc + Phase 1 landed 2026-05-11 (arborist/search/vec.pyVecBackend + chunk_vecs vec0 + embed_documents + arborist embed / search --backend vec; [vec] extra = sqlite-vec + fastembed; v1 tuning: bge-small-en-v1.5 / 384 / float32 / cosine / flat / top_k 20; 7 tests; demonstrated on crawl_appliedcombinatorics_org.db — 168 chunks, semantic hits topically correct, chain-check 0). Phase 2 (RRF hybrid fusion) gated on ≥5pp recall lift 2026-05-09
#000038 Phase 4 content acquisition — proprietary textbook license decisions for warrant coverage closed · obviated 2026-05-10 by alias-substitution sprint under #000031 (74 rows in #000041 + 13 rows in #000042); 92/92 records now resolve. Residue (multilingual PD, Hilbert-Ackermann OCR, Knuth permission, personal-copy path B) preserved as design log §8 2026-05-09
#000037 Prometheus-Σ recursive falsification controller (bicameral substrate) in progress · Phases 0 + 1 + 1.b + 1.c + 2 landed 2026-05-10; §12 Trigger 2 fired (divergence variance 0.575 / N=37); §22 Findings 2 + 3 RESOLVED (kernel/llm cost split + sweep_weights §15.4 + per-mode τ_qa); controller_events carries 4 event kinds (decision · difficulty · budget_allocation · falsification_proposal) feeding arborist controller-events inspector + live-harvest third bucket in bench/scripts/harvest_falsification_proposals.py; §12 Trigger 1 probe wired 2026-05-11 (trigger_1_branch_density reads fork_score_branches — measurable, not yet fired); Phase 3 sleep-sweep scheduler tracked under #000045 (gating ticket) 2026-05-09
#000036 T3 per-window covert-channel budget bound closed · 2026-05-11 · Phase 1 + dav1d review → Tier-1 + Tier-2 (Option B = b1_model=max_envelope default, in v1, no v2 fork) + KAT-regen tooling (scripts/generate_t3_bound_kat.py) all landed 2026-05-11; baseline 625.87 → 6183.02 (max_envelope), NOT_CERTIFIED_BY_BOUND at W=10000; 53 → 83 tests; 12-entry active KAT; both dav1d closure blockers cleared, all §5 acceptance criteria met. Continuation: empirical C_B* tightening under #000043 (parks on v7) 2026-05-09
#000035 PRG choice for φ_PRG (HMAC-SHA-512 expansion) closed · 2026-05-11 — Phase 1 (arborist/substrate/anchor_prg.py HMAC-SHA-512 counter-mode KDF, pure stdlib, PHI_PRG_VERSION = phi-prg-v1-hmac-sha512-le, dim_h ≤ 16·2^32 guard + bool-reject; 31 tests; 10 KATs) + Phase 2 (§ 9.10 + § 9.10.1 folded into the in-repo v7 substrate spec docs/_source/merkle-agi-dag-v7.rst — the v7 formal substrate was brought into the repo for this; supersedes the prior "edit upstream ~/Downloads/merkle-agi-dag_v7.txt" plan, and is now version-controlled with docs/_source/index.rst toctree entry). dav1d-reviewed-final 2026-05-11: little-endian per v7 §A1 (counter + uint32-word reads flipped, KATs regenerated), manifest field anchor_prg_seed, float-map prose "negligible mean 2^32" not "unbiased", seed-independence/single-purpose, M1-enablement-policy separation; HMAC-SHA-512 / 32-byte seed / SHALL-replace all LOCKED. #000018 §9.2 ("which PRG?") resolved → HMAC-SHA-512 + 32-byte committed seed. (Full upstream v7 spec revision stays exogenous; this lands the amendment into the tracked in-repo copy.) 2026-05-09
#000034 Hessian alignment under φ_linear in progress · Phase 1a landed 2026-05-10 (synthetic-ablation probe + KAT fixture); Phase 1b parks for v7 ramp-up 2026-05-09
#000033 Claim-pack pillar VII (combinatorics) closed · landed 2026-05-09 (live in shard 000.db; lift verified) 2026-05-09
#000032 combinatorics@v1 π* (integer counting kernel) closed · landed 2026-05-09 2026-05-09
#000031 Surface-ingest cited textbooks for claim-pack warrant promotion closed · Phase 1 + 2 + 2.5 (92/92) + 3 (verifier wiring) + B-1 (via_citation_alias attribution: 18 direct + 74 +alias) + B-2 (source-side title-from-author backfill) all landed 2026-05-09 / 2026-05-10 2026-05-09
#000030 Math π* expansion: SymPy substrate (algebra · calculus · linalg) closed · all 7 phases + 1b landed 2026-05-09 (abe5988) 2026-05-09
#000029 Claim-pack source (axiom/theorem JSON bundles) closed · landed 2026-05-09 2026-05-09
#000028 Multi-modality witness for canonical shapes closed · landed 2026-05-09 + follow-ups (capital ledger · sample rate) 2026-05-08
#000027 Canonical projections persist to providence_cache closed · landed 2026-05-09 2026-05-08
#000026 Real-shard workload baseline + search latency closed · Phase 1 + 2 + 3 landed 2026-05-09 (Phase 3 in 60b5748) 2026-05-08
#000025 5F battery (Function · Finetuning · Falsification · Formulate · Feedback Loop) closed · 2026-05-11 — Phase 1a1f landed 2026-05-09/10; Phase 1g (§10.13 feedback latency + persisted-footprint efficiency), 1h (§10.14 ForkScore threshold-calibration handoff → #000012 §8), 1i (§10.11 persistent SelfModel-chain lineage: bench-5f-selfmodel-snapshot grows the chain, run_finetuning shard-chain mode reads the 2 latest snapshots) all landed 2026-05-11. Every §10 closure criterion met 2026-05-07
#000024 5T Phase 1b + Dav1DPrometheus vocabulary alignment closed · landed 2026-05-08 2026-05-07
#000023 5S Phase 1b: Syllogism · Synthesis · Semiotics closed · landed 2026-05-08 2026-05-07
#000022 Adapter LossReport (PRD I9 analogue) closed · landed 2026-05-07 2026-05-07
#000021 5S/5T/5R benchmark fixtures + harness closed · Phase 1a + 2 landed 2026-05-08 2026-05-07
#000020 Capital-cost ledger (8-capital queues) closed · landed 2026-05-07 2026-05-07
#000019 Specification methodology for π* and V closed · landed 2026-05-07 2026-05-07
#000018 Adversarial soft-hash covert-channel analysis closed · landed 2026-05-09 (analysis doc) 2026-05-07
#000017 Memory-root: lifelong learning audit chain closed · landed 2026-05-07 2026-05-07
#000016 ZK Phase-2 frontier proof (concretize) parked · bench-plan + wire-protocol landed 2026-05-09 (sibling-repo measurement pending) 2026-05-07
#000015 π* domain library + cross-domain composition closed · landed 2026-05-07 2026-05-07
#000014 SelfModel: schema, falsification, integration closed · landed 2026-05-07 2026-05-07
#000013 Spatial-temporal substrate (Merkle-AGI v7-W) closed · landed 2026-05-09 (substrate paper + frontier catalog + namespace stub) 2026-05-07
#000012 Selection & consensus protocol (Merkle-AGI v8) in progress · Phase 1a (ForkScore) landed 2026-05-08; Phase 1b (consensus paper, docs/_source/merkle-agi-v8-consensus.rst 834 lines) landed 2026-05-10; Phase 1c (branch-set persistence — fork_score_branches sibling table, persist_branch_score + branch_set_density, 6 new CLI flags on arborist substrate score, default-off) landed 2026-05-10 — feeds #000037 §12 Trigger 1; Trigger 1 probe wired 2026-05-11 (trigger_1_branch_density reads fork_score_branches via branch_set_density() — measurable, not yet fired: no branch sets persisted) 2026-05-07
#000011 SOFT_PREFLIGHT_HINT model-assisted sidecar closed · landed 2026-05-04 (zero-shot full impl) 2026-05-04 D1 (preserves)
#000010 Meta-Cognition Preflight Guard (M0 / MCTL) closed · landed 2026-05-03 (Phases 14); DAG binding shipped via #000009 2026-05-03 D1, D3
#000009 Preflight run-DAG node binding (#000008+#000010) closed · re-landed 2026-05-04 (§8 corrections: reject-path DAG, nested CTI clauses) 2026-05-03 D3, D4
#000008 Broad-quantifier preflight guard closed · landed in 4f2b5a6; Phase 5 DAG binding split into #000009 2026-05-02
#000007 Query-layer hyphen folding closed · 2026-05-02 2026-05-02
#000006 Bench-emergent findings (rolling research log) open · rolling research log (verifier-ladder healthy at 134 cycles; emergent gems still being mined) 2026-05-02
#000005 Label ladder migration (POINTER-LINKED → …) closed · 2026-05-02 2026-05-01 D7
#000004 Directive coverage in bench summary closed · acd1f9c 2026-05-01 D8
#000003 Anchor-class warrant generalization (Module H+) closed · 2026-05-02 2026-05-01 D6
#000002 Reference-Frame Polarity Contract (Module L) closed · 2026-05-02 2026-05-01 D3
#000001 Retrieval-keywords audit gap closed · 2026-05-02 2026-05-01 D4

Next ID

000049