docs/calculator-test-patterns.md: codify the 2026-05-10 pattern bench

Captures the testing patterns surfaced from the three-module
pattern bench (anchor_prg / phi_alignment_probe /
t3_bound_calculator) as a docs/ architecture reference. Future
blackops shifts find it via grep instead of having to read commit
messages from de997f7 + a4b3056 + 35aeebb to reconstruct the
discipline.

Eight patterns codified:

  1. KAT (known-answer-test) regression — pin (input, expected)
     fixtures under bench/fixtures/<module>/
  2. Hand-computed formula assertion — compute the spec's math
     in the test file from first principles, assert agreement
  3. Monotonicity per input axis — scaling one knob while
     holding others fixed verifies the function responds in
     the expected direction
  4. Closure / sum-of-parts invariants — A + B + C ≡ total;
     no missing terms, no double-counting
  5. Parametrized invalid-input cones — collapse N
     test_rejects_X functions into one @parametrize per
     validation class
  6. CLI subprocess end-to-end — argparse + main() drift the
     import-only tests miss
  7. Doc parity — pin the spec's worked-example numbers
     against the function's actual output (catches doc drift,
     which surfaced today as fox's calibration-pass note on the
     T3 bound doc §11)
  8. Module-export shape — version constants, asdict round-trip,
     JSON-serializability, "v1" version-string discipline

Includes:

- "When to use" header that explicitly limits applicability to
  calculator-style code (math-heavy, closed-form output, JSON-
  serializing CLI)
- "When NOT to use" disclaimer that verifier / parser /
  state-machine tests have their own appropriate structures
  and shouldn't be retrofitted (saves future shifts the same
  realization I had today walking warrant_resolver tests)
- A 9-item checklist for new calculator-style modules to land
  with consistent test density
- Concrete code excerpts from the three exemplar test files

Architecture-reference convention followed (per docs/TICKETS.md
§ "Distinction from other docs"): docs/<filename>.md, no ticket
number, listed in the architecture-reference block. Joins
concept-relations-design.md, warrant-substrate-cookbook.md, and
the cti-architecture.md / mesh.md / etc. set.

docs/TICKETS.md amended with the new entry under "Distinction
from other docs".
This commit is contained in:
russell@unturf.com 2026-05-10 13:13:19 -04:00
parent 28b57da381
commit 0725eb418b
No known key found for this signature in database
2 changed files with 292 additions and 0 deletions

View file

@ -56,6 +56,14 @@ proposing change. Examples that stay un-numbered:
record, written 2026-05-10 after the 18/92 → 92/92 push under
`#000031`. Per-pillar substrate map, ingest patterns, alias
discipline, cascade tuning, honest tier breakdown.)
- `calculator-test-patterns.md` (architecture reference for
testing calculator / probe-style code: KAT regression,
hand-computed formula assertions, monotonicity, closure /
sum-of-parts invariants, parametrized invalid-input cones, CLI
subprocess, doc parity, module-export shape. Codified
2026-05-10 from the three-module pattern bench across
`t3_bound_calculator` / `phi_alignment_probe` / `anchor_prg`.
Explicit "doesn't apply to verifier-style code" caveat.)
If a doc proposes change AND awaits a decision AND has scoped
implementation cost, it's a ticket. Otherwise it's reference.

View file

@ -0,0 +1,284 @@
# Calculator / probe test patterns
Architecture reference for testing **calculator-style** Python
modules: pure functions over numeric / cryptographic inputs that
emit JSON-serializable structured output and (often) a CLI surface.
The patterns codified here surfaced from the 2026-05-10 pattern
bench across three modules:
- ``arborist/substrate/anchor_prg.py`` — HMAC-SHA-512 PRG for #000035
- ``bench/scripts/phi_alignment_probe.py`` — Lanczos alignment probe
for #000034 Phase 1a
- ``bench/scripts/t3_bound_calculator.py`` — T3 channel-capacity
closed-form bound for #000036
The exemplar test file is ``tests/test_t3_bound_calculator.py``
(51 cases, written by fox 2026-05-10). The other two test files
(``tests/test_anchor_prg.py`` and
``tests/test_phi_alignment_probe.py``) were retroactively
backfilled with the same patterns in commits
``de997f7`` and ``a4b3056``.
---
## When to use these patterns
**Apply to**: pure functions over numeric input, math-heavy code,
modules that expose a structured-output dataclass / JSON-emitting
CLI, anything with closed-form formulas the test can hand-compute.
**Don't apply to**: parser / state-machine code (e.g.
``arborist/qa/warrant_resolver.py``), verifier code (e.g.
``arborist/qa/warrant_chain.py``), retrieval code, anything where
inputs and outputs lack a hand-derivable mathematical relationship.
The verifier-style tests in ``tests/test_warrant_resolver.py``
(33 tests) and ``tests/test_warrant_chain.py`` (10 tests) are
appropriately structured for their domain — descriptive
function-name-per-input-shape works better than parametrized
tables for parsers, and there's no formula to hand-check.
---
## The patterns
### 1. KAT (known-answer-test) regression
Pin a small fixture of (input, expected_output) tuples into
``bench/fixtures/<module>/known-answer-tests.jsonl``. Tests replay
every fixture row against the live module; mismatches fail. Catches
unintentional algorithm changes between releases.
**Examples**:
- ``bench/fixtures/phi-prg/known-answer-tests.jsonl`` (10 vectors)
- ``bench/fixtures/phi-alignment/synthetic-checkpoints.jsonl`` (30 vectors)
**Discipline**: bump the module's ``VERSION`` constant + emit a new
fixture file when the algorithm changes. Old runs replay against
old data; new runs replay against new data.
**Caveat**: KAT regression alone can't catch bugs in the original
implementation that were baked into the fixture. Always pair with
hand-computed formula tests below.
### 2. Hand-computed formula assertion
For closed-form math, compute the expected output from first
principles inside the test file using stdlib (``math``, ``hashlib``)
or numpy. Assert agreement with the function output.
**Why**: catches algorithm drift that KAT regression misses (KAT
could have been generated against a buggy version).
**Example** — fox's ``test_b1_exact_formula`` in
``test_t3_bound_calculator.py``:
```python
def test_b1_exact_formula():
g, gnm, sigma, W = 0.1, 2.0, 0.5, 5000
snr = g * gnm / sigma # 0.4
expected = 1.0 * g * W * math.log2(snr + 1)
r = t3_bound_bits(**{**_BASELINE, "gradient_fraction": g, ...})
assert r["B1_contribution"] == pytest.approx(expected, abs=1e-3)
```
The test computes B1 by hand using the §3 formula from the source
spec, then asserts the function returns the same value. Catches a
sign error or constant drop that no KAT regression would.
**Example** — ``test_phi_prg_first_block_matches_direct_hmac`` in
``test_anchor_prg.py``: directly invokes
``hmac.new(seed, h + counter_bytes, sha512).digest()`` and asserts
the function's first 64 bytes match.
**Example** — ``test_uniform_baseline_matches_analytical_formula``
in ``test_phi_alignment_probe.py``: hand-computes
``(1/dim_d) Σ 1/(λ_j+ε)`` from the §2.1 derivation and asserts
agreement with the function's ``a_uniform`` field.
### 3. Monotonicity in each input axis
For each input the function takes, assert that scaling it
monotonically moves the output in the expected direction.
Holding others fixed.
**Why**: catches sign errors, dropped terms, missed dependencies.
**Examples**:
- T3 calculator: doubling ``window_length`` doubles ``B1`` exactly
(linear); raising ``gradient_fraction`` raises ``I_window``
while leaving ``B2`` and ``B3`` untouched.
- phi_prg: increasing ``dim_h`` from N to N+k yields output whose
first N entries are byte-identical to the dim_h=N output
(prefix-extension / streaming-counter invariant).
- phi_alignment_probe: tighter W concentration on low-λ subspace
produces strictly larger ratio.
**Pattern code**:
```python
def test_monotone_in_X():
base = func(**baseline)
higher_x = func(**{**baseline, "X": baseline["X"] * 2})
assert higher_x["output"] > base["output"]
# Bonus: exact relationship if the formula is linear in X
assert higher_x["output"] == pytest.approx(2 * base["output"], rel=1e-6)
```
### 4. Closure / sum-of-parts invariants
If the function returns a structured output with multiple
contributions, assert the contributions sum / compose to the
total. No missing terms, no double-counting.
**Examples**:
- T3: ``I_window ≡ B1 + B2 + B3`` (test_total_equals_sum_of_three_contributions)
- phi_alignment_probe: when ``k_top + k_bot = dim_d`` (full
spectrum), ``a_top + a_bot`` must equal full-spectrum A from
dense ``numpy.linalg.eigh`` decomposition
(test_full_spectrum_a_top_plus_a_bot_covers_full_isotropic_baseline)
- phi_alignment_probe: top-k eigenvalue min ≥ bot-k eigenvalue
max (no overlap; test_eigenvalue_ordering_top_dominates_bot)
**Why**: these are mechanical sanity invariants. A pure
implementation passes them automatically; a buggy implementation
that drops a term, double-counts, or has overlapping ranges fails
loudly.
### 5. Parametrized invalid-input cones
Use ``@pytest.mark.parametrize`` to spawn one test function per
invalid-input class.
**Why**: cleaner than N separate ``test_rejects_X`` /
``test_rejects_Y`` functions for the same validation surface.
**Pattern**:
```python
@pytest.mark.parametrize("g", [0.0, -0.1, 1.5, 2.0])
def test_invalid_gradient_fraction_rejected(g):
args = {**_BASELINE, "gradient_fraction": g}
with pytest.raises(ValueError, match="gradient_fraction"):
func(**args)
```
**Don't over-parametrize**: if each case has a unique error
message that the test asserts specifically, separate functions
remain readable. The cone form works best when the assertion
shape is identical across cases.
### 6. CLI subprocess end-to-end
For modules that expose a CLI surface, test ``argparse + main()``
via ``subprocess.run``. Parse the JSON output; check exit codes
on validation errors.
**Why**: import-only tests miss argparse + main() drift. Refactors
that break the CLI surface (rename flag, change exit code) won't
surface until someone runs the command. fox's
``test_cli_baseline_runs_clean`` and ``test_cli_invalid_input_exits_2``
catch exactly this.
**Example**:
```python
def test_cli_baseline_runs_clean(tmp_path):
cmd = [
sys.executable, "-m", "bench.scripts.t3_bound_calculator",
"--gradient-fraction", "0.05", ...
]
out = subprocess.run(cmd, capture_output=True, text=True,
check=True, cwd="/home/fox/git/arborist")
j = json.loads(out.stdout)
assert j["calculator_version"] == CALCULATOR_VERSION
assert j["I_window_bits_upper_bound"] == pytest.approx(625.87, abs=0.05)
```
**Caveat**: the ``cwd`` argument matters — invoking
``python -m bench.scripts.<X>`` requires the project root as cwd
so the package import resolves.
### 7. Doc parity (and what fox's pattern caught)
Test that the function's actual output matches the numbers in the
spec doc's worked-example sections. Catches **doc drift**
when a doc was hand-written before the implementation crystallized
and approximate rounding got baked into prose.
**Example finding (2026-05-10)**: fox's
``test_baseline_matches_section_11_doc`` flagged that
``soft-hash-channel-t3-bound.md §11`` said 622.7 / 290.0 / 32.7
bits but the calculator produced 625.87 / 292.48 / 33.39. The
test's docstring (lines 56-61) explicitly tracks the
calibration-pass follow-up. Refresh landed in commit ``de997f7``.
**Why**: docs go stale faster than code. A test that pins the
spec-quoted numbers fails noisily when the doc drifts; the prose
in the doc gets refreshed (or the function gets fixed, depending
on which side has the bug).
### 8. Module-export shape
Test the public exports: version constants, threshold constants,
dataclass round-trip via ``asdict``, JSON-serializability.
**Example** — fox's ``test_returns_calculator_version_token``:
```python
def test_returns_calculator_version_token():
r = func(**_BASELINE)
assert r["calculator_version"] == CALCULATOR_VERSION
assert "v1" in CALCULATOR_VERSION
```
The "v1" assertion encodes the discipline: when the algorithm
changes, the version string MUST change too. A future contributor
who refactors the algorithm without bumping ``CALCULATOR_VERSION``
gets caught here.
**Example** — ``test_alignment_report_round_trips_via_asdict``:
asserts ``asdict(report)`` is JSON-serializable. Catches
non-stdlib types leaking into the dataclass that would break
downstream consumers.
---
## Checklist for new calculator-style code
When opening a new ``bench/scripts/<X>.py`` or substrate primitive:
1. [ ] KAT fixture under ``bench/fixtures/<module>/`` (≥ 5 cases)
2. [ ] Module-level ``VERSION`` constant; tests assert its
presence + the ``v1`` discipline
3. [ ] Hand-computed formula tests — at least one per
independent contribution / output field
4. [ ] Monotonicity test per input axis (or invariance assertion
if the function is meant to be insensitive to that axis)
5. [ ] Closure / sum-of-parts invariant if the output decomposes
6. [ ] Parametrized invalid-input cone (one ``@parametrize`` per
validation class — wrong dim, out-of-range value, wrong type)
7. [ ] CLI subprocess test if the module exposes a CLI
8. [ ] Doc-parity test if the spec doc has worked-example numbers
9. [ ] Module-export shape test (asdict round-trip,
JSON-serializability, version token presence)
The 51 tests fox shipped for ``t3_bound_calculator`` are the
exemplar; new modules should aim for similar coverage density on
their own surface.
---
## What this doc is NOT
- A general testing tutorial. Pytest mechanics + fixture scoping
+ ``tmp_path`` use are documented in pytest's own docs.
- A code-style guide. Ruff / black / type-hints discipline lives
in CLAUDE.md.
- A list of every pattern in the repo. Verifier tests, parser
tests, retrieval tests have their own structures appropriate
to their domains; this doc is calculator-style only.
The "When to use" caveat at the top is the load-bearing line:
**don't retrofit these patterns onto verifier-style tests**.
Calculator code has math; verifier code has state machines.
Different shapes need different test patterns.