Walking6aca7d9(cookbook test-coverage appendix) surfaced two findings: (1) two stale test counts since fox wrote the appendix at 2026-05-10 13:11 EDT; (2) the appendix and docs/calculator-test- patterns.md are complementary lenses but had no explicit cross-reference. Both fixed in this docs-only commit. Stale counts refreshed ====================== - `tests/test_phi_alignment_probe.py` "14 tests" → "23 tests". Drift cause: my `a4b3056` (2026-05-10 14:08 EDT) added 9 pattern-backfill tests after fox's appendix snapshot at 13:11 EDT (~57 min lag). - `tests/test_t3_bound_calculator.py` "51 tests" → "53 tests". Drift cause: my `581ad90` (2026-05-10 ~13:50 EDT) added 2 KAT-fixture-gap closures after fox's appendix snapshot. Both refreshes preserve the trajectory by noting the ``+9 froma4b3056`` / ``+2 from581ad90`` provenance inline. Same durability pattern fox used in `018a2a1` for the alias-count refresh + my `6f1dbed` ditto. Per the appendix-author's own ``unit-test density`` heuristic, the refreshed counts confirm both files keep their ≥1× test/code ratio. ``test_phi_alignment_probe.py`` jumps from 268/200 ≈ 1.34 to 419/200 ≈ 2.10× (closer to the contract-defining-foundation ratio fox flagged for warrant_chain.py at 3.6×). Missing entry added ==================== `tests/test_substrate_fork_score.py` (renamed from `test_v8_fork_score.py` in `a4058a4` per the 2026-05-10 v-prefix retirement) wasn't listed in fox's appendix. The file is the ``arborist substrate score`` CLI surface coverage — adapter tests + 4 in-process build_parser CLI tests + 1 real subprocess invocation. Distinct from `test_fork_score.py` (fox's pure-function unit tests for ScoredFork at 18 tests). Added under "Substrate-paper-spec'd primitives" section alongside test_fork_score.py. Reciprocal cross-reference ========================== `docs/calculator-test-patterns.md` (the per-pattern CHECKLIST for new tests) and `docs/warrant-substrate-cookbook.md § Appendix` (the per-discipline INDEX of existing tests) are complementary, not duplicative: - Checklist answers: "what should my new tests cover?" - Index answers: "where are the tests for X?" Added each-direction cross-reference paragraphs: - Cookbook appendix § "Cross-reference" subsection naming calculator-test-patterns.md as the checklist for new code. When adding a new substrate-paper-spec'd primitive: walk the checklist to design the test file, then add a row to the appendix under the matching discipline. - calculator-test-patterns.md § "What this doc is NOT" expanded with a bullet pointing readers at the cookbook appendix as the existing-test inventory. Closes the gap where future shifts might find one without the other and miss half the discipline. Hygiene ======= - make test → 1986 passed, 45 skipped. - Both docs are reference-only; no test or code surface change.
291 lines
11 KiB
Markdown
291 lines
11 KiB
Markdown
# Calculator / probe test patterns
|
|
|
|
Architecture reference for testing **calculator-style** Python
|
|
modules: pure functions over numeric / cryptographic inputs that
|
|
emit JSON-serializable structured output and (often) a CLI surface.
|
|
The patterns codified here surfaced from the 2026-05-10 pattern
|
|
bench across three modules:
|
|
|
|
- ``arborist/substrate/anchor_prg.py`` — HMAC-SHA-512 PRG for #000035
|
|
- ``bench/scripts/phi_alignment_probe.py`` — Lanczos alignment probe
|
|
for #000034 Phase 1a
|
|
- ``bench/scripts/t3_bound_calculator.py`` — T3 channel-capacity
|
|
closed-form bound for #000036
|
|
|
|
The exemplar test file is ``tests/test_t3_bound_calculator.py``
|
|
(51 cases, written by fox 2026-05-10). The other two test files
|
|
(``tests/test_anchor_prg.py`` and
|
|
``tests/test_phi_alignment_probe.py``) were retroactively
|
|
backfilled with the same patterns in commits
|
|
``de997f7`` and ``a4b3056``.
|
|
|
|
---
|
|
|
|
## When to use these patterns
|
|
|
|
**Apply to**: pure functions over numeric input, math-heavy code,
|
|
modules that expose a structured-output dataclass / JSON-emitting
|
|
CLI, anything with closed-form formulas the test can hand-compute.
|
|
|
|
**Don't apply to**: parser / state-machine code (e.g.
|
|
``arborist/qa/warrant_resolver.py``), verifier code (e.g.
|
|
``arborist/qa/warrant_chain.py``), retrieval code, anything where
|
|
inputs and outputs lack a hand-derivable mathematical relationship.
|
|
The verifier-style tests in ``tests/test_warrant_resolver.py``
|
|
(33 tests) and ``tests/test_warrant_chain.py`` (10 tests) are
|
|
appropriately structured for their domain — descriptive
|
|
function-name-per-input-shape works better than parametrized
|
|
tables for parsers, and there's no formula to hand-check.
|
|
|
|
---
|
|
|
|
## The patterns
|
|
|
|
### 1. KAT (known-answer-test) regression
|
|
|
|
Pin a small fixture of (input, expected_output) tuples into
|
|
``bench/fixtures/<module>/known-answer-tests.jsonl``. Tests replay
|
|
every fixture row against the live module; mismatches fail. Catches
|
|
unintentional algorithm changes between releases.
|
|
|
|
**Examples**:
|
|
- ``bench/fixtures/phi-prg/known-answer-tests.jsonl`` (10 vectors)
|
|
- ``bench/fixtures/phi-alignment/synthetic-checkpoints.jsonl`` (30 vectors)
|
|
|
|
**Discipline**: bump the module's ``VERSION`` constant + emit a new
|
|
fixture file when the algorithm changes. Old runs replay against
|
|
old data; new runs replay against new data.
|
|
|
|
**Caveat**: KAT regression alone can't catch bugs in the original
|
|
implementation that were baked into the fixture. Always pair with
|
|
hand-computed formula tests below.
|
|
|
|
### 2. Hand-computed formula assertion
|
|
|
|
For closed-form math, compute the expected output from first
|
|
principles inside the test file using stdlib (``math``, ``hashlib``)
|
|
or numpy. Assert agreement with the function output.
|
|
|
|
**Why**: catches algorithm drift that KAT regression misses (KAT
|
|
could have been generated against a buggy version).
|
|
|
|
**Example** — fox's ``test_b1_exact_formula`` in
|
|
``test_t3_bound_calculator.py``:
|
|
|
|
```python
|
|
def test_b1_exact_formula():
|
|
g, gnm, sigma, W = 0.1, 2.0, 0.5, 5000
|
|
snr = g * gnm / sigma # 0.4
|
|
expected = 1.0 * g * W * math.log2(snr + 1)
|
|
r = t3_bound_bits(**{**_BASELINE, "gradient_fraction": g, ...})
|
|
assert r["B1_contribution"] == pytest.approx(expected, abs=1e-3)
|
|
```
|
|
|
|
The test computes B1 by hand using the §3 formula from the source
|
|
spec, then asserts the function returns the same value. Catches a
|
|
sign error or constant drop that no KAT regression would.
|
|
|
|
**Example** — ``test_phi_prg_first_block_matches_direct_hmac`` in
|
|
``test_anchor_prg.py``: directly invokes
|
|
``hmac.new(seed, h + counter_bytes, sha512).digest()`` and asserts
|
|
the function's first 64 bytes match.
|
|
|
|
**Example** — ``test_uniform_baseline_matches_analytical_formula``
|
|
in ``test_phi_alignment_probe.py``: hand-computes
|
|
``(1/dim_d) Σ 1/(λ_j+ε)`` from the §2.1 derivation and asserts
|
|
agreement with the function's ``a_uniform`` field.
|
|
|
|
### 3. Monotonicity in each input axis
|
|
|
|
For each input the function takes, assert that scaling it
|
|
monotonically moves the output in the expected direction.
|
|
Holding others fixed.
|
|
|
|
**Why**: catches sign errors, dropped terms, missed dependencies.
|
|
|
|
**Examples**:
|
|
- T3 calculator: doubling ``window_length`` doubles ``B1`` exactly
|
|
(linear); raising ``gradient_fraction`` raises ``I_window``
|
|
while leaving ``B2`` and ``B3`` untouched.
|
|
- phi_prg: increasing ``dim_h`` from N to N+k yields output whose
|
|
first N entries are byte-identical to the dim_h=N output
|
|
(prefix-extension / streaming-counter invariant).
|
|
- phi_alignment_probe: tighter W concentration on low-λ subspace
|
|
produces strictly larger ratio.
|
|
|
|
**Pattern code**:
|
|
|
|
```python
|
|
def test_monotone_in_X():
|
|
base = func(**baseline)
|
|
higher_x = func(**{**baseline, "X": baseline["X"] * 2})
|
|
assert higher_x["output"] > base["output"]
|
|
# Bonus: exact relationship if the formula is linear in X
|
|
assert higher_x["output"] == pytest.approx(2 * base["output"], rel=1e-6)
|
|
```
|
|
|
|
### 4. Closure / sum-of-parts invariants
|
|
|
|
If the function returns a structured output with multiple
|
|
contributions, assert the contributions sum / compose to the
|
|
total. No missing terms, no double-counting.
|
|
|
|
**Examples**:
|
|
- T3: ``I_window ≡ B1 + B2 + B3`` (test_total_equals_sum_of_three_contributions)
|
|
- phi_alignment_probe: when ``k_top + k_bot = dim_d`` (full
|
|
spectrum), ``a_top + a_bot`` must equal full-spectrum A from
|
|
dense ``numpy.linalg.eigh`` decomposition
|
|
(test_full_spectrum_a_top_plus_a_bot_covers_full_isotropic_baseline)
|
|
- phi_alignment_probe: top-k eigenvalue min ≥ bot-k eigenvalue
|
|
max (no overlap; test_eigenvalue_ordering_top_dominates_bot)
|
|
|
|
**Why**: these are mechanical sanity invariants. A pure
|
|
implementation passes them automatically; a buggy implementation
|
|
that drops a term, double-counts, or has overlapping ranges fails
|
|
loudly.
|
|
|
|
### 5. Parametrized invalid-input cones
|
|
|
|
Use ``@pytest.mark.parametrize`` to spawn one test function per
|
|
invalid-input class.
|
|
|
|
**Why**: cleaner than N separate ``test_rejects_X`` /
|
|
``test_rejects_Y`` functions for the same validation surface.
|
|
|
|
**Pattern**:
|
|
|
|
```python
|
|
@pytest.mark.parametrize("g", [0.0, -0.1, 1.5, 2.0])
|
|
def test_invalid_gradient_fraction_rejected(g):
|
|
args = {**_BASELINE, "gradient_fraction": g}
|
|
with pytest.raises(ValueError, match="gradient_fraction"):
|
|
func(**args)
|
|
```
|
|
|
|
**Don't over-parametrize**: if each case has a unique error
|
|
message that the test asserts specifically, separate functions
|
|
remain readable. The cone form works best when the assertion
|
|
shape is identical across cases.
|
|
|
|
### 6. CLI subprocess end-to-end
|
|
|
|
For modules that expose a CLI surface, test ``argparse + main()``
|
|
via ``subprocess.run``. Parse the JSON output; check exit codes
|
|
on validation errors.
|
|
|
|
**Why**: import-only tests miss argparse + main() drift. Refactors
|
|
that break the CLI surface (rename flag, change exit code) won't
|
|
surface until someone runs the command. fox's
|
|
``test_cli_baseline_runs_clean`` and ``test_cli_invalid_input_exits_2``
|
|
catch exactly this.
|
|
|
|
**Example**:
|
|
|
|
```python
|
|
def test_cli_baseline_runs_clean(tmp_path):
|
|
cmd = [
|
|
sys.executable, "-m", "bench.scripts.t3_bound_calculator",
|
|
"--gradient-fraction", "0.05", ...
|
|
]
|
|
out = subprocess.run(cmd, capture_output=True, text=True,
|
|
check=True, cwd="/home/fox/git/arborist")
|
|
j = json.loads(out.stdout)
|
|
assert j["calculator_version"] == CALCULATOR_VERSION
|
|
assert j["I_window_bits_upper_bound"] == pytest.approx(625.87, abs=0.05)
|
|
```
|
|
|
|
**Caveat**: the ``cwd`` argument matters — invoking
|
|
``python -m bench.scripts.<X>`` requires the project root as cwd
|
|
so the package import resolves.
|
|
|
|
### 7. Doc parity (and what fox's pattern caught)
|
|
|
|
Test that the function's actual output matches the numbers in the
|
|
spec doc's worked-example sections. Catches **doc drift** —
|
|
when a doc was hand-written before the implementation crystallized
|
|
and approximate rounding got baked into prose.
|
|
|
|
**Example finding (2026-05-10)**: fox's
|
|
``test_baseline_matches_section_11_doc`` flagged that
|
|
``soft-hash-channel-t3-bound.md §11`` said 622.7 / 290.0 / 32.7
|
|
bits but the calculator produced 625.87 / 292.48 / 33.39. The
|
|
test's docstring (lines 56-61) explicitly tracks the
|
|
calibration-pass follow-up. Refresh landed in commit ``de997f7``.
|
|
|
|
**Why**: docs go stale faster than code. A test that pins the
|
|
spec-quoted numbers fails noisily when the doc drifts; the prose
|
|
in the doc gets refreshed (or the function gets fixed, depending
|
|
on which side has the bug).
|
|
|
|
### 8. Module-export shape
|
|
|
|
Test the public exports: version constants, threshold constants,
|
|
dataclass round-trip via ``asdict``, JSON-serializability.
|
|
|
|
**Example** — fox's ``test_returns_calculator_version_token``:
|
|
|
|
```python
|
|
def test_returns_calculator_version_token():
|
|
r = func(**_BASELINE)
|
|
assert r["calculator_version"] == CALCULATOR_VERSION
|
|
assert "v1" in CALCULATOR_VERSION
|
|
```
|
|
|
|
The "v1" assertion encodes the discipline: when the algorithm
|
|
changes, the version string MUST change too. A future contributor
|
|
who refactors the algorithm without bumping ``CALCULATOR_VERSION``
|
|
gets caught here.
|
|
|
|
**Example** — ``test_alignment_report_round_trips_via_asdict``:
|
|
asserts ``asdict(report)`` is JSON-serializable. Catches
|
|
non-stdlib types leaking into the dataclass that would break
|
|
downstream consumers.
|
|
|
|
---
|
|
|
|
## Checklist for new calculator-style code
|
|
|
|
When opening a new ``bench/scripts/<X>.py`` or substrate primitive:
|
|
|
|
1. [ ] KAT fixture under ``bench/fixtures/<module>/`` (≥ 5 cases)
|
|
2. [ ] Module-level ``VERSION`` constant; tests assert its
|
|
presence + the ``v1`` discipline
|
|
3. [ ] Hand-computed formula tests — at least one per
|
|
independent contribution / output field
|
|
4. [ ] Monotonicity test per input axis (or invariance assertion
|
|
if the function is meant to be insensitive to that axis)
|
|
5. [ ] Closure / sum-of-parts invariant if the output decomposes
|
|
6. [ ] Parametrized invalid-input cone (one ``@parametrize`` per
|
|
validation class — wrong dim, out-of-range value, wrong type)
|
|
7. [ ] CLI subprocess test if the module exposes a CLI
|
|
8. [ ] Doc-parity test if the spec doc has worked-example numbers
|
|
9. [ ] Module-export shape test (asdict round-trip,
|
|
JSON-serializability, version token presence)
|
|
|
|
The 51 tests fox shipped for ``t3_bound_calculator`` are the
|
|
exemplar; new modules should aim for similar coverage density on
|
|
their own surface.
|
|
|
|
---
|
|
|
|
## What this doc is NOT
|
|
|
|
- A general testing tutorial. Pytest mechanics + fixture scoping
|
|
+ ``tmp_path`` use are documented in pytest's own docs.
|
|
- A code-style guide. Ruff / black / type-hints discipline lives
|
|
in CLAUDE.md.
|
|
- A list of every pattern in the repo. Verifier tests, parser
|
|
tests, retrieval tests have their own structures appropriate
|
|
to their domains; this doc is calculator-style only.
|
|
- A coverage inventory. The per-discipline test-file index lives
|
|
at `docs/warrant-substrate-cookbook.md` § "Appendix —
|
|
test-coverage cross-reference". This doc is the **checklist
|
|
for new tests**; the cookbook appendix is the **index of
|
|
existing tests** keyed by substrate discipline. Walk both:
|
|
this doc designs the test file, the cookbook appendix records
|
|
it under the matching discipline.
|
|
|
|
The "When to use" caveat at the top is the load-bearing line:
|
|
**don't retrofit these patterns onto verifier-style tests**.
|
|
Calculator code has math; verifier code has state machines.
|
|
Different shapes need different test patterns.
|