bootstrap-nli-only: a venv + [nli] only — no [dev] extras (no crawler/
hessian/vec/sympy). On a CUDA host PyPI's torch wheel is the CUDA build,
so ShadowNLI auto-detects cuda and bench-nli-shadow / export-nli-onnx /
the candidate bench all run on the GPU with no further wiring. Verified
on the ai box (RTX 4090): torch 2.11+cu130, cuda True; 82M cross-encoder
batched ≈ 0.09 ms/pair (512 pairs in 0.047s — ~350x onnx-int8-cpu,
~1300x torch-cpu-batch1); 62-record synthetic shadow sweep in 2.4s wall;
24/24 tests pass. The GPU only accelerates the NLI half — make bench-qa
(Hermes LLM) still runs wherever the shard corpus is.