Step 4 of bend-gpu.lsp now dispatches (cuda-sim-ops-bin path 141)
against /tmp/ecdsa-queue/out.bin on 3090-ai — the foxhop kickmix
circuit simulator at 9024 shots over 141 batches. The HTTP payload
is ~70 bytes; the worker does ~1.5 seconds of real Toffoli-heavy
QECC simulation, the GPU kernel comes back in ~360 ms, byte-identical
to the CPU reference. Speedup ~4x on this workload, scales further
with bigger circuits.
Why this works over HTTP where 1M-input SHAKE wouldn't: the request
references a bin file that already lives on the worker disk, so we
never have to stream the bytes through the wire. The response is a
structured (cuda-sim-result …) S-expression with timing-ms,
gates-sum, status, byte-identity, and the speedup-kernel field —
everything you need to see the GPU win directly in the playground
output panel.
Pairs with the existing small-payload (cuda-shake-fanout (...) 32)
step so the demo shows both the round-trip pattern and the real
"slow on CPU, fast on GPU" headline that the original demo header
promised but couldn't deliver until cuda-sim-ops-bin was wired up.