The S-expression wire format was the bottleneck at huge payload sizes
-- 23.8 s end-to-end for 1M x 16 B inputs on the Python tier, while
the actual CUDA kernel finishes the same workload in ~47 ms. The
hex-S-exp parser ate everything between.
New binary wire mode (magic 'BSHK' prefix; payload is the daemon's
binary portal format verbatim) bypasses S-expression parsing entirely.
Worker writes the blob to disk, calls daemon process-bin, reads result,
prepends 'BSHR' magic, replies.
Measured 3090-ai, daemon warm, localhost:
workload Py S-exp Py binary C S-exp C binary
100 x 16 B 3.43 ms 0.74 ms 0.40 ms 0.15 ms
1k x 16 B 23.24 ms 0.76 ms 2.77 ms 0.22 ms
10k x 16 B 218.82 ms 1.27 ms CLIFF 0.88 ms
100k x 16 B 2,219 ms 10.18 ms CLIFF 10.35 ms
1M x 16 B 23,811 ms 159 ms CLIFF 157 ms
150x speedup at 1M inputs on Python tier. C tier S-exp CLIFFs
between 1k and 10k inputs (reader payload limit); binary mode
bypasses the CLIFF entirely. At 100k+ inputs both tiers converge
since file I/O + CUDA kernel dominates over wire framing.
Host comparison: hashlib.shake_256 over 1M tiny inputs takes ~2 s
on a single Python core. Bend via binary worker = 157 ms = 12x
faster than host. Bend now wins at huge workloads, not just heavy
ones.
Implementation:
lumbda.py
* tcp-send/tcp-recv switched to latin-1 (1:1 byte mapping)
so binary payloads pass through cleanly. UTF-8 was mangling
bytes with replacement chars.
* write-binary-file / read-binary-file primitives.
c/builtins.c
* write-binary-file / read-binary-file matching Python tier.
examples/cuda-fanout/wire.lsp
* wire-send-raw / wire-recv-raw helpers that frame a raw
payload string without S-expression serialization.
examples/cuda-fanout/gpu-worker.lsp
* handle-binary-shake: write portal blob, daemon process-bin,
read result, wire-send 'BSHR' + bytes.
* handle-one dispatches on first 4 bytes of payload: 'BSHK'
goes to binary path, anything else stays S-exp.
examples/cuda-fanout/bench_tiers.py
* make_payload_binary builds the BSHK protocol payload.
* --binary flag in CLI.
www/index.html
* full S-exp + binary comparison table.
* 'bend now beats host hashlib at huge workloads' headline finding.
|
||
|---|---|---|
| asm | ||
| c | ||
| docs | ||
| examples | ||
| proof | ||
| tests | ||
| whitepaper | ||
| www | ||
| .gitignore | ||
| .gitlab-ci.yml | ||
| bench.py | ||
| cl-compat.lsp | ||
| CLAUDE.md | ||
| friction.sh | ||
| lumbda.py | ||
| Makefile | ||
| README.md | ||
| stdlib.lsp | ||
| tests.py | ||
Lumbda
A Lisp/Scheme-derived, just-in-time lambda language. Four implementation tiers with MOAD defect isolation. Workloads migrate across basic UNIX systems.
Four implementation tiers sharing one wire format — Scheme source itself:
- Python bytecode VM — reference, full first-class continuations
- C tree-walker + bytecode VM — portable C, JSON portal
- C + x86_64 JIT — pattern-matched native code, 7–10× faster than CPython
- Pure x86_64 assembly — 22 KB stripped, zero libc, 14 syscalls
Feedback is the primitive across four scopes: continuations within a process, portal files across processes, S-expressions across implementations, TCP sockets across machines.
Home: lumbda.com
λ> (define (fib n)
(let loop ((a 0) (b 1) (i 0))
(if (= i n) a (loop b (+ a b) (+ i 1)))))
λ> (map fib (iota 10))
(0 1 1 2 3 5 8 13 21 34)
Usage
python3 lumbda.py # interactive REPL
python3 lumbda.py script.lsp # run a file
python3 lumbda.py -e '(+ 1 2)' # eval an expression
python3 lumbda.py --fast script.lsp # auto-compile (7-19x faster)
Bytecode compiler
Lumbda includes a stack-based bytecode compiler and VM. Enable it with --fast or (auto-compile! #t):
python3 lumbda.py --fast examples/fibonacci.lsp
(auto-compile! #t)
(define (ack m n)
(cond ((= m 0) (+ n 1))
((= n 0) (ack (- m 1) 1))
(else (ack (- m 1) (ack m (- n 1))))))
(compiled? ack) ; => #t
(ack 3 4) ; => 125
The compiler handles: if, begin, and, or, when, unless, cond, define, set!, lambda, let, named-let, let*, letrec, do, call/cc, function calls with tail-call optimization. Macros are expanded at compile time. 20 specialized opcodes for hot builtins (+, -, *, =, <, car, cdr, cons, null?, etc.) avoid function call overhead.
Features:
- Explicit frame stack — compiled-to-compiled calls don't grow the Python stack
- Full continuations —
call/ccsupports upward continuations; generators work - Constant folding —
(+ 1 2)folds to3at compile time - Peephole optimizer — eliminates dead code (VOID+POP, JUMP-to-next)
(disassemble proc)— inspect generated bytecode
What's implemented
Core language
- Full lexical scoping and closures
- Tail-call optimization (TCO) — deep recursion never blows the stack
- Hygienic macros via
syntax-ruleswith ellipsis (...) support define-macro/defmacrofor procedural macroscall/cc— full continuations (escape + upward) in compiled codevalues/call-with-valuesdynamic-wind,guard,with-exception-handlerquasiquote/unquote/unquote-splicingwith proper nesting- R7RS internal defines with letrec* body semantics
- R7RS error objects
- Exact rational arithmetic —
(/ 1 3)→1/3,(+ 1/4 3/4)→1 - String ports —
open-input-stringopen-output-stringreadon ports - Mutable strings —
string-set!string-fill!string-copy! - Module system —
module/importwith export lists define-record-typewith(inherit parent)for single-inheritance- Pretty-print —
pp/pretty-print - Tracing —
(trace fn)/(untrace fn)
Special forms
define set! lambda λ if cond case and or when unless
begin let let* letrec letrec* named-let do
quasiquote define-macro define-syntax syntax-rules
let-syntax letrec-syntax apply eval values call/cc
dynamic-wind guard parameterize load error
module import define-record-type
Built-ins
- Arithmetic:
+-*/quotientremaindermoduloexptsqrtabsfloorceilingroundtruncateminmaxgcdlcmlogexptrig functions,numeratordenominator - Rationals:
(/ 1 3)→1/3, literal1/3syntax,exact/inexactconversion - Comparison:
=<><=>=zero?positive?negative?odd?even? - Pairs & lists:
conscarcdrlistlengthappendreversemapfor-eachfilterfold-leftfold-rightreduceanyeverysortpartitionfindtakedropzipflattenand more - SRFI-1:
lastfirst–fifthdeletelset-unionlset-intersectionlset-differenceunfoldlist-tabulate - Strings:
string-lengthstring-refstring-set!substringstring-appendstring-copystring-copy!string-fill!string->liststring->numberformatand more - Characters:
char->integerinteger->charchar-alphabetic?char-upcasechar-downcase - Vectors:
make-vectorvectorvector-refvector-set!vector-copyvector-copy! - Hash tables:
make-hash-tablehash-table-set!hash-table-refhash-table-keyshash-table-valueshash-table-walkand more - I/O:
displaywritenewlinereadread-charread-lineopen-input-stringopen-output-stringwith-output-to-string - File system:
file-exists?delete-filerename-filedirectory-filescurrent-directory - System:
command-lineget-environment-variablecurrent-timeexit - Python interop:
py-evalpy-execpy-importpy-callpy-attr - Compiler:
compilecompiled?disassembleauto-compile!
Standard library (stdlib.lsp)
Additional macros, string/list/numeric/tree utilities, alist/hash helpers, simple object system, SRFI-2/8/64 test framework.
Examples
python3 lumbda.py --fast examples/fibonacci.lsp
python3 lumbda.py --fast examples/generator.lsp
python3 lumbda.py --fast examples/mergesort.lsp
python3 lumbda.py examples/objects.lsp
;; Generator using full continuations
(auto-compile! #t)
(define (make-gen thunk)
(let ((k #f) (done #f))
(lambda ()
(if done 'done
(call/cc (lambda (return)
(if k (k return)
(begin (thunk (lambda (val)
(call/cc (lambda (next)
(set! k next) (return val)))))
(set! done #t) (return 'done)))))))))
(define counter (make-gen (lambda (yield)
(let loop ((i 0)) (yield i) (loop (+ i 1))))))
(counter) ; => 0
(counter) ; => 1
(counter) ; => 2
Running tests & benchmarks
make test # run 529 tests
make test-verbose # verbose output
make bench # compare interpreter vs bytecode vs CPython
make lint # syntax check all Python files
Portal — machine state migration
Serialize a running VM mid-computation, transfer to another machine, resume:
# Machine A: start a long computation with checkpoints
python3 lumbda.py --fast examples/portal-prime.lsp
# saves prime-state.portal at checkpoint
# Machine B: resume from checkpoint
python3 lumbda.py --portal-resume prime-state.portal
# continues from exact instruction
The portal captures the full env chain, compiled procedures, continuations, and frame stack as JSON. 16KB for a primality test in progress.
EML universality proof
The proof/ directory contains a formal verification that eml(x,y) = exp(x) - ln(y)
with constant 1 generates all elementary functions (arXiv:2603.21852v2).
Three approaches, benchmarked:
| Approach | Time | Guarantee |
|---|---|---|
| Python (numerical) | 0.04s | 1e-10 tolerance |
| Lumbda (numerical) | 59s | 1e-10 tolerance |
| Lean 4 (formal proof) | 1.5s | kernel-verified |
The formal proof is 40x faster than brute-force search with infinitely stronger
guarantees. See proof/benchmark_results.md for the full analysis — including
why this is MOAD-0001 (the sedimentary defect) at the proof methodology layer.
File layout
lumbda.py interpreter + bytecode compiler (one file, ~3200 lines)
stdlib.lsp extended standard library
tests.py test suite (571 tests)
bench.py benchmarks vs CPython
examples/ example programs
proof/ EML universality proof (Python, Scheme, Lean 4)
Makefile make test / make bench / make repl