Adds §6.6.3 "Collaborative Meta-GC: From Greedy to Adaptive" with
the three-workload benchmark (friendly / hostile / mixed × greedy
/ adaptive). Honest read of the numbers:
friendly greedy 732 ms 1000 resets, 0 escapes
friendly adaptive 691 ms 1000 resets, 0 escapes (-6%)
hostile greedy 568 ms 0 resets, 1000 escapes
hostile adaptive 607 ms 0 resets, 1000 escapes, 11 skipped (+7%)
mixed greedy 1981 ms 17 resets, 1983 escapes
mixed adaptive 1694 ms 14 resets, 1986 escapes, 2 skipped (-17%)
Adaptive wins on friendly (-6%) and mixed (-17%, the policy's
design target). On fully hostile workloads implicit GC fires 982
of 1000 arenas before the dispatcher sees them, so the signal is
drowned and greedy happens to edge adaptive by ~7%. Section
explicitly calls out the collaborative-but-local structure
(shared state on arena_active + EMA + countdown, decisions made
locally by each component) and credits the benchmark work with
surfacing two real correctness bugs in the conservative stack
scan — 24-byte strings misread as env nodes, 40-byte strings
misread as 25-element vectors — both now fixed.
Also:
- meta-gc-policy.dot rewritten to show the adaptive gate
(rate > 50% + probe countdown) before the greedy verify path;
new skip branch, new EMA annotations on edges.
- §6 reproducibility list + Makefile bench-gc-adaptive target.
- PDF rebuilt at 2.64 MB.
137 asm no-GC + 137 asm GC + 189 shared functional tests pass
against the new asm.
|
||
|---|---|---|
| .. | ||
| asm-architecture.dot | ||
| asm-architecture.png | ||
| benchmark-ack.dot | ||
| benchmark-ack.png | ||
| benchmark-binary-size.dot | ||
| benchmark-binary-size.png | ||
| benchmark-fib.dot | ||
| benchmark-fib.png | ||
| benchmark-gc.dot | ||
| benchmark-gc.png | ||
| benchmark-speedup.dot | ||
| benchmark-speedup.png | ||
| benchmark-sumto.dot | ||
| benchmark-sumto.png | ||
| c-architecture.dot | ||
| c-architecture.png | ||
| gpu-architecture.md | ||
| jit-pipeline.dot | ||
| jit-pipeline.png | ||
| meta-gc-policy.dot | ||
| meta-gc-policy.png | ||
| python-architecture.dot | ||
| python-architecture.png | ||
| README.md | ||
uncommonlisp Architecture Documentation
"A diagram is worth 10,000 words." — russell@unturf.com
Three implementations of the same Scheme language, sharing the same .lsp test files.
Python Implementation (uncommonlisp.py)
3,324 lines. Bytecode compiler + stack VM + full continuations + portal.
Execution tiers:
- Tree-walker (
leval): default, handles all forms including macros - Bytecode VM (
--fast): 40 opcodes + superinstructions, 7-19x faster - Python JIT prototype: exec()-based transpilation (labeled as prototype)
Key features:
- Full multi-shot continuations via explicit frame stack
- Portal: serialize VM state to JSON, resume on another machine
- Inline cache, constant folding, peephole optimizer
- Source maps for error reporting with line numbers
- Bytecode serialization (.lspc files)
Tests: 571 unit + integration tests (tests.py)
C Implementation (c/)
8,120 lines. Tree-walker + bytecode VM + x86_64 JIT.
Execution tiers:
- Tree-walker: default, full special form support
- Bytecode VM (
--fast): matching Python's opcodes - x86_64 JIT (
--jit): 10-24x faster than CPython
Key features:
- NaN-boxed 64-bit values (zero-alloc numbers)
- Hash-map environments with parent chain + global shortcut
- Interned symbols
- Real JIT: mmap(PROT_EXEC) + raw x86_64 bytes
Tests: 76 unit + integration + JIT tests (test.c)
Assembly Implementation (asm/)
2,592 lines of GNU assembler. 13KB binary. Zero dependencies.
Design:
- No C. No libc. Only Linux syscalls (read, write, mmap, exit)
- Tag-in-low-3-bits value representation
- Bump allocator on 64MB mmap'd page
- TCO via
jmp .eval_top(never grows the stack) - 34 builtins, all special forms
Tests: 75 unit + integration + functional tests (test.sh)
JIT Pipeline (c/jit.c)
1,309 lines. Compiles Scheme AST directly to x86_64 machine code.
What gets JIT'd:
- if, cond, and, or (conditional jumps)
- +, -, *, =, <, >, <=, >= (native integer ops)
- let, let* (stack-allocated locals)
- Named-let loops (native jmp, zero call overhead)
- car, cdr, cons, null?, pair? (NaN-box pointer ops)
- Self-recursive calls (call/ret) and tail calls (jmp)
What falls back to interpreter:
- call/cc, macros, syntax-rules, quasiquote, modules
- String/vector/hash-table operations
- Any form the AST analyzer can't verify as integer-safe
Performance Summary
All benchmarks measured in-process (no startup overhead) on the same machine.
| Implementation | ack(3,4) | fib(35) | sum-to(50k) | Binary |
|---|---|---|---|---|
| C + x86_64 JIT | 0.19ms | 0.09ms | 0.55ms | 171KB |
| CPython (native) | 1.3ms | 0.006ms | 5.5ms | ~5MB |
| C interpreter | 20ms | 0.06ms | 109ms | 171KB |
| Python bytecode VM | 149ms | 0.75ms | 437ms | 3,324 lines |
| Assembly (13KB) | ~8ms* | ~0.6ms* | ~43ms* | 13KB |
*Assembly times include process startup + tokenizer + parser.
The JIT runs Scheme faster than CPython runs Python on recursive workloads: ack(3,4) is 7x faster, sum-to(50k) is 10x faster. The JIT compiles Scheme AST directly to x86_64 machine code via mmap(PROT_EXEC).
Test Coverage
943 verified assertions across all implementations:
make test-all
Python unit/integration: 571 tests
C unit/integration/JIT: 83 tests
Assembly unit/int/func: 108 tests
Shared functional: 181 tests (Python + C)
Total: 943 assertions
The language is R7RS Scheme. uncommonlisp is the project name — a play on Common Lisp, since this is decidedly uncommon.



