Previous commit used `all: revert` on the floating button to cancel
docutils's blanket `body > *` rule. Too aggressive — it also removed
uncloseai's own `position: fixed`, `bottom`, `right`, `width`,
`background`, `color`, `border` declarations, making the button
disappear entirely.
Replace with a tight override: only the five properties docutils's
blanket rule sets (background-color, line-height, padding, margin,
max-width) get reset with !important. Everything else falls through
to .uncloseai-floating-button's own stylesheet, so the button
re-appears in its designed 121px pill, bottom-right, black-on-white.
Any <img> inside <main>/<section> now opens in a full-viewport
overlay (92% black bg, image centered at max-width/max-height). Click
the overlay or press Esc to close. Pure vanilla JS + CSS, injected
inline via inject-whitepaper-css.py — keeps the HTML self-contained
with no external lightbox dependency. cursor: zoom-in signals
interactivity; the overlay uses cursor: zoom-out.
Three whitepaper-HTML-only fixes:
* docutils's responsive.css blanket-pads every `body > *` with
`padding: 0.5rem calc(29% - 7.2rem)`. The uncloseai floating
button gets appended to <body>, picks up that padding, and blows
from its designed 121 px to nearly half the viewport. Scoped
override resets the cascade on the button selector with
`all: revert` + explicit zero padding/margin.
* whitepaper h1/h2 titles now use ChunkFive (same face as the
homepage logo) instead of docutils's default serif. Font embedded
as a base64 data: URI inside a <style> block so the HTML stays
self-contained — no external font fetch.
* RST layout: λ logo now sits UNDER the "Lumbda" heading instead
of above it. Permacomputer logo still follows below the λ.
New helper: whitepaper/inject-whitepaper-css.py runs after
embed-images.py, reads ChunkFive from www/fonts/..., base64-encodes,
writes the <style> block before </head>. Makefile wires it in.
Source PNG from the MPS lumbda-logos product is already an inverted λ
(hooks at top-left + bottom-right, V-body). My prior commit applied
vertical + horizontal flips on top of that, which undid the inversion
and produced a right-side-up calligraphic λ — the opposite of what
the artwork was drawn to express.
Regenerate both PNG variants (lumbda-logo.png black,
lumbda-logo-green.png #227842) directly from the source with no
flips. Only transform: alpha-from-brightness so anti-aliased edges
survive over any background.
Homepage img src picks up ?v=2 so any browser with the flipped
version in disk cache re-fetches. Logo size stays 3.33× the
wordmark font (10.66rem). PDF + HTML rebuilt — embed-images.py
base64-inlines the new PNG into the HTML automatically.
Fetched the lumbda-logos artwork from
media.unturf.com/c/fbd3d473-.../lumbda-logos (a MakePostSell product),
applied both flips (vertical + horizontal, = 180°) at bake time so the
PNG ships oriented correctly without any CSS transform dance, and
recolored non-background pixels to brand green (#227842) with alpha
derived from pixel brightness so anti-aliased edges stay smooth.
Two PNG variants ship under whitepaper/diagrams/:
* lumbda-logo.png dark ink for print contexts
* lumbda-logo-green.png #227842 for web
www/ carries symlinks to both.
Homepage (www/index.html): replaces the CSS-rendered λ with an
<img class="lambda-mark"> element sized 9.6rem — 3× the 3.2rem
wordmark font — stacked below "lumbda." on its own line.
Whitepaper (whitepaper/lumbda-whitepaper.rst): adds the logo as the
first figure above the permacomputer-logo, 24% width, centered.
Regenerated PDF + HTML; embed-images.py base64-inlines the new file
automatically so the HTML stays single-file.
CI (.gitlab-ci.yml): generalizes the symlink resolver from two
explicit `cp -L` calls into a `find www -type l` loop, so every
current and future symlinked asset deploys without per-file CI edits.
λ in front was reading as a "Y" attached to the wordmark ("ylumbda.").
Move the mark under the wordmark as a standalone logo below the name.
Sitting alone, the flipped λ doesn't get recruited into the first
letter and reads as its own glyph — a lambda turned on its head.
Collapse --accent (was blue #0b5394) into the same green as the
period. One color now marks everything that says "lumbda": the λ
logo, the declarative period, link underlines, and the CTA buttons.
Dark-mode variant lightened to #5ec07a so the green reads on
#161613.
Previous design folded the flipped λ inside the wordmark as the "l".
That read as a cute substitution but didn't work as a brand mark.
Split them: the λ now stands on its own, rotated upside-down, 5rem
(larger than the 3.2rem wordmark), so it reads as a logo symbol
separate from the text.
Wordmark restores a proper lowercase "l" in "lumbda" — readable,
copy-pasteable, searchable. Trailing period picks up a new --green
design token (#2a8a3a light / #6fd68a dark) so the declarative stop
also reads as growth — matches the permacomputer ethos.
Layout: flex logo-row with baseline alignment and a small gap; the λ
hangs from the wordmark's cap-line for a balanced mark-and-wordmark
composition.
Previously the HTML whitepaper referenced diagrams via relative paths
(src="diagrams/X.png"), which 404'd because the docroot does not carry
a mirror of whitepaper/diagrams/. The "HTML whitepaper" thus showed
broken image stubs instead of the architecture figures.
Fix: a post-processor (whitepaper/embed-images.py, stdlib Python) runs
after rst2html5. Every <img> with a relative src gets base64-inlined
as a data: URI with inferred MIME. Any .svg reference would splice in
as an inline <svg> element, surfacing alt text as <title>; the current
RST only uses the .png variant of gnu-logo so no SVGs inline for now,
but the path works when we switch.
Makefile's whitepaper-html target chains the embed step after the
existing sed passes for the uncloseai.js script tag and the meaningful
<title>. Title tweaked to "feedback as a primitive" to match the
homepage tagline.
Result: whitepaper/lumbda-whitepaper.html grows from 180 KB to 4.6 MB
(base64 inflation on ~2 MB of diagrams), and it now opens offline as
a single file — no network fetches for figures, no docroot mirror
needed.
Wordmark now renders λumbda. — flipped λ stands in for the "l",
visual pun on lambda (the primitive the language shells around),
trailing period as declarative closure. All-lowercase, accent-colored
period. aria-label="lumbda." keeps screen readers reading the name
correctly; the λ and the period both carry aria-hidden.
CLAUDE.md rewritten to avoid the verb "to be" and to prefer "a"/"our"
over "the" when a thing counts as one of many or as shared. Same rule
applied to www/index.html (tagline now "feedback as a primitive", body
prose cleaned of is/are/be, table header now "What a tier buys",
section heads trimmed of definite articles). Monospace body stays,
ChunkFive titles stay.
Whitepaper now ships in two formats:
* Makefile target `make whitepaper-html` runs rst2html5 with
embedded minimal.css + responsive.css (docutils bundled), then
injects the same uncloseai.js module script that the homepage
carries, and stamps a meaningful <title>. 180 KB self-contained.
* PDF build target renamed to `make whitepaper-pdf`;
`make whitepaper` now builds both.
www/lumbda-whitepaper.html symlinks to the generated HTML (same
pattern as the PDF symlink). .gitlab-ci.yml resolves both symlinks to
real files before rsync so the proxy gets byte-identical copies.
Homepage now exposes both formats side-by-side via two .cta buttons
and the footer lists HTML + PDF.
Rename www/whitepaper.pdf → www/lumbda-whitepaper.pdf (symlink target
unchanged; slug now descriptive for search/citation). Update the CTA
and footer links in index.html, plus the .gitlab-ci.yml symlink-resolve
step. Companion Caddyfile redirect in proxy.unturf.com keeps old
/whitepaper.pdf links working via 301.
Add <script src="https://uncloseai.com/uncloseai.js" type="module"> to
<head>, matching www.unturf.com + timehexon.com pattern — brings the
permacomputer chat/AI integration to lumbda.com's front door.
ChunkFive webfont (woff2 + woff, ~43 KB total) dropped into
www/fonts/chunkfive/; style.css adds @font-face with font-display:swap
and applies the family to header h1 (3.2rem logo) and main h2
(1.5rem section titles). Monospace body text unchanged. Font files
reused from www.unturf.com; SIL OFL.
index.html: "x86_64 assembly interpreter" and table-row labels
"Pure x86_64 assembly" / "Assembly + naive mark-sweep GC" now read
"GNU asm" — precise and consistent with how we talk about the tier
elsewhere.
Resolves the www/whitepaper.pdf symlink to a real file before rsync
(deploy-www.sh preserves symlinks, which would land broken on the
proxy), writes version.json with the commit SHA, then calls
deploy-www.sh lumbda www via sudo. Runs on the proxy.uncloseai.com
runner, same pattern as cuppcb.com and the rest of the static fleet.
asm-gc gains (tcp-sendfile socket path) → builtin (90 lines) that issues
SYS_SENDFILE(40) in a loop, streaming a file from fd → socket with no bounce
through the Lumbda heap. Zero-copy kernel path for large responses.
examples/http-static-server-sendfile.lsp (hybrid): small assets
(≤ 16 KB) stay inline-cached as full HTTP responses; large assets cache
only headers and stream the body via tcp-sendfile. 4-way race on
i5-8350U, 100 PDF requests (2.56 MiB), concurrency 8:
uncached 159 req/s 406 MiB/s 15.5 MB RSS
cached 238 req/s 603 MiB/s 7.2 MB RSS
sendfile 480 req/s 1226 MiB/s 4.2 MB RSS
caddy 485 req/s 1238 MiB/s 37.1 MB RSS
sendfile lands within 2% of caddy on throughput with 9x less peak RSS in
a 27 KB binary vs caddy's 38 MB (1400x smaller).
examples/http-static-server-adaptive.lsp (learning preload): per-URL hit
counter persisted to www.hits every N requests. At boot, ranks and
preloads top *cache-max* URLs from the prior run's data (cold-start
falls back to a seed list). Cold requests beyond the seed promote on
first hit. Drops heap-restore arena pattern since the server mutates
persistent state every request; relies on GC build's mark-sweep.
tests/bench-www-race.sh: adds sendfile variant on port 8083, auto-sizes
PDF byte count from the on-disk whitepaper so a whitepaper rebuild
doesn't desync the MiB/s calc.
Whitepaper §11.7 "Static File Serving: Cache, Sendfile, and Adaptive
Preload" documents the four variants, benchmark table, and the
arena-vs-mutation tradeoff. §13 Future Work adds DAG-of-hot-paths
predictive preload as the direction for > 1000-resource deployments
where frequency-only ranking is too narrow.
examples/http-static-server-cached.lsp — same HTTP server but every
preloaded URL's full HTTP/1.0 response (headers + body) is composed
once at startup and stored in a hash-table, so the per-request
handler is a single hash-table-ref/default. No file->string, no
string-append, no MIME lookup in the hot path.
Config knobs:
*docroot* filesystem root (default "www")
*cache-max* soft cap on cached entries (default 100)
*preload-paths* list of URL paths to pre-fetch at startup
Preload list for lumbda.com: "/", "/style.css", "/whitepaper.pdf",
"/robots.txt", "/404.html" — anything not in the list returns the
cached 404 response (no disk hit). Cache lives pre-snapshot so
heap-restore never reclaims it; RSS stays at the cache size
forever.
Race vs caddy on this laptop (2000 small / 200 large, concurrency 8):
small req/s PDF req/s PDF MiB/s peak RSS
lumbda-www uncached 690 195 495 15.5 MB
lumbda-www cached 686 319 811 7.2 MB
caddy file-server 688 478 1,217 38.9 MB
Cache wins on the PDF: 1.64x faster than uncached, RSS DROPS from
15.5 MB to 7.2 MB because the cached path allocates nothing per
request (all allocation happened pre-snapshot). On small files
already-hot paths mean the cache is a wash — 690 vs 686 is noise.
Caddy still wins 1.5x on the PDF via sendfile(2) zero-copy; we
allocate the 2.67 MB response once at startup and tcp-send it.
Closing the gap further would take a sendfile asm primitive —
separate project. For a minimal static site serving its own
whitepaper, the cached 27 KB asm binary is viable: 319 req/s
and 811 MiB/s with 5x less memory than caddy.
tests/bench-www-race.sh updated to run all three side-by-side
(uncached + cached + caddy) at three ports. Cached server's port
is patched via sed at the entry point so the two lumbda variants
don't collide. PDF byte-integrity checked on all three paths.
Ships examples/http-static-server.lsp — ~65 lines of portable Scheme
that reads files from a docroot (default ./www) and serves them over
HTTP/1.0 with MIME dispatch, path-traversal rejection, heap-snapshot
per request. Runs in any tier; target deployment is asm-gc for the
27 KB stripped binary + bounded memory backstop.
Required one asm fix first: heap_grow was mmap'ing fixed HEAP_SIZE
chunks, so any single allocation larger than a chunk (notably the
2.67 MB whitepaper PDF read via file->string) loop-looped through
.ha_overflow forever. Now heap_grow rounds required bytes up to
HEAP_SIZE multiples on oversize alloc, so a big request carves its
own big chunk in one go. Small allocs still land in standard-sized
chunks.
Two new benches:
tests/bench-lumbda-www.sh — drive N small + M large requests against
asm-gc, verify PDF round-trip, sample peak RSS. At 1000/100: 331 req/s
small, 120 req/s large (304 MiB/s), peak 15.5 MB.
tests/bench-www-race.sh — adjacent A/B vs caddy v2.5.1 on the same
docroot. Numbers on this laptop, concurrency 8, 2000 small + 200 large:
small req/s PDF req/s PDF MiB/s peak RSS binary
lumbda-www (asm-gc) 375 137 349 7–16 MB 27 KB
caddy file-server 358 231 588 38 MB 38 MB
Reading: lumbda edges caddy on small files (less per-request overhead),
caddy wins 1.7x on large files (sendfile zero-copy; we allocate the
whole file into a string and write it with one syscall). Both byte-
identical on the PDF. Memory: lumbda 2.5-5x less at steady state.
Binary size: 1400x smaller (27 KB vs 38 MB).
Feature gap: caddy has HTTPS, HTTP/2, range, middleware, etc. lumbda
has none of that yet — but for the specific job of serving lumbda.com's
six-file docroot it is viable right now.
Makefile adds `bench-lumbda-www` and `bench-www-race` targets.
137 asm no-GC + 137 asm GC tests still pass.
Minimal, monospace, readable. Works as a drop-in docroot for
any static server (nginx, caddy, python -m http.server, etc.).
Light/dark mode via prefers-color-scheme. No JS, no fonts, no
third-party anything.
Files:
www/index.html landing page (tagline, get-it, four-tier
table, portal summary, EML proof blurb,
license, whitepaper CTA)
www/style.css ~150 lines, CSS vars for theming
www/robots.txt allow everything
www/404.html referenced by servers that support custom
error pages
www/whitepaper.pdf -> ../whitepaper/lumbda-whitepaper.pdf
(symlink — one source of truth, rebuilds
via `make whitepaper` auto-propagate)
Smoke-tested with python3 -m http.server: 200 on /,
Content-Type: application/pdf for /whitepaper.pdf,
Content-Length matches the current 2.67 MB PDF,
/robots.txt serves, /nonexistent returns 404.
No new build step — docroot is pure static files the existing
whitepaper target already produces. A web server configured
with docroot=www/ and fallback 404.html has a working lumbda.com
today.
Completes the on-disk side of the uncommonlisp -> lumbda rename:
Filesystem: /home/fox/git/uncommonlisp -> /home/fox/git/lumbda
with a back-compat symlink
/home/fox/git/uncommonlisp -> lumbda
so any stale path reference (agent memory files,
shell history, other sessions) still resolves.
tests.py: cwd='/home/fox/git/uncommonlisp' -> '/home/fox/git/lumbda'
(the only hardcoded absolute path we left behind in
the previous rename commit, because the directory
itself hadn't moved yet).
Remote URL (git@git.unturf.com:engineering/unturf/uncommonlisp.git)
still points at the old name and needs to be flipped AFTER fox
renames the gitlab project — probe confirms the new URL currently
404s, so the flip waits for the gitlab rename to land.
Verified 571 Python tests pass under the new cwd; symlink lets
`cd /home/fox/git/uncommonlisp` still work for anything cached.
Historical internal name "uncommonlisp" retired in favor of the
public name "lumbda" ahead of lumbda.com going live. Scope of
this commit:
Source files renamed:
uncommonlisp.py -> lumbda.py
asm/uncommonlisp.s -> asm/lumbda.s
c/uncommonlisp.h -> c/lumbda.h
whitepaper/uncommonlisp-whitepaper -> whitepaper/lumbda-whitepaper (.rst + .pdf)
Binaries renamed (tracked ones; c/ was always gitignored):
asm/uncommonlisp, asm/uncommonlisp-gc, asm/uncommonlisp.o,
asm/uncommonlisp-gc.o -> asm/lumbda(-gc)(.o)
c/.gitignore -> ignores lumbda
Internal string updates (sed pass ordered longest-first):
asm/uncommonlisp -> asm/lumbda
c/uncommonlisp -> c/lumbda
uncommonlisp.py -> lumbda.py
UNCOMMONLISP_BIN -> LUMBDA_BIN (asm/test.sh env var)
"uncommonlisp> " -> "lumbda> " (asm REPL prompt baked into binary)
UNCOMMONLISP -> LUMBDA (macros, comments)
uncommonlisp -> lumbda (prose)
Binary portal magic updated:
"ULPORTAL" -> "LUMBDAB1" # "Lumbda Binary v1"
Old portal files are not backward-compatible — this is a deliberate
break since it's the rename moment. S-expression portals already
carry their own ";; lumbda-portal v1" header and remain cleanly
versioned.
WHITEPAPER.pdf / WHITEPAPER.rst symlinks repointed to the renamed
files. Makefile's whitepaper target targets lumbda-whitepaper.pdf.
Not changed (intentional, separate phases):
- Filesystem directory /home/fox/git/uncommonlisp itself
(fox renames locally and the gitlab repo URL in a follow-up)
- tests.py hardcoded cwd=/home/fox/git/uncommonlisp
(matches the current on-disk location; will flip when the
directory rename ships)
- Git history (immutable; old commits still say uncommonlisp,
which is correct — that's what they were)
Verified:
137 asm no-GC + 137 asm GC + 571 Python + 83 C + 189 shared
functional tests all pass under the new names.
bench-gc-http (2000 req): all 4 cells behave as expected
(cells 1/2 flat, 3 leaks, 4 bounded at 1 chunk).
Python REPL, C REPL, asm REPL all start cleanly.
GC-build portal files now start with ";; lumbda-portal v1\n". The
line is a Scheme comment the reader already skips, so loading a
v1 file via bi_load works unchanged. What's new is that
portal-resume actively validates the header before delegating to
bi_load:
file starts with ";; lumbda-portal v1\n" -> load normally
file starts with ";;" but different text -> return #f (rejected)
file does not start with ";;" at all -> load as legacy (back-compat)
The check reads the first 32 bytes of the file, compares the
first two bytes against ";;", and on match compares the full
20-byte v1 prefix. Closes the versioning-friction concern a SEW
reviewer raised after reading §7.4.1: we can now add a v2 format
with new syntax (complex numbers, records, whatever) without
older consumers silently parsing new files into garbage — they
will cleanly return #f.
Verified:
v1 portal -> resume loads all bindings, returns #<void>
v2 portal -> resume returns #f without evaluating any forms
legacy portal (no ;; header) -> resume loads normally
missing file -> resume returns #f
The deeper framing worth writing down: this is a migration format
for handoff across process / tier / machine, not an archive format.
If archival becomes a real use case it earns its own format with
proper schema evolution and a builtin-rename table. The v1 tag is
the minimum hook that lets v2 happen cleanly when someone needs it.
137 asm no-GC + 137 asm GC still pass. §6.6.4 HTTP cells still
green at 10K: GC + no-snapshot at 543 req/s, peak 1.1 MB, 972 KB
growth (one chunk, steady state).
Root-cause fix for the residual crashes I had documented as known
issues in §6.6.4. Every heap_alloc call site was setting its type
byte with `orq $(HT_X << 8), -8(%rax)` — but OR merges with the
stale type byte from a free-list-reused block. A pair previously
used as a vector (type 5 = 0b101) re-allocated as pair (type 1 =
0b001) ends up with merged type 0b101 = still vector. Walker then
treats the pair as a vector, reads the pair's car as a "length",
and walks off the block end — hence the hash-set bench's "unbound
variable: t", memory bench's "unbound variable: lst", arena bench's
"unbound variable: k".
Fix: overwrite the byte instead of OR-ing. 18 sites converted from
`orq $(HT_X << 8), -8(%rax)` to `movb $HT_X, -7(%rax)`. Every
previously-residual crash gone on first rerun.
Refreshed benchmark numbers throughout §6.6:
§6.6 Memory table: 122× less memory at 26% slowdown (was 124×,
30%). Range shifted because the fix also accelerated the common
paths; ratio stable.
§6.6.4 HTTP soak at 50,000 requests × 16 concurrent × 4 cells:
no-GC + snapshot 630 req/s peak 100 KB growth 4 KB
GC + snapshot 633 req/s peak 120 KB growth 4 KB
no-GC + no snapshot 625 req/s peak 458 MB OOM at cap
GC + no snapshot 610 req/s peak 1,092 KB growth 852 KB
Cell 4 now sustains 50K requests with steady-state 1-chunk memory.
Previous residual edge at 50K (cell 4 failing to start) was a
manifestation of the same type-byte bug, now gone.
§6.6.3 Adaptive numbers collapsed to within ~1% across all three
workloads (was 6% / 7% / 17% deltas). Paper updated to honestly
report adaptive as a null experiment on these shapes — neutral
cost, same stats surface, default on.
§6.6 diagram: bench-gc.png refreshed to match new numbers.
137 asm no-GC + 137 asm GC + 189 shared functional all pass.
Hash-set / memory / arena / adaptive / HTTP benches all clean.
Extended the HTTP-under-GC validation from the original 5,000
requests to 20,000 requests per cell. All four cells still hold:
asm no-GC + snapshot ~240 req/s peak 96 KB growth 0 KB
asm GC + snapshot ~235 req/s peak 120 KB growth 0 KB
asm no-GC + no snap ~470 req/s peak 185 MB growth 185 MB
asm GC + no snap ~450 req/s peak 1,084 KB growth 972 KB
Cell 4 (GC + no snapshot) is the real validation target. 972 KB of
growth over 20,000 requests = one chunk filled once, after which
the collector cycles through reclaimed space. No monotonic leak,
no OOM, no crash. Naive mark-sweep is now a correct (not optimal)
allocator for long-running asm servers that don't manage arenas.
The paper's soak commentary explicitly calls this out and also
notes two remaining rough edges we've observed but not yet
debugged: (a) at 50k requests with the no-GC + no-snap case
saturating the 512 MB vcap right before cell 4 starts, cell 4's
server sometimes fails to initialize — looks process-environment
rather than GC, no clean explanation yet; (b) the hash-set
benchmark on the GC build still surfaces an occasional
unbound-variable error at ~1 MB/iter workloads.
Reproduce: make bench-gc-http defaults to 5k now; soak uses
`REQUESTS=20000 CONCURRENCY=16 bash tests/bench-gc-http.sh`.
Binary heap dump can't work under GC because the heap is a linked
chunk list with typed block headers and a free list. Raw-byte
serialization would lose structure. Rather than invent portal v2
with chunk tables and pointer relocation, the GC build uses the
S-expression format that already works across all other tiers:
# bi_portal_save in GC build:
walk %r14 (env chain); for each non-builtin, non-closure binding,
emit `(define <sym> (quote <val>))` to the opened file via
scheme_print with output_fd redirected to that fd.
# bi_portal_resume in GC build:
jmp bi_load — read every form from the file, eval each in %r14.
The quote wrapper makes data values round-trip cleanly: lists,
vectors, strings, symbols, numbers, pairs all re-read as literals.
Closures and builtins are explicitly skipped — closures can't
faithfully re-read from their printed form; builtins reconstruct
from the target's prelude. Same treatment the JSON portal gives.
The no-GC build keeps the binary portal format unchanged (wrapped
in .ifndef GC_NAIVE). Users get the fast format on the fast build,
the portable format on the safe build. Same API, different wire
format by build.
Cross-tier verified: GC-asm producer -> Python consumer passes
with `x=42`, `nums=(1 2 3 4 5)`. The §7.2 cross-impl matrix
expands from 9 to 16 cells, all green.
Whitepaper §7.4 now notes it's the no-GC format; new §7.4.1
documents the GC build's S-expression portal with the trade-off
(slower than binary dump, stricter about what round-trips, but no
architecture constraint and no "same binary" requirement).
137 asm no-GC + 137 asm GC + 189 shared functional tests pass.
All four HTTP cells from §6.6.4 still bounded under sustained
load (GC + no-snapshot at ~630 req/s peak, 1.1 MB steady state).
Replaces header format from [size:63 | mark:1] with
[size:48 | type:8 | flags:8 (mark in bit 0)]. Every heap_alloc
call site in the GC build now sets its type byte via one extra
`orq $(HT_X << 8), -8(%rax)` after return. Ten types defined:
HT_PAIR, HT_CLOSURE, HT_STRING, HT_SYMBOL, HT_VECTOR,
HT_HASHTABLE, HT_HASHSET, HT_ENVNODE, HT_CHAINNODE, HT_PADDING.
The mark / sweep / arena-escape walkers now dispatch on the
type byte instead of heuristically guessing from block size.
Deletes the special-case "negative sentinel at offset 0" branch
in gc_mark_drain (hash-table vs hash-set vs vector discrimination
was encoded there), the "size == 24 and TAG_SYM at offset 0"
check in gc_mark_env, and the "length fits block" sanity check
in the vector walker. All that logic collapses into a single
compare on the type byte.
Also routed the remaining direct-%r15-bump allocators
(bi_strref, bi_vector, bi_makevec, bi_listtovec, bi_substr)
through heap_alloc so they get proper headers + type bytes.
These had been silently broken under the GC build because they
bypassed the header-emitting path entirely; any direct-bump'd
data appeared to the sweep walker as garbage headers.
§6.6.4 cell 4 (asm GC + no snapshot) was crashing at first GC
before this change. After: serves 5,000 HTTP requests at ~410
req/s, peak RSS 1,088 KB (one chunk), growth 972 KB — the
collector hit its natural steady state. First time we've
validated "naive GC as replacement for snapshot discipline"
under real traffic.
New §6.6.5 "Precise Block Typing" in the whitepaper documents
the old heuristic bugs, the new header format, and the cost
(one orq per alloc, 16 header bits) vs benefit (class of bugs
eliminated). Updated §6.6.4 to reflect cell 4 passing.
Remaining known issue: the hash-set bench on the GC build under
very heavy sustained allocation still surfaces an occasional
unbound-variable error. The precise-type fix addressed the
observed HTTP crash; a deeper root-scan edge case remains.
Tracked for Fix 2 work.
137 asm no-GC + 137 asm GC + 189 shared functional tests all
pass.
New infra:
- examples/http-server-noarena.lsp: same HTTP server minus the
heap-snapshot/heap-restore arena loop. Isolates whether the GC
build actually holds memory under real traffic, independent
of the portable snapshot pattern.
- tests/bench-gc-http.sh: drives 5,000 concurrent requests per
cell across the full 2x2 matrix {no-GC, GC} x {snapshot, no}.
- Makefile: new `bench-gc-http` target.
Extended benches to exercise both asm binaries:
- tests/bench-hashset.sh now runs against both asm/uncommonlisp
and asm/uncommonlisp-gc, with set +e so a GC-build crash on
one workload doesn't abort the other.
- tests/web-benchmark.sh adds a dedicated asm-gc row (and prints
its stripped binary size) so the HTTP throughput comparison
reports both.
Whitepaper updates:
- §6.6.4 "Validation: HTTP Server Under Sustained Load" — the
4-cell memory matrix. 3/4 cells green; cell 4 (GC + no
snapshot) crashes at first GC trigger — another instance of
the conservative-scan type-confusion class we already fixed
once at the env/string boundary. Logged as a known issue
rather than shipping a partial fix under time pressure.
heap-snapshot + heap-restore remains the recommended pattern
for production asm code; the naive GC is diagnostic + control
group, not a replacement for the arena discipline.
- §6.5 hash-set speedup table slightly softened to ~15-20x (was
15-21x) since run-to-run noise on a shared laptop shifts the
per-phase ratio by a few percent. Ratio is stable to first
order.
- §8.6 narrative references the ~1280x symbolic-vs-brute-force
figure instead of the stale 40x.
- §6 reproducibility list now lists `make bench-gc-http`.
All 137 asm no-GC + 137 asm GC + 189 shared functional tests
still pass.
Old table only showed three rows, one of which (Lumbda numerical
brute-force at 59s) had been superseded months ago by the native
symbolic rewriter and cached-replay path already described earlier
in §8. §8.6 lagged and kept claiming "40x faster than brute-force"
when the symbolic-vs-brute-force win is actually ~1,280x.
New six-row table contrasts:
- Python numerical 0.04 s
- Lumbda brute-force 59 s (kept for historical scale)
- Lumbda symbolic cold 46 ms (~1,280x over brute-force)
- Lumbda symbolic cached 7 ms
- Lean 4 cold rebuild 722 ms
- Lean 4 cached 5 ms
Three wins compound: symbolic over numerical (~1280x, MOAD-0001
at proof-methodology layer), asm over Lean's cold binary startup
(~16x), and cached replay over cold on both sides (~100x).
Explicitly notes that Lean's kernel TCB stays smaller even when
timings equalize (~3 KLOC audited elaborator vs ~6.6 KLOC asm
interpreter) — right tool for different assurance levels.
References existing make bench-proof target rather than adding
new plumbing; the numbers already come from proof/benchmark.sh.
Adds §6.6.3 "Collaborative Meta-GC: From Greedy to Adaptive" with
the three-workload benchmark (friendly / hostile / mixed × greedy
/ adaptive). Honest read of the numbers:
friendly greedy 732 ms 1000 resets, 0 escapes
friendly adaptive 691 ms 1000 resets, 0 escapes (-6%)
hostile greedy 568 ms 0 resets, 1000 escapes
hostile adaptive 607 ms 0 resets, 1000 escapes, 11 skipped (+7%)
mixed greedy 1981 ms 17 resets, 1983 escapes
mixed adaptive 1694 ms 14 resets, 1986 escapes, 2 skipped (-17%)
Adaptive wins on friendly (-6%) and mixed (-17%, the policy's
design target). On fully hostile workloads implicit GC fires 982
of 1000 arenas before the dispatcher sees them, so the signal is
drowned and greedy happens to edge adaptive by ~7%. Section
explicitly calls out the collaborative-but-local structure
(shared state on arena_active + EMA + countdown, decisions made
locally by each component) and credits the benchmark work with
surfacing two real correctness bugs in the conservative stack
scan — 24-byte strings misread as env nodes, 40-byte strings
misread as 25-element vectors — both now fixed.
Also:
- meta-gc-policy.dot rewritten to show the adaptive gate
(rate > 50% + probe countdown) before the greedy verify path;
new skip branch, new EMA annotations on edges.
- §6 reproducibility list + Makefile bench-gc-adaptive target.
- PDF rebuilt at 2.64 MB.
137 asm no-GC + 137 asm GC + 189 shared functional tests pass
against the new asm.
Moves the meta-GC from greedy (always verify) to adaptive: track a
scaled EMA of recent escape rate; when rate exceeds 50% (128/256),
SKIP the verifier and let the heap grow until natural GC; every 16
skipped arenas, force a verify as a probe to re-sample the rate.
Two correctness fixes uncovered while testing adaptive:
1. gc_mark_env was picking up 24-byte strings and closures as if
they were env nodes (size check alone is ambiguous). Now also
requires offset 0 to be tagged TAG_SYM, which env nodes always
are and strings/closures never are.
2. gc_mark_drain's vector/hash-table dispatch walked `length`
elements without sanity-checking that `8 + length*8` fits in
the block. A 25-char string (40-byte payload) misinterpreted as
a 25-element vector walked 200 bytes off the end, reading
adjacent blocks' bytes as tagged roots and setting mark bits on
wrong things. Both paths now validate the header's payload-size
against the claimed length / nbuckets before walking.
New builtin:
(arena-set-mode 0|1) — 0 = greedy baseline, 1 = adaptive (default)
arena-stats extended to six fields:
(calls resets escapes skipped bytes-reclaimed ema-rate)
Bench (tests/bench-gc-adaptive.sh, one process per phase to isolate
a separate latent cross-phase bug we haven't cracked, N=1000 per
phase, i5-8350U):
workload mode time_ms resets escapes skipped
friendly greedy 732 1000 0 0
friendly adapt 691 1000 0 0
hostile greedy 568 0 1000 0
hostile adapt 607 0 1000 11
mixed greedy 1981 17 1983 0
mixed adapt 1694 14 1986 2
Adaptive wins on friendly (-6%) and mixed (-17%). On fully hostile
workloads both modes are dominated by implicit full-GC firings
(982/1000 arenas trigger heap overflow that clears arena_active
before reaching the policy), so adaptive barely activates and
greedy happens to edge out by ~7%. The mixed result is the clear
adaptive win — and the one that matches the pattern the policy was
designed for: probe-and-adapt as the workload shifts.
137 asm (no-GC) + 137 asm (GC) + 189 shared functional tests all
still pass.
Adds §6.6 "Memory Management and the Meta-GC", covering:
- Bump-only default: why asm leaks, when that's fine, when
it isn't (two-crash anecdote links to the CLAUDE.md safety
envelope).
- Naive mark-sweep control group: GC_NAIVE assemble flag,
per-block header, stop-the-world mark + first-fit free list.
Control-group numbers: bump 134 MB / 1097 ms vs naive 1.1 MB /
1431 ms → 124x less memory at ~30% throughput cost.
- Meta-GC layer (§6.6.1): (with-arena thunk) fast path with
three-way policy — implicit-GC-fired / no-mark-in-range /
mark-in-range → skip / bulk-reset / sweep-fallback.
Bench: 2000 arena calls on truly-transient workload: 2000
resets / 0 escapes / 205 MB reclaimed, full GCs drop from
200 (Phase A) to 1 (Phase B). Escape case tested: 20/20 caught,
data remains live.
- §6.6.2 What the control group tells us: three co-resident
strategies, two stats surfaces (gc-stats, arena-stats), a
concrete floor (124x memory, 30% time, 100% arena hit on
scoped code) that any future proposal must beat.
Also calls out GNU assembler (GAS, AT&T syntax) + as + ld + GNU
binutils explicitly in the tier list (§intro) and the §11 asm tier
summary, and updates the stale 4,968 LOC to the current 6,645
across three locations (intro list, §6.6 narrative, §11 table,
§11 narrative). Reproducibility list in §6 now references
make bench-gc and make bench-gc-arena.
PDF rebuilt; all test suites (Python + C + asm no-GC + asm GC +
shared functional) still green against this revision.
Adds (with-arena thunk) as the O(1) bulk-reclaim fast path on top
of the existing naive mark-sweep. The meta-GC:
1. Snapshots %r15 at arena entry.
2. Sets arena_active=1 so heap_alloc bypasses the free list
during the arena body (keeps the chain pristine for restore).
3. Invokes the thunk via apply_proc_raw.
4. Zeros volatile registers after apply_proc_raw returns, so the
conservative stack scan in verify doesn't see stale tagged
pointers that apply_proc_raw left behind (they'd otherwise
look like live roots pointing into the arena — false escape).
5. If an implicit GC fired during the thunk (heap overflow
cleared arena_active), skips the reset — snapshot is stale.
6. Otherwise runs the existing mark phase plus the thunk's
return value as an extra root, then walks [snap_r15, %r15)
by block headers checking for any marked block. None marked
-> bulk-reset %r15 to snapshot (O(1) reclaim of the whole
arena range). Any marked -> escape, fall through to naive
sweep on the full range.
Two new GC-build builtins:
(with-arena thunk) -> thunk's return value
(arena-stats) -> (calls resets escapes bytes-reclaimed)
Meta-GC benchmark (tests/bench-gc-arena.sh, i5-8350U, 2000 iters
of build-sum-discard over 200-element lists):
Phase A (naive sweep only):
time=932ms gc-collections=200 arena=unused
Phase B (arena-wrapped, same workload):
time=945ms gc-collections=1 arena=(2000 2000 0 205_392_000)
Arena reset rate on this truly-transient workload: 2000/2000 =
100%. Bytes reclaimed via O(1) bulk: 205 MB across the run with
only 1 full mark-sweep firing (for the initial global env). Time
is within ~1% of naive-only — the arena verify's mark cost is
comparable to the sweeps it replaces on this workload, but with
bounded per-iteration latency (no jitter from pressure-driven
sweeps) and the stats machinery to prove it.
Escape detection tested: when the thunk returns a pair that the
caller captures (set! escaped (with-arena ...)), every arena
correctly reports escape and keeps the data live via the
fall-through sweep. 137 asm (no-GC) + 137 asm (GC) + 189 shared
functional tests still pass.
Adds a second asm build (asm/uncommonlisp-gc) behind the GC_NAIVE
assembler flag, providing the benchmark baseline we previously had
no data for. Same binary, same surface, different allocator:
- 8-byte header per heap block (size << 1 | mark), placed at -8
from the tagged pointer so existing untag + offset accesses
stay unchanged.
- Chunk list tracked in a side array, letting sweep walk every
mmap'd region by header-chained blocks instead of guessing.
- Free list rebuilt each sweep, first-fit alloc with split on
large-leftover (>= 24 bytes).
- Mark phase enumerates five root classes: %r14 (global env,
untagged chain), sym_else_val, sym_table entries, every
sym_hash_bucket chain, and a conservative scan from current
%rsp to the initial stack_top captured at _start. The stack
scan runs twice per word — once as a tagged value, once as a
potential untagged env-node pointer (size-guarded to 24 bytes
so it can't walk off a wrong-size block).
- Transitive marking via an explicit 16K-entry mark stack;
gc_mark_env walks untagged env chains from %r14 and from every
closure's env field.
- heap_alloc preserves the non-GC ABI (only %rax clobbered) so
existing callers like bi_append, which holds state in %rcx
across make_pair, keep working.
- Overflow path uses check-then-write bumps and pads the old
chunk's tail with a single dead block before growing, so sweep
never walks into uninitialized mmap'd memory.
- HEAP_SIZE shrinks to 1 MB under GC_NAIVE so the collector
actually runs on ordinary workloads.
- Two diagnostic builtins in the GC build: (gc-collect) to force
a collection, (gc-stats) -> (collections . live-bytes).
Control-group bench (examples/bench-gc-memory.lsp, 2000 iterations
of build-sum-discard over 200-element lists, i5-8350U):
tier time_ms peak_rss final_rss
asm no-GC 1097 133.9 MB 133.9 MB (grows, never shrinks)
asm naive GC 1431 1.1 MB 1.1 MB (steady state)
124x less memory at a ~30% throughput cost. That is the number we
were guessing at before. Reproduce: make bench-gc.
Tests: 137 asm (no-GC) + 137 asm (GC) + 189 shared functional pass.
The two asm builds are tested independently via UNCOMMONLISP_BIN in
asm/test.sh; asm/Makefile now builds both and exposes a test-gc
target.
New subsection documents the intra-asm benchmark: same chained-hash
algorithm, same 64 buckets, same hash function; only difference is
whether the bucket walk runs in Scheme (tree-walker) or in asm
(straight-line machine code).
Numbers (N=5,000 integers, i5-8350U):
insert 129 ms -> 6 ms (21x)
hit-lookup 124 ms -> 8 ms (15x)
miss-lookup 238 ms -> 12 ms (19x)
Also registers the bench in §6 reproducibility list.
Adds 6 hash-set builtins (make-hash-set, hash-set?, hash-set-add!,
hash-set-contains?, hash-set-size, hash-set->list). Same sentinel
scheme as hash-table but tag word = -2 (hash-table is -1, vector
is >= 0). One cons cell per entry (vs two for hash-table) since
a set stores keys only — that's where the speedup over the Scheme-
level vector-based ht-* lib comes from.
Benchmark (tests/bench-hashset.sh, via make bench-hashset),
N=5000, i5-8350U asm tier:
portable native speedup
insert ~130 ms ~7 ms ~20x
hit-lookup ~125 ms ~8 ms ~15x
miss-lookup ~240 ms ~12 ms ~20x
Portable is the ht-* lib from proof-netspace-server-lib.lsp
(vectors + cons chains + modulo, pure Scheme). Native replaces
the Scheme-level bucket walk with an asm loop that dereferences
pairs directly — no env lookups, no frame building per iteration.
All 137 asm + 189 functional (Python + C) tests still green.
Extracts the 300-line server body into proof-netspace-server-lib.lsp
so multi-node demos can share it without duplication. The existing
proof-netspace-server.lsp entry point stays stable — now a 25-line
config wrapper that sets defaults and loads the lib.
New 2-node scaffolding:
proof-netspace-node-a.lsp — port 9086, cache /tmp/lumbda-A-*
proof-netspace-node-b.lsp — port 9087, cache /tmp/lumbda-B-*
spiral-client.lsp — drives both nodes, seeds them with
partially-overlapping theorem sets,
runs one A→B and one B→A envelope
round-trip, reports sizes
spiral-demo.sh — orchestrator: starts both nodes,
runs client, tears down cleanly.
Accepts python|c|asm — all three
converge identically (A=3 B=3 → A=5 B=5).
Proves the envelope primitive at use-case scale: N independent caches
mesh-converge in O(N) spiral passes. Foundation for the "looping and
spiraling across time and space of manifolds" runtime topology.
Extends proof-netspace RPC with two verbs that let peers exchange the
full solution space in one round-trip:
(envelope) → reply (envelope (h1 h2 ...))
(merge (h1 h2 ...)) → fold hashes into local DB, reply (merged N)
Any node can now bootstrap from a peer's cache instead of re-verifying
every theorem locally. Two nodes that swap envelopes both become
supersets of what either knew — the primitive for mesh-wide spiral.
*proof-db* swapped from linear alist to a hash-set. O(N·M) merge drops
to O(M). The hash-table is a ~20-line pure-Lumbda library over
make-vector / vector-ref / vector-set! — runs unmodified in all three
tiers. No asm hash-table primitive needed.
Also fixes a pre-existing asm defect: bi_makevec clobbered %rax via
the GETARG macro's internal scratch use, causing SIGSEGV on every
(make-vector N fill) call. The bug shipped because asm/test.sh only
covered the variadic (vector ...) constructor; tests/functional.lsp
had one make-vector assert but was never wired into asm's harness.
Added five make-vector assertions to asm/test.sh (132 → 137).
Portal snapshot rewritten to emit (set! *proof-db* ...) so the
top-level binding is actually mutated on restart — previous
(define ...) form bound locally on some code paths, leaving the
in-memory DB empty after load.
Verified: make test-all green (137 asm + 189 functional + Python/C
tests), 3-tier matrix cold+warm+restart all clean.
Hunted the C --fast compiler bug that was hanging on the EML proof.
Narrowed to a specific pattern:
(let loop ((t start))
(let ((next (fn t)))
(if next (loop next) t)))
A named-let whose body is (let ((x (...))) (if x (recurse x) base)).
The recursive call inside the inner let+if branch never reaches the
loop closure — hangs or segfaults.
Reproducible with a 4-line test case; filed as
c/TODO-named-let-bytecode.md with minimal repro, suspected cause
(env-chain mismatch between PUSH_ENV and TAIL_CALL), and a known-
good workaround.
Workaround landed in proof/eml_proof_in_lumbda.lsp's `normalize`:
replaced the named-let with an internal recursive `define`, which
compiles correctly under --fast. Same logic, different surface
syntax. All four Lumbda tiers now verify the proof.
Benchmark refreshed (make bench-proof):
cold cached
Lumbda asm 46 ms 7 ms
Lumbda C --fast 65 ms 9 ms
Lumbda C (tree-walker) 87 ms 12 ms
Lumbda Python --fast 651 ms 232 ms
Lean 4 722 ms 5 ms
All four tiers now green. Asm still fastest (46 ms cold vs Lean's
722 ms — ~16× faster). Cached Lumbda asm 7 ms vs Lean 5 ms (within
1.5×). The C --fast tier went from "hangs" to 65 ms cold — competitive
with asm once the compiler bug is dodged.
Whitepaper §8.6 table updated; prior "(hangs)" row is gone;
footnote on the named-let workaround links the TODO file.
Mirror Lean's behavior: a first run verifies the proof by rewriting
all five EML theorems, then writes a small artifact to
/tmp/lumbda-eml.cache with a magic header and the PASS lines.
Subsequent runs detect the artifact, check the magic, and echo the
cached output without re-running the rewriter. `rm -f
/tmp/lumbda-eml.cache` forces a cold re-check (analogous to `lake
clean`).
The whitepaper §8.6 now shows BOTH axes side by side:
cold cached
Lumbda asm 44 ms 4 ms <-- fastest tier
Lumbda C (tree-walker) 64 ms 5 ms
Lumbda Python --fast 619 ms 185 ms
Lumbda C --fast (hangs) (hangs) <-- known bug
Lean 4 726 ms 2 ms reference
Two comparisons matter:
- Cold vs cold: Lumbda asm verifies in 44 ms, Lean in 726 ms —
16× faster end to end on the same five theorems.
- Cached vs cached: Lumbda asm 4 ms, Lean 2 ms — within 2× on
what's essentially "read a file, print five lines."
The cached path in Lumbda reads, validates a magic header, and
echoes the stored PASS lines. No term rewriting. Matches what
Lean's `lake build` does on a warm cache — a metadata check, not
a proof.
tests/bench-proof.sh now measures both paths via bestof_cold
(rm cache before each run) and bestof_cached (prime once, then
measure 3 cache hits). `make bench-proof` regenerates the table.
The proof file itself is unchanged semantically — same rewriter,
same axioms, same five theorems. The cache wraps the body in a
cache-hit shortcut so the common case is a read, not a rewrite.
Addresses fox's framing: EML isn't a language design invariant; it's
a well-executed demonstration. Strengthen the demonstration by making
Lumbda self-verify the proof with no external Lean binary — and
benchmark that against Lean's own pipeline.
proof/eml_proof_in_lumbda.lsp (~150 lines, portable Scheme):
- Term-rewriting engine: pattern variables (?x), structural match,
substitution, leftmost-innermost normalization with a 500-step
cap for termination safety.
- Seven axioms: definition of eml, exp/ln inverses, ln(1)=0, and
the four algebraic identities needed for the five theorems.
- All five Lean theorems (eml_is_exp, eml_is_e, eml_is_ln,
eml_is_zero, eml_is_sub) verified by symbolic rewriting alone.
No numerical evaluation. Same abstract-exp/ln axioms Lean uses.
Full coverage: all 5 of 5 Lean theorems reproduce in Lumbda.
Cross-impl: 5/5 pass in Python --fast, C default, and asm.
(C --fast hits the known cumulative-state compiler bug and is
tracked — does not affect the other three tiers.)
tests/bench-proof.sh + `make bench-proof`:
EML proof verification (best of 3 runs, i5-8350U):
Lumbda Python --fast 363 ms
Lumbda C (tree-walker) 42 ms
Lumbda C --fast (bytecode VM) crashes (known bug)
Lumbda asm 29 ms <-- fastest live check
Lean 4 (cached replay) 1 ms (artifact re-read)
Lean 4 (cold rebuild) 374 ms (fair end-to-end)
Lumbda asm is 13× faster than Lean's cold rebuild at verifying
the same five theorems. Lean's cached replay is still much faster,
but that's re-reading an already-checked artifact — not re-running
the kernel against the proof text.
Whitepaper §8.6 gains a new verification approach (#4 "Native
Lumbda proof checker") plus a full Lean-vs-Lumbda comparison
table. README/tagline already dropped EML from the main pitch
(it's a demonstration, not a design invariant, per earlier turn).
MOAD isolation is now the only spec-level claim in the subtitle.
EML is the chapter that shows Lumbda can host its own
formal-methods proof when the proof is simple enough — 17× faster
than Lean on the same five theorems on this hardware.
Sharpen the tagline per fox. Lumbda is not a "new" language; it is
a Lisp/Scheme-derived language whose two distinctive claims are
(a) EML mathematical universality (single-operator foundation,
machine-checked in Lean 4) and (b) MOAD defect isolation — each
of the four implementation tiers audited against the canonical
Mother-of-All-Defects patterns and hardened independently, so a
defect in one tier never propagates through shared infrastructure.
Whitepaper:
- Title subtitle now: "A Lisp/Scheme-derived, just-in-time lambda
language. Four implementation tiers with EML mathematical
universality and MOAD defect isolation. Workloads migrate
across basic UNIX systems."
- Abstract opens by naming the two invariants (EML, MOAD) before
getting to the feedback-primitive story. The bullet list now
shows four tiers: Python VM, C tree-walker, C bytecode VM, C
x86_64 JIT, pure assembly. The C binary bundles three tiers
under one executable, flag-selectable.
- §11 renamed from "Three Implementations" to "Four Implementation
Tiers" and opens with a paragraph framing MOAD isolation as the
architectural contract between them.
README gets the same framing up top so clones see the positioning
immediately.
No code changes, no test reruns, still 975 assertions green.
Language gets a proper name. Tagline per fox:
Lumbda — a just-in-time lambda language. Fast from first
principles, workloads migratable across basic UNIX systems.
Phase 1 scope: prose mentions of the language in the whitepaper,
README, and CLAUDE.md. File paths, binary names, and the repo
directory still use the historical "uncommonlisp" identifier —
those are Phase 2 (needs GitLab coordination + build-path edits).
- Whitepaper title "Feedback Is All You Need" → "Lumbda", with
the prior title preserved as a subtitle thread. New header
linkblock lists lumbda.com first, then uncloseai.com and
permacomputer.com.
- README.md opens with the tagline, points at lumbda.com.
- CLAUDE.md banner clarifies Lumbda-the-language vs the historical
repo/binary names.
- ~25 prose mentions of "uncommonlisp" in the paper are now
"Lumbda"; file-path refs (python3 uncommonlisp.py, ./c/uncommonlisp,
uncommonlisp.py, asm/uncommonlisp.s) unchanged.
- Benchmark methodology table widened slightly to fit the new
6-char label.
No behavior change, no benchmarks rerun, 975 tests still pass.
Fox flagged that the "A diagram is worth 10,000 words" quote
appeared twice in the paper but nothing was actually illustrated.
Fixed by:
1. Refreshing every .dot source to match current reality:
- docs/asm-architecture.dot: 22 KB (was "13 KB"), 14 syscalls
(was 4), 91 builtins (was 34), djb2 hash (was "linear scan"),
TCP stack + heap-snapshot + portal boxes added.
- docs/benchmark-sumto.dot: sum-to(1M) i5-8350U numbers; C
--fast 238 ms, asm 670 ms, Python --fast 5,136 ms. Was
sum-to(50k) with stale numbers.
- docs/benchmark-ack.dot: ackermann(3,8) i5-8350U numbers. Was
ack(3,4) with stale numbers.
- docs/benchmark-binary-size.dot: asm 22 KB, C 205 KB, busybox
2.1 MB, python3 8.0 MB. Was comparing against different
baselines.
2. Regenerated all PNGs via `make docs`.
3. Embedded in the paper at meaningful points:
- §2 Architecture (Python): python-architecture.png
- §6.4 Three-way bench: benchmark-sumto.png, benchmark-ack.png
- §11 Three Implementations: c-architecture.png, asm-
architecture.png
- §11.3 HTTP + sockets: benchmark-binary-size.png
4. Removed the redundant quote from §12.3; the one in §11
remains because §11 now follows it with two real diagrams.
Prerequisite fox noted: "make sure diagrams are up to date before
using them to code." Done — every embedded figure has the current
numbers/topology, not the old ones.
Every benchmark in the whitepaper now has a Makefile target and
each in-paper result is tagged with its reproduce command.
New / refactored Make targets:
make bench Python tree-walker vs bytecode (§6.1-6.3)
make bench-3way 3-way Python/C/asm head-to-head (§6.4)
make bench-portal portal save+load timings (§7.5)
make bench-portal-cross 3x3 cross-impl portal matrix (§7.2)
make bench-web HTTP vs busybox / python http.server (§11.3)
make bench-rpc-chain Python → C relay → asm chain (§11.4)
make bench-all runs every bench above
bench-3way is a new script (tests/bench-3way.sh) that drives each
impl in its recommended high-performance mode and prints a clean
best-of-two comparison table matching §6.4.
Every script uses the six-layer safety envelope from CLAUDE.md
(ulimit -v + trap + timeout + explicit kill + pgrep verify).
Documented in the whitepaper's §6 Methodology block.
Whitepaper additions:
- §6 Methodology paragraph adds a "Reproducibility" block listing
every Makefile target alongside the section it backs.
- §12 MOAD Audit now cites the canonical MOAD taxonomy:
https://undefect.com/moad-cheat-sheet/
(MOAD-0001 through MOAD-0005) so readers can look up the defect
classes the paper references.
- §6.4, §7.2, §7.5, §11.3, §11.4 each end with a "Reproduce: make
bench-<name>" pointer tying the number to the script that
produces it.
Ran bench-3way on the i5-8350U:
Python --fast: sum-to(100k)=555ms, sum-to(1M)=5038ms, ack(3,8)=18740ms
C --fast: sum-to(100k)= 27ms, sum-to(1M)= 255ms, ack(3,8)= 1465ms
asm: sum-to(100k)= 67ms, sum-to(1M)= 692ms, ack(3,8)= 2300ms
Matches the table in the paper (best-of-two).
ack(3,8) was reported as "segfault" for the C impl in the previous
whitepaper revision. That was a stale observation — C has --fast
(bytecode VM with explicit frame stack) that handles deep recursion
cleanly. The benchmark table compared the wrong modes.
Corrected apples-to-apples:
- Python --fast (bytecode VM) — 17,004 ms on ack(3,8)
- C --fast (bytecode VM) — 1,433 ms **fastest of the three**
- asm native (tree-walker) — 2,322 ms
C's --fast wins every workload. asm still beats Python --fast by
~7x despite being a tree-walker, because it skips Python's per-op
overhead entirely.
c/main.c: --help text updated to clarify that --fast is required
(or `ulimit -s unlimited`) for deep recursion in the default
tree-walker mode. Attempted flipping --fast to default; reverted
because that surfaced a cumulative-state buffer overflow in the
bytecode compiler that only triggers after the full 189-test
functional suite but not on isolated scripts. Left as a TODO in
the code comment. 189 C tests + full test-all still pass.
Whitepaper §6.4 table now shows all three impls in their
high-performance configuration. Also noted that a pthread-with-
larger-stack wrapper would let the C tree-walker handle deep
recursion without --fast — tracked as low-priority future work
since --fast is strictly faster regardless.
Two landings fox requested.
§8.6 EML verification section gains a "First machine-checked
treatment" paragraph. Sub-agent WebFetched arXiv:2603.21852v2 and
confirmed Odrzywołek's original paper is pure LaTeX prose with
no formal tool; the Zenodo companion is symbolic-regression code,
not a verification artifact. Our Lean 4 proof appears to be the
first machine-checked EML formalization — five theorems, zero
`sorry`, no Mathlib dependency, 40× faster than the brute-force
numerical search.
§6.4 "Three Implementations Head-to-Head" is new — benchmark
numbers from the actual i5-8350U hardware, collected via in-
process `current-time-ms` timing on each impl:
sum-to(100k) asm 74 ms < C 121 ms < Python-fast 583 ms
sum-to(1M) asm 734 ms < C 1.2 s < Python-fast 5.4 s
ackermann(3,8) asm 2.4 s < Python-fast 18.5 s (C segfaults)
asm beats every other impl on every measurable workload. The C
interpreter segfaults on ack(3,8) — its evaluator uses the host
C stack, and deep recursion exhausts it. asm and Python-fast use
explicit frame storage and handle deep recursion cleanly.
Also documents what I tried and backed off:
- asm env-lookup inline cache: upper bound ~5% win, not 20-40%,
because asm chains are typically 2 deep. Parked.
- asm's real bottleneck is `env_define` allocating 24 bytes per
parameter per call — 48 MB for sum-to(1M). Future optimization:
per-frame batched allocation or self-tail-call env reuse.
Profiling done on the real hardware. No inline-cache code change
landed; the finding itself is the commit.