| This is an update. I posted LocalMind here many moons ago from another account, when it was a Gemma chat in a tab. LocalMind is a static web page that runs models on your GPU through WebGPU. It has no server, no account and no install. The new part: two engines that stream mixture-of-experts weights from disk while they generate. That lets a tab run models bigger than the machine's RAM. Live: https://localmind.naklitechie.com · Code (MIT): https://github.com/NakliTechie/LocalMind All numbers are from one MacBook M4 Pro (24 GB) in Chrome. How it works - On first load the GGUF is copied into OPFS, the browser's private file system.
- Dense weights, routers and the KV cache go to the GPU.
- Routed experts stay on disk. A pool of workers reads them on demand with sync access handles into a GPU slot cache (LRU, two layers of prefetch).
- The trunk kernels are hand-written WGSL that follow llama.cpp's graphs. That lets me test against llama.cpp on the exact same GGUF.
Gemma 4 26B-A4B (Google's QAT Q4_0, 14.4 GB) - Same output as llama.cpp b9830 Metal: the live site's chat replies were character-identical on 9/9 test conversations (capped at 64 tokens). 15/16 fresh prompts matched token for token. The 16th split on a 0.00009-nat near tie, where llama.cpp's own two attention paths also disagree.
- Memory: the Chrome GPU process sits at 6.9 GB with a 4 GB expert cache. About 8.6 GB of experts stay on disk.
- Speed: 23.6 tok/s decode, 55 tok/s prompt processing. llama.cpp Metal does 70.6 and 204 on the same Mac, so the tab is ~3× slower at decode. Per token: ~22.5 ms GPU compute, ~11 ms routing round trips, ~8–13 ms SSD reads.
- First load from the site: 11.5 min (14.4 GB download). After that: 1.6 s.
Qwen3.6 35B-A3B (unsloth Q8_0, 36.9 GB, on a 24 GB Mac) — experimental - The file is bigger than the machine's memory. The GPU process measured 7.3 GB with a 4 GB expert cache.
- Live site: 9.9 tok/s decode, 2.2 s to first token. First load is 36 min (download plus the OPFS copy).
- Output matches llama.cpp Metal 8/8 on 4- and 16-layer cuts. On the full model it matches llama.cpp CPU 5/8; the other 3 swap near-tie tokens. I can't run the full file on llama.cpp Metal on this Mac, so full-model parity is still open.
- Per token (~99 ms): ~23 ms GPU compute, ~39 ms routing round trips, ~35 ms expert reads from the SSD. Moving routing onto the GPU gave no gain (10.3 vs 10.3 tok/s): the misses are experts nobody predicted.
Also - Gemma 4 E2B can keep its 1.2 GB per-layer embedding table on disk: GPU process 4.27 → 2.07 GB, identical output, 3–8% slower decode. It's a setting, off by default.
- The whole app is one
index.html again (854 KB with brotli). Engines, workers and the disk tier are rolled into it, and the tab builds them from blob URLs. - The disk tier is also a standalone library: diskformer.js.
Prior art As far as I can find (searched 6 Oct 2026), no earlier browser engine reads weights from disk during generation. wllama and LlamaWeb stream from OPFS only at load. Pooled runs Qwen3.6-35B-A3B in a browser with experts paged from system RAM. On-demand disk reads exist in native runtimes: llama.cpp's --moe-stream PR (#25294) and Google's LiteRT-LM for Gemma's per-layer embeddings. Corrections welcome. The Gemma 4 E2B kernels are webml-community's (Xenova and the Transformers.js team). My part there is the disk path. Limits - Chrome or Edge with WebGPU. Tested on one M4 Pro 24 GB only; 8 and 16 GB machines are untested.
- Not faster than native: llama.cpp is ~3× faster on Gemma 26B. The point is that a tab can run these at all, with the same output.
- Parity covers greedy decoding, the prompts listed above, and 64 tokens each.
- I haven't tried llama.cpp's expert-offload flags (
-ot exps=CPU) for comparison. If you have an NVIDIA/AMD GPU or a 32–64 GB Mac, I'd like your tok/s numbers. A bigger expert cache should move the Qwen3.6 number the most. submitted by /u/naklitechie [link] [留言] |