跳到正文
r/LocalLLaMA· /u/fechyyy·· 4 小时前AI 评分54

21M 模型配 6.4B 参数查表:追平 114M 稠密模型,表可放 SSD

I gave a 21M model a 6.4B-parameter lookup table. It matches a 114M dense model and runs with the table on an SSD (RX 9070)

AI 导读

作者公开了一个稀疏记忆实验:21M 模型外挂 16.8M 行、6.4B 参数的查表,每个 token 只用 33M 参数,效果与同样 500M Wikipedia token 训练的 114M 稠密模型相当。

正文

I spent the last few weeks on a hobby research project and just made it public.

The idea isn't new (product-key memory, Lample et al. 2019, and Meta's "Memory Layers at Scale"): give a model a huge table of learned vectors and let it read only a few hundred of them per token. I wanted to know what that's actually worth on a small model, what it costs, and whether the table even has to sit in VRAM.

What came out:

- A 21M model with a 16.8M-row table (6.4B parameters in the table, 33M used per token) is about as good as a 114M dense model trained on the same 500M Wikipedia tokens.

- The table doesn't need VRAM. With the 4-bit table memory-mapped from an NVMe SSD the model still writes ~140 tok/s on my RX 9070, using 0.4 GB of VRAM. Reading long prompts from the SSD is slow though, every missed row costs a whole 4 KB page.

- I wrote Triton kernels for it. They run unchanged on my Radeon, an MI350X and H100/H200.

- Bolting a table onto a finished model (Qwen3.5-0.8B) didn't work: no better than a small dense add-on with the same compute.

Caveats: it's tiny, one seed for the big runs, and the text it writes is fluent Wikipedia English with made-up facts. I wrote down the success criteria before every run, and the stuff that didn't work is in there too.

Most of it ran on my gaming PC, the big runs cost about 70 dollars on Runpod. I built it together with Claude Code (you'll see it in the commits), the ideas, decisions and money were mine.

Repo: https://github.com/re133/sparse-memory-lm

Click a word and see which table entries the model reads: https://re133.github.io/sparse-memory-lm/explorer/

Model: https://huggingface.co/fechyy/sparse-memory-lm-B-16M

Feedback welcome, especially if I got something wrong. And if anyone has bigger GPUs to spare, I'd love to try this at 1B scale.

submitted by /u/fechyyy
[link] [留言]

来源:r/LocalLLaMA · reddit.com