跳到正文
r/LocalLLaMA· /u/IceFog72·· 5 小时前AI 评分36

k_llama.cpp 分支为 MoE 模型加入专家驻留与 CPU/GPU 混合执行优化

k_llama.cpp MoE Optimizations: Expert Residency, Hybrid CPU/GPU Execution, Q2_0 Support

AI 导读

开发者 IceFog72 发布 k_llama.cpp 分支,为 MoE 模型加入带宽自适应 CPU/GPU 混合执行、可配置专家驻留显存上限(--moe-resident-mib N)和 Q2_0 量化支持。

正文

Finally finished my fork:
https://github.com/IceFog72/ik_llama.cpp

Nothing else I wanted to add/try currently works

In short, it now has:

Basic usage:

-cmoe --moe-resident auto --moe-resident-mib N

I don't know how -ncmoe behaves because I can't properly test it on my hardware.

Without --moe-resident-mib N, --moe-resident auto will fill all available free VRAM with resident experts.

With something like:

--moe-resident auto --moe-resident-mib 2048

you can cap how much VRAM the resident cache uses and intentionally leave some free. There a useful cap how much helps to improve speed. If you set the cap too low, performance will drop too.

And gaze upon the magic of higher generation speed Kek

The important part: this only helps when the full MoE does not fit in VRAM and the GPU still has unused compute capacity, and free pci buss speed.

If your GPU was already fully loaded, this fork probably won't improve anything.

If your GPU is sitting around ~75% while you have many layers in VRAM, it may be worth trying 1-2 fewer regular GPU layers and using:

--moe-resident auto --moe-resident-mib 1024/2048

Adjust the cap depending on your GPU and available VRAM. The goal is to use resident experts to fill otherwise-idle GPU capacity rather than simply maximizing the number of fully offloaded layers.

On my setup — RTX 2060 6GB + Ryzen 7 2700X + 40 GB DDR 4 2993Mhz using arch — I have too little VRAM to offload enough complete expert layers for useful acceleration, so I use -cmoe.

Before these changes, generation could leave my GPU at only around 25-35% utilization, with roughly 1.5-2GB VRAM still free with fully loaded cpu.

With Qwen3.6-35B-A3B-UD-Q4_K_M.gguf at around 15-30k context, default ik_llama.cpp gives me roughly 23 t/s, while this fork gives me around 26-30 t/s.

So on my hardware I'm seeing roughly 20-30% speedup.

People with better GPUs and more VRAM may see better results, depending on where their bottleneck is.

The two experimental options still need more testing:

--moe-resident-profiler new/old

gives me a more balanced CPU/GPU work split, with somewhat more work left on the CPU and lower GPU load, but no clear speed difference for my setup

--moe-resident-grouping off/layout

also needs more testing, especially on better systems.

I sometimes see around 1-2 t/s difference from these options, but on my PC a browser tab sneezing can cause +/-2-4 t/s, so I don't consider that conclusive.

My current command:

./llama-server \ -m /mnt/Kingstone_SSD/GGUF/Qwen3.6-35B-A3B-UD-Q4_K_M.gguf \ --alias "hz" \ --host 0.0.0.0 \ -ctk q4_0 \ -ctv q4_0 \ -ctv-first q8_0,4 \ -ctv-last q8_0,4 \ -cmoe \ -b $((6 * 512)) \ -ub $((3 * 512)) \ --ctx-size $((64 * 1024)) \ --jinja \ -fa on \ --no-mmap \ --no-context-shift \ --temp 0.6 \ --top-k 24 \ --top-p 0.95 \ --min-p 0.00 \ -ngl 999 \ -np 1 \ --samplers "penalties;dry;top_n_sigma;top_k;typ_p;top_p;min_p;xtc;temperature" \ --moe-resident auto \ --moe-resident-mib $((2 * 512)) \ --k-cache-hadamard \ --v-cache-hadamard \ --moe-resident-profiler old \ --moe-resident-grouping off 

I plan to keep the fork updated with the main ik_llama.cpp branch for my own use.

If more people test it and provide feedback, especially on systems where the model still doesn't fully fit in VRAM, I may eventually make a PR to merge it upstream.

submitted by /u/IceFog72
[link] [留言]

来源:r/LocalLLaMA · reddit.com