为加速 MoE 模型预填充与解码,给 DDR5 内存超频
Overclocking DDR5 For Faster MoE Prefill and Decode
在 Intel B70(32GB)上用 llama.cpp 的定制 SYCL 后端运行放不进 VRAM 的 MoE 模型时,将 AMD 7950X 平台的 4 条双 rank DDR5-5600(128GB)从 3600MHz 超到 4800MHz,内存带宽实测提升 36%。
With llama.cpp using a customized SYCL backend on Intel B70 (32GB), overclocking my DDR5 memory gave modest gains for MoE models which do not fit in VRAM.
Both PP and TG increased after overclocking DDR5.
I've never been one to overclock my system, but on the advice of my agent, I overclocked the DDR5 RAM on my AMD 7950X (4x dual-rank DDR5-5600, 128GB) from 3600MHz, the AMD safe default for that memory configuration, to 4800MHz, with a measured 36% increase in memory bandwidth.
What was the improvement? PP increased by about 10% and TG increased about 5%. And the prefill numbers increase the deeper the context gets.
llama-benchy numbers, including prefix caching tests.
| test | 3600 base | 4800 avg (r1/r2) | delta |
|---|---|---|---|
| pp2048 @ d0 | 682.2 | 713.7 (716.6/710.9) | +4.6% |
| tg128 @ d0 | 30.3 | 31.2 (31.0/31.5) | +3.0% |
| ctx_pp @ d8192 | 660.6 | 715.0 | +8.2% |
| ctx_tg @ d8192 | 26.9 | 27.8 | +3.3% |
| pp2048 @ d8192 | 554.3 | 632.0 (632.4/631.6) | +14.0% |
| tg128 @ d8192 | 28.7 | 31.1 (31.7/30.5) | +8.2% |
Because people seem to want this level of detail:
llama-server -m Qwen3.8-Flash-Next-GSQ-RCO-Q2_0-00001-of-00002.gguf --alias qwen38-flash-next --mmproj Qwen-3.8-Flash-Next-mmproj-BF16.gguf --model-draft mtp-Qwen3.8-Flash-Next-shared-Q4_K_M.gguf --spec-type draft-mtp --spec-draft-n-max 3 --host 0.0.0.0 --port 8081 -ngl all -ncmoe 34 -c 262144 -fitc 786432 --kv-unified --lazy-mode off -lm none -ub 2048 -b 4096 -fa on -ctk q8_0 -ctv q8_0 --ctx-checkpoints 32 --checkpoint-min-step 2048 -t 12 -tb 12 --jinja --reasoning on --reasoning-preserve --temp 1.0 --top-p 0.95 --top-k 20 --min-p 0.0 --presence-penalty 0.0 --repeat-penalty 1.0 --chat-template-kwargs {"reasoning_effort":"medium"} --parallel 3 -cram 10240 --slot-save-path /var/tmp/kv-cache --log-file /tmp/qwen38-pristine.log -lv 3 submitted by /u/EvolvingDior
[link] [留言]
来源:r/LocalLLaMA · reddit.com