为什么说 ik_llama.cpp 比主线 llama.cpp 更快?在混合多 GPU 机器上主线反而大幅胜出
Why is ik_llama.cpp said to be faster than Mainline? On my hybrid multi-GPU rig, Mainline easily beats it
有用户在 4 卡混合多 GPU 机器上实测,主线 llama.cpp 以 19.70 tokens/sec 的 prompt eval 和 10.86 tokens/sec 的生成速度,大幅领先 ik_llama.cpp 的 5.48 和 8.43 tokens/sec。
I constantly see recommendations saying that ik_llama.cpp (ikawrakow's fork) is the undisputed king of hybrid CPU/GPU offloading and MoE performance. However, every time I benchmark it against mainline ggml-org, mainline consistently beats it by a wide margin.
Am I missing specific flags, or is ik_llama simply not designed for multi-GPU layer splitting?
My Rig & Hardware Constraints:
- Host CPU: Intel Core i9-10920X (12 physical cores, AVX-512 & VNNI enabled).
- GPUs: 4x asymmetric setup:
- GPU 0, 1, 3: NVIDIA P102-100 (10GB Pascal, PCIe 1.0 bus bottleneck).
- GPU 2: RTX 3060 12GB (Ampere, acts as Master node via -mg 2).
- Known Hardware Laws / Workarounds:
- I run layer splitting (-sm layer) across the 4 cards with asymmetric tensor splits (-ts).
- Pascals must strictly stay under 9.7 GB VRAM; exceeding that triggers PCIe micro-paging and tanks speed.
- I use --poll 100 on mainline to prevent AVX-512 CPU threads from dropping into low-power sleep states between GPU layer handoffs.
The Test:
- Model: Qwen 3.8 Flash-Next 177B Uncensored (IQ3_XXS, ~89 GB) with multimodal vision (mmproj).
- Offload: 35 layers offloaded to the 4 GPUs (-ngl 35), remaining 13 layers computed on the AVX-512 CPU. 32k context.
The Head-to-Head Benchmark:
- Mainline (ggml-org/llama.cpp):
- Prompt Eval: 19.70 tokens/sec
- Token Generation: 10.86 tokens/sec
- CUDA Graphs: 2,490 CUDA graphs reused across the GPUs.
- ik_llama.cpp:
- Prompt Eval: 5.48 tokens/sec (72% drop)
- Token Generation: 8.43 tokens/sec (22% drop)
- Observations: 0 CUDA graphs engaged. It spent time taking context checkpoints during generation (100ms+ pauses), and --poll is unsupported.
The Question:
Is ik_llama.cpp's speed advantage strictly meant for pure CPU inference or single-GPU systems?
Does its custom CPU threadpool fall apart when coordinating pipelined layer splits across heterogeneous GPUs over PCIe, where mainline's CUDA graph caching takes over? Would love to hear from anyone running hybrid multi-GPU setups.
submitted by /u/vulcan4d
[link] [留言]
来源:r/LocalLLaMA · reddit.com