Infermeld:在 AMD + NVIDIA 混合 GPU 上用 llama.cpp 跑单个 GGUF 的 Linux 工具包
Infermeld: a Linux kit for running one GGUF across AMD + NVIDIA GPUs with llama.cpp
开发者发布开源 Linux 工具包 Infermeld v0.1.0,可在 AMD 与 NVIDIA 混合 GPU 上通过 llama.cpp 运行同一个 GGUF 模型,提供 Vulkan + CUDA 设备选择、可复现构建说明和只读温度保护。
Following on from club-5060ti and club-rdna16, I’ve put together Infermeld: a small, open-source Linux companion kit for running one GGUF across an AMD GPU and an NVIDIA GPU, powered by llama.cpp.
I’m the maintainer. This is an experimental v0.1.0 release, and I’m looking for people with other mixed GPU combinations to help reproduce the setup and find the rough edges.
The idea is practical: if you already have cards from both vendors, can you put them to work together without buying a matching pair?
What Infermeld adds
The inference engine is llama.cpp. Infermeld isn’t a new backend, and I’m not claiming to have invented mixed-GPU inference.
It packages the supporting pieces around that setup:
- Explicit AMD/Vulkan + NVIDIA/CUDA device selection and runtime preflight.
- Reproducible build instructions and inspectable launch arguments.
- A read-only thermal guard, with shutdown limited to the server process it started.
- Documentation and a results site that keep configurations, failures and limitations visible.
The release is source-only. You build the documented llama.cpp revision separately and supply your own model weights. It’s intended for people comfortable with an experimental Linux setup, not as a one-click installer.
Current tested setup
| Component | Tested configuration |
|---|---|
| AMD GPU | RX 6900 XT, 16GB |
| NVIDIA GPU | RTX 3080, 10GB |
| Model | Qwen3.6-35B-A3B, UD-Q4_K_M GGUF |
| Backends | Vulkan + CUDA |
| Split mode | Layer |
| Context reservation | 8,192 tokens |
The acceptance checks include loading and short completions with MTP off and on.
That’s a narrow result on one hardware pair, not broad compatibility testing. An 8K context reservation is not the same as testing a filled 8K prompt, and a short successful response is not a sustained performance benchmark.
Important limitations
- Sustained Q4 throughput and full-length high-context results are not yet qualified.
- Historical measurements are labelled with their original configurations. They should not be read as performance numbers for the current Q4 setup.
- There’s no promise that combining cards is faster than using one.
- Adding the advertised VRAM capacities does not guarantee that all of it is usable for the model and its runtime allocations.
I’d rather make those boundaries clear than present a successful load as a complete benchmark.
Looking for other AMD/NVIDIA combinations
Successful runs and failures are both useful. If you try it, please include:
- Both GPU models and their VRAM sizes.
- OS, driver versions and llama.cpp revision.
- Model and quantization.
- Launch settings, including the split and context reservation.
- How far it got: preflight, loading, first completion or a longer workload.
There’s a hardware/result issue form in the repository. Please sanitize paths and keep credentials and private logs out of reports.
Repository and setup instructions:
https://github.com/5p00kyy/infermeld
Results and evidence:
https://5p00kyy.github.io/infermeld/
Anyone already using an AMD/NVIDIA pair for local inference? I’d be interested in what works for you, and where this setup breaks on different hardware.
submitted by /u/do_u_think_im_spooky
[link] [留言]
来源:r/LocalLLaMA · reddit.com