RDNA4 上跑 Qwen 3.8 的当前最佳推理方案是什么?
What's the current meta for RDNA4 with Qwen 3.8?
有用户在双 R9700 上通过 vllm-radiance 运行 Qwen 3.8 27B FP8(Swift 1.5 FP8),实测 prefill 达 5000t/s、代码生成 130t/s+。
As it says, really - I'm currently running vllm-radiance on dual R9700s, with Qwen 3.8 27B FP8 (or, rather, Swift 1.5 FP8). Performance is great an' all (5000t/s prefill, 130t/s+ code gen), but I'm just wondering...with all the architecture-specific inference engines popping up all over the place...is there anything I'm missing out on? I couldn't find anything that could give better performance on RDNA4 when I looked, so...over to you guys?
I'm particularly interested in anything that could potentially get up and running with Qwen 3.8 Flash Next - vllm-radiance doesn't support it yet, but I don't particularly want to regress to the performance of llama.cpp after having experienced vllm-radiance performance levels.
来源:r/LocalLLaMA · reddit.com