跳到正文
r/LocalLLaMA· /u/N34257·· 3 小时前AI 评分27

RDNA4 上跑 Qwen 3.8 的当前最佳推理方案是什么?

What's the current meta for RDNA4 with Qwen 3.8?

AI 导读

有用户在双 R9700 上通过 vllm-radiance 运行 Qwen 3.8 27B FP8(Swift 1.5 FP8),实测 prefill 达 5000t/s、代码生成 130t/s+。

正文

As it says, really - I'm currently running vllm-radiance on dual R9700s, with Qwen 3.8 27B FP8 (or, rather, Swift 1.5 FP8). Performance is great an' all (5000t/s prefill, 130t/s+ code gen), but I'm just wondering...with all the architecture-specific inference engines popping up all over the place...is there anything I'm missing out on? I couldn't find anything that could give better performance on RDNA4 when I looked, so...over to you guys?

I'm particularly interested in anything that could potentially get up and running with Qwen 3.8 Flash Next - vllm-radiance doesn't support it yet, but I don't particularly want to regress to the performance of llama.cpp after having experienced vllm-radiance performance levels.

submitted by /u/N34257
[link] [留言]

来源:r/LocalLLaMA · reddit.com