跳到正文
r/LocalLLaMA· /u/BraceletGrolf·· 3 小时前AI 评分12

如何真正学会使用 vLLM?从 Voxtral 3B 到量化部署的困惑

Ok how to actually learn vLLM ?

AI 导读

一位用户在 r/LocalLLaMA 发帖求助如何系统学习 vLLM,称其生态难以理解,官方文档也分不清 server 与 client library 的区别。他目前用单张 GPU 跑 Voxtral 3B(无需量化),但在量化及更高级功能上毫无头绪,此前曾在 RX 7900 XTX 上用 llama.cpp 全 GPU 运行量化版 Qwen 3.8 27B。

正文

Said in title, I find the ecosystem difficult to understand, and RTFMing doesn't help me as it's never clear what is the server vs their client library ? I'm using it for voxtral 3B on one GPU, but it's because I can run that with no quantization, I'm lost on learning to run with quantization / more advanced features.

I think it makes sense, because I'm running Qwen 3.8 27B quantized on llama.cpp but with everything on the GPU (RX 7900 XTX).

submitted by /u/BraceletGrolf
[link] [留言]

来源:r/LocalLLaMA · reddit.com