跳到正文
r/LocalLLaMA· /u/bodhi371·· 2 小时前AI 评分67

Qwen3.8-27B 在 12GB 显存加 8GB 内存下跑到约 18 tok/sec

Qwen3.8-27B with just 12GB VRAM + 8GB RAM - ~18 tok/sec

AI 导读

作者用 ISTA-DASLab 的 Qwen3.8-27B-GSQ-RCO IQ3_S 量化(约 11GB),在 12GB 显存加 8GB 内存的机器上把 Qwen3.8-27B 跑到约 18 tok/sec 解码、64k 上下文下约 500-600 tok/sec 预填充。

正文

I got Qwen3.8-27B running at ~18 tok/sec decode & ~500-600 tok/sec prefill (at 64k context) on just 12GB VRAM + 8GB RAM, using ISTA-DASLab Qwen3.8-27B-GSQ-RCO IQ3_S quant (~11GB). This recipe also works with 0bserverx’ Qwen3.8-27B-Heretic-GSQ-RCO IQ3_S quant, achieving similar speeds (about a 7% loss).

This is the best Qwen3.8-27B quant I’ve tested so far (and I’ve tried everything), and for it to fit in such limited RAM/VRAM is wild. GSQ-RCO quantization is magic, it performs very close to the full precision weights in all of my testing.

The reason it fits at all is Qwen3.8 is hybrid, so only 16 of the 64 layers need KV cache. With q4_0 for cache the full 64k is only about 1.1GB instead of 4GB for f16.

I'm on a 9900X + 4070S 12GB + 32GB RAM for reference, using stock llama.cpp. Settings are -ngl 58 -ot token_embd=CPU -ctk q4_0 -ctv q4_0 -c 64000.

Full build + serve scripts and all my numbers are here if you’d like to reproduce yourselves: https://github.com/bodhi37/Qwen3.8-27B-12GBVRAM-Recipe

submitted by /u/bodhi371
[link] [留言]

来源:r/LocalLLaMA · reddit.com