跳到正文
r/LocalLLaMA· /u/lkarlslund·· 4 小时前AI 评分56

NInfer6000 分支在 RTX6000 上跑 Qwen 3.8 Flash Next,达 400 tg/s

NInfer6000 - Qwen 3.8 Flash Next @ 400 tg/s & 13K pp/s

AI 导读

作者发布 NInfer 的 fork 版本 NInfer6000,面向 RTX6000 96GB 显卡运行 Qwen 3.8 Flash Next,理想条件下解码速度达到 400 tg/s。

正文

I've tinkered some more with my fork of NInfer for the Qwen 3.8 Flash Next model on RTX6000, and I've just hit 400tg/s under ideal conditions with it, so I thought it was worth a share.

Decode MTP3 with --lm-head-draft

Context 16-bit 8-bit Change
512 197.1 tok/s 274.8 tok/s +39%
8K 303.1 tok/s 401.3 tok/s +32%
64K 291.7 tok/s 380.3 tok/s +30%
128K 282.6 tok/s 368.0 tok/s +30%
256K (maximum) 277.4 tok/s 360.6 tok/s +30%

Prefill

Prompt length 16-bit 8-bit Change
512 6,709 tok/s 5,905 tok/s -12%
8K 13,903 tok/s 13,908 tok/s 0%
64K 13,043 tok/s 12,171 tok/s -7%
128K 11,927 tok/s 11,032 tok/s -8%
256K (maximum) 9,941 tok/s 9,154 tok/s -8%

More benchmark variants in the readme in the repo

The original NInfer is for 5090 cards 32GB and variants below that, but I was both missing Qwen 3.8 Flash Next in it (when I started the fork) and something that could properly use a RTX6000 96GB card. The performance and options in VLLM and llama.cpp offerings just didn't really cut it for me, so I've vibed on this for some weeks now.

This fork supports both the NVFP4 quants from "radixark" and the "Swift 1.5" variant with 'less thinking but same results' post-training. With non-experts downsampled from 16-bit to 8-bit, MTP3 and smaller drafting head you get up to 400 tokens per second. You can also opt not to do the downsampling at a performance cost, but a bit higher quality.

Vision is also supported. Have fun.

https://github.com/lkarlslund/ninfer6000

submitted by /u/lkarlslund
[link] [留言]

来源:r/LocalLLaMA · reddit.com