跳到正文
r/LocalLLaMA· /u/KnownAd4832·· 3 小时前AI 评分44

Qwen3.8-Flash-Next 在 Strata 上支持 Strix Halo

Qwen3.8-Flash-Next on Strata

AI 导读

Strata 现已官方支持在 Strix Halo 机器上运行 Qwen3.8-Flash-Next,长上下文解码与 ppts 表现最佳,使用 Unsloth 的 Q4 和 GSQ-RCO 权重。该支持可扩展至 1M 上下文长度且速度损失不大,目前标记为实验性,仅在 Linux 上完成。

正文
Qwen3.8-Flash-Next on Strata

Hey! 👋

I have released an official support for Strix Halo machines on Strata for Qwen3.8-Flash-Next.

Currently numbers are the best on long context decode and ppts using typical Unsloth’s Q4 and GSQ-RCO model weights.

Can go up to 1M context length without big speed loss. Currently support is marked as experimental and was done on Linux only.

https://github.com/Niko1221/Strata/

Will be happy for any feedback and pull requests you could give! 👀

submitted by /u/KnownAd4832
[link] [留言]

来源:r/LocalLLaMA · reddit.com