跳到正文
r/LocalLLaMA· /u/DankpawsDev·· 4 小时前AI 评分22

Swift 1.5 Qwen3.8 Flash Next:面向 96GB Mac Studio M5 Ultra 的 MLX 量化版

Swift1.5 Qwen3.8 Flash Next - Tailored for the 96GB Mac Studio with M5 Ultra

AI 导读

开发者针对 96GB Mac Studio M5 Ultra 发布了 UkisAI Swift 1.5 的 MLX 量化版本,称其为目前校准效果最好的尝试。

正文

https://huggingface.co/Dankpaws/Swift1.5-Qwen3.8-Flash-Next-MLX-4.7bpw

I've had the 96GB Mac Studio with M5 Ultra for about a week now and wasn't satisfied with the results I was getting from the limited number of models available to me. It was a combination of speed, memory headroom, and/or output quality.

This is my best attempt at a calibrated MLX quantization of UkisAI’s Swift 1.5. Hope those of you with the hardware enjoy it!

Measurement This pack Swift llama.cpp IQ3_XXS
Prefill · 25k prompt 3,191 tok/s 1,427 tok/s
Prefill · 95k prompt 2,928 tok/s 1,307 tok/s
Decode · after 4k prompt 113.7 tok/s 62.7 tok/s
Decode · after 95k prompt 81.1 tok/s 44.7 tok/s
Top-1 agreement with Swift BF16 91.0% 84.1%

91% is next-token agreement with BF16 across 680 common held-out positions, not task accuracy.

Results above are simply from my own machine. mlx-serve 26.10.1. ~107GB download, text-only, 179,200-token tested context.

submitted by /u/DankpawsDev
[link] [留言]

来源:r/LocalLLaMA · reddit.com