Swift 1.5 Qwen3.8 Flash Next:面向 96GB Mac Studio M5 Ultra 的 MLX 量化版
Swift1.5 Qwen3.8 Flash Next - Tailored for the 96GB Mac Studio with M5 Ultra
开发者针对 96GB Mac Studio M5 Ultra 发布了 UkisAI Swift 1.5 的 MLX 量化版本,称其为目前校准效果最好的尝试。
https://huggingface.co/Dankpaws/Swift1.5-Qwen3.8-Flash-Next-MLX-4.7bpw
I've had the 96GB Mac Studio with M5 Ultra for about a week now and wasn't satisfied with the results I was getting from the limited number of models available to me. It was a combination of speed, memory headroom, and/or output quality.
This is my best attempt at a calibrated MLX quantization of UkisAI’s Swift 1.5. Hope those of you with the hardware enjoy it!
| Measurement | This pack | Swift llama.cpp IQ3_XXS |
|---|---|---|
| Prefill · 25k prompt | 3,191 tok/s | 1,427 tok/s |
| Prefill · 95k prompt | 2,928 tok/s | 1,307 tok/s |
| Decode · after 4k prompt | 113.7 tok/s | 62.7 tok/s |
| Decode · after 95k prompt | 81.1 tok/s | 44.7 tok/s |
| Top-1 agreement with Swift BF16 | 91.0% | 84.1% |
91% is next-token agreement with BF16 across 680 common held-out positions, not task accuracy.
Results above are simply from my own machine. mlx-serve 26.10.1. ~107GB download, text-only, 179,200-token tested context.
submitted by /u/DankpawsDev
[link] [留言]
来源:r/LocalLLaMA · reddit.com