Unsloth 新增 AMD GPU 训练支持并重做 Studio
Unsloth x AMD
Unsloth 与 AMD 合作,支持在 AMD Radeon、Instinct、Ryzen 及数据中心 GPU 上本地训练和运行 500 多个 LLM,覆盖 Windows、WSL 和 Linux,官方称仅需 3GB VRAM 即可训练,速度提升 2 倍、显存占用减少 70% 且无精度损失,并提供针对 GGUF 和 Safetensors 推理优化的 ROCm 构建。
Unsloth 新增 AMD GPU 本地训练支持并重做 Studio,读者可据此判断本地微调的硬件选择面。
Hey there! It’s been a while since our last newsletter and we have a lot of updates. We’ve added AMD training support and completely reimagined Unsloth Studio, introducing a new Hub, expanded customization, more language support, many new features, models, and much more.
💚 Unsloth x AMD
We collaborated with AMD to enable you to train and run LLMs locally on AMD GPUs across 500+ models. Works on AMD Radeon, Instinct, Ryzen and data center GPUs across Windows, WSL, and Linux. You can train models 2× faster with 70% less VRAM and no accuracy loss on just 3GB VRAM. We also have optimized ROCm builds for GGUF and Safetensors inference. GitHub repo
🧭 New Model Hub
We’re introducing a new Model Hub which makes it easier to discover, download and run models, with a trending feed, search, README previews and a resumable download manager. We also rebuilt model context auto-fitting, delivering up to ~3× longer context in several scenarios.
⚡ Dynamic NVFP4 quants + Better export
We’ve released new Dynamic NVFP4 quants for NVIDIA Blackwell GPUs, covering Qwen3.6 and Gemma 4. They enable faster and more accurate 4-bit inference. Unsloth can now also create merged FP8 and NVFP4 exports after training, plus imatrix-assisted GGUF quantization.
💎 Gemma 4 improvements
The full Gemma 4, QAT family now runs and trains in Unsloth, including the new 12B model alongside E2B, E4B, 26B-A4B and 31B. We also uploaded new Dynamic NVFP4 quants. Recently, Google also added better tool calling and MTP for faster and improved local inference.
🎨 Make Unsloth yours
Unsloth Studio’s appearance is now fully customizable, palettes, custom accent/background colors, fonts and accessibility settings, all synced across devices. We also now support 12 languages including Japanese, Chinese, Portuguese, Hindu, Arabic and others. Local transcription is also coming this week!
🐳 Better model and GPU control
We now have much more customization for model loading. Loading now exposes Flash Attention and tensor-parallel options. A new GPU Memory panel lets you choose which GPUs to use, set GPU layers and control MoE expert offload, or leave everything to Unsloth’s automatic fit mode.
🔮 New models
DeepSeek-V4-Flash, GLM-5.2, Qwen3.6 - run the new agentic models
Inkling, DiffusionGemma, Qwen-AgentWorld, Ornith, Kimi K2.7 Code and MiniMax-M3 - are all now available and supported in Unsloth.
Hope you have a lovely rest of July! We’ve got even more hardware, Unsloth and model updates on the way. 🦥
来源:Unsloth Blog · unslothai.substack.com