Hugging Face Blog·· 2026-08-25精选AI 评分60
Multiverse Computing 提出 Quantization-Aware Healing:4-bit 压缩模型在 9 项基准中 7 项超过其全精度版本
Quantization-Aware Healing: a compressed, 4-bit model that outperforms its full-precision original
AI 导读
Multiverse Computing 提出 Quantization-Aware Healing(QAH),把 GPT-OSS 120B 压缩到 60B 参数并量化为 MXFP4 后,在 9 项基准中有 7 项超过该 60B 模型的 bfloat16 版本。
推荐理由
论文给出压缩加量化后从原始模型蒸馏的修复配方,并附 9 项基准对比,可据此判断 4-bit 部署的精度取舍。
来源:Hugging Face Blog · huggingface.co