Hugging Face Blog·· 2024-10-29精选AI 评分62
Intel Labs 与 Hugging Face 提出 Universal Assisted Generation,可用任意助手模型加速解码
Universal Assisted Generation: Faster Decoding with Any Assistant Model
AI 导读
Intel Labs 与 Hugging Face 提出 Universal Assisted Generation(UAG),通过双向 tokenizer 转换让目标模型与任意模型族的助手模型配对,从而加速推理。
推荐理由
Intel Labs 与 Hugging Face 提出跨 tokenizer 的通用辅助生成方法,读者可了解其如何让任意小模型充当助手并带来 1.5x-2.0x 推理加速。
来源:Hugging Face Blog · huggingface.co