跳到正文
Unsloth Blog·· 2024-04-24精选AI 评分59

Unsloth 支持 Llama 3 微调,8B 训练提速 2 倍

Llama 3 is now on Unsloth 🦙

AI 导读

Unsloth 宣布支持 Meta 的 Llama 3 模型微调,称 Llama 3 8B 训练速度比 Flash Attention 2 加 Hugging Face 快 2 倍、内存占用减少 63%,70B 版本快 1.8 倍、显存减少 68%。

推荐理由

Unsloth 给出 Llama 3 微调的加速与显存数据,并附免费 Colab 笔记本,可据此评估训练成本。

正文 · 原文

Meta's new Llama 3 models are the most capable open LLMs to date - outperforming many open models on industry standard benchmarks.

Unsloth makes Llama 3 (8B) model training 2x faster and use 63% less memory than Flash Attention 2 + Hugging Face. Llama 3 (70B) is 1.8x faster and uses 68% less VRAM.

To train your own Llama 3 model for free, we uploaded a Google Colab notebook to finetune Llama 3 (8B): Notebook.

We also uploaded pre-quantized 4bit models for 4x faster downloading to our Hugging Face page. On one A100 80GB GPU, Llama 3 (70B) with Unsloth can now fit 48K total tokens (8192 * bsz of 5) vs 7K tokens without Unsloth. That's 6x longer context lengths!

P.S. Don't forget to ⭐Star us on GitHub and join our Discord server ❤️

来源:Unsloth Blog · unslothai.substack.com