跳到正文
Google Developers Blog·· 13 天前精选AI 评分65

Google 在 MaxText 中复现 Ai2 的 Olmo 3 7B 预训练

Reproducing Olmo 3 7B Pre-training in MaxText: case study of large scale training on TPUs

AI 导读

Google 团队在 Google Cloud TPU 上用 MaxText 从零复现了 Ai2 的 Olmo 3 7B 预训练与中期退火,覆盖约 5.93T token、1.41M 步的 stage-1 和 47,684 步的 stage-2,并在留出集指标上验证匹配。

推荐理由

完整记录了一次跨框架复现实验,其中数据加载 bug 伪装成性能提升的过程对训练验证方法有参考价值。

来源:Google Developers Blog · developers.googleblog.com