Hugging Face Blog·· 2023-05-16AI 评分34
BigCode 大规模近重复数据去重背后的技术
Large-scale Near-deduplication Behind BigCode
AI 导读
Hugging Face 博客详解 BigCode 项目采用 MinHash + LSH 进行大规模文档级近重复去重的方法,参数为 (256, 0.7, 5)。该方案已在 The Stack 代码数据集上验证,去重能提升代码模型性能并减少训练步数。文章还对比了 InCoder、CodeGen、AlphaCode 等模型所用的精确匹配与 MinHash 去重策略。
来源:Hugging Face Blog · huggingface.co