Overmind 开源了其平台,可用智能体实际工作记录构建训练与评测数据集,并微调开源模型。作者引用其自测数据称,在 100 多份合同、4000 多个问题的测试中,训练后模型编造引文的比率为 0.12%,对比的前沿模型为 3.36%,低 28 倍。代码已发布在 GitHub,用户可保留训练模型的权重并自行部署运行。
Overmind has open-sourced its platform for training smaller models on the work your agents actually do.
In its own test of 4,000+ questions across 100+ contracts, Overmind reports that its trained model invented quotes on 0.12% of questions, compared with 3.36% for the frontier model it tested. That's a 28x lower rate.
It uses records of your agents' real work to build training data and evals, then fine-tunes open models. For an agent pulling clauses out of contracts all day, training a model for that specific job makes sense.
I'm a big fan of open source, so its cool that they released the code. You keep the trained model's weights and can run the platform yourself.
Overmind turns anyone into an AI lab. • Builds a context graph from code and traces • Curates training and eval datasets • Evaluates prompts and models • Trains smaller, specialized models https://github.com/overmind-core/overmind在 X 查看被引用的帖子
来源:Chubby♨️ · x.com