跳到正文
Rohan Paul· @rohanpaul_ai · X·· 3 小时前AI 评分53
AI 导读

Overmind 将用户自己的生产 trace 用于微调小型开源模型,再用同一批真实任务构建的评测与现有模型对比。其公布的对比为经 Overmind 微调的 Qwen3.5 9B 对 GPT5.6 Luna,公司称在法律合同中幻觉条款减少 20 至 30 倍,逐字引用条款的准确率提升 7 倍。产出的模型权重归用户所有,可托管在 Overmind 或自行运行。

正文

A frontier model reading a contract will sometimes cite a clause that isn't there.

Overmind fine-tunes a small open model on your own production traces, then scores it against your current model on evals built from those same real tasks.

Their published comparison is Qwen3.5 9B tuned through Overmind against GPT5.6 Luna. The company reports 20 to 30x fewer phantom clauses in legal contracts and 7x better accuracy at quoting a clause word for word.

Here the observability layer and the training layer share the same data. The traces that show you where the agent fails become the dataset, and the evals are built from real tasks rather than a public benchmark.

The output is a smaller open model with weights you own. You can host it on Overmind or run it yourself.

引用Overmind@OvermindLab
Overmind turns anyone into an AI lab. • Builds a context graph from code and traces • Curates training and eval datasets • Evaluates prompts and models • Trains smaller, specialized models https://github.com/overmind-core/overmind
在 X 查看被引用的帖子

来源:Rohan Paul · x.com