Sakana AI 提出 MASS,用多智能体自监督替代外部验证器,让无检查器的开放式任务也能进入自我改进循环。同一基座模型提出并运行多智能体工作流、自行打分,再用进化搜索保留高分工作流,并在自身轨迹上微调,改进后的模型进入下一轮循环。
Recommended paper from Sakana AI on recursive self-improvement.
They propose an interesting way to scale recursive self-improvement through multi-agent self-supervision.
In this line of research, self-improvement loops usually need an external verifier, so open-ended tasks without a checker are left out.
MASS removes that requirement.
One base model proposes multi-agent workflows, runs them and grades them, and an evolutionary search keeps the workflows that score best. The model is then fine-tuned on its own traces, and the improved model starts the next cycle as a better optimizer and grader.
Two cycles on Qwen3.6-27B raise performance per output token from 1.2 to 1.6x on four open-ended benchmarks.
A student trained on multi-agent traces also beats a single-agent student trained on 1.4x more tokens.
Paper: https://arxiv.org/abs/2610.12176
Chat with Paper: https://academy.dair.ai/papers/recursive-self-improvement-through-multi-agent-self-supervision-2610.12176
来源:elvis · x.com