Microsoft 及合作者发布论文 ScholarEvolve,提出依据已发表的智能体研究而非智能体自身失败日志来演化 agent harness。该方法把 harness 拆分为工具使用、记忆管理和任务执行三个模块,对近期论文做主题建模以找出各模块的改进策略,再实现并测试不同组合,且可持续加入新论文。
New paper from Microsoft and colleagues on evolving agent harnesses.
It's a really cool idea to evolve a harness from published research. Something I have also been testing for the past couple of months.
ScholarEvolve proposes harness changes based on published agent research rather than the agent's own failure logs.
It splits the harness into modules for tool use, memory management, and task execution.
It runs topic modeling over recent papers to identify distinct improvement strategies for each module, then implements and tests combinations.
New papers can be added over time.
With the model held fixed, Qwen3.5-27B goal completion on AppWorld Challenge rises from 49.6% to 63.6%, and GPT-5.4-mini on Tau2-Bench Telecom rises from 72.7% to 81.9%.
Paper: https://arxiv.org/abs/2609.40169
Chat with Paper: https://academy.dair.ai/papers/learning-from-research-toward-lifelong-agent-harness-evolution-2609.40169
来源:elvis · x.com