跳到正文
Rohan Paul· @rohanpaul_ai · X·· 4 小时前AI 评分52
AI 导读

微软一篇新论文发现,编码智能体出错主要发生在需要理解大量代码时,而不是需要修改大量代码时,因此建议用阅读和比较代码来测试它们,而非看 diff 大小。研究者构建了 CABRA,可生成合成编码任务并每次只提高一种难度,在 6,840 个任务上跑了 8 个 LLM 和 6 个智能体,并把每次工具调用标注为阅读、分析、搜索、编辑或测试。

正文

New Microsoft paper finds that coding agents trip up when they have to understand a lot of code, not when they have to edit a lot of it, so test them on reading and comparing code instead of diff size.

Microsoft researchers built CABRA, which generates synthetic coding tasks and raises 1 kind of difficulty at a time. They ran 8 LLMs and 6 agents on 6,840 tasks and labeled each tool call as reading, analyzing, searching, editing, or testing.

Plain LLMs got worse as tasks grew, but agents stayed near-perfect by using tools like grep. On SWE-bench Verified, the count of reading and analysis calls tracked agent failures better than lines edited, with correlations of -0.200 versus -0.159.

来源:Rohan Paul · x.com