Anthropic 详解 Claude 3.5 Sonnet 如何在 SWE-bench Verified 达到 49%
Raising the bar on SWE-bench Verified with Claude 3.5 Sonnet
Anthropic 升级版 Claude 3.5 Sonnet 在 SWE-bench Verified 上取得 49%,超过此前最优模型的 45%。文章介绍了围绕模型搭建的智能体脚手架,只用一个提示词、Bash Tool 和 Edit Tool 两个通用工具,尽量把控制权交给模型本身。
Anthropic 公开了让 Claude 3.5 Sonnet 在 SWE-bench Verified 拿到 49% 的智能体脚手架与工具设计细节,可供开发者复现。
我们最新的模型,升级版 Claude 3.5 Sonnet,在软件工程评估 SWE-bench Verified 上取得了 49% 的成绩,超过了此前最先进模型的 45%。本文介绍了我们围绕该模型构建的“agent”,旨在帮助开发者充分发挥 Claude 3.5 Sonnet 的性能。
SWE-bench 是一个 AI 评估基准,用于衡量模型完成真实世界软件工程任务的能力。具体而言,它测试模型如何解决来自热门开源 Python 仓库的 GitHub issue。对于基准中的每个任务,AI 模型会获得一个配置好的 Python 环境,以及该仓库在 issue 被解决之前的检出(本地工作副本)。模型随后需要理解、修改并测试代码,然后提交其解决方案。
每个解决方案都会根据关闭原始 GitHub issue 的 pull request 中的真实单元测试进行评分。这测试了 AI 模型是否能够实现与 PR 原始人类作者相同的功能。
SWE-bench 并非孤立地评估 AI 模型,而是评估整个“agent”系统。在此语境下,“agent”指的是 AI 模型与其周围软件脚手架的组合。该脚手架负责生成输入模型的提示、解析模型输出以采取行动,以及管理交互循环,将模型先前行动的结果纳入其下一次提示中。即使使用相同的底层 AI 模型,agent 在 SWE-bench 上的表现也可能因脚手架的不同而有显著差异。
还有许多其他用于评估大型语言模型编码能力的基准,但 SWE-bench 因以下几个原因而日益流行:
- 它使用来自实际项目的真实工程任务,而非竞赛或面试风格的问题;
- 它尚未饱和——仍有很大的改进空间。目前还没有模型在 SWE-bench Verified 上突破 50% 的完成率(不过截至撰写本文时,更新后的 Claude 3.5 Sonnet 已达到 49%);
- 它衡量的是整个“agent”,而非孤立的模型。开源开发者和初创公司已成功优化脚手架,从而在同一模型上大幅提升性能。
请注意,原始 SWE-bench 数据集包含一些任务,若没有 GitHub issue 之外的额外上下文(例如关于要返回的特定错误消息),则无法解决。SWE-bench-Verified 是 SWE-bench 的一个包含 500 个问题的子集,已经过人工审查以确保其可解决,因此能最清晰地衡量编码 agent 的性能。本文中我们将引用这一基准。
实现最先进水平
工具使用型 Agent
我们为更新版 Claude 3.5 Sonnet 优化 agent 脚手架时的设计理念是,尽可能将控制权交给语言模型本身,并保持脚手架最小化。该 agent 有一个提示、一个用于执行 bash 命令的 Bash Tool,以及一个用于查看和编辑文件及目录的 Edit Tool。我们持续采样,直到模型决定完成,或超出其 200k 上下文长度。这种脚手架允许模型自行判断如何解决问题,而不是被硬编码到特定模式或工作流程中。
提示词概述了建议模型采用的方法,但对于这项任务来说,它并不过长或过于详细。模型可以自由选择如何从一个步骤过渡到下一个步骤,而不是进行严格且离散的转换。如果你对 token 不敏感,明确鼓励模型生成较长的回复可能会有所帮助。
以下代码展示了我们 agent 脚手架中的提示词:
<uploaded_files>
{location}
</uploaded_files>
I've uploaded a python code repository in the directory {location} (not in /tmp/inputs). Consider the following PR description:
<pr_description>
{pr_description}
</pr_description>
Can you help me implement the necessary changes to the repository so that the requirements specified in the <pr_description> are met?
I've already taken care of all changes to any of the test files described in the <pr_description>. This means you DON'T have to modify the testing logic or any of the tests in any way!
Your task is to make the minimal changes to non-tests files in the {location} directory to ensure the <pr_description> is satisfied.
Follow these steps to resolve the issue:
1. As a first step, it might be a good idea to explore the repo to familiarize yourself with its structure.
2. Create a script to reproduce the error and execute it with `python <filename.py>` using the BashTool, to confirm the error
3. Edit the sourcecode of the repo to resolve the issue
4. Rerun your reproduce script and confirm that the error is fixed!
5. Think about edgecases and make sure your fix handles them as well
Your thinking should be thorough and so it's fine if it's very long.模型的第一个工具执行 Bash 命令。其 schema 很简单,只接收要在环境中运行的命令。然而,工具的描述承载了更多分量。它包含给模型的更详细指令,包括转义输入、无互联网访问,以及如何在后台运行命令。
接下来,我们展示 Bash 工具的规范:
{
"name": "bash",
"description": "Run commands in a bash shell\n
* When invoking this tool, the contents of the \"command\" parameter does NOT need to be XML-escaped.\n
* You don't have access to the internet via this tool.\n
* You do have access to a mirror of common linux and python packages via apt and pip.\n
* State is persistent across command calls and discussions with the user.\n
* To inspect a particular line range of a file, e.g. lines 10-25, try 'sed -n 10,25p /path/to/the/file'.\n
* Please avoid commands that may produce a very large amount of output.\n
* Please run long lived commands in the background, e.g. 'sleep 10 &' or start a server in the background.",
"input_schema": {
"type": "object",
"properties": {
"command": {
"type": "string",
"description": "The bash command to run."
}
},
"required": ["command"]
}
}模型的第二个工具(Edit 工具)要复杂得多,包含了模型查看、创建和编辑文件所需的一切。同样,我们的工具描述中包含了给模型的关于如何使用该工具的详细信息。
我们在各种 agentic 任务中为这些工具的描述和规范投入了大量精力。我们测试它们,以发现模型可能误解规范的任何方式,或使用这些工具可能存在的陷阱,然后编辑描述以预先防范这些问题。我们认为,应该像为人类设计工具界面时投入大量注意力那样,为模型设计工具界面投入更多注意力。
以下代码展示了我们 Edit 工具的描述:
{
"name": "str_replace_editor",
"description": "Custom editing tool for viewing, creating and editing files\n
* State is persistent across command calls and discussions with the user\n
* If `path` is a file, `view` displays the result of applying `cat -n`. If `path` is a directory, `view` lists non-hidden files and directories up to 2 levels deep\n
* The `create` command cannot be used if the specified `path` already exists as a file\n
* If a `command` generates a long output, it will be truncated and marked with `<response clipped>` \n
* The `undo_edit` command will revert the last edit made to the file at `path`\n
\n
Notes for using the `str_replace` command:\n
* The `old_str` parameter should match EXACTLY one or more consecutive lines from the original file. Be mindful of whitespaces!\n
* If the `old_str` parameter is not unique in the file, the replacement will not be performed. Make sure to include enough context in `old_str` to make it unique\n
* The `new_str` parameter should contain the edited lines that should replace the `old_str`",
...我们改进性能的一种方式是对工具进行“防错”处理。例如,有时在 agent 移出根目录后,模型可能会弄错相对文件路径。为防止这种情况,我们直接让工具始终要求使用绝对路径。
我们尝试了多种不同的策略来指定对现有文件的编辑,其中字符串替换的可靠性最高,即模型指定 `old_str`,在给定文件中用 `new_str` 替换它。只有当 `old_str` 恰好匹配一次时才会执行替换。如果匹配次数多于或少于一次,模型会看到相应的错误消息以便重试。
我们 Edit 工具的规范如下所示:
...
"input_schema": {
"type": "object",
"properties": {
"command": {
"type": "string",
"enum": ["view", "create", "str_replace", "insert", "undo_edit"],
"description": "The commands to run. Allowed options are: `view`, `create`, `str_replace`, `insert`, `undo_edit`."
},
"file_text": {
"description": "Required parameter of `create` command, with the content of the file to be created.",
"type": "string"
},
"insert_line": {
"description": "Required parameter of `insert` command. The `new_str` will be inserted AFTER the line `insert_line` of `path`.",
"type": "integer"
},
"new_str": {
"description": "Required parameter of `str_replace` command containing the new string. Required parameter of `insert` command containing the string to insert.",
"type": "string"
},
"old_str": {
"description": "Required parameter of `str_replace` command containing the string in `path` to replace.",
"type": "string"
},
"path": {
"description": "Absolute path to file or directory, e.g. `/repo/file.py` or `/repo`.",
"type": "string"
},
"view_range": {
"description": "Optional parameter of `view` command when `path` points to a file. If none is given, the full file is shown. If provided, the file will be shown in the indicated line number range, e.g. [11, 12] will show lines 11 and 12. Indexing at 1 to start. Setting `[start_line, -1]` shows all lines from `start_line` to the end of the file.",
"items": {
"type": "integer"
},
"type": "array"
}
},
"required": ["command", "path"]
}
}结果
总体而言,升级后的 Claude 3.5 Sonnet 在推理、编码和数学能力上均优于我们之前的模型,以及此前的最先进模型。它还展现出更强的 agentic 能力:这些工具和脚手架有助于将其改进后的能力发挥到最佳。
| 模型 | Claude 3.5 Sonnet(新版) | 此前的最先进 | Claude 3.5 Sonnet(旧版) | Claude 3 Opus |
|---|---|---|---|---|
| SWE-bench Verified 得分 | 49% | 45% | 33% | 22% |
agent 行为示例
为运行该基准测试,我们使用 SWE-Agent 框架作为 agent 代码的基础。在下面的日志中,我们将 agent 的文本输出、工具调用和工具响应分别呈现为 THOUGHT、ACTION 和 OBSERVATION,尽管我们并未将模型限制为固定的顺序。
下面的代码块将展示 Sonnet 3.5 解决一个 SWE-bench 问题的典型案例。
在第一个代码块中,你可以看到给模型的初始提示词的一部分,其中 `{pr_description}` 已填入来自某个 SWE-bench 任务的真实值。重要的是,该任务包含重现问题的步骤,这将为模型提供一个有价值的调查起点。
<uploaded_files>
/repo
</uploaded_files>
I've uploaded a python code repository in the directory /repo (not in /tmp/inputs). Consider the following PR description:
<pr_description>
linear_model.RidgeClassifierCV's Parameter store_cv_values issue
#### Description
Parameter store_cv_values error on sklearn.linear_model.RidgeClassifierCV
#### Steps/Code to Reproduce
import numpy as np
from sklearn import linear_model as lm
#test database
n = 100
x = np.random.randn(n, 30)
y = np.random.normal(size = n)
rr = lm.RidgeClassifierCV(alphas = np.arange(0.1, 1000, 0.1), normalize = True,
store_cv_values = True).fit(x, y)
#### Expected Results
Expected to get the usual ridge regression model output, keeping the cross validation predictions as attribute.
#### Actual Results
TypeError: __init__() got an unexpected keyword argument 'store_cv_values'
lm.RidgeClassifierCV actually has no parameter store_cv_values, even though some attributes depends on it.模型作出响应,首先使用 Edit Tool 查看仓库结构。你可以在 THOUGHT 和 ACTION 下看到模型的文本输出和工具调用参数,并在 OBSERVATION 下看到部分工具输出:
THOUGHT:
I'll help you implement the changes to satisfy the PR description. Let's follow the steps you outlined.
1. First, let's explore the repo structure:
ACTION (str_replace_editor):
command:
view
path:
/repo
OBSERVATION:
Here's the files and directories up to 2 levels deep in /repo, excluding hidden items:
/repo
/repo/PULL_REQUEST_TEMPLATE.md
/repo/ISSUE_TEMPLATE.md
/repo/sklearn
/repo/sklearn/decomposition
/repo/sklearn/kernel_ridge.py
/repo/sklearn/model_selection
/repo/sklearn/covariance
/repo/sklearn/cluster
/repo/sklearn/base.py
/repo/sklearn/dummy.py
/repo/sklearn/naive_bayes.py
...现在模型对仓库结构有了更好的理解,它使用 Edit Tool 创建一个新脚本,用来复现问题并测试其修复:
THOUGHT:
Based on the repository structure, we'll need to modify the Ridge Classifier in the `sklearn/linear_model` directory.
Let's create a script to reproduce the error:
ACTION (str_replace_editor):
command:
create
path:
/repo/reproduce_error.py
file_text:
import numpy as np
from sklearn import linear_model as lm
#test database
n = 100
x = np.random.randn(n, 30)
y = np.random.normal(size = n)
rr = lm.RidgeClassifierCV(alphas=np.arange(0.1, 1000, 0.1), normalize=True,
store_cv_values=True).fit(x, y)
OBSERVATION:
File created successfully at: /repo/reproduce_error.py随后,模型使用 Bash Tool 执行它编写的脚本,并成功复现了任务中的问题:
THOUGHT:
Let's verify that we can reproduce the error:
ACTION (bash):
command:
python3 /repo/reproduce_error.py
OBSERVATION:
Traceback (most recent call last):
File "/repo/reproduce_error.py", line 10, in <module>
store_cv_values=True).fit(x, y)
TypeError: __init__() got an unexpected keyword argument 'store_cv_values'从这里开始,模型使用 Edit Tool 修改仓库中的源代码,并重新运行脚本以验证修改是否解决了问题:
THOUGHT:
The error is reproducible. Looking at the code, I can see that the issue is that while the `RidgeClassifierCV` class inherits from `_BaseRidgeCV`, it fails to pass the `store_cv_values` parameter from its constructor through to the base class constructor.
Let's fix this:
ACTION (str_replace_editor):
command:
str_replace
path:
/repo/sklearn/linear_model/ridge.py
old_str:
def __init__(self, alphas=(0.1, 1.0, 10.0), fit_intercept=True,
normalize=False, scoring=None, cv=None, class_weight=None):
super(RidgeClassifierCV, self).__init__(
alphas=alphas, fit_intercept=fit_intercept, normalize=normalize,
scoring=scoring, cv=cv)
self.class_weight = class_weight
new_str:
def __init__(self, alphas=(0.1, 1.0, 10.0), fit_intercept=True,
normalize=False, scoring=None, cv=None, class_weight=None,
store_cv_values=False):
super(RidgeClassifierCV, self).__init__(
alphas=alphas, fit_intercept=fit_intercept, normalize=normalize,
scoring=scoring, cv=cv, store_cv_values=store_cv_values)
self.class_weight = class_weight在这个具体示例中,模型工作了 12 步后才决定可以提交。随后任务的测试成功运行,验证了模型的解决方案解决了问题。有些任务在模型提交解决方案前需要超过 100 轮;在其他任务中,模型一直尝试,直到上下文耗尽。
通过审查更新版 Claude 3.5 Sonnet 与旧模型的尝试对比,更新版 3.5 Sonnet 更频繁地自我纠正。它还展现出尝试多种不同解决方案的能力,而不是反复陷入同一个错误。
挑战
SWE-bench Verified 是一个强大的评估,但运行起来也比简单的单轮评估更复杂。以下是我们在使用它时面临的一些挑战——其他 AI 开发者也可能遇到这些挑战。
- 时长和高 token 成本。 上述示例来自一个在 12 步内成功完成的案例。然而,许多成功运行需要模型花费数百轮才能解决,并且消耗 >100k tokens。更新版 Claude 3.5 Sonnet 很有韧性:只要有足够时间,它通常能找到绕过问题的办法,但这可能代价高昂;
- 评分。 在检查失败任务时,我们发现了模型行为正确、但存在环境设置问题,或安装补丁被应用两次的问题的案例。解决这些系统问题对于准确了解 AI agent 的表现至关重要。
- 隐藏测试。 由于模型看不到用于评分的测试,它常常“认为”自己已经成功,而任务实际上失败了。其中一些失败是因为模型在错误的抽象层级上解决了问题(打补丁而不是进行更深入的 refactor)。其他失败则显得不那么公平:它们解决了问题,但与原始任务中的单元测试不匹配。
- 多模态。 尽管更新版 Claude 3.5 Sonnet 具备出色的视觉和多模态能力,我们并未实现让它查看保存到文件系统的文件或作为 URL 引用的文件的方式。这使得调试某些任务(尤其是来自 Matplotlib 的任务)特别困难,也容易导致模型幻觉。开发者在这里显然有容易改进的空间——而且 SWE-bench 已经推出了一个专注于多模态任务的新 评估。我们期待在不久的将来看到开发者用 Claude 在此评估上取得更高分数。
升级后的 Claude 3.5 Sonnet 在 SWE-bench Verified 上达到了 49%,击败了之前的最先进水平(45%),仅使用一个简单的提示和两个通用工具。我们相信,使用新版 Claude 3.5 Sonnet 进行开发的开发者将很快找到新的、更好的方法来提高 SWE-bench 分数,超越我们在此初步展示的成果。
致谢
Erik Schluntz 优化了 SWE-bench 智能体并撰写了这篇博客文章。Simon Biggs、Dawn Drain 和 Eric Christiansen 帮助实现了该基准测试。Shauna Kravec、Dawn Drain、Felipe Rosso、Nova DasSarma、Ven Chandrasekaran 以及许多其他人为训练 Claude 3.5 Sonnet 使其擅长智能体编程做出了贡献。
来源:Anthropic Engineering · anthropic.com