跳到正文
r/MachineLearning· /u/heyitsdannyle·· 5 小时前AI 评分53

SWE-Race:188 个真实并发缺陷的编码智能体基准,含三个模型结果

SWE-Race: a coding-agent benchmark of 188 real concurrency bugs, with results from three models [P]

AI 导读

SWE-Race 是一个由约 100 个 Python 项目已合并 PR 中真实并发缺陷(竞态、死锁、取消问题)构成的编码智能体基准,共 188 个任务,每个任务用项目自带测试在无网络容器中评分,仓库被裁剪到单个 commit 以防从 git 历史找回修复。

正文
SWE-Race: a coding-agent benchmark of 188 real concurrency bugs, with results from three models [P]

We've been building a benchmark out of real concurrency bugs (race conditions, deadlocks, cancellation issues) taken from merged PRs in about 100 Python projects. Each task gets graded by the project's own tests, in a container with no network, and the repo is cut down to a single commit so the agent can't recover the fix from git history.

https://preview.redd.it/0uopztlmpsth1.png?width=1200&format=png&auto=webp&s=59c0698ca0a76c17be41e41362324d6a262f1f25

Some findings:

With one attempt per task GLM-5.3 Flash scored 85%. With two to three attempts it scored 82%, within the margin of error of GPT-5.6 Luna (81%). The leaderboard now shows the number of attempts and the interval for every score.

https://preview.redd.it/9dqmk9uppsth1.png?width=1200&format=png&auto=webp&s=006cd587d11c03e47baf50c9954d95d937d09005

About half the tasks are easy for every model (near 100%). The other half is where they actually differ: 50%, 45% and 23% on the hard ones. Most of the difference between models comes from the hard half.

Since every fix is public on GitHub, we reviewed all 11k commands the agents ran. 69 tried to access the network and all failed. GLM tried 50 times to pip download the already-fixed release of the library it was fixing.

We also checked contamination by comparing older bugs (pre-2026) with newer ones of similar size. Older ones are solved about 9 points more often, but the confidence interval crosses zero, so we can't say much yet.

Half the tasks are private. So far public and private scores line up for all three models.

Results and every agent run: https://labs.evaligo.com/swe-race?utm_source=reddit&utm_medium=ml&utm_campaign=launch

Tasks: https://huggingface.co/datasets/evaligo/swe-race

The protocol follows DeepSWE (100 steps). Feedback on it, and suggestions for which models to run next, are welcome.

submitted by /u/heyitsdannyle
[link] [留言]

来源:r/MachineLearning · reddit.com