SWE-Race:188 个真实并发缺陷的编码智能体基准,含三个模型结果
SWE-Race: a coding-agent benchmark of 188 real concurrency bugs, with results from three models [P]
SWE-Race 是一个由约 100 个 Python 项目已合并 PR 中真实并发缺陷(竞态、死锁、取消问题)构成的编码智能体基准,共 188 个任务,每个任务用项目自带测试在无网络容器中评分,仓库被裁剪到单个 commit 以防从 git 历史找回修复。
| We've been building a benchmark out of real concurrency bugs (race conditions, deadlocks, cancellation issues) taken from merged PRs in about 100 Python projects. Each task gets graded by the project's own tests, in a container with no network, and the repo is cut down to a single commit so the agent can't recover the fix from git history. Some findings: With one attempt per task GLM-5.3 Flash scored 85%. With two to three attempts it scored 82%, within the margin of error of GPT-5.6 Luna (81%). The leaderboard now shows the number of attempts and the interval for every score. About half the tasks are easy for every model (near 100%). The other half is where they actually differ: 50%, 45% and 23% on the hard ones. Most of the difference between models comes from the hard half. Since every fix is public on GitHub, we reviewed all 11k commands the agents ran. 69 tried to access the network and all failed. GLM tried 50 times to pip download the already-fixed release of the library it was fixing. We also checked contamination by comparing older bugs (pre-2026) with newer ones of similar size. Older ones are solved about 9 points more often, but the confidence interval crosses zero, so we can't say much yet. Half the tasks are private. So far public and private scores line up for all three models. Results and every agent run: https://labs.evaligo.com/swe-race?utm_source=reddit&utm_medium=ml&utm_campaign=launch Tasks: https://huggingface.co/datasets/evaligo/swe-race The protocol follows DeepSWE (100 steps). Feedback on it, and suggestions for which models to run next, are welcome. submitted by /u/heyitsdannyle[link] [留言] |
来源:r/MachineLearning · reddit.com