ThinkingBox:507 个有状态工作流评测智能体任务单次成功与 20/20 全通过
ThinkingBox: Solving an agent task once vs. solving it 20/20: 507 stateful workflows graded on terminal database state [R]
微软作者发布 ThinkingBox-Bench,包含 5 个领域的 507 个策略条件业务工作流,每个任务独立执行 20 次,共 10,140 次试验,按终端后端状态和副作用评分。
| Disclosure: I'm one of the authors (Microsoft). The paper, code, dataset are public and ThinkingBox is on Hugging Face OpenEnv as well. Raw evaluation trajectories are not released Links at the bottom. We wanted to know how much of a single agent success rate survives repetition, and whether "the agent finished the task" means the backend/database actually ended up in the correct state. Setup
Three metrics, because these get conflated constantly
Two findings Discovery and repeatability rank models very differently. The attached figure plots all three metrics for 9 models, and the orange-to-green spread is the whole point. Kimi-K3 has the broadest coverage we measured, solving 93.89% of tasks at least once (476/507), but only 13.41% (68/507) on all 20. Claude Opus 5 discovers fewer (79.09%) and repeats far more (47.53%, or 241 tasks). Qwen3.8-27B: 89.35% at least once, 7.50% every time. Ranking by pass@20 and ranking by all-20 give you nearly reversed leaderboards. Failures often look clean. In a retrospective ablation over 121,680 valid trials across 12 models, 79,853 failed the executable checks. 67.24% of those failures still terminated cleanly, invoked a state changing tool, and ended without a final tool error, so a completion-style proxy would have scored them as done. Of those clean terminating failures, the state checks found (categories overlap):
What it does not show
You can run any of the 507 tasks against your own model through HF OpenEnv environment, which is the fastest way to disagree with us using your own numbers. We report both the observed all-20 count and a plug-in passk estimate in the appendix. They answer different questions and can differ substantially: one records what happened in a fixed 20-attempt campaign, while the other estimates repeated success under additional assumptions. Which would you want emphasized on a reliability leaderboard or should both be shown?
[link] [留言] |
来源:r/MachineLearning · reddit.com