跳到正文
r/MachineLearning· /u/tuhin_k·· 5 小时前AI 评分71

ThinkingBox:507 个有状态工作流评测智能体任务单次成功与 20/20 全通过

ThinkingBox: Solving an agent task once vs. solving it 20/20: 507 stateful workflows graded on terminal database state [R]

AI 导读

微软作者发布 ThinkingBox-Bench,包含 5 个领域的 507 个策略条件业务工作流,每个任务独立执行 20 次,共 10,140 次试验,按终端后端状态和副作用评分。

正文
ThinkingBox: Solving an agent task once vs. solving it 20/20: 507 stateful workflows graded on terminal database state [R]

Figure 1b from our paper

Disclosure: I'm one of the authors (Microsoft). The paper, code, dataset are public and ThinkingBox is on Hugging Face OpenEnv as well. Raw evaluation trajectories are not released Links at the bottom.

We wanted to know how much of a single agent success rate survives repetition, and whether "the agent finished the task" means the backend/database actually ended up in the correct state.

Setup

  • Thinkingbox-Bench includes 507 policy conditioned business workflows across 5 domains (retail, travel/hospitality, auto insurance, neobank internal IT, consulting IT/HR)
  • Each task is run in 20 independently executed attempts, each starting from an identical clean backend: 10,140 trials per model
  • A simulated user holds private context and only reveals it when asked
  • Grading compares the terminal backend state and side effects against the required end state. Any trajectory that produces the right outcome passes; wrong, missing or extra effects fail. 477 of 507 tasks are graded on state alone; 30 also check a narrow property of the final response

Three metrics, because these get conflated constantly

  • pass@1 — fraction of all attempts that succeed
  • pass@20 — fraction of tasks solved at least once across the 20 attempts
  • all-20 — fraction of tasks solved on every one of the 20 attempts. This is an observed count on a fixed trial budget, not an estimator

Two findings

Discovery and repeatability rank models very differently. The attached figure plots all three metrics for 9 models, and the orange-to-green spread is the whole point. Kimi-K3 has the broadest coverage we measured, solving 93.89% of tasks at least once (476/507), but only 13.41% (68/507) on all 20. Claude Opus 5 discovers fewer (79.09%) and repeats far more (47.53%, or 241 tasks). Qwen3.8-27B: 89.35% at least once, 7.50% every time. Ranking by pass@20 and ranking by all-20 give you nearly reversed leaderboards.

Failures often look clean. In a retrospective ablation over 121,680 valid trials across 12 models, 79,853 failed the executable checks. 67.24% of those failures still terminated cleanly, invoked a state changing tool, and ended without a final tool error, so a completion-style proxy would have scored them as done. Of those clean terminating failures, the state checks found (categories overlap):

  • wrong field values: 77.61%
  • unintended extra effects: 43.30%
  • missing required effects: 25.36%

What it does not show

  • Tasks are synthetic reconstructions of enterprise workflow patterns, not production traffic
  • 20/20 on our trial budget is an observed count, not a guarantee of future reliability
  • The simulated user is a fixed LLM; that is a source of variance we discuss in the appendix
  • Failure categories are deterministic diagnostic labels, not causal explanations

You can run any of the 507 tasks against your own model through HF OpenEnv environment, which is the fastest way to disagree with us using your own numbers.

We report both the observed all-20 count and a plug-in passk estimate in the appendix. They answer different questions and can differ substantially: one records what happened in a fixed 20-attempt campaign, while the other estimates repeated success under additional assumptions. Which would you want emphasized on a reliability leaderboard or should both be shown?

submitted by /u/tuhin_k
[link] [留言]

来源:r/MachineLearning · reddit.com