跳到正文
elvis· @omarsar0 · X·· 3 小时前AI 评分64
AI 导读

Parsewave 对 Zapier AutomationBench 全部 600 个公开任务进行验证器审计,智能体标记出 323 个可疑验证器,人工复核确认 206 个真实错误并全部修复。

正文

We need more efforts like this.

Every agent benchmark should audit its verifiers.

Parsewave went through all 600 public tasks in Zapier's AutomationBench.

Agents wrote realistic wrong answers to try to fool each verifier, and human review confirmed 206 real bugs. AutomationBench Verified fixed all 206.

Regarding 1,235 Kimi K3 runs, the fixed verifiers changed 27.9% of the grades.

I just started looking into this benchmark for some independent eval work I am doing, so this is good timing to see this audit.

引用shaped@shaped
AutomationBench Verified is out. We went through all 600 public tasks in @Zapier's AutomationBench and checked every verifier. • Agents flagged 323 verifiers as suspicious, human review confirmed 206 real bugs, all 206 are fixed. • We replayed 1,235 Kimi K3 runs on the old and fixed verifiers and 344 of them (27.9%) got a different grade • Where verifiers were too strict, pass rate went from 18.8% to 43.8% • Where they were too lenient, it dropped from 60.2% to 49.7% Audit and dataset links below.
在 X 查看被引用的帖子

来源:elvis · x.com