Google 一篇论文提出 VeriHarness,把同一基座模型改造成智能体验证器,分别处理多轮采样结果不一致和全部一致两类情况。作者称在五个长时程基准上,它的选择得分优于所测基线,配合有证据支撑的修订,在 Gemini 3.5 Flash 上比单次采样提升 6.2 分,在 Claude Opus 4.8 上提升 6.4 分。
Banger paper from Google.
It's standard practice to sample several agent rollouts and trust the answers they agree on.
This Google paper shows that agreement can hide shared errors, while disagreement often points to the correct alternative.
VeriHarness turns the same base model into an agentic verifier with two jobs.
One resolves claims where rollouts disagree by checking workspace evidence.
The other challenges claims that every rollout agrees on and looks for requirements they all missed.
Across five long-horizon benchmarks, it gives the best selection scores among the baselines tested. With evidence-backed revision, it adds 6.2 points over a single rollout with Gemini 3.5 Flash and 6.4 points with Claude Opus 4.8.
The authors also release about 26,000 rollouts.
Paper: https://arxiv.org/abs/2610.00972
Chat with Paper: https://academy.dair.ai/papers/veriharness-scaling-agentic-verification-for-long-horizon-tasks-2610.00972
来源:elvis · x.com