一篇论文提出 HERMES harness,在整仓库迁移任务上让 GPT-5.6 Sol 的通过率从 6.5% 提升到 31.0%,模型和 effort 设置保持不变。
Build your own harness, folks.
Reading papers like this makes me realize how underexplored harness engineering really is.
The authors find that on whole-repository migration, GPT-5.6 Sol goes from 6.5% to 31.0% when Codex is replaced with the HERMES harness, with the same model and effort setting.
The gain comes from the harness.
HERMES pairs each repository component with a resident LLM that knows its own code and dependencies. A dependency-aware step decides which components to activate, and a diagnosis step maps test failures back to the components that need changes.
Across four software engineering benchmarks, it beats matched baseline harnesses by 12.4 points on average.
With strong activation and diagnosis models, Qwen3-8B components come within 4.5 points of an all-GPT-5.6 Sol setup and cut Terminal-Bench 4.0 inference cost by 26.2%.
Paper: https://arxiv.org/abs/2610.07832
Chat with Paper: https://academy.dair.ai/papers/harness-engineering-for-software-engineering-via-modular-executable-dev-primitiv-2610.07832
来源:elvis · x.com