跳到正文
elvis· @omarsar0 · X·· 2 天前AI 评分60
AI 导读

CMU 一篇论文提出 harness learning,用 RL 训练 proposer 模型读取任务、当前 harness 和执行报告,然后写出对 harness 的代码编辑,奖励是修改后 harness 的得分,solver 模型始终不变。

正文

Banger paper from CMU on harness learning.

(bookmark it)

Also, pay attention to this important new AI engineering skill of improving agents by editing their harness code instead of their weights.

Seeing a huge shift towards this.

The authors train a proposer model with RL to read a task, the current harness and an execution report, then write a code edit to the harness.

The reward is the score of the revised harness. The solver model never changes.

A trained 4B proposer beats its 35B teacher at single-step revision on Reasoning Gym, including task families it never saw in training. A proposer trained on HotpotQA keeps improving harnesses on MuSiQue and 2WikiMultihopQA.

Paper: https://arxiv.org/abs/2609.35738

Chat with Paper: https://academy.dair.ai/papers/harness-learning-enables-generalizable-test-time-adaptation-2609.35738

来源:elvis · x.com