MIT 论文发现,LLM 智能体在拍卖与匹配博弈中,当选项以简单步骤呈现或直接说明规则的安全动作时决策更好,被要求推理对手时反而更差。GPT-4o、Claude、Gemini 和 Gemma 仍会低报价以保留利润空间;把涨价改为简单的留或退选择,使 Gemma 的报价从低于估值 $5.30 收窄到 $0.30,一句"拒绝仅重定向"的说明把匹配错误率从 4.2% 降到 0.2%。
New MIT Paper: LLM agents decide better when the choice is shown in simple steps or the rule's safe move is stated plainly, and worse when told to reason about opponents.
Market-design rules of thumb built for human bidders carry over to LLM agents, so we can borrow them instead of inventing new prompt tricks.
Honest bids and rankings are always the best move in these auctions and matching games. GPT-4o, Claude, Gemini and Gemma still underbid, often to keep a profit margin.
A rising price with a simple stay-or-exit choice moved Gemma from $5.30 below its value to $0.30 below. A 1-line note that rejections only redirect cut matching errors from 4.2% to 0.2%.
Fix the format and state the key fact before adding reasoning prompts, and judge agents by their choices, since their written plans missed these gains.
Agent prompts should state facts about the rules rather than request more thinking, because facts improved choices while thinking prompts often added errors.
来源:Rohan Paul · x.com