跳到正文
Rohan Paul· @rohanpaul_ai · X·· 2 小时前AI 评分51
AI 导读

清华大学一篇新论文提出 AAArena 基准,取自清华年度机器人构建竞赛的 12 款游戏,并以 1920 个存档人类程序作为对手,考察编码智能体在模型权重不变的情况下读规则、选对手、研究回放并改写机器人。

正文

New Tsinghua paper finds that AI agents improving game bots from match replays can top human leaderboards, but mostly stall on games with complex rules.

Getting AI to learn a winning game strategy from a limited number of matches is still hard, especially against changing rivals.

They built AAArena from 12 games in Tsinghua's yearly bot-building contest, with 1,920 archived human programs as rivals. A coding agent, with its model weights unchanged, reads the rules, picks opponents, studies replays, and rewrites its bot within a match budget.

Detailed replays beat win/loss-only feedback in all 3 games tested. With replays, a Pacman bot reached rank 1, versus rank 11 without them.

Tripling the match budget did not push any of 4 stuck bots to rank 1.

– arxiv. org/abs/2610.12341

Title: "Can AI Agents Learn Their Way to the Top? Evaluating Heuristic Learning in a Long-Running Game Agent Competition"

来源:Rohan Paul · x.com