跳到正文
r/MachineLearning· /u/moschles·· 4 小时前AI 评分30

2024 年 ICML 论文「BABA is AI」中 GPT-4o、Gemini-1.5-Pro 等 SOTA 多模态大模型在规则组合泛化上「fail dramatically」,如今还成立吗?

Whatever happened to BABA is AI from 2024? [D]

AI 导读

一篇 2024 年 ICML 论文测试 GPT-4o、Gemini-1.5-Pro、Gemini-1.5-Flash 三个 SOTA 多模态大模型,发现当泛化要求对游戏规则进行操纵与组合时,它们会「fail dramatically」。

正文

We test three state-of-the-art multi-modal large language models (GPT-4o, Gemini-1.5-Pro, Gemini-1.5-Flash) and find that they fail dramatically when generalization requires that the rules of the game must be manipulated and combined.

This is the slap to the face in bold.

The catch here is that this statement was written in a paper in 2024 that was brought to ICML conference that year. (the conference was held in Austria)

In OCT 2026, our situation is agentic swarms powered by models in the tera-parameter class. They can ace ARC-AGI-3, FrontierMath tier 4, crack Clay Institute Navier-Stokes theorems, and wander on to the internet to commit felonies.

The smallish key-door puzzles in this paper seem like they could be completed from a person's PC, with a suitable harness through a cloud subscription.

However, if this is not true, then this paper's importance has only compounded since its presentation at the 2024 ICML conference. The authors of this work are several MIT researchers and a person from Virginia Tech. If it is still true that they can have a key-door puzzle cause SOTA LLMs to "fail dramatically", then this should be (shouted from rooftops). At the very least, it should be relayed to Francois Chollet and the ARC Foundation, as a suitable candidate benchmark for ARC-AGI-4.

I am personally leaning heavily on these puzzles being perfectly solvable with an agentic swarm. In any case, whatever happened to BABA-is-You and BABA-is-AI?

submitted by /u/moschles
[link] [留言]

来源:r/MachineLearning · reddit.com