跳到正文
Chubby♨️· @kimmonismus · X·· 3 小时前AI 评分28
AI 导读

3/ 在同一个 Python 任务上,Step 5 Preview 和 Claude Opus 5.5 通过了 45/45 项本地检查;Gemini 3.1 Pro Preview 通过了 41/45。每个模型只有一次尝试机会。对这个任务来说表现不错,但测试规模太小,无法确立整体能力相当。不过初步来看相当不错。

正文

3/ On the same Python task, Step 5 Preview and Claude Opus 5.5 passed 45/45 local checks; Gemini 3.1 Pro Preview passed 41/45. Each model got one attempt. Promising for this task, but too small a test to establish equal overall capability. But initially it looks pretty good.

来源:Chubby♨️ · x.com