跳到正文
Artificial Analysis· @ArtificialAnlys · X·· 2 小时前AI 评分52
AI 导读

Artificial Analysis 让四款前沿图像编辑模型对同一张照片连续执行 30 次编辑,测试多轮编辑一致性。

正文

We gave four frontier image editing models the same photo and 30 edits in a row: Ideogram 4.5 keeps most of the room intact, while the others drift, GPT Image 2.5 Sunburst most visibly.

We've seen some interesting demos of multi-turn editing consistency from the latest image editing models, so we ran our own test: GPT Image 2.5 Sunburst, #1 on our Image Editing leaderboard, against Ideogram 4.5, FLUX 3 and Nano Banana 2.1. Our leaderboard scores single edits; this tests what happens when 30 consecutive changes stack up, with each model editing its own previous output through a 30-step real estate staging sequence: light the fire, add a sofa, repaint the walls, swap day for twilight, and more.

Why are the final results so different?

@ideogram_ai's Ideogram 4.5 and @bfl_ml's FLUX 3 edit locally. On small edits like adding a vase of tulips, we measured that they left 95% or more of the image essentially untouched. GPT Image 2.5 (Sunburst) re-renders most of the entire image on every edit, leaving only about a fifth of the image unchanged, so small shifts in colour and detail compound over turns. Nano Banana 2.1 sits in between: its edits stay local, but the rest of the image shifts slightly and gradually darkens.

来源:Artificial Analysis · x.com