Artificial Analysis 用 AA-Video-T2V v2.0 评测 Grok Imagine Video 1.5 Lite,该基准覆盖 10 项能力与 10 类用例。Lite 在多场景叙事、光照材质和文字渲染上最接近前沿,在对话口型同步和人体结构上差距最大。对比 Grok Imagine Video 1.5,Lite 仅在物理表现上持平,其余九项能力均落后,多场景叙事差距最小。
Where Grok Imagine Video 1.5 Lite comes closest to the frontier on AA-Video-T2V v2.0: Multi-Scene & Narrative, Lighting & Materials and Text Rendering
AA-Video-T2V v2.0 measures 10 capabilities and 10 use cases, each with its own leaderboard. Capabilities are the model skills a prompt is written to test. They draw on lab and academic research, and on how creators and businesses push video models today.
Grok Imagine Video 1.5 Lite sits closest to the frontier in Multi-Scene & Narrative, Lighting & Materials and Text Rendering. It sits furthest from it in Dialogue & Lip Sync and Human Anatomy. Against Grok Imagine Video 1.5, Lite matches it in Physics and trails it on the other nine capabilities, by the least in Multi-Scene & Narrative.
➤ Multi-Scene & Narrative covers multi-event stories, shot transitions, montage, and identity consistency across shots.
➤ Lighting & Materials covers lighting, exposure, shadows, color temperature, texture, cloth, and hair and fur.
➤ Text Rendering covers short and long text, 2D layout, non-lexical strings, and text on deforming surfaces.
来源:Artificial Analysis · x.com