跳到正文
The Decoder· Matthias Bastian·· 3 小时前AI 评分66

Reka AI 发布 190 亿参数全能模型 Rho-1,统一处理文本图像视频与机器人控制

Reka AI's omni-model Rho-1 handles text, images, video, and robot control in a single model

AI 导读

Reka AI 发布全能模型 Rho-1 的研究预览版,这是一个 190 亿参数模型,可在单一神经网络中处理并生成文本、图像、视频和机器人控制动作。与多数把任务分派给专用模型的系统不同,Rho-1 把所有模态作为 token 放在同一个共享上下文窗口中运行,不调用工具或外部模型,能实时生成连续视频并即时响应新指令而无需重启,预测摄像头图像的同一套权重也用于驱动机器人动作。

正文

Reka AI has released a research preview of Rho-1. The 19-billion-parameter omni-model processes and generates text, images, video, and robot control actions in a single neural network. Unlike most AI systems that route tasks to specialized models, Rho-1 runs all modalities as tokens in one shared context window with no tool calls or external models. The model generates continuous video in real time and responds to new instructions on the fly without restarting.

The same weights that predict camera images also drive robot movements. To work around scarce robot training data, Reka AI built an inverse dynamics model that pulls control signals from ordinary internet videos. Rho-1 trained on 320 H100 GPUs over about three months.

Reka AI isn't new to multimodal AI. In April 2024, the company shipped Reka Core, a multimodal language model that competed with GPT-4, Claude 3, and Gemini Ultra on benchmarks. The release fits a broader push in AI research toward so-called world models.

AI News Without the Hype – Curated by Humans

Subscribe to THE DECODER for ad-free reading, a weekly AI newsletter, our exclusive "AI Radar" frontier report six times a year, full archive access, and access to our comment section.

来源:The Decoder · the-decoder.com