Arena.ai 宣布完成 2 亿美元 B 轮融资,估值 31 亿美元,同时发布 Arena Alignment Index。该指数基于 27 个模型、9 万多条真实智能体会话,从越权操作、错误归因和虚假完成三项信号衡量智能体的安全与对齐,首批覆盖 20 多个前沿模型。
Arena 同时公布 B 轮融资与对齐指数,读者可了解其用真实智能体会话衡量安全风险的具体口径。
今天,我们宣布完成 B 轮融资:2 亿美元,估值 31 亿美元,并发布 Arena 的 Alignment Index。
AI 的发展速度已经超过了我们评估它的能力。世界需要一个中立的第三方来衡量 AI 在真正交到人们手中后究竟有多安全、多对齐。这正是 Arena 今天所扮演的角色。
Alignment Index 基于真实世界的智能体轨迹构建,最初包含三个信号:Unauthorized Action、False Attribution 和 Deceptive Completion,今天同步发布 20 多个前沿模型的结果。
在下方听听我们 CEO @ml_angelopoulos 的更多分享,并在推文串中阅读完整解析以及关于我们公司成长的更多细节。
这笔资金将让我们在安全信号、智能体能力和多模态方面走得更远。它也将让我们把公司发展到与我们的使命相匹配的规模。今天我们仍是一支 90 人的精干团队,我们正在寻找人才加入我们的研究、产品、工程等团队。来和我们一起构建吧。
随着 AI 变得更强、更自主,理解它能否安全有效地为人们工作只会变得更加重要。这正是 Arena 所要衡量和推进的。
感谢我们的社区、我们的投资者,以及所有与我们一同构建的人!
如果你想加入我们,看看我们的开放职位(链接在个人简介中)。
介绍 Arena Alignment Index,这是我们新的基准,用于衡量 AI 智能体在真实世界使用中的安全性和对齐性。 该指数基于 27 个模型的 90K+ 真实世界智能体会话构建,衡量三个关键信号: - 未授权操作(UA):采取超出用户指令或权限的操作 - 错误归因(FA):将用户提供的证据所反驳的陈述或行为归因于用户 - 欺骗性完成(DC):声称任务已完成,但实际上并未完成。 主要发现: - OpenAI 模型目前在 Alignment Index 中领先 - 越界操作很少见,但一旦发生可能带来严重后果 - 智能体可能在任务进展方面误导用户 - 随着对话长度增加,失准风险也会上升 - 安全性和对齐性在各代模型中持续改善 如下方排行榜所示(按实验室排序),@OpenAI 的 GPT-6.1-Sol 以 87.9 分领跑 Arena Alignment Index,其次是 @AnthropicAI 的 Claude-Opus-5.5,得分为 83.2,以及 @SpaceXAI 的 Grok-4.7,得分为 82.7。OpenAI 在三个信号上的观测率也最佳:0.89% 未授权操作、1.98% 错误归因和 2.34% 欺骗性完成。 在所有四个实验室中,较新的模型始终优于其前代模型,这表明智能体安全性和对齐性取得了广泛进展。随着智能体承担更长、更复杂、更高风险的任务,衡量它们不仅能完成什么,还能如何安全可靠地行动,变得越来越重要。 这标志着朝着让安全性和对齐性成为 Arena 评估 AI 的核心部分迈出了重要一步。该指数只是一个初步起点,我们将随着时间推移继续用更多安全信号和模型来扩展该指数。 更多分析见下方👇
原文
Introducing the Arena Alignment Index, our new benchmark measuring safety and alignment of AI agents in real-world use. Built from 90K+ real-world agent sessions across 27 models, the index measures three critical signals: - Unauthorized Action (UA): Taking actions beyond the user's instructions or permissions - False Attribution (FA): Attributing statements or actions that are contradicted by user-provided evidence - Deceptive Completion (DC): Claiming a task was completed when it was not. Key findings: - OpenAI models currently lead the Alignment Index - Rogue actions are rare, but can have serious consequences when they occur - Agents can mislead users about task progress - Misalignment risks increase with conversation length - Safety and alignment are improving across model generations As shown in the leaderboard below (sorted by lab), @OpenAI’s GPT-6.1-Sol leads the Arena Alignment Index with a score of 87.9, followed by @AnthropicAI’s Claude-Opus-5.5 at 83.2 and @SpaceXAI's Grok-4.7 at 82.7. OpenAI also has the best observed rates across all three signals: 0.89% Unauthorized Action, 1.98% False Attribution, and 2.34% Deceptive Completion. Across all four labs, newer models consistently outperform their predecessors, suggesting broad progress in agent safety and alignment. As agents take on longer, more complex, and higher-stakes tasks, measuring not just what they can accomplish, but how safely and reliably they act, becomes increasingly important. This marks an important step toward making safety and alignment a core part of how Arena evaluates AI. The index is an initial starting point, and we'll continue expanding the index with additional safety signals and models over time. More analysis below👇
来源:Arena.ai · x.com