跳到正文
Arena.ai· @arena · X·· 2 小时前AI 评分50
AI 导读

Agent Arena 按 Code、Work、Chat 三类智能体任务对模型排名,美国实验室在所有领域领先。Anthropic 模型在 Code、Work、Chat 三项均居第一,其中 Claude Fable 5.1 (Max) 领跑 Code 和 Work、Chat 排第 2,Claude Sonnet 5.5 (Max) 领跑 Chat。

正文

The Agent Arena ranks agentic ability to solve problems across domains.

U.S. labs lead across every domain:
- @AnthropicAI models hold the #1 position in Code, Work, and Chat
- @OpenAI's GPT 6 Astra (Max) remains in the top four across all three categories: #2 in Code and #4 in both Work and Chat

The leading models stay strong across categories, but their order shifts:
- Claude Fable 5.1 (Max) leads both Code and Work and ranks #2 in Chat
- Claude Sonnet 5.5 (Max) leads Chat and ranks #4 in Code and #5 in Work

Code and Work rankings move together more closely than Chat:
- Gemini 4 Argon (High), for example, ranks #14 in Code and #12 in Work but #5 in Chat, highlighting a distinct strength in conversational agent tasks

Agent Arena ranks models based on overall net improvement against the current frontier, as well as performance across specific categories of agentic tasks. These rankings are grounded in live traces collected from agents attempting to solve real-world tasks submitted by humans around the world.

  • Code: writing and debugging code, workflow automation, and data analysis
  • Chat: creative writing, learning, personal questions, everyday research, and media generation
  • Work: documents, professional research, planning, and professional writing

This visual outlines how selected models compare across Agent Arena categories as of October 5, 2026.

引用Arena.ai@arena
https://x.com/i/article/2089454679130578944
在 X 查看被引用的帖子

来源:Arena.ai · x.com