跳到正文
Arena.ai· @arena · X·· 3 天前精选AI 评分66
AI 导读

Arena.ai 数据显示,Anthropic 的 Claude Sonnet 5.5 在 Agent Arena 首次登场即排名第三,净提升 +12.5%,单任务中位成本 2.74 美元。相比排名第 13 的 Claude Sonnet 5(High,+4.4%),其净提升高出 8.1 个百分点,并在 Chat 分类以 +15.6% 位列第一。不过其成本比排名第二的 Claude Opus 5.5(High,1.58 美元)高约 73%,因此未能进入 Agent Arena 的帕累托前沿。

推荐理由

榜单数据给出了 Claude Sonnet 5.5 的排名、净提升与单任务成本,可据此比较同门模型的性价比取舍。

正文 · 原文

Claude Sonnet 5.5 by @AnthropicAI just landed at #3 in the Agent Arena. This model has a median cost per task of $2.74, and a +12.5% net improvement score.

Claude Sonnet 5.5 delivers top-tier performance, but at a cost premium: #2 Claude Opus 5.5 costs $1.58 per task while achieving a higher score. That tradeoff keeps Sonnet 5.5 just off the Agent Arena Pareto frontier.

引用Arena.ai@arena
Exciting news: Claude Sonnet 5.5 (Max) by @AnthropicAI has debuted at #3 in the Agent Arena with +12.5% net improvement! This release is a 8.1 percentage-point increase over Claude Sonnet 5 (High), which ranks #13 with +4.4% net improvement. By category, Claude Sonnet 5.5 secured the #1 spot in Chat (+15.6%) above both Fable 5.1 (+11.49%) and Opus 5.5 (+10.29%). This performance comes with a higher cost: Claude Sonnet 5.5 (Max) has a median cost of $2.74 per task, about 73% higher than #2 Claude Opus 5.5 (High) at $1.58. @AnthropicAI models now hold all three top positions in Agent Arena. Congrats to the team!
在 X 查看被引用的帖子

来源:Arena.ai · x.com