Arena.ai 公布 GPT-6.1 Sol (Max) 在 Agent Arena 排名第 5,较此前提升 11.23%,并重塑帕累托前沿。其每任务中位成本为 $0.56,比 GPT-6 Sol 低 39% 且得分高 1.52 分,比 GPT-6 Astra 低 81% 且差距在 1.04 分以内。
Arena 榜单给出 GPT-6.1 Sol 在 Agent Arena 的排名与每任务成本,可据此比较它与同价位模型的性价比位置。
Exciting news: GPT-6.1 Sol (Max) by @OpenAi just landed in the Agent Arena at #5 (+11.23%) and reshaped the Pareto frontier!
At a $0.56 median cost per task, it delivers performance within 2 percentage points of GPT-6 Sol and GPT-6 Astra for substantially less cost:
- 39% lower cost than GPT-6 Sol, while scoring +1.52 pts higher
- 81% lower cost than GPT-6 Astra, while landing within 1.04 pts
GPT-6.1 Sol also delivers top-five performance at substantially lower cost compared to:
- 88% lower cost than Claude Fable 5.1 (Max), while landing within 3.08 pts (ranked #1)
- 65% lower cost than Claude Opus 5.5 (High), while landing within 2.59 pts (ranked #2)
- 80% lower cost than Claude Sonnet 5.5 (Max), while landing within 1.29 pts (ranked #3)
Congrats to the team @OpenAI on this release!
Exciting news: GPT-6.1 Sol (Max) by @OpenAI just landed the Code Arena: WebDev at #3 with 1759 pts, and at a blended $8/MToken it reshapes the Pareto frontier! GPT-6.1 Sol (Max) marks a clear improvement in cost efficiency: it gained 70 points over GPT-6 Sol (Max) for the same price. It landed within 30 points of GPT-6 Astra (Max) at 80% lower blended token cost, and 59 points from Claude Opus 5.5 (Max) at 50% of the price. See position on the Pareto frontier for the Code Arena: WebDev in the post below. Overall, GPT-6.1 Sol improved from GPT-6 Sol by 4 rankings! It also improved in every category: - Consumer Product: #5 → #1 - Simulations: #6 → #3 - Data & Analytics: #4 → #3 - Content Creation Tools: #4 → #3 - Gaming: #6 → #4 - Reference-Based Design: #6 → #4 - Brand & Marketing: #10 → #6 Congrats to the @OpenAI team on the release!在 X 查看被引用的帖子
来源:Arena.ai · x.com