Artificial Analysis· @ArtificialAnlys · X·· 3 小时前AI 评分49
AI 导读
生成更多输出 token 并不一定带来更高分数。GPT-6 Astra(max)每任务约 81k 输出 token,得分 8.6%,不到 Grok 4.7(xhigh)约 180k 的一半。三个 Claude 模型生成的输出 token 最多(每任务约 202k 至 562k),得分 2.8% 至 6.4%。
正文
Generating more output tokens doesn’t necessarily translate to a higher score. GPT-6 Astra (max) scores 8.6% on ~81k output tokens per task, under half the ~180k of Grok 4.7 (xhigh). Three Claude models generated the most output tokens (~202k to ~562k per task) and score 2.8% to 6.4%.
来源:Artificial Analysis · x.com