跳到正文
Arena.ai· @arena · X·· 2 小时前精选AI 评分65
AI 导读

Mistral Large 4 在 Code Arena WebDev 榜单排名第 45,得分 1534,比 Mistral Large 3 的第 130 名提升 304 分,也比 Mistral Medium 3.5 的第 122 名高 271 分。

推荐理由

Mistral Large 4 在 Code Arena WebDev 的排名与价格对比,可帮助读者判断其相对前沿闭源模型的位置。

正文 · AI 翻译

Mistral Large 4 由 @MistralAI 刚刚登陆 Code Arena:WebDev 排名第 45。凭借 1534 分,这相比排名第 130 的 Mistral Large 3 提升了 +304 分!

此次发布还标志着相比其表现次佳的变体 Mistral Medium 3.5(排名第 122)提升了 271 分。

Mistral Large 4 的性能仅比 Claude Opus 4.8 (High) 低 2 分,而混合价格却低了近 6 倍:$3.47/M 对比 $20/M tokens。附近但性能更高的模型可以以更低价格获得,这使其刚好落在 Code Arena:WebDev 帕累托前沿之外。

@MistralAI 宣布其开放权重将于十月底发布。敬请期待它在 Arena 开放模型中的得分,并祝贺团队此次发布!

引用Mistral AI@MistralAI
认识 Mistral Large 4,又名 Le Chonk。 • 1T 参数,原生多模态。49B 活跃。 在综合基准测试上,它是来自美国或欧洲的最佳开放权重模型。 • 在关键工作负载上达到最先进水平,包括网络防御、制造业和金融,并且在视觉定位上超越了闭源前沿模型。 • 全程在欧洲打造,并可通过我们自己的 Mistral Cloud 基础设施从欧洲部署。 • 今天即可通过 API 向所有人开放。正与网络安全合作伙伴私下合作。 开放权重版本将于十月底发布。
原文

Meet Mistral Large 4, aka Le Chonk. • 1T parameters, natively multimodal. 49B active. It is the best open weights model from US or Europe on aggregated benchmarks. • State-of-the-art on critical workloads, including cyber defense, manufacturing and finance and it surpasses closed frontier models on visual grounding. • Forged in Europe end-to-end and is deployable from Europe via our own Mistral Cloud infrastructure. • Available to all via API today. Working with cybersecurity partners privately. Open weights release end of October.

在 X 查看被引用的帖子

来源:Arena.ai · x.com