Arena.ai· @arena · X·· 2 小时前精选AI 评分62
AI 导读
Mistral Large 4 在 Agent Arena 的 5000 多场真实智能体会话中录得 -6.6% 净提升分,总排名第 43,比上一版本 Mistral Medium 3.5(-12.60%)高 11 位;按当前分数,它会在开放模型中排第 13。
推荐理由
Mistral Large 4 在 Agent Arena 的实测排名与开放权重计划,可对照其与前一版本及开源模型的差距。
正文 · 原文
Mistral Large 4 by @MistralAI has landed in the Agent Arena top 15 labs!
Across +5K real-world agentic sessions, this preview model records a -6.6% net improvement score. It is 11 rankings above the previous variant, Mistral Medium 3.5 (-12.60%).
Mistral Large 4 is #43 overall in the Agent Arena. Open weights are expected at the end of October, but at its current score Mistral Large 4 would rank #13 among open models.
@MistralAI is also the only European lab in the Agent Arena top 15. Congrats to the team!
Meet Mistral Large 4, aka Le Chonk. • 1T parameters, natively multimodal. 49B active. It is the best open weights model from US or Europe on aggregated benchmarks. • State-of-the-art on critical workloads, including cyber defense, manufacturing and finance and it surpasses closed frontier models on visual grounding. • Forged in Europe end-to-end and is deployable from Europe via our own Mistral Cloud infrastructure. • Available to all via API today. Working with cybersecurity partners privately. Open weights release end of October.在 X 查看被引用的帖子
来源:Arena.ai · x.com