Arena.ai 宣布 Mistral Large 4 已进入 Arena,可在 Agent Arena 中实测,投票将影响其评测结果,分数稍后公布。Agent Arena 基于数百万真实长周期智能体任务,模型可调用网页搜索、文件系统和终端工具,排行榜用因果追踪方法衡量相对平均模型的结果表现。
Mistral Large 4 进入 Agent Arena 与 Code Arena 评测,读者可了解其评测入口与开放权重时间安排。
Mistral Large 4 by @MistralAI is now in the Arena!
Head to Agent Arena to test it out, and your votes will shape its evaluation. Scores coming soon.
In Agent Arena, we measure models on millions of real-world, long-horizon agentic tasks. Models can access web search, filesystem, and terminal tools to complete complex workflows. The leaderboard measures model performance on outcomes relative to the average model using a causal tracing methodology.
Mistral Large 4 is also available in Code Arena: WebDev, Text, and Vision.
Meet Mistral Large 4, aka Le Chonk. • 1T parameters, natively multimodal. 49B active. It is the best open weights model from US or Europe on aggregated benchmarks. • State-of-the-art on critical workloads, including cyber defense, manufacturing and finance and it surpasses closed frontier models on visual grounding. • Forged in Europe end-to-end and is deployable from Europe via our own Mistral Cloud infrastructure. • Available to all via API today. Working with cybersecurity partners privately. Open weights release end of October.在 X 查看被引用的帖子
来源:Arena.ai · x.com