跳到正文
r/LocalLLaMA· /u/Gold-Bat-3225·· 4 小时前AI 评分51

InferBench 基准测试:MiMo V2.6 Pro 推断用户意图接近 GPT-6 Astra

MiMo V2.6 Pro almost matched GPT-6 Astra at figuring out what a user actually wants

AI 导读

作者发布 InferBench 基准,测试前沿 LLM 从指令中推断用户优先级的能力,覆盖 12 个模型、20 个场景、2.8k 组对话。用户优先级明确给出时模型选对最佳选项的比例为 89%,部分优先级未说明时降至 60%;GPT-6 Astra 为 76%,MiMo V2.6 Pro 为 75%,Gemini 3.1 Pro 为 69%。

正文
MiMo V2.6 Pro almost matched GPT-6 Astra at figuring out what a user actually wants

Introducing InferBench: A benchmark testing how well frontier LLMs infer a user's priorities from their instructions.

We tested 12 LLMs across 20 scenarios with 2.8k conversations to see which models understand the user's goals.

Each conversation has a simulated user with a private profile of their priorities. The assistant decides on the best option or can first ask clarifying questions.

When a user's priorities were clearly stated up front, the models picked the best option 89% of the time. If some priorities were initially unstated then accuracy went down to 60%.

The open weight models did better than I expected:

GPT-6 Astra: 76%

MiMo V2.6 Pro: 75%

Gemini 3.1 Pro: 69%

Astra always asked clarifying questions when preferences weren't stated and never did when they were (extremely impressive). By comparison, Grok 4.6 asked to clarify in 18/64 conversations.

However LLMs are still just as confident when they're wrong. 105/288 wrong decisions were submitted with >90% confidence.

The full report is linked here: https://laugh.so/research/inferbench/

What surprised you the most?

submitted by /u/Gold-Bat-3225
[link] [留言]

来源:r/LocalLLaMA · reddit.com