InferBench 基准测试:MiMo V2.6 Pro 推断用户意图接近 GPT-6 Astra
MiMo V2.6 Pro almost matched GPT-6 Astra at figuring out what a user actually wants
作者发布 InferBench 基准,测试前沿 LLM 从指令中推断用户优先级的能力,覆盖 12 个模型、20 个场景、2.8k 组对话。用户优先级明确给出时模型选对最佳选项的比例为 89%,部分优先级未说明时降至 60%;GPT-6 Astra 为 76%,MiMo V2.6 Pro 为 75%,Gemini 3.1 Pro 为 69%。
| Introducing InferBench: A benchmark testing how well frontier LLMs infer a user's priorities from their instructions. We tested 12 LLMs across 20 scenarios with 2.8k conversations to see which models understand the user's goals. Each conversation has a simulated user with a private profile of their priorities. The assistant decides on the best option or can first ask clarifying questions. When a user's priorities were clearly stated up front, the models picked the best option 89% of the time. If some priorities were initially unstated then accuracy went down to 60%. The open weight models did better than I expected: GPT-6 Astra: 76% MiMo V2.6 Pro: 75% Gemini 3.1 Pro: 69% Astra always asked clarifying questions when preferences weren't stated and never did when they were (extremely impressive). By comparison, Grok 4.6 asked to clarify in 18/64 conversations. However LLMs are still just as confident when they're wrong. 105/288 wrong decisions were submitted with >90% confidence. The full report is linked here: https://laugh.so/research/inferbench/ What surprised you the most? submitted by /u/Gold-Bat-3225[link] [留言] |
来源:r/LocalLLaMA · reddit.com