跳到正文
Artificial Analysis· @ArtificialAnlys · X·· 3 小时前AI 评分50
AI 导读

引入幻觉门控后,Muse Spark 1.3 (max) 从 26.7% 降至 8.9%,Grok 4.7 (xhigh) 以 9.4% 升至榜首。

正文

Hallucination gating changes the picture for most models. Muse Spark 1.3 (max) from @AIatMeta ranks first before the hallucination gate at 26.7%, but falls to 8.9% after accounting for hallucinations, leaving Grok 4.7 (xhigh) from @SpaceXAI first at 9.4% on the headline metric. GLM-5.3 (max) passes all criteria for 13.9% of tasks but just 0.3% without a material hallucination, Kimi K3 (max) from @Kimi_Moonshot falls from 16.7% to 5.3%, and Claude Sonnet 5.5 (max with fallback) from 11.7% to 2.8%. GPT-6.1 Sol (max) falls relatively little, from 7.5% to 6.9%.

来源:Artificial Analysis · x.com