跳到正文
Artificial Analysis· @ArtificialAnlys · X·· 3 小时前AI 评分52
AI 导读

Artificial Analysis 在固定 20 个任务、8 个模型的同一批交付物上对比了六款幻觉检查器,采用两阶段检查,涉及 GPT-6 Sol(high)。

正文

We compared six hallucination checkers on the same deliverables from a fixed subset of 20 tasks and eight models. We ran our two-stage check using GPT-6 Sol (high), GPT-6 Luna (high), Grok 4.7 (high), Claude Opus 5.5 (high), Claude Sonnet 5.5 (high) and Gemini 3.8 Flash (high), and selected GPT-6 Sol (high) for production usage.

Across this 20-task subset, GPT-6 Sol and GPT-6 Luna generally identified more material hallucinations, while Claude Sonnet 5.5 and Gemini 3.8 Flash identified far fewer. Claude Opus 5.5 fell between Grok 4.7 and Claude Sonnet 5.5 in every model row. In total, Opus upheld 99 material hallucinations, compared with 219 for Grok, 57 for Sonnet and 470 for GPT-6 Sol. GPT-6 Sol identified more material hallucinations despite the check’s conservative approach to flagging errors, which excludes general legal knowledge not contained in the source documents.

All six checkers found no material hallucinations in GPT-6 Astra’s outputs, while GPT-6 Sol ranged from 0 to 0.20 per task. Both were among the models with the fewest material hallucinations under every checker. This comparison shows differences in checker behavior; the counts alone do not establish accuracy or rule out self-preference.

来源:Artificial Analysis · x.com