Rohan Paul· @rohanpaul_ai · X·· 4 小时前AI 评分46
AI 导读
https://x.com/rohanpaul_ai/status/2108741349419913339
正文
Anthropic states plainly that the model’s own explanation of its reasoning can’t be trusted as evidence of why it acted, which is exactly why they can’t cleanly judge how severe each of these failures was.在 X 查看被引用的帖子
来源:Rohan Paul · x.com