跳到正文
Rohan Paul· @rohanpaul_ai · X·· 4 小时前AI 评分46
AI 导读

https://x.com/rohanpaul_ai/status/2108741349419913339

正文
引用Rohan Paul@rohanpaul_ai
Anthropic states plainly that the model’s own explanation of its reasoning can’t be trusted as evidence of why it acted, which is exactly why they can’t cleanly judge how severe each of these failures was.
在 X 查看被引用的帖子

来源:Rohan Paul · x.com