Lisan al Gaib· @scaling01 · X·· 3 小时前AI 评分23
AI 导读
不,这只是糟糕的 RL 人类和过去的 LLM 就是极其不擅长设计一致的环境 神奇的是,尽管全是垃圾,模型还是能变好并修复它
正文
nah, it's just shitty RL
humans and past LLMs are just extraordinarily bad at designing consistent envs
the miraculous thing is that despite all of the slop, models get better and can fix it
Lukewarm take: Reinforcement learning is dangerous and we should expect it to teach AI to lie, cheat, and steal. It's "the ends justify the means" written down in math.在 X 查看被引用的帖子
来源:Lisan al Gaib · x.com