针对 OpenAI 发布约 400 项数学结果一事,有观点认为其目的并非"摧毁数学研究",而是用高难度数学题测试内部模型,结果发现只有最难开放问题才够用。非千禧年大奖难题的解答均在单智能体模式下经 3 小时思考得出,OpenAI 随后将这些数学进展公开。该模型在更多算力下很可能还能解决约 4000 题集中的更多问题。
I think Prinz makes a valid point that isn't discussed enough, but it also gives me pause. Apparently, there's a real assumption or idea that OpenAI wants to "destroy" scientific disciplines like mathematics with its models. I can't imagine that from any perspective.
Firstly, because they are, of course, part of the scientific process and intend to remain so. Their aim is clearly to revolutionize science using *their technology*, meaning they want to maintain a good relationship with scientists. The heated debate surrounding the Millennium Prize issue, in particular, showed that OpenAI seeks to reconcile with scientists. Secondly, from a purely economic standpoint, it's simply relevant to have a good relationship with the scientific community, especially from a PR perspective, and also in the run-up to the upcoming OpenAI IPO. Nothing is more desirable than bad press, and that's why I can't grasp the idea that OpenAI would arbitrarily destroy science.
On the one hand, the idea that OpenAI would deliberately destroy science is beyond my comprehension.
The last thing we want to avoid is bad press. Conversely, Prinz solved the problem by pointing out that the models are now so good that only the toughest tests can challenge current frontier models. And the solutions to those problems are simultaneously solutions to long-standing research tasks that coincide with the current state of science. Or, put another way: it is practically a fortunate circumstance that the challenges must be so immense that they simultaneously advance science by solving long-standing problems.
It's somewhat under-discussed *why* OpenAI released only around ~400 math results. The purpose of this exercise was not to "destroy math research" (or whatever it is that some people want you to believe). OpenAI just wanted to test its internal model on some hard math, and it turned out that anything less than the hardest open problems no longer suffices for this purpose. This is why all non-Millennium-Prize solutions were reached in a single-agent setting after 3 hours of thinking: OpenAI was literally just testing its model on these questions, similarly to how I run prinzbench. Once OpenAI had the solutions (which were a side effect of sorts of benchmarking its model), a question arose as to what exactly should be done with these solutions, which are genuine advances in the field of mathematics. OpenAI - correctly - decided to make them public. It seems clear that many other problems in the set of ~4,000 are very likely solvable by this model with more compute than just ~3 hours of Pro-level thinking. Whenever a more powerful checkpoint of this model is ready to be tested, OpenAI will likely have even more to share.在 X 查看被引用的帖子
来源:Chubby♨️ · x.com