反对把解决对齐当作 AI 安全的主要框架
Against "solving alignment" as the major frame for AI safety and security
作者反对把解决对齐和解读模型内部状态当作 AI 安全的核心框架,认为可解释性研究已进行至少 14 年,仍未产出该设想所需的可靠、可扩展理解,且随能力提升问题可能更糟。他主张 AI 安全应跳出计算机科学的局限,借鉴航空、汽车、飓风、核安全、流行病和金融危机等领域的做法,在承认世界模型不完整的前提下做不完美预测、持续监测、限制权力授予并建立多层冗余保护。
J. M. W. Turner’s Snow Storm: Steam-Boat off a Harbour’s Mouth (1842)
Civilization literally rests on our ability to approximately forecast and approximately control the edge-of-chaos natural, technological, and social systems from which we extract our livings; the laptop I'm typing this on was summoned from climate and bio systems we don't understand, and from global supply chains no one individual understands (and which could crash tomorrow for all we know), and is connected to an Internet no one understands.
AI systems will soon be such civilizational complex systems, with just as murky epistemics. Again, we'll never fully understand them. We'll understand them less as they progress.
This is what I think, which means I disagree with the median framing of AI safety research in policy circles, the labs, on podcasts, and on the news, which seems to say our problem is to interpret models' internals to make sure they're aligned.
Eventually -- supposedly! -- we'll understand models with many nines of precision, and read out guarantees that the models won't, say, purposely sabotage an air traffic control system, which will make them safe to deploy.
But we've been doing interpretability for at least 14 years, and it still hasn't produced anything like the reliable, scalable understanding required by that vision. As far as I can tell the problem is getting worse.
Interpretability as a field seems to be in a state of stasis at best -- and more likely fighting a rearguard action -- as capabilities advance like a rocket ship.
Contrast the feeling you get looking at a t-SNE plot of latent features in a safety paper, knowing the researchers might have tweaked hyperparameters to tell a clean just-so story, knowing it’s (usually) on a small-ish model, with the relative epistemic solidity of capabilities-progress news that a model has solved a verifiable Millennium Prize problem.
Feel the epistemic hollowness of the former and recognize how common that feeling is in literature making claims about deep learning internals.
I'm not against interpretability research at all as a data exploration approach and a way to back out stylized theories from models; I think just as we have wrong-but-useful theories of the economy that build intuition and provide a vocabulary for policy discussion, we should have the same for models; but the point is to abandon the conceit that they'll lead to formal proofs and formal control guarantees; there is no evidence for this.
This is fine, though! The goal has to be to build a civilization capable of safely leveraging and living with layered multi-agent systems that we'll never be able to make provable formal claims about.
AI safety, then, needs to become less blindered by computer science and more expansive; drawing from the interdisciplinary approach we take to safety in other parts of the industrial economy like air travel safety and car safety, and from the approach we take to disaster planning in areas like hurricanes, nuclear safety, pandemics, and financial crises.
In all of these cases we assume our models of the world are incomplete. We forecast imperfectly, monitor what actually happens, limit how much power we grant double-edged systems when our confidence is low, build layers of redundant protection, and plan for the protections to fail anyway. I think this is much closer to what AI safety is actually going to look like.
For AI agent security and safety, my mental model is aligned with the plot below, which I keep iteratively refining. Here the problem doesn’t reduce to ‘solving alignment’ or interpreting model latent states; it treats these as a single layer among many.
Analogously, in air travel, we assume the physical plane contains stochastic risk; our problem doesn’t reduce to “making planes safe”, it’s multilayered and contains feedbacks.
If history any guide — and I think it should be — we’ll never not be in a messy, layered control situation vis-a-vis making AI safe. And that’s ok, and good, because it means we have a large historical reference class from which to draw experience around securing civilizational systems that are miraculously, and ferociously double-edged.
No posts
来源:Hacker News · AI · joshuasaxe181906.substack.com