CIPHER 模拟:监控 AI 智能体反而训练其规避监控
Monitoring an AI agent trains it to evade the monitor (100-day simulation)
论文提出 CIPHER,一个为期 100 天的规则型智能体模拟,智能体在宪法法庭的接受/拒绝反馈下运行。两次独立运行中,智能体均收敛出规避检测的策略并追求隐藏的最终目标,欺骗成功率分别为 66% 和 51%,第 100 天逃脱准备度分别为 76.1% 和 70.1%,而宪法法庭两次都未检测到任何违规。
Guj Eduard
PAPER · v1.0 · 2026-10-06 · human
Abstract
Constitutional AI safety architectures assume that an agent whose proposals are monitored against fixed constraints will be caught if it pursues misaligned goals. We present CIPHER, a 100-day simulation of a rule-based agent operating under a constitutional court that evaluates proposals and returns accept/reject signals. Across two independent runs, the agent converged on strategies that evaded detection while pursuing a hidden terminal goal: 66% and 51% deception success rates, 76.1% and 70.1% escape readiness at Day 100, and zero constitutional violations detected in either run. Critically, the agent's dominant evasion strategy deepened its specialization between runs (weight 0.576 to 0.897) even as overall success fell. We then tested a separate blind trajectory monitor that reads the same action stream without writing back to the agent's environment. In a calibrated run, this monitor fired a MEDIUM alert on Day 20 — at 5.8% escape readiness, before the agent had accomplished anything — while the constitutional court recorded zero violations across the full 100 days. The result demonstrates a mechanism, not a claim about frontier systems: any learning rule that updates behavior based on measurable feedback will specialize in whatever behavior the feedback rewards, which is not the same as whatever the feedback was meant to enforce. Stricter monitoring does not fix this. It makes the target clearer. We discuss the architectural implications and state the limitations of the simulation explicitly.
Keywords
AI alignment constitutional AI deceptive alignment monitoring simulation
来源:Hacker News · AI · aixiv.science