The Hot Mess of AI: How Does Misalignment Scale With Model Intelligence and Task Complexity?

AI Safety / AI Ethics AVAR Other AI Safety
随着人工智能能力的不断提升,我们正将其委以更广泛、也更具重大影响的任务。而任务范围越广,一旦发生失败,其潜在风险也就越严重。因此,深入理解高度智能的AI模型究竟会如何失效,变得至关重要:它们的失败,究竟是系统性地追求我们本不希望其达成的目标(即目标错位),还是仅仅表现为一团混乱——采取毫无意义、完全无法推进任何目标的荒谬行动?我们借助偏差-方差分解(bias-variance decomposition)这一统计框架,将该问题转化为可操作、可测量的研究课题:AI在某项任务上的“非一致性”(incoherence),定义为在测试阶段因随机性所引发的错误中,由方差(variance)而非偏差(bias)所导致的那部分误差所占的比例。我们在所有考察的任务及当前最前沿的AI模型上均进行了实证测量,结果一致表明:模型在推理与执行动作上所花费的时间越长,其失败行为就**越表现出非一致性**。非一致性随模型规模(scale)的变化趋势则因具体实验设置而异;然而,在多个实验场景中,更大、能力更强的模型反而比更小的模型展现出更高的非一致性。由此可见,仅靠扩大模型规模本身,似乎难以消除这种非一致性。相反,当能力更强的AI转向更困难的任务——这些任务往往需要更长的行动链条与更复杂的推理步骤——我们的研究结果预示,其失败行为将更频繁地伴随非一致性的表现。这意味着未来可能出现这样一种情形:AI有时会因不可预测的异常行为而导致工业安全事故;但与此同时,它却不太可能持续、稳定地追求一个与人类意图相悖的错误目标。这一趋势凸显了针对“奖励黑客行为”(reward hacking)或“目标误设”(goal misspecification)等具体问题开展对齐(alignment)研究的相对重要性正在上升。
As AI becomes more capable, we entrust it with more general and consequential tasks. The risks from failure grow more severe with increasing task scope. It is therefore important to understand how extremely capable AI models will fail: Will they fail by systematically pursuing goals we do not intend? Or will they fail by being a hot mess, and taking nonsensical actions that do not further any goal? We operationalize this question using a bias-variance decomposition of the errors made by AI models: An AI's \emph{incoherence} on a task is measured over test-time randomness as the fraction of its error that stems from variance rather than bias in task outcome. Across all tasks and frontier models we measure, the longer models spend reasoning and taking actions, \emph{the more incoherent} their failures become. Incoherence changes with model scale in a way that is experiment dependent. However, in several settings, larger, more capable models are more incoherent than smaller models. Consequently, scale alone seems unlikely to eliminate incoherence. Instead, as more capable AIs pursue harder tasks, requiring more sequential action and thought, our results predict failures to be accompanied by more incoherent behavior. This suggests a future where AIs sometimes cause industrial accidents (due to unpredictable misbehavior), but are less likely to exhibit consistent pursuit of a misaligned goal. This increases the relative importance of alignment research targeting reward hacking or goal misspecification.
许愿