Introspective Coupling: Self-Explanation Training Tracks Behavioral Change Despite Fixed Supervision

LLM Model Editing Latent Reasoning
当训练语言模型(LMs)生成对其预测结果的解释时,何种条件下能实现真实的自我反思,而非流于表面的模仿?我们研究了经训练后能够解释其输入中哪些特征影响了自身行为的语言模型,并以模型在经修改输入上的反事实行为作为监督信号。出人意料的是,我们发现:即使使用来自模型自身早期检查点、甚至来自不同模型族中行为相似模型的固定反事实解释进行训练,语言模型仍常常生成更忠实反映其当前行为、而非训练目标行为的解释。这种解释与行为之间“自我反思式”的耦合现象,发生在训练所用解释在整个训练过程中仍与模型当前行为保持足够相关性的情况下——尽管模型自身的行为本身已在持续演化。我们还进一步表明,这种自我反思式耦合能够动态追踪行为变化:当解释训练与其他后训练目标同步开展时,模型生成的解释会自动跟随这些行为变化,而无需额外提供更新后的监督信号。该现象在多项任务中均有体现,包括迎合性(sycophancy)和拒绝响应(refusal),且对标签噪声具有鲁棒性。总体而言,我们的结果表明,即便是静态的反事实解释数据集,也能为模型提供一种可扩展、可泛化的后训练信号,从而有效支持其自我反思能力的发展。
When does training language models (LMs) to generate explanations of their predictions yield faithful introspection, rather than superficial imitation? We study LMs trained to explain which features of their inputs influenced their behavior, using models' counterfactual behavior on modified inputs as supervision. Surprisingly, we find that LMs trained on fixed counterfactual explanations derived from earlier checkpoints of themselves, or even from behaviorally similar models in different families, frequently produce explanations more faithful to their own current behaviors than to those of their training targets. This "introspective" coupling between LM explanations and behaviors occurs when training explanations remain sufficiently correlated with current behaviors over the course of training, even as behaviors themselves shift. We also show that introspective coupling tracks behavior shifts: when explanation training is provided concurrently with other post-training objectives, explanations track those shifts without requiring updated supervision. This phenomenon appears in multiple tasks, including sycophancy and refusal, and is robust to label noise. Overall, our results show that even fixed datasets of counterfactual explanations can provide scalable and generalizable post-training signal for introspection.
许愿