详细信息
- 来源站点
- ArXiv CS.AI
- 作者
- Maxime Rich\'e, Daniel Tan, Vili Kohonen, Niels Warncke
- 文章类型
- PAPER
- 语言
- en
- 发布日期
- 2026-06-30
摘要
arXiv:2606.30252v1 Announce Type: new Abstract: Inoculation prompting is a selective generalization technique used against Emergent Misalignment. We introduce inoculation adapters (IA), which similarly diminish the optimization pressure to learn undesired traits by strengthening the trait at train time. Inoculation adapters are LoRAs that are trained and used over three steps: 1) trained on undesired traits; 2) attached frozen while a separate task adapter is trained on data exhibiting both desired and undesired traits; 3) at deployment, the IA is discarded, and only the task adapter is kept. We show across six model families and several undesired traits including emergent misalignment, that inoculation adapters are more effective at suppressing undesired traits, while avoiding two drawbacks of inoculation prompting: inoculation adapters can suppress capabilities and traits that cannot be reliably elicited by a prompt, and they introduce fewer surprising backdoors than inoculation…
摘要可能不完整,可查看原文
相关事件
暂无数据
相关人物
暂无数据