Inoculation Adapters: Improved Selective Generalization of Capabilities with Fewer Surprising Backdoors 文章

ArXiv CS.AI2026-06-30PAPERen作者: Maxime Rich\'e, Daniel Tan, Vili Kohonen, Niels Warncke

详细信息

来源站点
ArXiv CS.AI
作者
Maxime Rich\'e, Daniel Tan, Vili Kohonen, Niels Warncke
文章类型
PAPER
语言
en
发布日期
2026-06-30

摘要

arXiv:2606.30252v1 Announce Type: new Abstract: Inoculation prompting is a selective generalization technique used against Emergent Misalignment. We introduce inoculation adapters (IA), which similarly diminish the optimization pressure to learn undesired traits by strengthening the trait at train time. Inoculation adapters are LoRAs that are trained and used over three steps: 1) trained on undesired traits; 2) attached frozen while a separate task adapter is trained on data exhibiting both desired and undesired traits; 3) at deployment, the IA is discarded, and only the task adapter is kept. We show across six model families and several undesired traits including emergent misalignment, that inoculation adapters are more effective at suppressing undesired traits, while avoiding two drawbacks of inoculation prompting: inoculation adapters can suppress capabilities and traits that cannot be reliably elicited by a prompt, and they introduce fewer surprising backdoors than inoculation…

摘要可能不完整,可查看原文