Transplanting, inverting, and preventing a misalignment persona: method-conditional emergent misalignment in Qwen2.5 文章

ArXiv CS.CL2026-08-05PAPERen作者: Lyndon Drake (University of Oxford), Zandi Eberstadt (University of Oxford)

详细信息

来源站点
ArXiv CS.CL
作者
Lyndon Drake (University of Oxford), Zandi Eberstadt (University of Oxford)
文章类型
PAPER
语言
en
发布日期
2026-08-05

摘要

arXiv:2607.04510v2 Announce Type: replace Abstract: Emergent misalignment (EM) --- the broad misbehaviour a language model acquires after fine-tuning on narrow harmful data --- is mediated in Qwen2.5 models by a latent persona direction, and that direction is causal in open weights. Transplanting it into a model that shares only pretraining with its source induces broad EM (2.83 $\pm$ 0.26\% misaligned against a random-direction floor of $\sim$1.1\%), and ablating a model's own direction roughly halves an overt inducer's broadcast (21\% to 10\%). The transplant doubles as a measurement method, causally assaying directions that a source model represents but cannot itself express. Whether a fine-tune recruits this persona depends on method and capacity, and since low-rank PEFT is the cheaper regime at scale, the recruiting method is also the economical one. On Qwen2.5-32B, LoRA at low ranks on insecure code recruits it (3.4\% misaligned) while full SFT on identical data does not (0.

相关事件

暂无数据

相关人物

暂无数据