Data Attribution of Emergent Misalignment with Persona Features 文章

ArXiv CS.CL2026-08-12PAPERen作者: Clemens Vetter, David Kacz\'er, Lucie Flek, Florian Mai

详细信息

来源站点
ArXiv CS.CL
作者
Clemens Vetter, David Kacz\'er, Lucie Flek, Florian Mai
文章类型
PAPER
语言
en
发布日期
2026-08-12

摘要

arXiv:2608.11025v1 Announce Type: new Abstract: Emergent misalignment (EM) is the phenomenon where fine-tuning a language model on a narrow task leads to harmful behavior in unrelated domains. A leading mechanistic account attributes EM to persona features: latent directions acquired during pre-training that misaligned fine-tuning amplifies. We ask where these features come from: which pre-training documents activate them, and whether naturally occurring human-written text suffices to induce EM. Using Sparse Autoencoder (SAE) based model diffing across four open-weight models, we find that features related to jailbreak personas, sarcasm, deception, and manipulation are amplified by misalignment fine-tuning, while safety-relevant and assistant-identity features are suppressed.

相关事件

暂无数据

相关公司

暂无数据

相关人物

暂无数据

相关产品

暂无数据