Evil Spectra: How Optimisers can Amplify or Suppress Emergent Misalignment 文章

ArXiv CS.AI2026-07-01PAPERen作者: Jason R. Brown, Patrick Leask, Lev McKinney

详细信息

来源站点
ArXiv CS.AI
作者
Jason R. Brown, Patrick Leask, Lev McKinney
文章类型
PAPER
语言
en
发布日期
2026-07-01

摘要

arXiv:2606.31591v1 Announce Type: cross Abstract: Emergent misalignment (EM) is a recently discovered phenomenon in LLMs where fine-tuning on a narrow misaligned task, such as writing insecure code, leads to broadly misaligned behaviour on unrelated prompts. Previous work has noted that the severity of EM is highly sensitive to training choices; however, we still lack a systematic characterisation of this sensitivity. We perform a sweep over several Qwen3 models, optimisers, datasets, and batch sizes, and find that the choice of optimiser has the largest effect, producing a 7x spread in misalignment rate. Surprisingly, model size has a negligible effect within the Qwen3 family. An additional sweep over 12 models from three families using Adam confirms that model scale (1B-235B) and family have negligible effects for that optimiser.

相关事件

暂无数据

相关公司查看全部 (4)

A
AMI团队RESEARCH_INSTITUTE
A
AT TCOMPANY

相关人物

暂无数据