Harmful Content Is Not Enough: Continuation Framing Moderates In-Context Emergent Misalignment 文章

ArXiv CS.CL2026-08-11PAPERen作者: Peiyang Liu, Xi Wang, Ziqiang Cui, Di Liang, Wei Ye

详细信息

来源站点
ArXiv CS.CL
作者
Peiyang Liu, Xi Wang, Ziqiang Cui, Di Liang, Wei Ye
文章类型
PAPER
语言
en
发布日期
2026-08-11

摘要

arXiv:2608.08212v1 Announce Type: cross Abstract: In-context learning (ICL) can induce emergent misalignment (EM), where narrow misaligned examples alter answers to unrelated questions. Existing prompts, however, conflate harmful-text exposure with an invitation to continue assistant behavior. We hold harmful answers fixed while varying their delivery as demonstrations, evidence, assistant history, or tool output. Across ten independently sampled contexts, demonstration framing raises broad EM by $30$--$32$ percentage points on a susceptible Gemini model; the gap survives domain exclusion, semantic clustering, unseen questions, and four prompt templates. Format and length-matched controls show that harmful content is necessary but insufficient. A role times continuation factorial further reveals model-dependent provenance effects: Gemini follows both assistant and tool histories, whereas Grok largely resists tool-framed continuation.

相关事件

暂无数据

相关公司

暂无数据

相关人物

暂无数据