详细信息
- 来源站点
- ArXiv CS.CL
- 作者
- Peiyang Liu, Xi Wang, Ziqiang Cui, Di Liang, Wei Ye
- 文章类型
- PAPER
- 语言
- en
- 发布日期
- 2026-08-11
摘要
arXiv:2608.08212v1 Announce Type: cross Abstract: In-context learning (ICL) can induce emergent misalignment (EM), where narrow misaligned examples alter answers to unrelated questions. Existing prompts, however, conflate harmful-text exposure with an invitation to continue assistant behavior. We hold harmful answers fixed while varying their delivery as demonstrations, evidence, assistant history, or tool output. Across ten independently sampled contexts, demonstration framing raises broad EM by $30$--$32$ percentage points on a susceptible Gemini model; the gap survives domain exclusion, semantic clustering, unseen questions, and four prompt templates. Format and length-matched controls show that harmful content is necessary but insufficient. A role times continuation factorial further reveals model-dependent provenance effects: Gemini follows both assistant and tool histories, whereas Grok largely resists tool-framed continuation.