详细信息
- 来源站点
- ArXiv CS.CL
- 作者
- Deepanshu Mody, Samarth Agarwal, Utkarsh Mittal, Dipesh Mahato
- 文章类型
- PAPER
- 语言
- en
- 发布日期
- 2026-07-29
摘要
arXiv:2607.25907v1 Announce Type: cross Abstract: Activation steering controls model behavior by editing internal activations at inference time. We study its input-side dual: optimizing a fluent prompt so that a chosen internal latent is driven toward zero, with no inference-time model access. Our target is an "evaluation-awareness" latent-linearly readable and steerable in recent work-whose control would threaten the validity of safety evaluations if models behave differently when they detect being tested. Adapting Fluent Dreaming / EPO with a negated feature term (GCG-style token optimization plus a self-cross-entropy fluency regularizer, swept over a fluency weight), we suppress the latent under five target constructions-a CAA direction, a subspace norm, an SAE feature, a single MLP neuron, and a behavioral logit-on Llama-3.2-3B and Llama-3.1-8B.
相关事件
暂无数据
相关公司
暂无数据
相关人物
暂无数据