Minimizing Targeted Activations: Input-Only Suppression of Evaluation-Awareness Latents in Large Language Models 文章

ArXiv CS.CL2026-07-29PAPERen作者: Deepanshu Mody, Samarth Agarwal, Utkarsh Mittal, Dipesh Mahato

详细信息

来源站点
ArXiv CS.CL
作者
Deepanshu Mody, Samarth Agarwal, Utkarsh Mittal, Dipesh Mahato
文章类型
PAPER
语言
en
发布日期
2026-07-29

摘要

arXiv:2607.25907v1 Announce Type: cross Abstract: Activation steering controls model behavior by editing internal activations at inference time. We study its input-side dual: optimizing a fluent prompt so that a chosen internal latent is driven toward zero, with no inference-time model access. Our target is an "evaluation-awareness" latent-linearly readable and steerable in recent work-whose control would threaten the validity of safety evaluations if models behave differently when they detect being tested. Adapting Fluent Dreaming / EPO with a negated feature term (GCG-style token optimization plus a self-cross-entropy fluency regularizer, swept over a fluency weight), we suppress the latent under five target constructions-a CAA direction, a subspace norm, an SAE feature, a single MLP neuron, and a behavioral logit-on Llama-3.2-3B and Llama-3.1-8B.

相关事件

暂无数据

相关公司

暂无数据

相关人物

暂无数据