When Activation Oracles Learn Not to Read: Concept-Specific Blind Spots in Fine-Tuned Oracles 文章

ArXiv CS.CL2026-07-28PAPERen作者: Tobias Bersia, Tatiana Gaintseva

详细信息

来源站点
ArXiv CS.CL
作者
Tobias Bersia, Tatiana Gaintseva
文章类型
PAPER
语言
en
发布日期
2026-07-28

摘要

arXiv:2607.23379v1 Announce Type: new Abstract: Activation Oracles (AOs) are language models trained to answer natural-language questions about another model's internal activations. They offer a flexible interface for reading hidden information from model states, especially when relevant information is internally represented but absent or incomplete in visible behavior. However, AOs are themselves learned systems: their answers are shaped by training data, objectives, and learned reporting behavior, rather than being neutral readouts of represented information. We study this in a controlled Taboo Word Guessing setting, where subject models are fine-tuned to internally use a hidden concept while avoiding direct disclosure. Contrary to the expectation that an AO trained on such a subject becomes a specialist reader, we find that fine-tuned AOs can become concept-specific anti-readers: they selectively fail to recover the concept persistently present during their own training.

相关事件

暂无数据

相关公司

暂无数据

相关人物

暂无数据

相关产品

暂无数据