Retrieval is Enough: Training-Free Interpretability with a Tool-Using Agent 文章

ArXiv CS.AI2026-07-21PAPERen作者: Sriram Balasubramanian, Soheil Feizi

详细信息

来源站点
ArXiv CS.AI
作者
Sriram Balasubramanian, Soheil Feizi
文章类型
PAPER
语言
en
发布日期
2026-07-21

摘要

arXiv:2607.16448v1 Announce Type: cross Abstract: Interpretability methods for neural network activations span a wide cost spectrum, from cheap, training-free techniques (such as linear probes, PCA, SVD) to more expensive training-based ones (such as SAEs and activation oracles). Training-based methods are typically more powerful, in part because they leverage large activation datasets during training. This raises a natural question - do they actually surface insights that go beyond what is recoverable from the training dataset itself? To address this, we equip an LLM agent with a vector database of activations paired with their textual contexts, along with tools for manipulating activations - projecting out directions in latent space, computing activation differences and averages. The agent iteratively queries the database, forms hypotheses from the retrieved samples, and validates them by constructing linear probes.

相关事件

暂无数据

相关公司

暂无数据

相关人物

暂无数据

相关产品

暂无数据