Measuring Semantic Abstractness of SAE Features via Nonlocality 文章

ArXiv CS.AI2026-08-12PAPERen作者: Chuqiao Lin, Shivaji Sondhi, Xiao-Liang Qi

详细信息

来源站点
ArXiv CS.AI
作者
Chuqiao Lin, Shivaji Sondhi, Xiao-Liang Qi
文章类型
PAPER
语言
en
发布日期
2026-08-12

摘要

arXiv:2608.10537v1 Announce Type: new Abstract: Sparse autoencoders (SAEs) have helped uncover mechanistic explanations for LLM behaviours such as reasoning, jailbreaking etc., via understanding the corresponding task-relevant and causally effective features. To evaluate such mechanistic explanations, downstream studies must distinguish surface lexical features from genuinely high-level ones. However, neither an autointerp-based semantic description nor causal steering utility fully resolves the abstraction level of a feature. To this end, we introduce \emph{Feature Nonlocality} (FNL), defined as the entropy of the normalized per-position influence on an SAE feature's activation. We report that FNL correlates with existing LLM-based proxy metrics of feature semantic abstractness, and successfully distinguishes context-dependent reasoning features from token-driven ones, correctly assigning the higher FNL to the contextual feature in $73$--$84\%$ of randomly drawn pairs that consist…

摘要可能不完整,可查看原文

相关事件

暂无数据

相关公司

暂无数据

相关人物

暂无数据

相关产品

暂无数据