When Structured Sparse Autoencoders Learn Consistent Concepts Across Modalities 文章

ArXiv CS.CV2026-07-10PAPERen作者: Weiduo Liao, Yunqiao Yang, Ying Wei

详细信息

来源站点
ArXiv CS.CV
作者
Weiduo Liao, Yunqiao Yang, Ying Wei
文章类型
PAPER
语言
en
发布日期
2026-07-10

摘要

arXiv:2607.08605v1 Announce Type: new Abstract: Sparse autoencoders (SAEs) have emerged as a promising technique for mechanistic interpretability by learning a set of sparse latent features in large models, each of which encodes a distinct concept. However, in vision-language models (VLMs), vanilla SAEs struggle to learn modality-consistent concepts, with concepts often exhibiting fragmented coverage (i.e., disjoint regions) in the visual modality. To address this challenge, we propose a Structured Sparse AutoEncoder ($S^2AE$) that enforces concept consistency from both semantic and spatial perspectives in the visual modality. Specifically, we group image patches based on Transformer attention similarity and spatial proximity, and introduce a structured sparsity regularization when training the vanilla SAE.

相关事件

暂无数据

相关公司查看全部 (2)

A
ANICOMPANY

相关人物

暂无数据