Position: Use Sparse Autoencoders to Discover Unknowns 文章

ArXiv CS.CL2026-07-07PAPERen作者: Kenny Peng, Rajiv Movva, Jon Kleinberg, Emma Pierson, Nikhil Garg

详细信息

来源站点
ArXiv CS.CL
作者
Kenny Peng, Rajiv Movva, Jon Kleinberg, Emma Pierson, Nikhil Garg
文章类型
PAPER
语言
en
发布日期
2026-07-07

摘要

arXiv:2506.23845v2 Announce Type: replace-cross Abstract: While sparse autoencoders (SAEs) have generated significant excitement, a series of negative results have added to skepticism about their usefulness. Here, we establish a conceptual distinction that reconciles competing narratives surrounding SAEs. We argue that even if SAEs may be less effective for \textit{acting on known concepts}, SAEs are especially powerful tools for \textit{discovering unknown concepts}. This distinction separates existing negative results from positive results, and suggests several classes of SAE applications. Specifically, we outline use cases for SAEs in (i) ML interpretability, explainability, fairness, auditing, and safety, and (ii) social and health sciences.

相关事件

暂无数据

相关公司查看全部 (1)

相关人物

暂无数据