Cross-seed explainability using Procrustes-conditioned Joint End-to-end Top-K Sparse Autoencoders 文章

ArXiv CS.CL2026-07-10PAPERen作者: Bendeg\'uz V\'aradi, Zolt\'an Kmetty

详细信息

来源站点
ArXiv CS.CL
作者
Bendeg\'uz V\'aradi, Zolt\'an Kmetty
文章类型
PAPER
语言
en
发布日期
2026-07-10

摘要

arXiv:2607.08499v1 Announce Type: new Abstract: We present a Procrustes-conditioned Joint End-to-end Top-K Sparse Autoencoder (SAE) for extracting cross-seed universal features from independently trained BERT models. Cross-seed feature universality is a fundamental challenge in mechanistic interpretability: because dictionary learning is non-convex, independently trained networks learn misaligned feature spaces, so apparently identical features may differ by random initialization. We address this by computing an orthogonal Procrustes rotation between seeds' activation spaces before joint SAE training, combining Top-K sparsity, end-to-end downstream optimization, and an auxiliary dead-feature revival loss based on previous SAE literature. Evaluating on five independent seed pairs (ten BERT models) across three benchmark datasets (SST-2, Stanford Politeness, TweetEval Emotion), our full pipeline produces more universal features (Pearson r $\geq$ 0.

相关事件

暂无数据

相关公司查看全部 (2)

A
ANICOMPANY

相关人物

暂无数据