LangFIR: Discovering Sparse Language-Specific Features from Monolingual Data for Language Steering 文章

ArXiv CS.CL2026-08-04PAPERen作者: Sing Hieng Wong, Hassan Sajjad, A. B. Siddique

详细信息

来源站点
ArXiv CS.CL
作者
Sing Hieng Wong, Hassan Sajjad, A. B. Siddique
文章类型
PAPER
语言
en
发布日期
2026-08-04

摘要

arXiv:2604.03532v2 Announce Type: replace Abstract: Large language models (LLMs) show strong multilingual capabilities, yet reliably controlling the language of their outputs remains difficult. Representation-level steering addresses this by adding language-specific vectors to model activations at inference time, but identifying language-specific directions in the residual stream often relies on multilingual or parallel data that can be expensive to obtain. Sparse autoencoders (SAEs) decompose residual activations into interpretable, sparse feature directions and offer a natural basis for this search, yet existing SAE-based approaches face the same data constraint. We introduce LangFIR (Language Feature Identification via Random-token Filtering), a method that discovers language-specific SAE features using only a small amount of monolingual data and random-token sequences. Many SAE features consistently activated by target-language inputs do not encode language identity.

相关事件

暂无数据

相关公司

暂无数据

相关人物

暂无数据

相关产品

暂无数据