From Expressivity to Sample Complexity: Narrow Teachers for Transformers via C-RASP 文章

ArXiv CS.CL2026-07-14PAPERen作者: Michael Rizvi-Martel, Satwik Bhattamishra, Guillaume Rabusseau, Michael Hahn

详细信息

来源站点
ArXiv CS.CL
作者
Michael Rizvi-Martel, Satwik Bhattamishra, Guillaume Rabusseau, Michael Hahn
文章类型
PAPER
语言
en
发布日期
2026-07-14

摘要

arXiv:2607.11760v1 Announce Type: cross Abstract: A theoretical understanding of Transformers is crucial to better understand the capacities and limitations of large language models (LLMs). There is much work analyzing the expressivity of attention-based models. By proposing handcrafted weights or using computational complexity arguments, a large amount of past theoretical works have sought to characterize which tasks are and which are not in the hypothesis class of Transformer models. However, little work investigates the learnability of such solutions. In this work, we make progress towards this goal. Inspired by recent loss landscape analysis work, we propose preliminary sample complexity bounds for learning C-RASP constructions with Transformers.

相关事件

暂无数据

相关公司

暂无数据

相关人物

暂无数据