Learning without training: The implicit dynamics of in-context learning 文章

ArXiv CS.CL2026-06-03NEWSen作者: Benoit Dherin, Michael Munn, Hanna Mazzawi, Michael Wunder, Javier Gonzalvo

摘要

arXiv:2507.16003v4 Announce Type: replace Abstract: One of the most striking features of Large Language Models (LLMs) is their ability to learn in-context. Namely at inference time an LLM is able to learn new patterns without any additional weight update when these patterns are presented in the form of examples in the prompt, even if these patterns were not seen during training. The mechanisms through which this can happen are still largely unknown. In this work, we show that the stacking of a self-attention layer with an MLP allows the transformer block to implicitly modify the weights of the MLP layer according to the context. We argue through theoretical analysis and experimentation that this simple mechanism may help explain why LLMs demonstrate capabilities of in-context learning, beyond what is captured during training.

Learning without training: The implicit dynamics of in-context learning 文章

摘要

相关事件

相关公司

相关人物

相关产品

相关技术查看全部 (5)