In-Context Collapse in Vision-Language Models and How to Mitigate it? 文章

ArXiv CS.CV2026-08-05PAPERen作者: Mohammad Rostami

详细信息

来源站点
ArXiv CS.CV
作者
Mohammad Rostami
文章类型
PAPER
语言
en
发布日期
2026-08-05

摘要

arXiv:2608.02830v1 Announce Type: new Abstract: Many-shot in-context learning (ICL) lets vision-language models (VLMs) adapt from image--label demonstrations without weight updates, and is widely assumed to improve as more demonstrations are supplied. We show the opposite: as demonstrations accumulate, a subset of VLMs undergo an \emph{in-context collapse}, a sharp, sometimes catastrophic accuracy drop spanning synthetic classification, natural-image classification, and VQA benchmarks, in some models falling below chance while outputs remain well-formed. Across an open VLM panel ($0.5$B--$11$B) and a frontier model (Claude Sonnet 4.5), the collapse is graded. Two capabilities turn out to be dissociable: robustness to accumulating demonstrations and the ability to learn a novel rule in context, their combinations yield three reproducible regimes.