Repurposing CLIP to Localize at Pixel Level 文章

ArXiv CS.CV2026-07-07PAPERen作者: Jiaxiang Fang, Shiqiang Ma, Jing Wang, Siyu Chen, Fei Guo, Shengfeng He

详细信息

来源站点
ArXiv CS.CV
作者
Jiaxiang Fang, Shiqiang Ma, Jing Wang, Siyu Chen, Fei Guo, Shengfeng He
文章类型
PAPER
语言
en
发布日期
2026-07-07

摘要

arXiv:2607.05253v1 Announce Type: new Abstract: Large-scale Vision-Language Models like CLIP have demonstrated impressive open-set localization capabilities at the image level. However, adapting this capability to pixel-level dense prediction poses challenges due to global feature biases. In this paper, we introduce CLIPix, a simple yet effective framework that repurposes CLIP to perform pixel-level localization. By tracing back CLIP's classification process, CLIPix identifies object-specific attentive regions and repurposes them as pixel-level localization cues. To address noise introduced by global biases, we propose a Noise-Resistant Correction strategy, refining these cues for more precise segmentation. Additionally, we introduce a Localization Embedding strategy to integrate both localization and enriched detail information, enabling accurate, high-resolution segmentation.