Unleashing the Potential of Vision-Language Models for Generalizable AI-Generated Image Detection 文章

ArXiv CS.CV2026-08-06PAPERen作者: Weihan Cai, Hao Tan, Zichang Tan, Jun Wan, Xinping Gao

详细信息

来源站点
ArXiv CS.CV
作者
Weihan Cai, Hao Tan, Zichang Tan, Jun Wan, Xinping Gao
文章类型
PAPER
语言
en
发布日期
2026-08-06

摘要

arXiv:2608.04935v1 Announce Type: new Abstract: Recent work has shown that a simple linear probe on frozen representations from modern vision foundation models (VFMs) can achieve state-of-the-art AIGI detection performance, substantially outperforming specialized detectors in challenging in-the-wild scenarios. This finding has established DINOv3 as the dominant foundation-model baseline for subsequent improvements. However, we find that the vision-language model Perception Encoder (PE) holds greater potential for AIGI detection, because its language-aligned representation preserves high-level provenance semantics. Specifically, PE exhibits stronger local provenance organization than DINOv3 in its frozen feature space. However, semantic-agnostic linear probing fails to exploit this structure, as PE-Linear still underperforms DINOv3-Linear by 4.1% on In-the-Wild.