详细信息
- 来源站点
- ArXiv CS.CV
- 作者
- Weihan Cai, Hao Tan, Zichang Tan, Jun Wan, Xinping Gao
- 文章类型
- PAPER
- 语言
- en
- 发布日期
- 2026-08-06
摘要
arXiv:2608.04935v1 Announce Type: new Abstract: Recent work has shown that a simple linear probe on frozen representations from modern vision foundation models (VFMs) can achieve state-of-the-art AIGI detection performance, substantially outperforming specialized detectors in challenging in-the-wild scenarios. This finding has established DINOv3 as the dominant foundation-model baseline for subsequent improvements. However, we find that the vision-language model Perception Encoder (PE) holds greater potential for AIGI detection, because its language-aligned representation preserves high-level provenance semantics. Specifically, PE exhibits stronger local provenance organization than DINOv3 in its frozen feature space. However, semantic-agnostic linear probing fails to exploit this structure, as PE-Linear still underperforms DINOv3-Linear by 4.1% on In-the-Wild.