详细信息
- 来源站点
- ArXiv CS.CV
- 作者
- Pierre Gallin-Martel, Mika Feng, Koichi Ito, Takafumi Aoki
- 文章类型
- PAPER
- 语言
- en
- 发布日期
- 2026-07-24
摘要
arXiv:2607.20900v1 Announce Type: new Abstract: With the increasing diversity of spoofing attacks, there is a growing demand for unified Face Anti-Spoofing (FAS) models capable of detecting both physical and digital threats. While existing Vision-Language Models (VLMs) demonstrate high generalization in this context, they heavily rely on complex multimodal fusion and external text encoders. In this paper, we propose DINO-VPT, a lightweight, vision-only framework leveraging hierarchical visual prompt tuning. By dynamically injecting prompts conditioned on input features via a Prompt Routing Network (PRN), our method effectively disentangles diverse spoofing artifacts without requiring multimodal fusion. Evaluations on the UniAttackData benchmark demonstrate that DINO-VPT achieves higher accuracy than state-of-the-art VLM-based methods.
相关事件
暂无数据
相关公司
暂无数据
相关人物
暂无数据