Emerging Properties in Self-Supervised Vision Transformers 论文

20212021 IEEE/CVF International Conference on Computer Vision (ICCV)引用 4905

Domain Adaptation and Few-Shot LearningAdvanced Neural Network ApplicationsAdvanced Image and Video Retrieval Techniques

人工智能 Advanced Neural Network Applications Domain Adaptation and Few-Shot Learning Advanced Image and Video Retrieval Techniques

关系图谱

作者

摘要

In this paper, we question if self-supervised learning provides new properties to Vision Transformer (ViT) [16] that stand out compared to convolutional networks (convnets). Beyond the fact that adapting self-supervised methods to this architecture works particularly well, we make the following observations: first, self-supervised ViT features contain explicit information about the semantic segmentation of an image, which does not emerge as clearly with supervised ViTs, nor with convnets. Second, these features are also excellent k-NN classifiers, reaching 78.3% top-1 on ImageNet with a small ViT. Our study also underlines the importance of momentum encoder [26], multi-crop training [9], and the use of small patches with ViTs. We implement our findings into a simple self-supervised method, called DINO, which we interpret as a form of self-distillation with no labels. We show the synergy between DINO and ViTs by achieving 80.1% top-1 on ImageNet in linear evaluation with ViT-Base.

作者查看全部 (7)

Armand Joulin

Piotr Bojanowski

Julien Mairal

Hervé Jeǵou

Emerging Properties in Self-Supervised Vision Transformers 论文

摘要

作者查看全部 (7)

相关技术查看全部 (3)

相关事件

相关文章