Training data-efficient image transformers & distillation through attention
Recently, neural networks purely based on attention were shown to address image understanding tasks such as image classification. However, these visual transformers are pre-trained with hundreds of millions of images using an expensive infrastructure, thereby limiting their adoption.
Also cited · not yet reviewed (5)
- ViT2020 · cited 12דWe build upon the visual transformer architecture from Dosovitskiy et al. [15] and improvements included in the timm library [55].”From this paper · §Introduction
- Transformer2017 · cited 7דintroduced by Vaswani et al. [52] for machine translation are currently the reference model for all natural language processing (NLP) tasks.”From this paper · §Related work
- Knowledge Distillation2015 · cited 2ד(KD), introduced by Hinton et al. [24], refers to the training paradigm in which a student model leverages “soft” labels coming from a strong teacher network.”From this paper · §Related work
- BERT2018 · cited 2דMotivated by the success of attention-based models in Natural Language Processing [14, 52], there has been increasing interest in architectures leveraging attention mechanisms within convnets [2, 34, 61].”From this paper · §Introduction
Show 1 more
- EfficientNet2019 · cited 2דThe evolution of the state of the art on the ImageNet dataset [42] reflects the progress with convolutional neural network architectures and learning [32, 44, 48, 50, 51, 57].”From this paper · §Related work
Led to
- ViLT2021 · cited 2דRecent work (Dosovitskiy et al. 2020; Touvron et al. 2020) demonstrated that using a simple linear projection of a patch is effective enough to embed pixels before feeding them into transformers.”From ViLT · §Introduction
- Swin2021 · cited 8×, 1 in Method“Most related to our work is the Vision Transformer (ViT) [20] and its follow-ups [63, 72, 15, 28, 66].”From Swin · §Related Work
- BEiT2021 · cited 3דWe directly follow the most of hyperparameters of DeiT [38] in our fine-tuning experiments for a fair comparison.”From BEiT · §Experiments
- Video Swin2021 · cited 4דA shift in backbone architectures for computer vision, from CNNs to Transformers, began recently with Vision Transformer (ViT) [8, 34].”From Video Swin · §Related Works
- ALBEF2021 · cited 2×, 1 in Method“We use a 12-layer visual transformer ViT-B/16 [38] as the image encoder, and initialize it with weights pre-trained on ImageNet-1k from [31].”From ALBEF · §ALBEF Pre-training
- VLMo2021 · cited 3×, 1 in Method“Following vision Transformers [12, 44, 2], the 2D image 𝒗∈ℝH×W×C{\bm{v}}\in\mathbb{R}^{H\times W\times C} is split and reshaped into N=HW/P2N={HW}/{P^{2}} patches 𝒗p∈ℝN×(P2C){\bm{v}}^{p}\in\mathbb{R}^{N\times(P^{2}C)…”From VLMo · §Methods
- METER2021 · cited 4×, 4 in Method“ViT has become a popular research topic recently dosovitskiy2020image; touvron2020deit; touvron2020deit; touvron2021going; yuan2021volo; liu2021swin; bao2021beit, and has been introduced into VLP kim2021vilt; xue2021prob…”From METER · §The Meter Framework
Abstract
Recently, neural networks purely based on attention were shown to address image understanding tasks such as image classification. However, these visual transformers are pre-trained with hundreds of millions of images using an expensive infrastructure, thereby limiting their adoption. In this work, we produce a competitive convolution-free transformer by training on Imagenet only. We train them on a single computer in less than 3 days. Our reference vision transformer (86M parameters) achieves top-1 accuracy of 83.1% (single-crop evaluation) on ImageNet with no external data. More importantly, we introduce a teacher-student strategy specific to transformers. It relies on a distillation token ensuring that the student learns from the teacher through attention. We show the interest of this token-based distillation, especially when using a convnet as a teacher. This leads us to report results competitive with convnets for both Imagenet (where we obtain up to 85.2% accuracy) and when transferring to other tasks. We share our code and models.