Paper Lineage
Esc
MethodDec 2020arXiv 2012.12877cs.CV

Training data-efficient image transformers & distillation through attention

Hugo Touvron, Matthieu Cord, Matthijs Douze and 3 others

Recently, neural networks purely based on attention were shown to address image understanding tasks such as image classification. However, these visual transformers are pre-trained with hundreds of millions of images using an expensive infrastructure, thereby limiting their adoption.

From the abstract

Built on

5 papers · 0 verifiedSee as graph

Also cited · not yet reviewed (5)

  • ViT2020 · cited 12×
    “We build upon the visual transformer architecture from Dosovitskiy et al. [15] and improvements included in the timm library [55].”
    From this paper · §Introduction
  • Transformer2017 · cited 7×
    “introduced by Vaswani et al. [52] for machine translation are currently the reference model for all natural language processing (NLP) tasks.”
    From this paper · §Related work
  • Knowledge Distillation2015 · cited 2×
    “(KD), introduced by Hinton et al. [24], refers to the training paradigm in which a student model leverages “soft” labels coming from a strong teacher network.”
    From this paper · §Related work
  • BERT2018 · cited 2×
    “Motivated by the success of attention-based models in Natural Language Processing [14, 52], there has been increasing interest in architectures leveraging attention mechanisms within convnets [2, 34, 61].”
    From this paper · §Introduction
Show 1 more
  • EfficientNet2019 · cited 2×
    “The evolution of the state of the art on the ImageNet dataset [42] reflects the progress with convolutional neural network architectures and learning [32, 44, 48, 50, 51, 57].”
    From this paper · §Related work

Led to

  • ViLT2021 · cited 2×
    “Recent work (Dosovitskiy et al. 2020; Touvron et al. 2020) demonstrated that using a simple linear projection of a patch is effective enough to embed pixels before feeding them into transformers.”
    From ViLT · §Introduction
  • Swin2021 · cited 8×, 1 in Method
    “Most related to our work is the Vision Transformer (ViT) [20] and its follow-ups [63, 72, 15, 28, 66].”
    From Swin · §Related Work
  • BEiT2021 · cited 3×
    “We directly follow the most of hyperparameters of DeiT [38] in our fine-tuning experiments for a fair comparison.”
    From BEiT · §Experiments
  • Video Swin2021 · cited 4×
    “A shift in backbone architectures for computer vision, from CNNs to Transformers, began recently with Vision Transformer (ViT) [8, 34].”
    From Video Swin · §Related Works
  • ALBEF2021 · cited 2×, 1 in Method
    “We use a 12-layer visual transformer ViT-B/16 [38] as the image encoder, and initialize it with weights pre-trained on ImageNet-1k from [31].”
    From ALBEF · §ALBEF Pre-training
  • VLMo2021 · cited 3×, 1 in Method
    “Following vision Transformers [12, 44, 2], the 2D image 𝒗∈ℝH×W×C{\bm{v}}\in\mathbb{R}^{H\times W\times C} is split and reshaped into N=H​W/P2N={HW}/{P^{2}} patches 𝒗p∈ℝN×(P2​C){\bm{v}}^{p}\in\mathbb{R}^{N\times(P^{2}C)…”
    From VLMo · §Methods
  • METER2021 · cited 4×, 4 in Method
    “ViT has become a popular research topic recently dosovitskiy2020image; touvron2020deit; touvron2020deit; touvron2021going; yuan2021volo; liu2021swin; bao2021beit, and has been introduced into VLP kim2021vilt; xue2021prob…”
    From METER · §The Meter Framework
Abstract

Recently, neural networks purely based on attention were shown to address image understanding tasks such as image classification. However, these visual transformers are pre-trained with hundreds of millions of images using an expensive infrastructure, thereby limiting their adoption. In this work, we produce a competitive convolution-free transformer by training on Imagenet only. We train them on a single computer in less than 3 days. Our reference vision transformer (86M parameters) achieves top-1 accuracy of 83.1% (single-crop evaluation) on ImageNet with no external data. More importantly, we introduce a teacher-student strategy specific to transformers. It relies on a distillation token ensuring that the student learns from the teacher through attention. We show the interest of this token-based distillation, especially when using a convnet as a teacher. This leads us to report results competitive with convnets for both Imagenet (where we obtain up to 85.2% accuracy) and when transferring to other tasks. We share our code and models.