Paper Lineage
Esc
MethodOct 2020arXiv 2010.11929cs.CV

An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov and 9 others

While the Transformer architecture has become the de-facto standard for natural language processing tasks, its applications to computer vision remain limited. In vision, attention is either applied in conjunction with convolutional networks, or used to replace certain components of convolutional networks while keeping their overall structure in place.

From the abstract

Built on

8 papers · 0 verifiedSee as graph

Also cited · not yet reviewed (8)

  • Transformer2017 · cited 4×, 2 in Method
    “In model design we follow the original Transformer (Vaswani et al. 2017) as closely as possible.”
    From this paper · §Method
  • BiT2019 · cited 7×, 1 in Method
    “We de-duplicate the pre-training datasets w.r.t. the test sets of the downstream tasks following Kolesnikov et al. 2020.”
    From this paper · §Experiments
  • BERT2018 · cited 4×
    “However, much of their success stems not only from their excellent scalability but also from large scale self-supervised pre-training (Devlin et al. 2019; Radford et al. 2018).”
    From this paper · §Experiments
  • Noisy Student2019 · cited 3×
    “The use of additional data sources allows to achieve state-of-the-art results on standard benchmarks (Mahajan et al. 2018; Touvron et al. 2019; Xie et al. 2020).”
    From this paper · §Related Work
Show 4 more
  • ResNet2015 · cited 2×
    “In computer vision, however, convolutional architectures remain dominant (LeCun et al. 1989; Krizhevsky et al. 2012; He et al. 2016).”
    From this paper · §Introduction
  • “Moreover, Sun et al. 2017 study how CNN performance scales with dataset size, and Kolesnikov et al. 2020; Djolonga et al. 2020 perform an empirical exploration of CNN transfer learning from large scale datasets such as I…”
    From this paper · §Related Work
  • “The use of additional data sources allows to achieve state-of-the-art results on standard benchmarks (Mahajan et al. 2018; Touvron et al. 2019; Xie et al. 2020).”
    From this paper · §Related Work
  • GPT-32020 · cited 2×
    “Large Transformer-based models are often pre-trained on large corpora and then fine-tuned for the task at hand: BERT (Devlin et al. 2019) uses a denoising self-supervised pre-training task, while the GPT line of work use…”
    From this paper · §Related Work

Led to

  • DeiT2020 · cited 12×
    “We build upon the visual transformer architecture from Dosovitskiy et al. [15] and improvements included in the timm library [55].”
    From DeiT · §Introduction
  • ViLT2021 · cited 4×, 1 in Method
    “Recent work (Dosovitskiy et al. 2020; Touvron et al. 2020) demonstrated that using a simple linear projection of a patch is effective enough to embed pixels before feeding them into transformers.”
    From ViLT · §Introduction
  • ALIGN2021 · cited 2×
    “High-quality visual representations for classification or retrieval are usually pre-trained on large-scale labeled datasets (Mahajan et al. 2018; Kolesnikov et al. 2020; Dosovitskiy et al. 2021; Juan et al. 2020).”
    From ALIGN · §Related Work
  • CLIP2021 · cited 4×, 1 in Method
    “For the second architecture, we experiment with the recently introduced Vision Transformer (ViT) (Dosovitskiy et al. 2020).”
    From CLIP · §Approach
  • Perceiver2021 · cited 4×
    “This is the strategy taken by the Vision Transformer (ViT) (Dosovitskiy et al. 2021), which first reduces the input size to ∼200\sim 200 using a 2D convolutional layer (referred to as “linear projection of flattened patc…”
    From Perceiver · §Related Work
  • Swin2021 · cited 8×, 3 in Method
    “Most related to our work is the Vision Transformer (ViT) [20] and its follow-ups [63, 72, 15, 28, 66].”
    From Swin · §Related Work
  • Scaling ViTs (ViT-G)2021 · cited 10×, 4 in Method
    “Attention-based Transformer architectures vaswani2017attention have taken computer vision domain by storm dosovitskiy2020; carion2020endtoend and are becoming an increasingly popular choice in research and practice.”
    From Scaling ViTs (ViT-G) · §Introduction
  • CoAtNet2021 · cited 4×
    “In comparison, after the success of ViT and ResNet-ViT [13], another popular line of research starts with a Transformer backbone and tries to incorporate explicit convolution or some desirable properties of convolution i…”
    From CoAtNet · §Related Work
  • BEiT2021 · cited 7×, 3 in Method
    “Following ViT [12], we use the standard Transformer [42] as the backbone network.”
    From BEiT · §Methods
  • Video Swin2021 · cited 3×
    “A shift in backbone architectures for computer vision, from CNNs to Transformers, began recently with Vision Transformer (ViT) [8, 34].”
    From Video Swin · §Related Works
  • ALBEF2021 · cited 2×, 2 in Method
    “We use a 12-layer visual transformer ViT-B/16 [38] as the image encoder, and initialize it with weights pre-trained on ImageNet-1k from [31].”
    From ALBEF · §ALBEF Pre-training
  • SimVLM2021 · cited 4×
    “We follow the setup in ViT (Dosovitskiy et al. 2021) to explore 3 variants of SimVLM, namely “Base”, “Large”, and “Huge”, such that each variant follows the same setting as its corresponding ViT variant.”
    From SimVLM · §Experiments
  • VLMo2021 · cited 5×, 1 in Method
    “Following vision Transformers [12, 44, 2], the 2D image 𝒗∈ℝH×W×C{\bm{v}}\in\mathbb{R}^{H\times W\times C} is split and reshaped into N=H​W/P2N={HW}/{P^{2}} patches 𝒗p∈ℝN×(P2​C){\bm{v}}^{p}\in\mathbb{R}^{N\times(P^{2}C)…”
    From VLMo · §Methods
  • METER2021 · cited 4×, 3 in Method
    “Transformers vaswani2017attention are prevalent in natural language processing and have recently shown promising performance in computer vision dosovitskiy2020image; liu2021swin.”
    From METER · §Introduction
  • MAE2021 · cited 12×, 2 in Method
    “Motivated by the success in NLP, related recent methods Chen2020c; Dosovitskiy2021; Bao2021 are based on Transformers Vaswani2017. iGPT Chen2020c operates on sequences of pixels and predicts unknown pixels.”
    From MAE · §Related Work
  • LiT2021 · cited 5×
    “We verify LiT across three image-text datasets, with Vision Transformer vit, ResNet bit, and MLP-Mixer mixer architectures.”
    From LiT · §Introduction
  • Swin V22021 · cited 5×
    “The current common practice is to perform a bi-cubic interpolation of the position bias maps dosovitskiy2020vit; liu2021swin.”
    From Swin V2 · §Introduction
  • Florence2021 · cited 2×
    “While inheriting performance benefits of the transformer self-attention operations (Dosovitskiy et al. 2021b), these hierarchical architectures model the scale invariance nature of images and have linear computational co…”
    From Florence · §Introduction
  • ViTCAP2021 · cited 3×
    “ViTCAP is constructed on the basis of a vision transformer dosovitskiy2020image as the stem image encoder.”
    From ViTCAP · §Introduction
  • BLIP2022 · cited 2×, 1 in Method
    “We employ a visual transformer (Dosovitskiy et al. 2021) as our image encoder, which divides an input image into patches and encodes them as a sequence of embeddings, with an additional [CLS] token to represent the globa…”
    From BLIP · §Method
  • Simple end-to-end captioning2022 · cited 3×, 1 in Method
    “Visual Encoder.We follow the same setting as the visual transformer (ViT) which recently attracts much attention and shows remarkable performance in the computer vision area (Dosovitskiy et al. 2021; Radford et al. 2021;…”
    From Simple end-to-end captioning · §. Methodology
  • CoCa2022 · cited 2×, 2 in Method
    “Following a standard encoder-decoder architecture, the image encoder provides latent encoded features (e.g., using a Vision Transformer [39] or ConvNets [40]) and the text decoder learns to maximize the conditional likel…”
    From CoCa · §Approach
  • VL-BEiT2022 · cited 3×, 1 in Method
    “Following [9], we split the image 𝒗∈ℝH×W×C{\bm{v}}\in\mathbb{R}^{H\times W\times C} into a sequence of patches, so that the image can be encoded by standard Transformer.”
    From VL-BEiT · §Methods
  • BEiT v22022 · cited 2×, 1 in Method
    “The vision Transformers (ViTs; Dosovitskiy et al. 2020) are employed as the backbone networks to obtain image representations.”
    From BEiT v2 · §Methodology
  • BEiT-32022 · cited 2×
    “Rather than appending a task layer to the vision encoder [13, 3], we formulate the task as an image-to-text retrieval task.”
    From BEiT-3 · §Experiments on Vision and Vision-Language Tasks
  • EVA2022 · cited 7×
    “However, the most competitive billion-sized vision pre-trained models dosovitskiy2020vit; zhai2022scalingvit; pham2021combined; swinv2 still heavily rely on supervised or weakly-supervised training with hundreds of milli…”
    From EVA · §Introduction
Abstract

While the Transformer architecture has become the de-facto standard for natural language processing tasks, its applications to computer vision remain limited. In vision, attention is either applied in conjunction with convolutional networks, or used to replace certain components of convolutional networks while keeping their overall structure in place. We show that this reliance on CNNs is not necessary and a pure transformer applied directly to sequences of image patches can perform very well on image classification tasks. When pre-trained on large amounts of data and transferred to multiple mid-sized or small image recognition benchmarks (ImageNet, CIFAR-100, VTAB, etc.), Vision Transformer (ViT) attains excellent results compared to state-of-the-art convolutional networks while requiring substantially fewer computational resources to train.