An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
While the Transformer architecture has become the de-facto standard for natural language processing tasks, its applications to computer vision remain limited. In vision, attention is either applied in conjunction with convolutional networks, or used to replace certain components of convolutional networks while keeping their overall structure in place.
Also cited · not yet reviewed (8)
- Transformer2017 · cited 4×, 2 in Method“In model design we follow the original Transformer (Vaswani et al. 2017) as closely as possible.”From this paper · §Method
- BiT2019 · cited 7×, 1 in Method“We de-duplicate the pre-training datasets w.r.t. the test sets of the downstream tasks following Kolesnikov et al. 2020.”From this paper · §Experiments
- BERT2018 · cited 4דHowever, much of their success stems not only from their excellent scalability but also from large scale self-supervised pre-training (Devlin et al. 2019; Radford et al. 2018).”From this paper · §Experiments
- Noisy Student2019 · cited 3דThe use of additional data sources allows to achieve state-of-the-art results on standard benchmarks (Mahajan et al. 2018; Touvron et al. 2019; Xie et al. 2020).”From this paper · §Related Work
Show 4 more
- ResNet2015 · cited 2דIn computer vision, however, convolutional architectures remain dominant (LeCun et al. 1989; Krizhevsky et al. 2012; He et al. 2016).”From this paper · §Introduction
- JFT-300M (unreasonable effectiveness)2017 · cited 2דMoreover, Sun et al. 2017 study how CNN performance scales with dataset size, and Kolesnikov et al. 2020; Djolonga et al. 2020 perform an empirical exploration of CNN transfer learning from large scale datasets such as I…”From this paper · §Related Work
- Instagram hashtag pre-training2018 · cited 2דThe use of additional data sources allows to achieve state-of-the-art results on standard benchmarks (Mahajan et al. 2018; Touvron et al. 2019; Xie et al. 2020).”From this paper · §Related Work
- GPT-32020 · cited 2דLarge Transformer-based models are often pre-trained on large corpora and then fine-tuned for the task at hand: BERT (Devlin et al. 2019) uses a denoising self-supervised pre-training task, while the GPT line of work use…”From this paper · §Related Work
Led to
- DeiT2020 · cited 12דWe build upon the visual transformer architecture from Dosovitskiy et al. [15] and improvements included in the timm library [55].”From DeiT · §Introduction
- ViLT2021 · cited 4×, 1 in Method“Recent work (Dosovitskiy et al. 2020; Touvron et al. 2020) demonstrated that using a simple linear projection of a patch is effective enough to embed pixels before feeding them into transformers.”From ViLT · §Introduction
- ALIGN2021 · cited 2דHigh-quality visual representations for classification or retrieval are usually pre-trained on large-scale labeled datasets (Mahajan et al. 2018; Kolesnikov et al. 2020; Dosovitskiy et al. 2021; Juan et al. 2020).”From ALIGN · §Related Work
- CLIP2021 · cited 4×, 1 in Method“For the second architecture, we experiment with the recently introduced Vision Transformer (ViT) (Dosovitskiy et al. 2020).”From CLIP · §Approach
- Perceiver2021 · cited 4דThis is the strategy taken by the Vision Transformer (ViT) (Dosovitskiy et al. 2021), which first reduces the input size to ∼200\sim 200 using a 2D convolutional layer (referred to as “linear projection of flattened patc…”From Perceiver · §Related Work
- Swin2021 · cited 8×, 3 in Method“Most related to our work is the Vision Transformer (ViT) [20] and its follow-ups [63, 72, 15, 28, 66].”From Swin · §Related Work
- Scaling ViTs (ViT-G)2021 · cited 10×, 4 in Method“Attention-based Transformer architectures vaswani2017attention have taken computer vision domain by storm dosovitskiy2020; carion2020endtoend and are becoming an increasingly popular choice in research and practice.”From Scaling ViTs (ViT-G) · §Introduction
- CoAtNet2021 · cited 4דIn comparison, after the success of ViT and ResNet-ViT [13], another popular line of research starts with a Transformer backbone and tries to incorporate explicit convolution or some desirable properties of convolution i…”From CoAtNet · §Related Work
- BEiT2021 · cited 7×, 3 in Method“Following ViT [12], we use the standard Transformer [42] as the backbone network.”From BEiT · §Methods
- Video Swin2021 · cited 3דA shift in backbone architectures for computer vision, from CNNs to Transformers, began recently with Vision Transformer (ViT) [8, 34].”From Video Swin · §Related Works
- ALBEF2021 · cited 2×, 2 in Method“We use a 12-layer visual transformer ViT-B/16 [38] as the image encoder, and initialize it with weights pre-trained on ImageNet-1k from [31].”From ALBEF · §ALBEF Pre-training
- SimVLM2021 · cited 4דWe follow the setup in ViT (Dosovitskiy et al. 2021) to explore 3 variants of SimVLM, namely “Base”, “Large”, and “Huge”, such that each variant follows the same setting as its corresponding ViT variant.”From SimVLM · §Experiments
- VLMo2021 · cited 5×, 1 in Method“Following vision Transformers [12, 44, 2], the 2D image 𝒗∈ℝH×W×C{\bm{v}}\in\mathbb{R}^{H\times W\times C} is split and reshaped into N=HW/P2N={HW}/{P^{2}} patches 𝒗p∈ℝN×(P2C){\bm{v}}^{p}\in\mathbb{R}^{N\times(P^{2}C)…”From VLMo · §Methods
- METER2021 · cited 4×, 3 in Method“Transformers vaswani2017attention are prevalent in natural language processing and have recently shown promising performance in computer vision dosovitskiy2020image; liu2021swin.”From METER · §Introduction
- MAE2021 · cited 12×, 2 in Method“Motivated by the success in NLP, related recent methods Chen2020c; Dosovitskiy2021; Bao2021 are based on Transformers Vaswani2017. iGPT Chen2020c operates on sequences of pixels and predicts unknown pixels.”From MAE · §Related Work
- LiT2021 · cited 5דWe verify LiT across three image-text datasets, with Vision Transformer vit, ResNet bit, and MLP-Mixer mixer architectures.”From LiT · §Introduction
- Swin V22021 · cited 5דThe current common practice is to perform a bi-cubic interpolation of the position bias maps dosovitskiy2020vit; liu2021swin.”From Swin V2 · §Introduction
- Florence2021 · cited 2דWhile inheriting performance benefits of the transformer self-attention operations (Dosovitskiy et al. 2021b), these hierarchical architectures model the scale invariance nature of images and have linear computational co…”From Florence · §Introduction
- ViTCAP2021 · cited 3דViTCAP is constructed on the basis of a vision transformer dosovitskiy2020image as the stem image encoder.”From ViTCAP · §Introduction
- BLIP2022 · cited 2×, 1 in Method“We employ a visual transformer (Dosovitskiy et al. 2021) as our image encoder, which divides an input image into patches and encodes them as a sequence of embeddings, with an additional [CLS] token to represent the globa…”From BLIP · §Method
- Simple end-to-end captioning2022 · cited 3×, 1 in Method“Visual Encoder.We follow the same setting as the visual transformer (ViT) which recently attracts much attention and shows remarkable performance in the computer vision area (Dosovitskiy et al. 2021; Radford et al. 2021;…”From Simple end-to-end captioning · §. Methodology
- CoCa2022 · cited 2×, 2 in Method“Following a standard encoder-decoder architecture, the image encoder provides latent encoded features (e.g., using a Vision Transformer [39] or ConvNets [40]) and the text decoder learns to maximize the conditional likel…”From CoCa · §Approach
- VL-BEiT2022 · cited 3×, 1 in Method“Following [9], we split the image 𝒗∈ℝH×W×C{\bm{v}}\in\mathbb{R}^{H\times W\times C} into a sequence of patches, so that the image can be encoded by standard Transformer.”From VL-BEiT · §Methods
- BEiT v22022 · cited 2×, 1 in Method“The vision Transformers (ViTs; Dosovitskiy et al. 2020) are employed as the backbone networks to obtain image representations.”From BEiT v2 · §Methodology
- BEiT-32022 · cited 2דRather than appending a task layer to the vision encoder [13, 3], we formulate the task as an image-to-text retrieval task.”From BEiT-3 · §Experiments on Vision and Vision-Language Tasks
- EVA2022 · cited 7דHowever, the most competitive billion-sized vision pre-trained models dosovitskiy2020vit; zhai2022scalingvit; pham2021combined; swinv2 still heavily rely on supervised or weakly-supervised training with hundreds of milli…”From EVA · §Introduction
Abstract
While the Transformer architecture has become the de-facto standard for natural language processing tasks, its applications to computer vision remain limited. In vision, attention is either applied in conjunction with convolutional networks, or used to replace certain components of convolutional networks while keeping their overall structure in place. We show that this reliance on CNNs is not necessary and a pure transformer applied directly to sequences of image patches can perform very well on image classification tasks. When pre-trained on large amounts of data and transferred to multiple mid-sized or small image recognition benchmarks (ImageNet, CIFAR-100, VTAB, etc.), Vision Transformer (ViT) attains excellent results compared to state-of-the-art convolutional networks while requiring substantially fewer computational resources to train.