Swin Transformer: Hierarchical Vision Transformer using Shifted Windows
This paper presents a new vision Transformer, called Swin Transformer, that capably serves as a general-purpose backbone for computer vision. Challenges in adapting Transformer from language to vision arise from differences between the two domains, such as large variations in the scale of visual entities and the high resolution of pixels in images compared to words in text.
Also cited · not yet reviewed (8)
- ViT2020 · cited 8×, 3 in Method“Most related to our work is the Vision Transformer (ViT) [20] and its follow-ups [63, 72, 15, 28, 66].”From this paper · §Related Work
- DeiT2020 · cited 8×, 1 in Method“Most related to our work is the Vision Transformer (ViT) [20] and its follow-ups [63, 72, 15, 28, 66].”From this paper · §Related Work
- ResNet2015 · cited 4×, 1 in Method“These stages jointly produce a hierarchical representation, with the same feature map resolutions as those of typical convolutional networks, e.g., VGG [52] and ResNet [30].”From this paper · §Method
- Transformer2017 · cited 3×, 1 in Method“The standard Transformer architecture [64] and its adaptation for image classification [20] both conduct global self-attention, where the relationships between a token and all other tokens are computed.”From this paper · §Method
Show 4 more
- VGG2014 · cited 2×, 1 in Method“These stages jointly produce a hierarchical representation, with the same feature map resolutions as those of typical convolutional networks, e.g., VGG [52] and ResNet [30].”From this paper · §Method
- T52019 · cited 1×, 1 in Method“In computing self-attention, we follow [49, 1, 32, 33] by including a relative position bias B∈ℝM2×M2B\in\mathbb{R}^{M^{2}\times M^{2}} to each head in computing similarity:”From this paper · §Method
- ResNeXt2016 · cited 3דIn addition to these architectural advances, there has also been much work on improving individual convolution layers, such as depth-wise convolution [70] and deformable convolution [18, 84].”From this paper · §Related Work
- EfficientNet2019 · cited 3דSince then, deeper and more effective convolutional neural architectures have been proposed to further propel the deep learning wave in computer vision, e.g., VGG [52], GoogleNet [57], ResNet [30], DenseNet [34], HRNet […”From this paper · §Related Work
Led to
- CoAtNet2021 · cited 6×, 2 in Method“Enforce local attention, which restricts the global receptive field 𝒢G in attention to a local field ℒL just like in convolution [22, 21].”From CoAtNet · §Model
- Dynamic Head2021 · cited 3דWhen training our dynamic head with latest transformer backbone [19], extra data and increased input size, we can further improve the current SOTA on COCO benchmark.”From Dynamic Head · §Appendix
- Video Swin2021 · cited 10דSwin Transformer [28] further introduces the inductive biases of locality, hierarchy and translation invariance, which enable it to serve as a general-purpose backbone for various image recognition tasks.”From Video Swin · §Related Works
- METER2021 · cited 5×, 3 in Method“Transformers vaswani2017attention are prevalent in natural language processing and have recently shown promising performance in computer vision dosovitskiy2020image; liu2021swin.”From METER · §Introduction
- Swin V22021 · cited 10דThe current common practice is to perform a bi-cubic interpolation of the position bias maps dosovitskiy2020vit; liu2021swin.”From Swin V2 · §Introduction
- Florence2021 · cited 3×, 2 in Method“For the image encoder, we chose hierarchical Vision Transformers (e.g. , Swin (Liu et al. 2021a), CvT (Wu et al. 2021), Vision Longformer (Zhang et al. 2021a), Focal Transformer (Yang et al. 2021), and CSwin (Dong et al.…”From Florence · §Introduction
- UniCL2022 · cited 2דWith this goal, numerous works have pushed the image recognition performance from different directions, such as data scale from MNIST lecun1989handwritten to ImageNet-1K deng2009imagenet, model architectures from convolu…”From UniCL · §Related works
Abstract
This paper presents a new vision Transformer, called Swin Transformer, that capably serves as a general-purpose backbone for computer vision. Challenges in adapting Transformer from language to vision arise from differences between the two domains, such as large variations in the scale of visual entities and the high resolution of pixels in images compared to words in text. To address these differences, we propose a hierarchical Transformer whose representation is computed with \textbf{S}hifted \textbf{win}dows. The shifted windowing scheme brings greater efficiency by limiting self-attention computation to non-overlapping local windows while also allowing for cross-window connection. This hierarchical architecture has the flexibility to model at various scales and has linear computational complexity with respect to image size. These qualities of Swin Transformer make it compatible with a broad range of vision tasks, including image classification (87.3 top-1 accuracy on ImageNet-1K) and dense prediction tasks such as object detection (58.7 box AP and 51.1 mask AP on COCO test-dev) and semantic segmentation (53.5 mIoU on ADE20K val). Its performance surpasses the previous state-of-the-art by a large margin of +2.7 box AP and +2.6 mask AP on COCO, and +3.2 mIoU on ADE20K, demonstrating the potential of Transformer-based models as vision backbones. The hierarchical design and the shifted window approach also prove beneficial for all-MLP architectures. The code and models are publicly available at~\url{https://github.com/microsoft/Swin-Transformer}.