Paper Lineage
Esc
MethodNov 2021arXiv 2111.09883cs.CV

Swin Transformer V2: Scaling Up Capacity and Resolution

Ze Liu, Han Hu, Yutong Lin and 9 others

Large-scale NLP models have been shown to significantly improve the performance on language tasks with no signs of saturation. They also demonstrate amazing few-shot capabilities like that of human beings.

From the abstract

Built on

13 papers · 0 verifiedSee as graph

Also cited · not yet reviewed (13)

  • AdamW2017 · cited 13×
    “Following liu2021swin, we employ an AdamW loshchilov2017decoupled optimizer for 300 epochs using a cosine decay learning rate scheduler with 20 epochs of linear warm-up.”
    From this paper · §A1 Experimental Settings for Ablation
  • Swin2021 · cited 10×
    “The current common practice is to perform a bi-cubic interpolation of the position bias maps dosovitskiy2020vit; liu2021swin.”
    From this paper · §Introduction
  • ViT2020 · cited 5×
    “The current common practice is to perform a bi-cubic interpolation of the position bias maps dosovitskiy2020vit; liu2021swin.”
    From this paper · §Introduction
  • GPT-32020 · cited 4×
    “It significantly improves a model’s performance on language tasks devlin2018bert; radford2019language; raffel2019t5; Turing-17B; fedus2021switch; Megatron-Turing-530B and the model demonstrates amazing few-shot capabilit…”
    From this paper · §Introduction
Show 9 more
  • CoAtNet2021 · cited 4×
    “While it has long been recognized that larger vision models usually perform better on vision tasks simonyan2014vgg; he2015resnet, the absolute model size was just able to reach about 1-2 billion parameters very recently…”
    From this paper · §Introduction
  • Transformer2017 · cited 3×
    “Transformer has served the standard network since the pioneer work of vaswani2017attention.”
    From this paper · §Related Works
  • BERT2018 · cited 3×
    “It significantly improves a model’s performance on language tasks devlin2018bert; radford2019language; raffel2019t5; Turing-17B; fedus2021switch; Megatron-Turing-530B and the model demonstrates amazing few-shot capabilit…”
    From this paper · §Introduction
  • T52019 · cited 3×
    “It significantly improves a model’s performance on language tasks devlin2018bert; radford2019language; raffel2019t5; Turing-17B; fedus2021switch; Megatron-Turing-530B and the model demonstrates amazing few-shot capabilit…”
    From this paper · §Introduction
  • Switch Transformer2021 · cited 3×
    “It significantly improves a model’s performance on language tasks devlin2018bert; radford2019language; raffel2019t5; Turing-17B; fedus2021switch; Megatron-Turing-530B and the model demonstrates amazing few-shot capabilit…”
    From this paper · §Introduction
  • BEiT2021 · cited 3×
    “Specifically, it obtains 84.0% top-1 accuracy on the ImageNet-V2 image classification validation set recht2019imagenet, 63.1 / 54.4 box / mask AP on the COCO test-dev set of object detection, 59.9 mIoU on ADE20K semantic…”
    From this paper · §Introduction
  • MS COCO2014 · cited 2×
    “We conduct experiments on ImageNet-1K image classification (V1 and V2) deng2009imagenet; recht2019imagenet, COCO object detection lin2014coco, and ADE20K semantic segmentation zhou2018semantic.”
    From this paper · §Experiments
  • VGG2014 · cited 2×
    “While it has long been recognized that larger vision models usually perform better on vision tasks simonyan2014vgg; he2015resnet, the absolute model size was just able to reach about 1-2 billion parameters very recently…”
    From this paper · §Introduction
  • BiT2019 · cited 2×
    “While it has long been recognized that larger vision models usually perform better on vision tasks simonyan2014vgg; he2015resnet, the absolute model size was just able to reach about 1-2 billion parameters very recently…”
    From this paper · §Introduction

Led to

  • EVA2022 · cited 7×
    “However, the most competitive billion-sized vision pre-trained models dosovitskiy2020vit; zhai2022scalingvit; pham2021combined; swinv2 still heavily rely on supervised or weakly-supervised training with hundreds of milli…”
    From EVA · §Introduction
Abstract

Large-scale NLP models have been shown to significantly improve the performance on language tasks with no signs of saturation. They also demonstrate amazing few-shot capabilities like that of human beings. This paper aims to explore large-scale models in computer vision. We tackle three major issues in training and application of large vision models, including training instability, resolution gaps between pre-training and fine-tuning, and hunger on labelled data. Three main techniques are proposed: 1) a residual-post-norm method combined with cosine attention to improve training stability; 2) A log-spaced continuous position bias method to effectively transfer models pre-trained using low-resolution images to downstream tasks with high-resolution inputs; 3) A self-supervised pre-training method, SimMIM, to reduce the needs of vast labeled images. Through these techniques, this paper successfully trained a 3 billion-parameter Swin Transformer V2 model, which is the largest dense vision model to date, and makes it capable of training with images of up to 1,536$\times$1,536 resolution. It set new performance records on 4 representative vision tasks, including ImageNet-V2 image classification, COCO object detection, ADE20K semantic segmentation, and Kinetics-400 video action classification. Also note our training is much more efficient than that in Google's billion-level visual models, which consumes 40 times less labelled data and 40 times less training time. Code is available at \url{https://github.com/microsoft/Swin-Transformer}.