Paper Lineage
Esc
MethodJun 2021arXiv 2106.04560cs.CV

Scaling Vision Transformers

Xiaohua Zhai, Alexander Kolesnikov, Neil Houlsby, Lucas Beyer

Attention-based neural networks such as the Vision Transformer (ViT) have recently attained state-of-the-art results on many computer vision benchmarks. Scale is a primary ingredient in attaining excellent results, therefore, understanding a model's scaling properties is a key to designing future generations effectively.

From the abstract

Built on

9 papers · 0 verifiedSee as graph

Also cited · not yet reviewed (9)

  • ViT2020 · cited 10×, 4 in Method
    “Attention-based Transformer architectures vaswani2017attention have taken computer vision domain by storm dosovitskiy2020; carion2020endtoend and are becoming an increasingly popular choice in research and practice.”
    From this paper · §Introduction
  • BiT2019 · cited 5×, 2 in Method
    “For ReaL, ViT-G/14 outperforms ViT-H dosovitskiy2020 and BiT-L kolesnikov2019big by only a small margin, which indicates again that the ImageNet classification task is likely reaching its saturation point.”
    From this paper · §Core Results
  • Instagram hashtag pre-training2018 · cited 2×, 1 in Method
    “Inspired by instagram, we address this issue by exploring learning-rate schedules that, similar to the warmup phase in the beginning, include a cooldown phase at the end of training, where the learning-rate is linearly a…”
    From this paper · §Method details
  • JFT-300M (unreasonable effectiveness)2017 · cited 1×, 1 in Method
    “For this study, we use the proprietary JFT-3B dataset, a larger version of the JFT-300M dataset used in many previous works on large-scale computer vision models sun2017unreasonable; kolesnikov2019big; dosovitskiy2020.”
    From this paper · §Method details
Show 5 more
  • Set Transformer2018 · cited 1×, 1 in Method
    “In particular, we evaluate global average pooling (GAP) and multihead attention pooling (MAP) lee2019set to aggregate representation from all patch tokens.”
    From this paper · §Method details
  • Kaplan scaling laws2020 · cited 3×
    “Optimal scaling of Transformers in NLP was carefully studied in kaplan2020scaling, with the main conclusion that large models not only perform better, but do use large computational budgets more efficiently.”
    From this paper · §Introduction
  • GPT-32020 · cited 3×
    “The few-shot transfer evaluation protocol has also been adopted by previous large-scale pre-training efforts in NLP domain gpt3.”
    From this paper · §Introduction
  • ImageNetV22019 · cited 2×
    “In addition to ImageNet fine-tuning and linear 10-shot results on the public validation set, we also report results of the ImageNet fine-tuned model on the ImageNet-v2 test set recht2019imagenet as an indicator of robust…”
    From this paper · §Core Results
  • Noisy Student2019 · cited 2×
    “On ImageNet-v2, ViT-G/14 improves 3% over the Noisy Student model xie2019selftraining based on EfficientNet-L2.”
    From this paper · §Core Results

Led to

  • CoAtNet2021 · cited 2×
    “Finally, when JFT-3B is used for pre-training, CoAtNet exhibits better efficiency compared to ViT, and pushes the ImageNet-1K top-1 accuracy to 90.88% while using 1.5x less computation of the prior art set by ViT-G/14 [2…”
    From CoAtNet · §Introduction
  • BEiT2021 · cited 2×
    “The results suggest that BEiT tends to help more for extremely larger models (such as 1B, or 10B), especially when labeled data are insufficient22 2 [47] report that supervised pre-training of a 1.8B-size vision Transfor…”
    From BEiT · §Experiments
  • LiT2021 · cited 8×
    “With the pre-trained model ViT-g/14 vitg, LiT achieves 85.2% zero-shot transfer accuracy on ImageNet, halving the gap between previous best zero-shot transfer results clip; align and supervised fine-tuning results vitg;…”
    From LiT · §Introduction
  • BEiT-32022 · cited 2×, 1 in Method
    “BEiT-3 is a giant-size foundation model following the setup of ViT-giant [63].”
    From BEiT-3 · §BEiT-3: A General-Purpose Multimodal Foundation Model
  • EVA2022 · cited 5×
    “However, the most competitive billion-sized vision pre-trained models dosovitskiy2020vit; zhai2022scalingvit; pham2021combined; swinv2 still heavily rely on supervised or weakly-supervised training with hundreds of milli…”
    From EVA · §Introduction
Abstract

Attention-based neural networks such as the Vision Transformer (ViT) have recently attained state-of-the-art results on many computer vision benchmarks. Scale is a primary ingredient in attaining excellent results, therefore, understanding a model's scaling properties is a key to designing future generations effectively. While the laws for scaling Transformer language models have been studied, it is unknown how Vision Transformers scale. To address this, we scale ViT models and data, both up and down, and characterize the relationships between error rate, data, and compute. Along the way, we refine the architecture and training of ViT, reducing memory consumption and increasing accuracy of the resulting models. As a result, we successfully train a ViT model with two billion parameters, which attains a new state-of-the-art on ImageNet of 90.45% top-1 accuracy. The model also performs well for few-shot transfer, for example, reaching 84.86% top-1 accuracy on ImageNet with only 10 examples per class.