Paper Lineage
Esc
MethodDec 2019arXiv 1912.11370cs.CV

Big Transfer (BiT): General Visual Representation Learning

Alexander Kolesnikov, Lucas Beyer, Xiaohua Zhai and 4 others

Transfer of pre-trained representations improves sample efficiency and simplifies hyperparameter tuning when training deep neural networks for vision. We revisit the paradigm of pre-training on large supervised datasets and fine-tuning the model on a target task.

From the abstract

Built on

3 papers · 0 verifiedSee as graph

Also cited · not yet reviewed (3)

  • Domain adaptive transfer2018 · cited 5×
    “Rather than pre-train generic representations, recent works have shown strong performance by training task-specific representations [63, 38, 61].”
    From this paper · §Related Work
  • “Rather than pre-train generic representations, recent works have shown strong performance by training task-specific representations [63, 38, 61].”
    From this paper · §Related Work
  • Noisy Student2019 · cited 4×
    “Rather than pre-train generic representations, recent works have shown strong performance by training task-specific representations [63, 38, 61].”
    From this paper · §Related Work

Led to

  • Natural distribution shift robustness2020 · cited 1×, 1 in Method
    “This subset includes models trained on (i) Facebook’s collection of 1 billion Instagram images [56, 104], (ii) the YFCC 100 million dataset [104], (iii) Google’s JFT 300 million dataset [82, 102], (iv) a subset of OpenIm…”
    From Natural distribution shift robustness · §Experimental setup
  • ViT2020 · cited 7×, 1 in Method
    “We de-duplicate the pre-training datasets w.r.t. the test sets of the downstream tasks following Kolesnikov et al. 2020.”
    From ViT · §Experiments
  • ALIGN2021 · cited 5×, 1 in Method
    “Following Kolesnikov et al. 2020, we also evaluate the robustness of our model on Visual Task Adaptation Benchmark (VTAB) (Zhai et al. 2019) which consists of 19 diverse (covering subgroups of natural, specialized and st…”
    From ALIGN · §Pre-training and Task Transfer
  • CLIP2021 · cited 5×
    “Kolesnikov et al. 2019 and Dosovitskiy et al. 2020 have also demonstrated large gains on a broader set of transfer benchmarks by pre-training models to predict the classes of the noisily labeled JFT-300M dataset.”
    From CLIP · §Introduction and Motivating Work
  • Scaling ViTs (ViT-G)2021 · cited 5×, 2 in Method
    “For ReaL, ViT-G/14 outperforms ViT-H dosovitskiy2020 and BiT-L kolesnikov2019big by only a small margin, which indicates again that the ImageNet classification task is likely reaching its saturation point.”
    From Scaling ViTs (ViT-G) · §Core Results
  • LiT2021 · cited 6×
    “Transfer learning transfer_learning_survey has been a successful paradigm in computer vision imagenet_transfer_better; bit; instagram_resnext.”
    From LiT · §Introduction
  • Swin V22021 · cited 2×
    “While it has long been recognized that larger vision models usually perform better on vision tasks simonyan2014vgg; he2015resnet, the absolute model size was just able to reach about 1-2 billion parameters very recently…”
    From Swin V2 · §Introduction
Abstract

Transfer of pre-trained representations improves sample efficiency and simplifies hyperparameter tuning when training deep neural networks for vision. We revisit the paradigm of pre-training on large supervised datasets and fine-tuning the model on a target task. We scale up pre-training, and propose a simple recipe that we call Big Transfer (BiT). By combining a few carefully selected components, and transferring using a simple heuristic, we achieve strong performance on over 20 datasets. BiT performs well across a surprisingly wide range of data regimes -- from 1 example per class to 1M total examples. BiT achieves 87.5% top-1 accuracy on ILSVRC-2012, 99.4% on CIFAR-10, and 76.3% on the 19 task Visual Task Adaptation Benchmark (VTAB). On small datasets, BiT attains 76.8% on ILSVRC-2012 with 10 examples per class, and 97.0% on CIFAR-10 with 10 examples per class. We conduct detailed analysis of the main components that lead to high transfer performance.