Paper Lineage
Esc
MethodMay 2019arXiv 1905.00546cs.CV

Billion-scale semi-supervised learning for image classification

I. Zeki Yalniz, Hervé Jégou, Kan Chen and 2 others

This paper presents a study of semi-supervised learning with large convolutional networks. We propose a pipeline, based on a teacher/student paradigm, that leverages a large collection of unlabelled images (up to 1 billion).

From the abstract

Built on

5 papers · 0 verifiedSee as graph

Also cited · not yet reviewed (5)

  • “IG-1B-Targeted: Following [27], we collected a dataset of 1B public images with associated hashtags from a social media website.”
    From this paper · §Image classification: experiments & analysis
  • ResNet2015 · cited 4×
    “Models: For student and teacher models, we use residual networks [16], ResNet-d with dd = {18,50}\{18,50\} and residual networks with group convolutions [43], ResNeXt-101 32XCd with 101101 layers and group widths CC = {4…”
    From this paper · §Image classification: experiments & analysis
  • ResNeXt2016 · cited 3×
    “Models: For student and teacher models, we use residual networks [16], ResNet-d with dd = {18,50}\{18,50\} and residual networks with group convolutions [43], ResNeXt-101 32XCd with 101101 layers and group widths CC = {4…”
    From this paper · §Image classification: experiments & analysis
  • BatchNorm2015 · cited 2×
    “Each GPU processes 2424 images at a time and apply batch normalization [22] to all convolutional layers on each GPU.”
    From this paper · §Image classification: experiments & analysis
Show 1 more
  • Goyal large-batch SGD2017 · cited 2×
    “We set the learning rate following the linear scaling procedure proposed in [13] with a warm-up and overall minibatch size of 64×24=153664\times 24=1536.”
    From this paper · §Image classification: experiments & analysis

Led to

  • Noisy Student2019 · cited 3×
    “Other Techniques – Noisy Student Training also works better with an additional trick: data filtering and balancing, similar to uda; billion_large_scale.”
    From Noisy Student · §Noisy Student Training
  • BiT2019 · cited 4×
    “Rather than pre-train generic representations, recent works have shown strong performance by training task-specific representations [63, 38, 61].”
    From BiT · §Related Work
  • Natural distribution shift robustness2020 · cited 2×, 2 in Method
    “This subset includes models trained on (i) Facebook’s collection of 1 billion Instagram images [56, 104], (ii) the YFCC 100 million dataset [104], (iii) Google’s JFT 300 million dataset [82, 102], (iv) a subset of OpenIm…”
    From Natural distribution shift robustness · §Experimental setup
Abstract

This paper presents a study of semi-supervised learning with large convolutional networks. We propose a pipeline, based on a teacher/student paradigm, that leverages a large collection of unlabelled images (up to 1 billion). Our main goal is to improve the performance for a given target architecture, like ResNet-50 or ResNext. We provide an extensive analysis of the success factors of our approach, which leads us to formulate some recommendations to produce high-accuracy models for image classification with semi-supervised learning. As a result, our approach brings important gains to standard architectures for image, video and fine-grained classification. For instance, by leveraging one billion unlabelled images, our learned vanilla ResNet-50 achieves 81.2% top-1 accuracy on the ImageNet benchmark.