Paper Lineage
Esc
MethodNov 2015arXiv 1511.02251cs.CV

Learning Visual Features from Large Weakly Supervised Data

Armand Joulin, Laurens van der Maaten, Allan Jabri, Nicolas Vasilache

Convolutional networks trained on large supervised dataset produce visual features which form the basis for the state-of-the-art in many computer-vision problems. Further improvements of these visual features will likely require even larger manually labeled data sets, which severely limits the pace at which progress can be made.

From the abstract

Built on

4 papers · 0 verifiedSee as graph

Also cited · not yet reviewed (4)

  • word2vec2013 · cited 6×
    “We follow Mikolov et al. mikolov13 and sample instances uniformly per class.”
    From this paper · §Weakly Supervised Learning of Convnets
  • “Recent studies have shown that using visual features extracted from convolutional networks trained on large object recognition datasets krizhevsky12; simonyan15; szegedy15 can lead to state-of-the-art results on many vis…”
    From this paper · §Introduction
  • GoogLeNet (Inception)2014 · cited 3×
    “Recent studies have shown that using visual features extracted from convolutional networks trained on large object recognition datasets krizhevsky12; simonyan15; szegedy15 can lead to state-of-the-art results on many vis…”
    From this paper · §Introduction
  • YFCC100M2015 · cited 2×
    “This type of data is available in great abundance via photo-sharing websites: specifically, we use a publicly available dataset of 100 million Flickr images and captions thomee15 (see Figure 1 for six randomly picked Fli…”
    From this paper · §Introduction

Led to

  • “In order to overcome the bottleneck, there have been recent efforts on visual representation learning using web-supervision [5, 6, 9, 21, 2, 23, 27, 24] or unsupervised [34, 10, 11, 43, 31, 32, 42] paradigms.”
    From JFT-300M (unreasonable effectiveness) · §Related Work
  • Instagram hashtag pre-training2018 · cited 7×, 1 in Method
    “Motivated by this work, we perform experiments in which we evaluate three different types of data sampling in the Instagram pretraining: (1) a natural sampling in which we sample images and hashtags according to the dist…”
    From Instagram hashtag pre-training · §Experiments
  • ALIGN2021 · cited 2×
    “Joulin et al. 2015; Li et al. 2017; Desai & Johnson 2020; Sariyildiz et al. 2020; Zhang et al. 2020 show that a good visual representation can be learned by predicting the captions from images, which inspires our work.”
    From ALIGN · §Related Work
  • CLIP2021 · cited 2×, 1 in Method
    “Zhang et al. 2020, Gomez et al. 2017, Joulin et al. 2016, and Desai & Johnson 2020 all introduce methods which learn visual representations from text paired with images but describe their approaches as unsupervised, self…”
    From CLIP · §Approach
Abstract

Convolutional networks trained on large supervised dataset produce visual features which form the basis for the state-of-the-art in many computer-vision problems. Further improvements of these visual features will likely require even larger manually labeled data sets, which severely limits the pace at which progress can be made. In this paper, we explore the potential of leveraging massive, weakly-labeled image collections for learning good visual features. We train convolutional networks on a dataset of 100 million Flickr photos and captions, and show that these networks produce features that perform well in a range of vision problems. We also show that the networks appropriately capture word similarity, and learn correspondences between different languages.