Paper Lineage
Esc
MethodDec 2016arXiv 1612.09161cs.CV

Learning Visual N-Grams from Web Data

Ang Li, Allan Jabri, Armand Joulin, Laurens van der Maaten

Real-world image recognition systems need to recognize tens of thousands of classes that constitute a plethora of visual concepts. The traditional approach of annotating thousands of images per class for training is infeasible in such a scenario, prompting the use of webly supervised data.

From the abstract

Built on

1 paper · 0 verifiedSee as graph

Also cited · not yet reviewed (1)

  • “Relating Images and Captions: Additional Results As an addition to the image and caption retrieval results on COCO-5K and Flickr-30K presented in the paper, we also provide retrieval results on the COCO-1K dataset, a tes…”
    From this paper · §Introduction

Led to

  • ALIGN2021 · cited 2×
    “Joulin et al. 2015; Li et al. 2017; Desai & Johnson 2020; Sariyildiz et al. 2020; Zhang et al. 2020 show that a good visual representation can be learned by predicting the captions from images, which inspires our work.”
    From ALIGN · §Related Work
  • CLIP2021 · cited 4×
    “Li et al. 2017 then extended this approach to predicting phrase n-grams in addition to individual words and demonstrated the ability of their system to zero-shot transfer to other image classification datasets by scoring…”
    From CLIP · §Introduction and Motivating Work
Abstract

Real-world image recognition systems need to recognize tens of thousands of classes that constitute a plethora of visual concepts. The traditional approach of annotating thousands of images per class for training is infeasible in such a scenario, prompting the use of webly supervised data. This paper explores the training of image-recognition systems on large numbers of images and associated user comments. In particular, we develop visual n-gram models that can predict arbitrary phrases that are relevant to the content of an image. Our visual n-gram models are feed-forward convolutional networks trained using new loss functions that are inspired by n-gram models commonly used in language modeling. We demonstrate the merits of our models in phrase prediction, phrase-based image retrieval, relating images and captions, and zero-shot transfer.