Learning Visual N-Grams from Web Data
Real-world image recognition systems need to recognize tens of thousands of classes that constitute a plethora of visual concepts. The traditional approach of annotating thousands of images per class for training is infeasible in such a scenario, prompting the use of webly supervised data.
Also cited · not yet reviewed (1)
- Karpathy visual-semantic alignment2014 · cited 5דRelating Images and Captions: Additional Results As an addition to the image and caption retrieval results on COCO-5K and Flickr-30K presented in the paper, we also provide retrieval results on the COCO-1K dataset, a tes…”From this paper · §Introduction
Led to
- ALIGN2021 · cited 2דJoulin et al. 2015; Li et al. 2017; Desai & Johnson 2020; Sariyildiz et al. 2020; Zhang et al. 2020 show that a good visual representation can be learned by predicting the captions from images, which inspires our work.”From ALIGN · §Related Work
- CLIP2021 · cited 4דLi et al. 2017 then extended this approach to predicting phrase n-grams in addition to individual words and demonstrated the ability of their system to zero-shot transfer to other image classification datasets by scoring…”From CLIP · §Introduction and Motivating Work
Abstract
Real-world image recognition systems need to recognize tens of thousands of classes that constitute a plethora of visual concepts. The traditional approach of annotating thousands of images per class for training is infeasible in such a scenario, prompting the use of webly supervised data. This paper explores the training of image-recognition systems on large numbers of images and associated user comments. In particular, we develop visual n-gram models that can predict arbitrary phrases that are relevant to the content of an image. Our visual n-gram models are feed-forward convolutional networks trained using new loss functions that are inspired by n-gram models commonly used in language modeling. We demonstrate the merits of our models in phrase prediction, phrase-based image retrieval, relating images and captions, and zero-shot transfer.