Paper Lineage
Esc
MethodApr 2015arXiv 1504.06692cs.CV

Learning like a Child: Fast Novel Visual Concept Learning from Sentence Descriptions of Images

Junhua Mao, Wei Xu, Yi Yang and 3 others

In this paper, we address the task of learning novel visual concepts, and their interactions with other concepts, from a few images with sentence descriptions. Using linguistic context and visual features, our method is able to efficiently hypothesize the semantic meaning of new words and add them to its word dictionary so that they can be used to describe images which contain these novel concepts.

From the abstract

Built on

3 papers · 0 verifiedSee as graph

Also cited · not yet reviewed (3)

  • VGG2014 · cited 2×, 1 in Method
    “The vision component contains a 16-layer deep convolutional neural network (CNN [48]) pre-trained on the ImageNet classification task [45].”
    From this paper · §The Image Captioning Model
  • “Very recent works of image captioning includes [39, 25, 24, 55, 10, 15, 7, 32, 37, 26, 57, 35].”
    From this paper · §Related Work
  • MS COCO2014 · cited 2×
    “The first two datasets are derived from the MS-COCO dataset [34].”
    From this paper · §Introduction

Led to

  • FM-IQA2015 · cited 5×, 3 in Method
    “Similar to [25], we use the sigmoid function as the activation function of the three gates and adopt ReLU [30] as the non-linear function for the LSTM memory cells.”
    From FM-IQA · §The Multimodal QA (mQA) Model
Abstract

In this paper, we address the task of learning novel visual concepts, and their interactions with other concepts, from a few images with sentence descriptions. Using linguistic context and visual features, our method is able to efficiently hypothesize the semantic meaning of new words and add them to its word dictionary so that they can be used to describe images which contain these novel concepts. Our method has an image captioning module based on m-RNN with several improvements. In particular, we propose a transposed weight sharing scheme, which not only improves performance on image captioning, but also makes the model more suitable for the novel concept learning task. We propose methods to prevent overfitting the new concepts. In addition, three novel concept datasets are constructed for this new task. In the experiments, we show that our method effectively learns novel visual concepts from a few examples without disturbing the previously learned concepts. The project page is http://www.stat.ucla.edu/~junhua.mao/projects/child_learning.html