Paper Lineage
Esc
MethodDec 2018arXiv 1812.00500cs.CV

Multi-task Learning of Hierarchical Vision-Language Representation

Duy-Kien Nguyen, Takayuki Okatani

It is still challenging to build an AI system that can perform tasks that involve vision and language at human level. So far, researchers have singled out individual tasks separately, for each of which they have designed networks and trained them on its dedicated datasets.

From the abstract

Built on

5 papers · 0 verifiedSee as graph

Also cited · not yet reviewed (5)

  • Flickr30k Entities2015 · cited 9×
    “Recently, studies of multi-modal tasks of vision and language have made significant progress, such as image captioning Anderson_2018_CVPR; Lu_2018_CVPR, visual question answering Nguyen_2018_CVPR; Teney_2018_CVPR, visual…”
    From this paper · §Related Work
  • VQA v22016 · cited 4×
    “Since the recent successes of deep learning on single modality tasks, multi-modal tasks lying on the intersection of vision and language, such as image captioning mscoco; Young_2014_TACL, visual question answering (VQA)…”
    From this paper · §Introduction
  • MS COCO2014 · cited 2×
    “Since the recent successes of deep learning on single modality tasks, multi-modal tasks lying on the intersection of vision and language, such as image captioning mscoco; Young_2014_TACL, visual question answering (VQA)…”
    From this paper · §Introduction
  • “Following the standard procedure Karpathy_2015_CVPR, we use the 1,000 val images and the 1,000 or 5,000 test images, which are selected from the original 40,504 val images.”
    From this paper · §Experiments
Show 1 more
  • VQA2015 · cited 2×
    “Since the recent successes of deep learning on single modality tasks, multi-modal tasks lying on the intersection of vision and language, such as image captioning mscoco; Young_2014_TACL, visual question answering (VQA)…”
    From this paper · §Introduction

Led to

  • 12-in-12019 · cited 6×, 1 in Method
    “Different from previous observation mccann2018natural; nguyen2019multi, we found that using no curriculum leads to superior performance when combined with other strategies proposed in this section.”
    From 12-in-1 · §Approach
Abstract

It is still challenging to build an AI system that can perform tasks that involve vision and language at human level. So far, researchers have singled out individual tasks separately, for each of which they have designed networks and trained them on its dedicated datasets. Although this approach has seen a certain degree of success, it comes with difficulties of understanding relations among different tasks and transferring the knowledge learned for a task to others. We propose a multi-task learning approach that enables to learn vision-language representation that is shared by many tasks from their diverse datasets. The representation is hierarchical, and prediction for each task is computed from the representation at its corresponding level of the hierarchy. We show through experiments that our method consistently outperforms previous single-task-learning methods on image caption retrieval, visual question answering, and visual grounding. We also analyze the learned hierarchical representation by visualizing attention maps generated in our network.