Detecting Visual Relationships Using Box Attention
We propose a new model for detecting visual relationships, such as "person riding motorcycle" or "bottle on table". This task is an important step towards comprehensive structured image understanding, going beyond detecting individual objects.
Also cited · not yet reviewed (2)
- ResNet2015 · cited 3דIn our experiments we use the RetinaNet [19] model with ResNet50 [11] backbone as the base detection model.”From this paper · §Experiments
- The Open Images Dataset V42018 · cited 2דWe evaluate our approach on the V-COCO [10], Visual Relationships [20] and Open Images [16] datasets in Section 4.”From this paper · §Introduction
Led to
- The Open Images Dataset V42018 · cited 4×, 3 in Method“Recently, high-performing models based on deep convolutional neural networks are dominating the field Gupta and Malik 2015; Gkioxari et al. 2018; Gao et al. 2018; Kolesnikov et al. 2018.”From The Open Images Dataset V4 · §Performance of baseline models
Abstract
We propose a new model for detecting visual relationships, such as "person riding motorcycle" or "bottle on table". This task is an important step towards comprehensive structured image understanding, going beyond detecting individual objects. Our main novelty is a Box Attention mechanism that allows to model pairwise interactions between objects using standard object detection pipelines. The resulting model is conceptually clean, expressive and relies on well-justified training and prediction procedures. Moreover, unlike previously proposed approaches, our model does not introduce any additional complex components or hyperparameters on top of those already required by the underlying detection model. We conduct an experimental evaluation on three challenging datasets, V-COCO, Visual Relationships and Open Images, demonstrating strong quantitative and qualitative results.