LXMERT: Learning Cross-Modality Encoder Representations from Transformers
Vision-and-language reasoning requires an understanding of visual concepts, language semantics, and, most importantly, the alignment and relationships between these two modalities. We thus propose the LXMERT (Learning Cross-Modality Encoder Representations from Transformers) framework to learn these vision-and-language connections.
Also cited · not yet reviewed (13)
- BERT2018 · cited 10×, 7 in Method“For the cross-modality output, following the practice in Devlin et al. 2019, we append a special token [CLS] (denoted as the top yellow block in the bottom branch of Fig. 1) before the sentence words, and the correspondi…”From this paper · §Model Architecture
- Bottom-Up Top-Down attention2017 · cited 7×, 4 in Method“Instead of using the feature map output by a convolutional neural network, we follow Anderson et al. 2018 in taking the features of detected objects as the embeddings of images.”From this paper · §Model Architecture
- Transformer2017 · cited 6×, 4 in Method“We build our cross-modality model with self-attention and cross-attention layers following the recent progress in designing natural language processing models (e.g., transformers Vaswani et al. 2017).”From this paper · §Model Architecture
- Bilinear Attention Networks2018 · cited 3×, 2 in Method“The SotA result is BAN+Counter in Kim et al. 2018, which achieves the best accuracy among other recent works: MFH Yu et al. 2018, Pythia Jiang et al. 2018, DFAF Gao et al. 2019a, and Cycle-Consistency Shah et al. 2019.55…”From this paper · §Experimental Setup and Results
Show 9 more
- NLVR22018 · cited 3×, 2 in Method“NLVR2NLVR^{2} Suhr et al. 2019 is a challenging visual reasoning dataset where some existing approaches Hu et al. 2017; Perez et al. 2018 fail, and the SotA method is ‘MaxEnt’ in Suhr et al. 2019.”From this paper · §Experimental Setup and Results
- Faster R-CNN2015 · cited 2×, 2 in Method“For these reasons, we take detected labels output by Faster R-CNN Ren et al. 2015.”From this paper · §Pre-Training Strategies
- GNMT2016 · cited 2×, 2 in Method“A sentence is first split into words {w1,…,wn}\left\{w_{1},\ldots,w_{n}\right\} with length of nn by the same WordPiece tokenizer Wu et al. 2016 in Devlin et al. 2019.”From this paper · §Model Architecture
- MS COCO2014 · cited 2×, 1 in Method“As shown in Table. 1, we aggregate pre-training data from five vision-and-language datasets whose images come from MS COCO Lin et al. 2014 or Visual Genome Krishna et al. 2017.”From this paper · §Pre-Training Strategies
- VQA2015 · cited 2×, 1 in Method“Besides the two original captioning datasets, we also aggregate three large image question answering (image QA) datasets: VQA v2.0 Antol et al. 2015, GQA balanced version Hudson and Manning 2019, and VG-QA Zhu et al. 201…”From this paper · §Pre-Training Strategies
- Visual Genome2016 · cited 2×, 1 in Method“As shown in Table. 1, we aggregate pre-training data from five vision-and-language datasets whose images come from MS COCO Lin et al. 2014 or Visual Genome Krishna et al. 2017.”From this paper · §Pre-Training Strategies
- Bahdanau attention2014 · cited 1×, 1 in Method“Attention layers Bahdanau et al. 2014; Xu et al. 2015 aim to retrieve information from a set of context vectors {yj}\{y_{j}\} related to a query vector xx.”From this paper · §Model Architecture
- VQA v22016 · cited 1×, 1 in Method“We use three datasets for evaluating our LXMERT framework: VQA v2.0 dataset Goyal et al. 2017, GQA Hudson and Manning 2019, and NLVR2NLVR^{2}.”From this paper · §Experimental Setup and Results
- ELMo2018 · cited 2דIn terms of language understanding, last year, we witnessed strong progress towards building a universal backbone model with large-scale contextualized language model pre-training Peters et al. 2018; Radford et al. 2018;…”From this paper · §Introduction
Led to
- VL-BERT2019 · cited 3דIn ViLBERT (Lu et al. 2019) and LXMERT (Tan & Bansal 2019), which are under review or just got accepted, the network architectures are of two single-modal networks applied on input sentences and images respectively, foll…”From VL-BERT · §Related Work
- Unified VLP2019 · cited 6×, 1 in Method“Inspired by the recent success of pre-trained language models such as BERT [\citeauthoryearDevlin et al.2018] and GPT [\citeauthoryearRadford et al.2018, \citeauthoryearRadford et al.2019], there is a growing interest in…”From Unified VLP · §Introduction
- 12-in-12019 · cited 5×, 1 in Method“The recent rise of general architectures for vision-and-language lu2019vilbert; tan2019lxmert; li2019visualbert; alberti2019fusion; li2019unicoder; su2019vl; zhou2019unified reduces the architectural differences across t…”From 12-in-1 · §Introduction
- VILLA2020 · cited 3דInspired by the success of BERT [13] on natural language understanding, there has been a surging research interest in developing multimodal pre-training methods for vision-and-language representation learning (e.g., ViLB…”From VILLA · §Introduction
- ConVIRT2020 · cited 2דContrastive-Binary-Loss: This baseline differs from ConVIRT by contrasting the paired image and text representations with a binary classification head, as is widely done in visual-linguistic pretraining work Tan and Bans…”From ConVIRT · §Experiments
- VL-BERT meta-analysis2020 · cited 3×, 1 in Method“Each model is initialized with the parameters of BERT, following the approaches described in the original papers.88 8 Only Tan and Bansal 2019 reported slightly better performance when pretraining from scratch but they r…”From VL-BERT meta-analysis · §Experimental Setup
- VL-T52021 · cited 9×, 4 in Method“As shown in Fig.3 (a), existing methods (Tan & Bansal 2019; Lu et al. 2019; Chen et al. 2020) typically introduce a multi-layer perceptron (MLP) multi-label classifier head on top of h[CLS]xh^{x}_{\texttt{[CLS]}}, which…”From VL-T5 · §Model
- ViLT2021 · cited 4×, 2 in Method“Backbone: ResNet-101 (Lu et al. 2019; Tan & Bansal 2019; Su et al. 2019) and ResNext-152 (Li et al. 2019; Li et al. 2020a; Zhang et al. 2021) are two commonly used backbones.”From ViLT · §Background
- Conceptual 12M2021 · cited 5×, 1 in Method“To train the model’s parameters, we use a contrastive softmax loss, for which the original image-text pairs are used as positive examples, while all other image-text pairs in the mini-batch are used as negative examples…”From Conceptual 12M · §Evaluating Vision-and-Language Pre-Training Data
- ALBEF2021 · cited 3דThe first category focuses on modelling the interactions between image and text features with transformer-based multimodal encoders [10, 11, 12, 13, 1, 14, 15, 2, 3, 16, 8, 17, 18].”From ALBEF · §Related Work
- SimVLM2021 · cited 5דWhile a variety of approaches have been proposed, a large portion of them require object detection for image region feature regression or tagging as part of the pre-training objectives (Tan & Bansal 2019; Su et al. 2020;…”From SimVLM · §Related Work
- VLMo2021 · cited 3דPre-training with Transformer [45] backbone networks has substantially advanced the state of the art across natural language processing [34, 10, 28, 22, 11, 36, 1, 7, 8, 4, 5, 6, 31], computer vision [12, 44, 2] and visi…”From VLMo · §Related Work
- METER2021 · cited 7×, 5 in Method“Vision-and-language pre-training (VLP) has now become the de facto practice to tackle these tasks tan-bansal-2019-lxmert; li2019visualbert; lu2019vilbert; su2019vl; chen2020uniter; li2020oscar.”From METER · §Introduction
- VL-BEiT2022 · cited 4דVision-language pretraining [37, 24, 35, 46, 29, 21, 16, 20, 42, 41, 40, 1, 45] aims to learn multimodal representations from large-scale image-text pairs.”From VL-BEiT · §Related Work
Abstract
Vision-and-language reasoning requires an understanding of visual concepts, language semantics, and, most importantly, the alignment and relationships between these two modalities. We thus propose the LXMERT (Learning Cross-Modality Encoder Representations from Transformers) framework to learn these vision-and-language connections. In LXMERT, we build a large-scale Transformer model that consists of three encoders: an object relationship encoder, a language encoder, and a cross-modality encoder. Next, to endow our model with the capability of connecting vision and language semantics, we pre-train the model with large amounts of image-and-sentence pairs, via five diverse representative pre-training tasks: masked language modeling, masked object prediction (feature regression and label classification), cross-modality matching, and image question answering. These tasks help in learning both intra-modality and cross-modality relationships. After fine-tuning from our pre-trained parameters, our model achieves the state-of-the-art results on two visual question answering datasets (i.e., VQA and GQA). We also show the generalizability of our pre-trained cross-modality model by adapting it to a challenging visual-reasoning task, NLVR2, and improve the previous best result by 22% absolute (54% to 76%). Lastly, we demonstrate detailed ablation studies to prove that both our novel model components and pre-training strategies significantly contribute to our strong results; and also present several attention visualizations for the different encoders. Code and pre-trained models publicly available at: https://github.com/airsplay/lxmert