Learned in Translation: Contextualized Word Vectors
Computer vision has benefited from initializing multiple deep layers with weights pretrained on large supervised training sets like ImageNet. Natural language processing (NLP) typically sees initialization of only the lowest layer of deep models with pretrained word vectors.
Learned in Translation has no earlier papers in this dataset.
Led to
- ELMo2018 · cited 7×, 1 in Method“Unlike previous approaches for learning contextualized word vectors (Peters et al. 2017; McCann et al. 2017), ELMo representations are deep, in the sense that they are a function of all of the internal layers of the biLM…”From ELMo · §Introduction
- Megatron-LM2019 · cited 1×, 1 in Method“Later work advanced research in this area by learning and transferring neural models that capture contextual representations of words (Melamud et al. 2016; McCann et al. 2017; Peters et al. 2018; Radford et al. 2017; Rad…”From Megatron-LM · §Background and Challenges
- CLIP2021 · cited 1×, 1 in Method“Although early work wrestled with the complexity of natural language when using topic model and n-gram representations, improvements in deep contextual representation learning suggest we now have the tools to effectively…”From CLIP · §Approach
Abstract
Computer vision has benefited from initializing multiple deep layers with weights pretrained on large supervised training sets like ImageNet. Natural language processing (NLP) typically sees initialization of only the lowest layer of deep models with pretrained word vectors. In this paper, we use a deep LSTM encoder from an attentional sequence-to-sequence model trained for machine translation (MT) to contextualize word vectors. We show that adding these context vectors (CoVe) improves performance over using only unsupervised word and character vectors on a wide variety of common NLP tasks: sentiment analysis (SST, IMDb), question classification (TREC), entailment (SNLI), and question answering (SQuAD). For fine-grained sentiment analysis and entailment, CoVe improves performance of our baseline models to the state of the art.