Paper Lineage
Esc
MethodFeb 2018arXiv 1802.05365cs.CL

Deep contextualized word representations

Matthew E. Peters, Mark Neumann, Mohit Iyyer and 4 others

, to model polysemy). Our word vectors are learned functions of the internal states of a deep bidirectional language model (biLM), which is pre-trained on a large text corpus.

From the abstract

Built on

3 papers · 0 verifiedSee as graph

Also cited · not yet reviewed (3)

  • TagLM2017 · cited 9×, 3 in Method
    “Overall, this formulation is similar to the approach of Peters et al. 2017, with the exception that we share some weights between directions instead of using completely independent parameters.”
    From this paper · §ELMo: Embeddings from Language Models
  • Limits of Language Modeling2016 · cited 3×, 3 in Method
    “Recent state-of-the-art neural language models (Józefowicz et al. 2016; Melis et al. 2017; Merity et al. 2017) compute a context-independent token representation 𝐱kL​Mx^{LM}_{k} (via token embeddings or a CNN over chara…”
    From this paper · §ELMo: Embeddings from Language Models
  • Learned in Translation2017 · cited 7×, 1 in Method
    “Unlike previous approaches for learning contextualized word vectors (Peters et al. 2017; McCann et al. 2017), ELMo representations are deep, in the sense that they are a function of all of the internal layers of the biLM…”
    From this paper · §Introduction

Led to

  • BERT2018 · cited 7×
    “ELMo and its predecessor Peters et al. 2017; Peters et al. 2018a generalize traditional word embedding research along a different dimension.”
    From BERT · §Related Work
  • ReCoRD2018 · cited 2×
    “We also evaluate DocQA with ELMo Peters et al. 2018 to analyze the impact of largely pre-trained encoder on our dataset.”
    From ReCoRD · §Evaluation
  • nocaps2018 · cited 2×, 1 in Method
    “When using ELMo [34], we use a dynamic representation of wcw_{c}, h¯t1\bar{h}_{t}^{1} and h¯t2\bar{h}_{t}^{2} as the input word embedding wE​L​M​otw_{ELMo}^{t} for our caption model. wcw_{c} is the character embedding of…”
    From nocaps · §Additional Implementation Details for Baseline Models
  • Transformer-XL2019 · cited 2×, 1 in Method
    “Second, though it is possible to use padding to respect the sentence or other semantic boundaries, in practice it has been standard practice to simply chunk long text into fixed-length segments due to improved efficiency…”
    From Transformer-XL · §Model
  • UniLM2019 · cited 2×
    “Language model (LM) pre-training has substantially advanced the state of the art across a variety of natural language processing tasks [8, 29, 19, 31, 9, 1].”
    From UniLM · §Introduction
  • BoolQ2019 · cited 4×, 1 in Method
    “Our Recurrent +ELMo model uses the language model from Peters et al. 2018 to provide contextualized embeddings to the baseline model outlined above, as recommended by the authors.”
    From BoolQ · §Results
  • XLNet2019 · cited 2×
    “Unsupervised representation learning has been highly successful in the domain of natural language processing [7, 22, 27, 28, 10].”
    From XLNet · §Introduction
  • RoBERTa2019 · cited 2×
    “Pretraining methods have been designed with different training objectives, including language modeling Dai and Le 2015; Peters et al. 2018; Howard and Ruder 2018, machine translation McCann et al. 2017, and masked langua…”
    From RoBERTa · §Related Work
  • ViLBERT2019 · cited 3×
    “Self-supervised language models on the other hand have resulted in significant improvements over prior work [12, 14, 13, 44].”
    From ViLBERT · §Related Work
  • LXMERT2019 · cited 2×
    “In terms of language understanding, last year, we witnessed strong progress towards building a universal backbone model with large-scale contextualized language model pre-training Peters et al. 2018; Radford et al. 2018;…”
    From LXMERT · §Introduction
  • Megatron-LM2019 · cited 2×, 1 in Method
    “By finetuning these pretrained language models on downstream natural language tasks, one can achieve state of the art results as shown in recent work (Devlin et al. 2018; Peters et al. 2018; Howard & Ruder 2018; Radford…”
    From Megatron-LM · §Introduction
  • ALBERT2019 · cited 2×
    “Learning representations of natural language has been shown to be useful for a wide range of NLP tasks and has been widely adopted (Mikolov et al. 2013; Le & Mikolov 2014; Dai & Le 2015; Peters et al. 2018; Devlin et al.…”
    From ALBERT · §Related work
  • T52019 · cited 3×, 1 in Method
    “Early results on transfer learning for NLP leveraged recurrent neural networks (Peters et al. 2018; Howard and Ruder 2018), but it has recently become more common to use models based on the “Transformer” architecture (Va…”
    From T5 · §Setup
  • BART2019 · cited 2×
    “ELMo Peters et al. 2018 concatenates left-only and right-only representations, but does not pre-train interactions between these features.”
    From BART · §Related Work
  • Knowledge in LM parameters2020 · cited 1×, 1 in Method
    “The popularity of this form of “transfer learning” is attributable to its empirical success on many NLP tasks Peters et al. 2018; Devlin et al. 2018; Yang et al. 2019; Lan et al. 2019; Raffel et al. 2019.”
    From Knowledge in LM parameters · §Background
Abstract

We introduce a new type of deep contextualized word representation that models both (1) complex characteristics of word use (e.g., syntax and semantics), and (2) how these uses vary across linguistic contexts (i.e., to model polysemy). Our word vectors are learned functions of the internal states of a deep bidirectional language model (biLM), which is pre-trained on a large text corpus. We show that these representations can be easily added to existing models and significantly improve the state of the art across six challenging NLP problems, including question answering, textual entailment and sentiment analysis. We also present an analysis showing that exposing the deep internals of the pre-trained network is crucial, allowing downstream models to mix different types of semi-supervision signals.