Paper Lineage
Esc
MethodFeb 2016arXiv 1602.02410cs.CL

Exploring the Limits of Language Modeling

Rafal Jozefowicz, Oriol Vinyals, Mike Schuster and 2 others

In this work we explore recent advances in Recurrent Neural Networks for large scale Language Modeling, a task central to language understanding. We extend current models to deal with two key challenges present in this task: corpora and vocabulary sizes, and complex, long term structure of language.

From the abstract

Built on

0 papers · 0 verifiedSee as graph

Limits of Language Modeling has no earlier papers in this dataset.

Led to

  • TagLM2017 · cited 5×, 1 in Method
    “Recent state of the art neural language models (Józefowicz et al. 2016) use a similar architecture to our baseline sequence tagger where they pass a token representation (either from a CNN over characters or as token emb…”
    From TagLM · §Language model augmented sequence taggers (TagLM)
  • ELMo2018 · cited 3×, 3 in Method
    “Recent state-of-the-art neural language models (Józefowicz et al. 2016; Melis et al. 2017; Merity et al. 2017) compute a context-independent token representation 𝐱kL​Mx^{LM}_{k} (via token embeddings or a CNN over chara…”
    From ELMo · §ELMo: Embeddings from Language Models
  • T52019 · cited 2×
    “This is a natural fit for neural networks, which have been shown to exhibit remarkable scalability, i.e. it is often possible to achieve better performance simply by training a larger model on a larger data set (Hestness…”
    From T5 · §Introduction
  • METER2021 · cited 1×, 1 in Method
    “The model is trained to maximize its probability similar to noise contrastive estimation nce1; nce2.”
    From METER · §The Meter Framework
Abstract

In this work we explore recent advances in Recurrent Neural Networks for large scale Language Modeling, a task central to language understanding. We extend current models to deal with two key challenges present in this task: corpora and vocabulary sizes, and complex, long term structure of language. We perform an exhaustive study on techniques such as character Convolutional Neural Networks or Long-Short Term Memory, on the One Billion Word Benchmark. Our best single model significantly improves state-of-the-art perplexity from 51.3 down to 30.0 (whilst reducing the number of parameters by a factor of 20), while an ensemble of models sets a new record by improving perplexity from 41.0 down to 23.7. We also release these models for the NLP and ML community to study and improve upon.