Exploring the Limits of Language Modeling
In this work we explore recent advances in Recurrent Neural Networks for large scale Language Modeling, a task central to language understanding. We extend current models to deal with two key challenges present in this task: corpora and vocabulary sizes, and complex, long term structure of language.
Limits of Language Modeling has no earlier papers in this dataset.
Led to
- TagLM2017 · cited 5×, 1 in Method“Recent state of the art neural language models (Józefowicz et al. 2016) use a similar architecture to our baseline sequence tagger where they pass a token representation (either from a CNN over characters or as token emb…”From TagLM · §Language model augmented sequence taggers (TagLM)
- ELMo2018 · cited 3×, 3 in Method“Recent state-of-the-art neural language models (Józefowicz et al. 2016; Melis et al. 2017; Merity et al. 2017) compute a context-independent token representation 𝐱kLMx^{LM}_{k} (via token embeddings or a CNN over chara…”From ELMo · §ELMo: Embeddings from Language Models
- T52019 · cited 2דThis is a natural fit for neural networks, which have been shown to exhibit remarkable scalability, i.e. it is often possible to achieve better performance simply by training a larger model on a larger data set (Hestness…”From T5 · §Introduction
- METER2021 · cited 1×, 1 in Method“The model is trained to maximize its probability similar to noise contrastive estimation nce1; nce2.”From METER · §The Meter Framework
Abstract
In this work we explore recent advances in Recurrent Neural Networks for large scale Language Modeling, a task central to language understanding. We extend current models to deal with two key challenges present in this task: corpora and vocabulary sizes, and complex, long term structure of language. We perform an exhaustive study on techniques such as character Convolutional Neural Networks or Long-Short Term Memory, on the One Billion Word Benchmark. Our best single model significantly improves state-of-the-art perplexity from 51.3 down to 30.0 (whilst reducing the number of parameters by a factor of 20), while an ensemble of models sets a new record by improving perplexity from 41.0 down to 23.7. We also release these models for the NLP and ML community to study and improve upon.