Paper Lineage
Esc

Language model pre-training

From word vectors to BERT-style and seq2seq pre-training objectives for text.

Concepts in Language model pre-training
ConceptIntroduced byYearsPapers
Word embeddings (skip-gram / CBOW)
Learn dense word vectors by predicting nearby words in huge text corpora.
word2vec2013–20218
Large-scale RNN language models
Push LSTM language models to billions of words and huge vocabularies.
Limits of Language Modeling2016–20191
Contextual word vectors
Word vectors that depend on the sentence, taken from a pretrained encoder.
Learned in Translation2017–20181
Deep bidirectional LM features
Use all internal layers of a pretrained bidirectional LM as features for downstream tasks.
ELMo2018–20199
Masked language modelling
Hide some tokens and predict them from context on both sides.
BERT2018–20205
Pre-train, then fine-tune
Train one big model on unlabelled data, then fine-tune it briefly for each task.
—2018–202243
Cross-layer parameter sharing
Share weights across layers and factorise embeddings to shrink BERT.
ALBERT2019–20203
Denoising seq2seq pre-training
Corrupt text in many ways and train an encoder–decoder to reconstruct it.
BART2019–20214
Permutation language modelling
Autoregressive training over all factorisation orders, capturing bidirectional context without masks.
XLNet2019–20215
Robustly optimised BERT recipe
Train BERT longer, on more data, with bigger batches and dynamic masking.
RoBERTa2019–202210
Span-corruption denoising
Replace random spans with sentinels and train the decoder to fill them in.
T52019–202210
Unified LM via attention masks
One Transformer does bidirectional, left-to-right and seq2seq modelling by switching attention masks.
UniLM2019–20235
Mixture-of-denoisers objective
Train on a mix of short-span, long-span and prefix denoising tasks.
U-PaLM20221