Language model pre-training
From word vectors to BERT-style and seq2seq pre-training objectives for text.
| Concept | Introduced by | Years | Papers |
|---|---|---|---|
| Word embeddings (skip-gram / CBOW) Learn dense word vectors by predicting nearby words in huge text corpora. | word2vec | 2013–2021 | 8 |
| Large-scale RNN language models Push LSTM language models to billions of words and huge vocabularies. | Limits of Language Modeling | 2016–2019 | 1 |
| Contextual word vectors Word vectors that depend on the sentence, taken from a pretrained encoder. | Learned in Translation | 2017–2018 | 1 |
| Deep bidirectional LM features Use all internal layers of a pretrained bidirectional LM as features for downstream tasks. | ELMo | 2018–2019 | 9 |
| Masked language modelling Hide some tokens and predict them from context on both sides. | BERT | 2018–2020 | 5 |
| Pre-train, then fine-tune Train one big model on unlabelled data, then fine-tune it briefly for each task. | — | 2018–2022 | 43 |
| Cross-layer parameter sharing Share weights across layers and factorise embeddings to shrink BERT. | ALBERT | 2019–2020 | 3 |
| Denoising seq2seq pre-training Corrupt text in many ways and train an encoder–decoder to reconstruct it. | BART | 2019–2021 | 4 |
| Permutation language modelling Autoregressive training over all factorisation orders, capturing bidirectional context without masks. | XLNet | 2019–2021 | 5 |
| Robustly optimised BERT recipe Train BERT longer, on more data, with bigger batches and dynamic masking. | RoBERTa | 2019–2022 | 10 |
| Span-corruption denoising Replace random spans with sentinels and train the decoder to fill them in. | T5 | 2019–2022 | 10 |
| Unified LM via attention masks One Transformer does bidirectional, left-to-right and seq2seq modelling by switching attention masks. | UniLM | 2019–2023 | 5 |
| Mixture-of-denoisers objective Train on a mix of short-span, long-span and prefix denoising tasks. | U-PaLM | 2022 | 1 |