training techniqueIntroduced by RoBERTa · 2019
Robustly optimised BERT recipe
Train BERT longer, on more data, with bigger batches and dynamic masking.
Drafted by AI · not yet reviewed
How this idea evolved
Suggested from citations. No curator has recorded what this idea builds on yet. These ideas from the same theme come from papers that RoBERTa cites, directly or one step removed. Citation is a fact; the connection between the ideas is not verified.
Hide some tokens and predict them from context on both sides.
Autoregressive training over all factorisation orders, capturing bidirectional context without masks.
Use all internal layers of a pretrained bidirectional LM as features for downstream tasks.
Papers using this
- 2019Unicoder-VL
- 2019Unified VLP
- 2019FreeLB
- 2019ALBERT
- 2019BART
- 2021METER
- 2021Florence
- 2021Fairseq MoE LMs
- 2022Simple end-to-end captioning
- 2022OPT