objectiveIntroduced by XLNet · 2019
Permutation language modelling
Autoregressive training over all factorisation orders, capturing bidirectional context without masks.
Drafted by AI · not yet reviewed
How this idea evolved
Suggested from citations. No curator has recorded what this idea builds on yet. These ideas from the same theme come from papers that XLNet cites, directly or one step removed. Citation is a fact; the connection between the ideas is not verified.
Hide some tokens and predict them from context on both sides.
Use all internal layers of a pretrained bidirectional LM as features for downstream tasks.
Word vectors that depend on the sentence, taken from a pretrained encoder.
Papers using this
- 2018Set Transformer
- 2019RoBERTa
- 2019ALBERT
- 2019BART
- 2021Perceiver