objectiveIntroduced by UniLM · 2019
Unified LM via attention masks
One Transformer does bidirectional, left-to-right and seq2seq modelling by switching attention masks.
Drafted by AI · not yet reviewed
How this idea evolved
Suggested from citations. No curator has recorded what this idea builds on yet. These ideas from the same theme come from papers that UniLM cites, directly or one step removed. Citation is a fact; the connection between the ideas is not verified.
Hide some tokens and predict them from context on both sides.
Use all internal layers of a pretrained bidirectional LM as features for downstream tasks.
Word vectors that depend on the sentence, taken from a pretrained encoder.
Papers using this
- 2019Unified VLP
- 2019BART
- 2021VLMo
- 2022BEiT-3
- 2023BLIP-2