BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and Comprehension
We present BART, a denoising autoencoder for pretraining sequence-to-sequence models. BART is trained by (1) corrupting text with an arbitrary noising function, and (2) learning a model to reconstruct the original text.
Also cited · not yet reviewed (7)
- XLNet2019 · cited 7×, 4 in Method“Based on XLNet (Yang et al. 2019), we sample 1/6 of the tokens, and generate them in a random order autoregressively.”From this paper · §Comparing Pre-training Objectives
- BERT2018 · cited 6×, 4 in Method“Following BERT Devlin et al. 2019, random tokens are sampled and replaced with [MASK] elements.”From this paper · §Model
- RoBERTa2019 · cited 7×, 3 in Method“Recent work has shown that downstream performance can dramatically improve when pre-training is scaled to large batch sizes (Yang et al. 2019; Liu et al. 2019) and corpora.”From this paper · §Large-scale Pre-training Experiments
- Transformer2017 · cited 2×, 2 in Method“BART uses the standard sequence-to-sequence Transformer architecture from Vaswani et al. 2017, except, following GPT, that we modify ReLU activation functions to GeLUs (Hendrycks & Gimpel 2016) and initialise parameters…”From this paper · §Model
Show 3 more
- UniLM2019 · cited 3×, 1 in Method“As in UniLM (Dong et al. 2019), we train a Masked Language Model with additional self-attention masks.”From this paper · §Comparing Pre-training Objectives
- SQuAD2016 · cited 2×, 1 in Method“Rajpurkar et al. 2016a an extractive question answering task on Wikipedia paragraphs.”From this paper · §Comparing Pre-training Objectives
- ELMo2018 · cited 2דELMo Peters et al. 2018 concatenates left-only and right-only representations, but does not pre-train interactions between these features.”From this paper · §Related Work
Led to
- Prefix-Tuning2021 · cited 3×, 1 in Method“For summarization, we compare against fine-tuning BART Lewis et al. 2020.”From Prefix-Tuning · §Experimental Setup
- VL-T52021 · cited 4×, 3 in Method“We introduce VL-T5 and VL-BART based on two pretrained transformer language models: T5Base (Raffel et al. 2020) and BARTBase (Lewis et al. 2020).”From VL-T5 · §Model
- Prompt tuning2021 · cited 1×, 1 in Method“Their work builds on GPT-2 Radford et al. 2019 and BART Lewis et al. 2020, while ours focuses on T5 and examines changes in performance and robustness to design choices as model size increases.”From Prompt tuning · §Comparison to Similar Approaches
- Natural Instructions2021 · cited 3×, 2 in Method“We use BART (base) Lewis et al. 2019 which allows us to fine-tune its model parameters.”From Natural Instructions · §Problem Setup and Models
- Fairseq MoE LMs2021 · cited 1×, 1 in Method“Models are pretrained by hiding parts of the input: predicting the next word sequentially left-to-right, masking words in the text (Devlin et al. 2019; Liu et al. 2019), or perturbing and/or masking spans (Lewis et al. 2…”From Fairseq MoE LMs · §Background and Related Work
Abstract
We present BART, a denoising autoencoder for pretraining sequence-to-sequence models. BART is trained by (1) corrupting text with an arbitrary noising function, and (2) learning a model to reconstruct the original text. It uses a standard Tranformer-based neural machine translation architecture which, despite its simplicity, can be seen as generalizing BERT (due to the bidirectional encoder), GPT (with the left-to-right decoder), and many other more recent pretraining schemes. We evaluate a number of noising approaches, finding the best performance by both randomly shuffling the order of the original sentences and using a novel in-filling scheme, where spans of text are replaced with a single mask token. BART is particularly effective when fine tuned for text generation but also works well for comprehension tasks. It matches the performance of RoBERTa with comparable training resources on GLUE and SQuAD, achieves new state-of-the-art results on a range of abstractive dialogue, question answering, and summarization tasks, with gains of up to 6 ROUGE. BART also provides a 1.1 BLEU increase over a back-translation system for machine translation, with only target language pretraining. We also report ablation experiments that replicate other pretraining schemes within the BART framework, to better measure which factors most influence end-task performance.