Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer
Transfer learning, where a model is first pre-trained on a data-rich task before being fine-tuned on a downstream task, has emerged as a powerful technique in natural language processing (NLP). The effectiveness of transfer learning has given rise to a diversity of approaches, methodology, and practice.
Also cited · not yet reviewed (18)
- Transformer2017 · cited 6×, 5 in Method“Early results on transfer learning for NLP leveraged recurrent neural networks (Peters et al. 2018; Howard and Ruder 2018), but it has recently become more common to use models based on the “Transformer” architecture (Va…”From this paper · §Setup
- BERT2018 · cited 19×, 3 in Method“The Transformer was initially shown to be effective for machine translation, but it has subsequently been used in a wide variety of NLP settings (Radford et al. 2018; Devlin et al. 2018; McCann et al. 2018; Yu et al. 201…”From this paper · §Setup
- RoBERTa2019 · cited 10×, 2 in Method“However, we opted to create a new data set because prior data sets use a more limited set of filtering heuristics, are not publicly available, and/or are different in scope (e.g. are limited to News data (Zellers et al.…”From this paper · §Setup
- decaNLP2018 · cited 4×, 2 in Method“This approach is inspired by previous unifying frameworks for NLP tasks, including casting all text problems as question answering (McCann et al. 2018), language modeling (Radford et al. 2019), or span extraction Keskar…”From this paper · §Introduction
Show 14 more
- SQuAD2016 · cited 3×, 2 in Method“Natural language inference (MNLI (Williams et al. 2017), QNLI (Rajpurkar et al. 2016), RTE (Dagan et al. 2005), CB (De Marneff et al. 2019))”From this paper · §Setup
- XLNet2019 · cited 10×, 1 in Method“It has recently also become common to use models consisting of a single Transformer layer stack, with varying forms of self-attention used to produce architectures appropriate for language modeling (Radford et al. 2018;…”From this paper · §Setup
- ELMo2018 · cited 3×, 1 in Method“Early results on transfer learning for NLP leveraged recurrent neural networks (Peters et al. 2018; Howard and Ruder 2018), but it has recently become more common to use models based on the “Transformer” architecture (Va…”From this paper · §Setup
- Bahdanau attention2014 · cited 2×, 1 in Method“Self-attention is a variant of attention (Graves 2013; Bahdanau et al. 2015) that processes a sequence by replacing each element by a weighted average of the rest of the sequence.”From this paper · §Setup
- Seq2Seq2014 · cited 2×, 1 in Method“The original Transformer consisted of an encoder-decoder architecture and was intended for sequence-to-sequence (Sutskever et al. 2014; Kalchbrenner et al. 2014) tasks.”From this paper · §Setup
- ResNet2015 · cited 1×, 1 in Method“After layer normalization, a residual skip connection (He et al. 2016) adds each subcomponent’s input to its output.”From this paper · §Setup
- ReCoRD2018 · cited 1×, 1 in Method“Question answering (MultiRC (Khashabi et al. 2018), ReCoRD (Zhang et al. 2018), BoolQ (Clark et al. 2019))”From this paper · §Setup
- BoolQ2019 · cited 1×, 1 in Method“Question answering (MultiRC (Khashabi et al. 2018), ReCoRD (Zhang et al. 2018), BoolQ (Clark et al. 2019))”From this paper · §Setup
- ALBERT2019 · cited 6דRecent results suggest that this may hold true for transfer learning in NLP (Liu et al. 2019c; Radford et al. 2019; Yang et al. 2019; Lan et al. 2019), i.e. it has repeatedly been shown that scaling up produces improved…”From this paper · §Experiments
- UniLM2019 · cited 5דThis approach has recently been used to obtain state-of-the-art results in many of the most common NLP benchmarks (Devlin et al. 2018; Yang et al. 2019; Dong et al. 2019; Liu et al. 2019c; Lan et al. 2019).”From this paper · §Introduction
- Adapters2019 · cited 3דThis synergy has resulted in a great deal of recent work developing transfer learning methodology for NLP, which has produced a wide landscape of pre-training objectives (Howard and Ruder 2018; Devlin et al. 2018; Yang e…”From this paper · §Introduction
- Limits of Language Modeling2016 · cited 2דThis is a natural fit for neural networks, which have been shown to exhibit remarkable scalability, i.e. it is often possible to achieve better performance simply by training a larger model on a larger data set (Hestness…”From this paper · §Introduction
- Instagram hashtag pre-training2018 · cited 2דThis is a natural fit for neural networks, which have been shown to exhibit remarkable scalability, i.e. it is often possible to achieve better performance simply by training a larger model on a larger data set (Hestness…”From this paper · §Introduction
- GPipe2018 · cited 2דThis is a natural fit for neural networks, which have been shown to exhibit remarkable scalability, i.e. it is often possible to achieve better performance simply by training a larger model on a larger data set (Hestness…”From this paper · §Introduction
Led to
- 12-in-12019 · cited 2דAdvances in multi-task learning have been developed in the context of vision zhang2013robust; zhang2014facial; misra2016cross; kokkinos2017ubernet; strezoski2019many; bragman2019stochastic, language collobert2008unified;…”From 12-in-1 · §Related Work
- Knowledge in LM parameters2020 · cited 9×, 3 in Method“Big, deep neural language models that have been pre-trained on unlabeled text have proven to be extremely performant when fine-tuned on downstream Natural Language Processing (NLP) tasks Devlin et al. 2018; Yang et al. 2…”From Knowledge in LM parameters · §Introduction
- GLU variants2020 · cited 10דFollowing the T5 codebase [Raffel et al. 2019] 11 1 Also in the interest of ML fairness., we use a version with no bias:”From GLU variants · §Introduction
- GPT-32020 · cited 10×, 3 in Method“On tasks that involve binary classification, we give the options more semantically meaningful names (e.g. “True” or “False” rather than 0 or 1) and then treat the task like multiple choice; we also sometimes frame the ta…”From GPT-3 · §Approach
- GShard2020 · cited 2דWithout loss of generality, emerging approaches follow scaling Transformer by stacking more and more layers [49, 15], widening the governing dimensions of the network (i.e. model dimension, hidden dimension or number of…”From GShard · §Massively Multilingual, Massive Machine Translation (M4)
- Learning to summarize from human feedbac2020 · cited 2דFinally, there has been extensive research on modifying architectures [22, 59] and pre-training procedures [70, 36, 49, 60, 53, 14] for improving summarization performance.”From Learning to summarize from human feedbac · §unknown section
- MMLU2020 · cited 2דUnifiedQA uses the T5 (Raffel et al. 2019) text-to-text backbone and is fine-tuned on previously proposed question answering datasets (Lai et al. 2017), where the prediction is the class with the highest token overlap wi…”From MMLU · §Experiments
- The Pile2020 · cited 5דMore recently, language models (Radford et al. 2018; Radford et al. 2019; Brown et al. 2020; Rosset 2019; Shoeybi et al. 2019) and masked language models (Devlin et al. 2019; Liu et al. 2019; Raffel et al. 2019) have bee…”From The Pile · §Related Work
- Prefix-Tuning2021 · cited 2דFor table-to-text generation, Kale 2020 fine-tunes a sequence-to-sequence model (Raffel et al. 2020, T5;).”From Prefix-Tuning · §Related Work
- Switch Transformer2021 · cited 12×, 2 in Method“An approach followed in Radford et al. 2018; Raffel et al. 2019; Brown et al. 2020 expands the model size of a densely-activated Transformer (Vaswani et al. 2017).”From Switch Transformer · §Introduction
- VL-T52021 · cited 6×, 3 in Method“We introduce VL-T5 and VL-BART based on two pretrained transformer language models: T5Base (Raffel et al. 2020) and BARTBase (Lewis et al. 2020).”From VL-T5 · §Model
- Conceptual 12M2021 · cited 2דOn one hand, this is due to advances in architectures and modeling that are mainly inspired by BERT and similar models in natural language understanding and generation [25, 53, 82, 46, 26, 66].”From Conceptual 12M · §Introduction
- CLIP2021 · cited 2דPre-training methods which learn directly from raw text have revolutionized NLP over the last few years (Dai & Le 2015; Peters et al. 2018; Howard & Ruder 2018; Radford et al. 2018; Devlin et al. 2018; Raffel et al. 2019…”From CLIP · §Introduction and Motivating Work
- Swin2021 · cited 1×, 1 in Method“In computing self-attention, we follow [49, 1, 32, 33] by including a relative position bias B∈ℝM2×M2B\in\mathbb{R}^{M^{2}\times M^{2}} to each head in computing similarity:”From Swin · §Method
- Prompt tuning2021 · cited 6דFollowing the “text-to-text” approach of T5 Raffel et al. 2020, we cast all tasks as text generation.”From Prompt tuning · §Prompt Tuning
- RoFormer (RoPE)2021 · cited 4×, 3 in Method“Later work Raffel et al. 2020; He et al. 2020; Ke et al. 2020; Huang et al. 2020 followed these settings by only encoding the relative position information into the attention weights.”From RoFormer (RoPE) · §Background and Related Work
- CoAtNet2021 · cited 3×, 1 in Method“Interestingly, while the idea seems overly simplified, the pre-normalization version yprey^{pre} corresponds to a particular variant of relative self-attention [30, 31].”From CoAtNet · §Model
- BEiT2021 · cited 1×, 1 in Method“Moreover, blockwise (or n-gram) masking is also widely applied in BERT-like models [19, 2, 35].”From BEiT · §Methods
- Frozen2021 · cited 2×, 1 in Method“We use a 7 billion parameter transformer trained on the public dataset C4 [30] – previous work has shown that the multi-billion parameter scale is sufficient to exhibit the key capacities we are interested in studying [2…”From Frozen · §The Frozen Method
- Codex2021 · cited 2דFollowing the success of large natural language models (Devlin et al. 2018; Radford et al. 2019; Liu et al. 2019; Raffel et al. 2020; Brown et al. 2020) large scale Transformers have also been applied towards program syn…”From Codex · §Related Work
- SimVLM2021 · cited 4דSelf-supervised textual representation learning (Devlin et al. 2018; Radford et al. 2018; Radford et al. 2019; Liu et al. 2019; Yang et al. 2019; Raffel et al. 2019; Brown et al. 2020) based on Transformers (Vaswani et a…”From SimVLM · §Introduction
- FLAN2021 · cited 3דTo balance the different sizes of datasets, we limit the number of training examples per dataset to 30k and follow the examples-proportional mixing scheme (Raffel et al. 2020) with a mixing rate maximum of 3k.22 2 In thi…”From FLAN · §FLAN: Instruction Tuning Improves Zero-Shot Learning
- Recursive book summarization2021 · cited 2דWe also evaluate our models on the recently proposed BookSum dataset for book-length summarization (Kryściński et al., 2021) We compare to the best extractive (BertExt; Liu and Lapata, 2019b) and abstractive (T5; Raffel…”From Recursive book summarization · §unknown section
- T02021 · cited 6×, 3 in Method“All models we trained are based on T5, a Transformer-based encoder-decoder language model pretrained with a masked language modeling-style objective on 1T tokens from C4 (Raffel et al. 2020).”From T0 · §Experimental Setup
- LiT2021 · cited 2דWe consider four possible transformer-based text models transformer—the transformer from ViT-B vit which also resembles that used in CLIP clip, T5-base T5, mT5-base mt5, and the classic BERT-base bert—and whether to init…”From LiT · §Experiments
- Swin V22021 · cited 3דIt significantly improves a model’s performance on language tasks devlin2018bert; radford2019language; raffel2019t5; Turing-17B; fedus2021switch; Megatron-Turing-530B and the model demonstrates amazing few-shot capabilit…”From Swin V2 · §Introduction
- Gopher2021 · cited 2×, 2 in Method“The largest sampling subset is our curated web-text corpus MassiveWeb, which we find to improve downstream performance relative to existing web-text datasets such as C4 (Raffel et al. 2020b) in Figure A5.”From Gopher · §Method
- Fairseq MoE LMs2021 · cited 4×, 2 in Method“This is in contrast to the traditional approach of augmenting LMs with task-specific heads, followed by supervised fine-tuning (Devlin et al. 2019; Raffel et al. 2020).”From Fairseq MoE LMs · §Background and Related Work
- LaMDA2022 · cited 2×, 1 in Method“The Transformer has 64 layers, dmodel=8192d_{model}=8192, dff=65536d_{ff}=65536, h=128h=128, dk=dv=128d_{k}=d_{v}=128, relative attention as described in T5 [11], and gated-GELU activation as described in Raffel et…”From LaMDA · §LaMDA pre-training
- Megatron-Turing NLG2022 · cited 2דPrior work after GPT-2 produced dense transformer models at 8 billion [63], 11 billion [54], and 17 billion [4] parameters, and GPT-3 [9] at 175 billion parameters demonstrated for the first time that language models at…”From Megatron-Turing NLG · §Related Works
- PaLM2022 · cited 7דOn SuperGLUE, we compare with state-of-the-art models such as T5-11B (Raffel et al. 2020) and ST-MoE-32B (Zoph et al. 2022) and show that PaLM obtains competitive close-to-SOTA performance.”From PaLM · §Evaluation
- GPT-NeoX-20B2022 · cited 2דOver the past several years, there has been an explosion in research surrounding large language models (LLMs) for natural language processing, catalyzed largely by the impressive performance of Transformer-based language…”From GPT-NeoX-20B · §Introduction
- Super-NaturalInstructions2022 · cited 4דWe conduct our experiments and analysis based on the T5 model Raffel et al. 2020.”From Super-NaturalInstructions · §Tkk-Instruct: Learning to Follow Instructions at Scale
- CoCa2022 · cited 2דHowever, these models rely heavily on image annotations as labeled vectors and do not bake in knowledge of free-form human natural language, hindering their application to downstream tasks that involving both vision and…”From CoCa · §Introduction
- U-PaLM2022 · cited 6דThe majority of this prior work, however, requires additional data such as aggregating dozens or hundreds of NLP datasets (Raffel et al. 2019; Aghajanyan et al. 2021; Aribandi et al. 2022), writing additional templates o…”From U-PaLM · §Related Work
- Flan-T5 / Flan-PaLM2022 · cited 3דThese checkpoints have strong zero-shot, few-shot, and CoT abilities, outperforming prior public checkpoints such as T5 (Raffel et al. 2020).”From Flan-T5 / Flan-PaLM · §unknown section
- EVA2022 · cited 2דScaling up pre-trained language models (PLMs) liu2019roberta; gpt3; t5 has revolutionized natural language processing (NLP) in the past few years.”From EVA · §Introduction
Abstract
Transfer learning, where a model is first pre-trained on a data-rich task before being fine-tuned on a downstream task, has emerged as a powerful technique in natural language processing (NLP). The effectiveness of transfer learning has given rise to a diversity of approaches, methodology, and practice. In this paper, we explore the landscape of transfer learning techniques for NLP by introducing a unified framework that converts all text-based language problems into a text-to-text format. Our systematic study compares pre-training objectives, architectures, unlabeled data sets, transfer approaches, and other factors on dozens of language understanding tasks. By combining the insights from our exploration with scale and our new ``Colossal Clean Crawled Corpus'', we achieve state-of-the-art results on many benchmarks covering summarization, question answering, text classification, and more. To facilitate future work on transfer learning for NLP, we release our data set, pre-trained models, and code.