Scaling Instruction-Finetuned Language Models
Finetuning language models on a collection of datasets phrased as instructions has been shown to improve model performance and generalization to unseen tasks. In this paper we explore instruction finetuning with a particular focus on (1) scaling the number of tasks, (2) scaling the model size, and (3) finetuning on chain-of-thought data.
Also cited · not yet reviewed (12)
- FLAN2021 · cited 15דWe call this finetuning procedure Flan (Wei et al. 2021, Finetuning language models;) and prepend ‘‘Flan’’ to the resulting finetuned models (e.g., Flan-PaLM).22 2 We use “Flan” to refer to our finetuning procedure. “FLA…”From this paper · §unknown section
- T02021 · cited 10דFurther progress has been made by finetuning language models on a collection of tasks phrased as instructions, which enables models to respond better to instructions and reduces the need for few-shot exemplars (Ouyang et…”From this paper · §unknown section
- InstructGPT2022 · cited 8דFurther progress has been made by finetuning language models on a collection of tasks phrased as instructions, which enables models to respond better to instructions and reduces the need for few-shot exemplars (Ouyang et…”From this paper · §unknown section
- Super-NaturalInstructions2022 · cited 7דFor instance, Flan-PaLM’s improved reasoning abilities enable it to leverage CoT and self-consistency (Wang et al. 2022c) to achieve 75.2% on Massive Multi-task Language Understanding (Hendrycks et al. 2020, MMLU;).”From this paper · §unknown section
Show 8 more
- GPT-32020 · cited 6דIn natural language processing (NLP), pretrained language models have made significant progress towards this goal, as they can perform tasks given natural language descriptions (Brown et al. 2020, inter alia).”From this paper · §unknown section
- GSM8K verifiers2021 · cited 6דThese nine datasets include tasks such as arithmetic reasoning (Cobbe et al. 2021), multi-hop reasoning (Geva et al. 2021), and natural language inference (Camburu et al. 2020).”From this paper · §unknown section
- PaLM2022 · cited 5דAs an extended version of the PaLM 62B model, we instruction-finetune cont-PaLM, which is a 62B PaLM-model initialized from PaLM-62B and then pretrained for 500B more tokens (Chowdhery et al. 2022).”From this paper · §unknown section
- MMLU2020 · cited 4דInstead, we use the following challenging benchmarks, for which current language models still perform well below expert human raters. (1) MMLU (Hendrycks et al. 2020) includes exam questions from 57 tasks such as mathema…”From this paper · §unknown section
- BIG-bench2022 · cited 4דInstead, we use the following challenging benchmarks, for which current language models still perform well below expert human raters. (1) MMLU (Hendrycks et al. 2020) includes exam questions from 57 tasks such as mathema…”From this paper · §unknown section
- U-PaLM2022 · cited 4דFinally, we instruction-finetune U-PaLM, which is a 540B PaLM model initialized from PaLM-540B and then pretrained with an UL2 objective for 20k additional steps (Tay et al. 2022a; Tay et al. 2022b).”From this paper · §unknown section
- T52019 · cited 3דThese checkpoints have strong zero-shot, few-shot, and CoT abilities, outperforming prior public checkpoints such as T5 (Raffel et al. 2020).”From this paper · §unknown section
- GLaM2021 · cited 2דThese benchmarks were also used in the PaLM paper (Chowdhery et al. 2022), which did not find any meaningful data contamination with pre-training data, consistent with data contamination analyses in previous work (Brown…”From this paper · §unknown section
Led to
- U-PaLM2022 · cited 2דNormally, we would like to end with a cliche future work statement here but not today because we have done this already here (Chung et al. 2022).”From U-PaLM · §Conclusion and Future Work
- BLIP-22023 · cited 3×, 1 in Method“For the frozen language model, we explore the unsupervised-trained OPT model family (Zhang et al. 2022) for decoder-based LLMs, and the instruction-trained FlanT5 model family (Chung et al. 2022) for encoder-decoder-base…”From BLIP-2 · §Method
Abstract
Finetuning language models on a collection of datasets phrased as instructions has been shown to improve model performance and generalization to unseen tasks. In this paper we explore instruction finetuning with a particular focus on (1) scaling the number of tasks, (2) scaling the model size, and (3) finetuning on chain-of-thought data. We find that instruction finetuning with the above aspects dramatically improves performance on a variety of model classes (PaLM, T5, U-PaLM), prompting setups (zero-shot, few-shot, CoT), and evaluation benchmarks (MMLU, BBH, TyDiQA, MGSM, open-ended generation). For instance, Flan-PaLM 540B instruction-finetuned on 1.8K tasks outperforms PALM 540B by a large margin (+9.4% on average). Flan-PaLM 540B achieves state-of-the-art performance on several benchmarks, such as 75.2% on five-shot MMLU. We also publicly release Flan-T5 checkpoints, which achieve strong few-shot performance even compared to much larger models, such as PaLM 62B. Overall, instruction finetuning is a general method for improving the performance and usability of pretrained language models.