Training language models to follow instructions with human feedback
Making language models bigger does not inherently make them better at following a user's intent. For example, large language models can generate outputs that are untruthful, toxic, or simply not helpful to the user.
Also cited · not yet reviewed (11)
- Learning to summarize from human feedbac2020 · cited 10×, 6 in Method“Compared to earlier work that collects human preference data on the task of summarization (Ziegler et al., 2019; Stiennon et al., 2020; Wu et al., 2021), our inputs span a much broader range of tasks, and can occasionall…”From this paper · §Methods and experimental details
- RLHF for LMs (Ziegler)2019 · cited 5×, 2 in Method“Compared to earlier work that collects human preference data on the task of summarization (Ziegler et al., 2019; Stiennon et al., 2020; Wu et al., 2021), our inputs span a much broader range of tasks, and can occasionall…”From this paper · §Methods and experimental details
- GPT-32020 · cited 4×, 2 in Method“We start with the GPT-3 pretrained language models from Brown et al., 2020.”From this paper · §Methods and experimental details
- Recursive book summarization2021 · cited 4×, 2 in Method“Compared to earlier work that collects human preference data on the task of summarization (Ziegler et al., 2019; Stiennon et al., 2020; Wu et al., 2021), our inputs span a much broader range of tasks, and can occasionall…”From this paper · §Methods and experimental details
Show 7 more
- PPO2017 · cited 3×, 2 in Method“Once again following Stiennon et al., 2020, we fine-tuned the SFT model on our environment using PPO (Schulman et al., 2017).”From this paper · §Methods and experimental details
- FLAN2021 · cited 4×, 1 in Method“We additionally compare InstructGPT to fine-tuning 175B GPT-3 on the FLAN (Wei et al., 2021) and T0 (Sanh et al., 2021) datasets, which both consist of a variety of NLP tasks, combined with natural language instructions…”From this paper · §Methods and experimental details
- T02021 · cited 4×, 1 in Method“We additionally compare InstructGPT to fine-tuning 175B GPT-3 on the FLAN (Wei et al., 2021) and T0 (Sanh et al., 2021) datasets, which both consist of a variety of NLP tasks, combined with natural language instructions…”From this paper · §Methods and experimental details
- Switch Transformer2021 · cited 2×, 1 in Method“This is because the language modeling objective used for many recent large LMs—predicting the next token on a webpage from the internet—is different from the objective “follow the user’s instructions helpfully and safely…”From this paper · §Introduction
- Gopher2021 · cited 2×, 1 in Method“This is because the language modeling objective used for many recent large LMs—predicting the next token on a webpage from the internet—is different from the objective “follow the user’s instructions helpfully and safely…”From this paper · §Introduction
- LaMDA2022 · cited 2×, 1 in Method“This is because the language modeling objective used for many recent large LMs—predicting the next token on a webpage from the internet—is different from the objective “follow the user’s instructions helpfully and safely…”From this paper · §Introduction
- Codex2021 · cited 1×, 1 in Method“The definition of alignment has historically been a vague and confusing topic, with various competing proposals (Chen et al., 2021; Leike et al., 2018; Gabriel, 2020).”From this paper · §Methods and experimental details
Led to
- Super-NaturalInstructions2022 · cited 3דFinally, the well-adopted InstructGPT model Ouyang et al. 2022 is partially enabled by a large dataset of prompts that are collected via various synthetic data augmentation which, unfortunately, is not publicly available…”From Super-NaturalInstructions · §Related Work
- OPT2022 · cited 3דRecent efforts have shown gains by fine-tuning models to directly respond to instruction-style prompting Wei et al. 2021; Min et al. 2021; Sanh et al. 2021; Ouyang et al. 2022.”From OPT · §Related Work
- Emergent abilities2022 · cited 3דConsider the nascent direction of enabling language models to follow natural language instructions describing a task (Wei et al. 2022a; Sanh et al. 2022; Ouyang et al. 2022, inter alia).”From Emergent abilities · §Discussion
- U-PaLM2022 · cited 2דA range of prior work has shown that finetuning language models on a collection of NLP tasks can improve downstream performance on a broad range of downstream tasks (Aghajanyan et al. 2021; Aribandi et al. 2022; Wei et a…”From U-PaLM · §Related Work
- Flan-T5 / Flan-PaLM2022 · cited 8דFurther progress has been made by finetuning language models on a collection of tasks phrased as instructions, which enables models to respond better to instructions and reduces the need for few-shot exemplars (Ouyang et…”From Flan-T5 / Flan-PaLM · §unknown section
Abstract
Making language models bigger does not inherently make them better at following a user's intent. For example, large language models can generate outputs that are untruthful, toxic, or simply not helpful to the user. In other words, these models are not aligned with their users. In this paper, we show an avenue for aligning language models with user intent on a wide range of tasks by fine-tuning with human feedback. Starting with a set of labeler-written prompts and prompts submitted through the OpenAI API, we collect a dataset of labeler demonstrations of the desired model behavior, which we use to fine-tune GPT-3 using supervised learning. We then collect a dataset of rankings of model outputs, which we use to further fine-tune this supervised model using reinforcement learning from human feedback. We call the resulting models InstructGPT. In human evaluations on our prompt distribution, outputs from the 1.3B parameter InstructGPT model are preferred to outputs from the 175B GPT-3, despite having 100x fewer parameters. Moreover, InstructGPT models show improvements in truthfulness and reductions in toxic output generation while having minimal performance regressions on public NLP datasets. Even though InstructGPT still makes simple mistakes, our results show that fine-tuning with human feedback is a promising direction for aligning language models with human intent.