training techniqueIntroduced by InstructGPT · 2022
SFT → reward model → PPO pipeline
Supervised demonstrations, then a preference reward model, then RL: the recipe behind chat assistants.
Drafted by AI · not yet reviewed
How this idea evolved
Each step is the idea’s introducing paper. The line between steps is the citation link between those papers.
- 2017
A simple, stable policy-gradient method with a clipped surrogate objective.
Cites · not yet reviewedcited 1× · §Methods“Train π\pi via Proximal Policy Optimization (PPO, Schulman et al. 2017) with reward RR from eq. 2 on x∼𝒟x\simD.”
From RLHF for LMs (Ziegler) · §Methods - 2019
Learn a reward model from human comparisons, then optimise the LM against it.
Cites · not yet reviewedcited 5× · §Methods and experimental details“Compared to earlier work that collects human preference data on the task of summarization (Ziegler et al., 2019; Stiennon et al., 2020; Wu et al., 2021), our inputs span a much broader range of tasks, and can occasionally include controversial and sensitive topics.”
From InstructGPT · §Methods and experimental details - 2022
Supervised demonstrations, then a preference reward model, then RL: the recipe behind chat assistants.
Also draws on: Proximal policy optimisation (PPO)
Papers using this
- 2022Super-NaturalInstructions
- 2022OPT
- 2022Flan-T5 / Flan-PaLM