training techniqueIntroduced by RLHF for LMs (Ziegler) · 2019
RL from human feedback
Learn a reward model from human comparisons, then optimise the LM against it.
Drafted by AI · not yet reviewed
How this idea evolved
Each step is the idea’s introducing paper. The line between steps is the citation link between those papers.
- 2017
A simple, stable policy-gradient method with a clipped surrogate objective.
Cites · not yet reviewedcited 1× · §Methods“Train π\pi via Proximal Policy Optimization (PPO, Schulman et al. 2017) with reward RR from eq. 2 on x∼𝒟x\simD.”
From RLHF for LMs (Ziegler) · §Methods - 2019
Learn a reward model from human comparisons, then optimise the LM against it.