Paper Lineage
Esc
training techniqueIntroduced by RLHF for LMs (Ziegler) · 2019

RL from human feedback

Learn a reward model from human comparisons, then optimise the LM against it.

Drafted by AI · not yet reviewed

How this idea evolved

Each step is the idea’s introducing paper. The line between steps is the citation link between those papers.

  1. 2017

    A simple, stable policy-gradient method with a clipped surrogate objective.

    Cites · not yet reviewedcited 1× · §Methods
    “Train π\pi via Proximal Policy Optimization (PPO, Schulman et al. 2017) with reward RR from eq. 2 on x∼𝒟x\simD.”
    From RLHF for LMs (Ziegler) · §Methods
  2. 2019

    Learn a reward model from human comparisons, then optimise the LM against it.

Papers using this