Paper Lineage
Esc
training techniqueIntroduced by InstructGPT · 2022

SFT → reward model → PPO pipeline

Supervised demonstrations, then a preference reward model, then RL: the recipe behind chat assistants.

Drafted by AI · not yet reviewed

How this idea evolved

Each step is the idea’s introducing paper. The line between steps is the citation link between those papers.

  1. 2017

    A simple, stable policy-gradient method with a clipped surrogate objective.

    Cites · not yet reviewedcited 1× · §Methods
    “Train π\pi via Proximal Policy Optimization (PPO, Schulman et al. 2017) with reward RR from eq. 2 on x∼𝒟x\simD.”
    From RLHF for LMs (Ziegler) · §Methods
  2. 2019

    Learn a reward model from human comparisons, then optimise the LM against it.

    Cites · not yet reviewedcited 5× · §Methods and experimental details
    “Compared to earlier work that collects human preference data on the task of summarization (Ziegler et al., 2019; Stiennon et al., 2020; Wu et al., 2021), our inputs span a much broader range of tasks, and can occasionally include controversial and sensitive topics.”
    From InstructGPT · §Methods and experimental details
  3. 2022

    Supervised demonstrations, then a preference reward model, then RL: the recipe behind chat assistants.

    Also draws on: Proximal policy optimisation (PPO)

Papers using this