Instruction tuning & alignment
Teaching LMs to follow instructions and human preferences.
| Concept | Introduced by | Years | Papers |
|---|---|---|---|
| Proximal policy optimisation A simple, stable policy-gradient method with a clipped surrogate objective. | PPO | 2017–2022 | 3 |
| RL from human feedback Learn a reward model from human comparisons, then optimise the LM against it. | RLHF for LMs (Ziegler) | 2019–2022 | 5 |
| Crowdsourced instruction datasets Collections of tasks paired with human-written instructions, used to train and test instruction following. | Natural Instructions | 2021–2022 | 2 |
| Instruction tuning Fine-tune on many tasks phrased as instructions, so the model follows new instructions zero-shot. | FLAN | 2021–2022 | 2 |
| Prompted multitask training Train on many datasets rewritten into diverse natural-language prompts. | T0 | 2021–2022 | 4 |
| Recursive task decomposition Break a hard-to-judge task into smaller ones that humans can supervise. | Recursive book summarization | 2021 | 0 |
| Chain-of-thought in instruction tuning Include step-by-step reasoning examples in instruction tuning to unlock reasoning. | Flan-T5 / Flan-PaLM | 2022 | 1 |
| SFT → reward model → PPO pipeline Supervised demonstrations, then a preference reward model, then RL: the recipe behind chat assistants. | InstructGPT | 2022 | 3 |