Paper Lineage
Esc

Instruction tuning & alignment

Teaching LMs to follow instructions and human preferences.

Concepts in Instruction tuning & alignment
ConceptIntroduced byYearsPapers
Proximal policy optimisation
A simple, stable policy-gradient method with a clipped surrogate objective.
PPO2017–20223
RL from human feedback
Learn a reward model from human comparisons, then optimise the LM against it.
RLHF for LMs (Ziegler)2019–20225
Crowdsourced instruction datasets
Collections of tasks paired with human-written instructions, used to train and test instruction following.
Natural Instructions2021–20222
Instruction tuning
Fine-tune on many tasks phrased as instructions, so the model follows new instructions zero-shot.
FLAN2021–20222
Prompted multitask training
Train on many datasets rewritten into diverse natural-language prompts.
T02021–20224
Recursive task decomposition
Break a hard-to-judge task into smaller ones that humans can supervise.
Recursive book summarization20210
Chain-of-thought in instruction tuning
Include step-by-step reasoning examples in instruction tuning to unlock reasoning.
Flan-T5 / Flan-PaLM20221
SFT → reward model → PPO pipeline
Supervised demonstrations, then a preference reward model, then RL: the recipe behind chat assistants.
InstructGPT20223