training techniqueIntroduced by Recursive book summarization · 2021
Recursive task decomposition
Break a hard-to-judge task into smaller ones that humans can supervise.
Drafted by AI · not yet reviewed
How this idea evolved
Suggested from citations. No curator has recorded what this idea builds on yet. These ideas from the same theme come from papers that Recursive book summarization cites, directly or one step removed. Citation is a fact; the connection between the ideas is not verified.
Learn a reward model from human comparisons, then optimise the LM against it.
A simple, stable policy-gradient method with a clipped surrogate objective.
Papers using this
No other paper in this dataset is tagged with it yet.