Paper Lineage
Esc
training techniqueIntroduced by Flan-T5 / Flan-PaLM · 2022

Chain-of-thought in instruction tuning

Include step-by-step reasoning examples in instruction tuning to unlock reasoning.

Drafted by AI · not yet reviewed

How this idea evolved

Each step is the idea’s introducing paper. The line between steps is the citation link between those papers.

  1. 2019

    Cast every task as text in, text out: one model, one loss.

    Cites · not yet reviewedcited 3× · §FLAN: Instruction Tuning Improves Zero-Shot Learning
    “To balance the different sizes of datasets, we limit the number of training examples per dataset to 30k and follow the examples-proportional mixing scheme (Raffel et al. 2020) with a mixing rate maximum of 3k.22 2 In this mixing scheme, a mixing rate maximum of 3,000 means that a dataset does not receive additional sampling weight for examples in excess of 3,000.”
    From FLAN · §FLAN: Instruction Tuning Improves Zero-Shot Learning
  2. 2021

    Fine-tune on many tasks phrased as instructions, so the model follows new instructions zero-shot.

    Cites · not yet reviewedcited 15×
    “We call this finetuning procedure Flan (Wei et al. 2021, Finetuning language models;) and prepend ‘‘Flan’’ to the resulting finetuned models (e.g., Flan-PaLM).22 2 We use “Flan” to refer to our finetuning procedure. “FLAN” is a model in Wei et al. 2021.”
    From Flan-T5 / Flan-PaLM · §unknown
  3. 2022

    Include step-by-step reasoning examples in instruction tuning to unlock reasoning.

Papers using this