Paper Lineage
Esc

Training & optimization

Normalisation, initialisation, regularisation, distillation and optimiser fixes that make training work.

Concepts in Training & optimization
ConceptIntroduced byYearsPapers
Batch normalisation
Normalise activations per mini-batch to stabilise and speed up training.
BatchNorm2015–202010
Knowledge distillation
Train a small student to match a large teacher's softened predictions.
Knowledge Distillation2015–20225
Label smoothing
Soften one-hot targets to regularise the classifier.
Inception v320150
PReLU & He initialisation
Learnable leaky ReLUs plus an initialisation that lets very deep rectifier nets train from scratch.
PReLU / He init20150
Decoupled weight decay (AdamW)
Apply weight decay separately from Adam's adaptive step, fixing its regularisation.
AdamW2017–20237
Large-batch training recipe
Scale learning rate with batch size and warm up, so huge batches train like small ones.
Goyal large-batch SGD2017–20181
Adversarial training in embedding space
Perturb word or region embeddings adversarially during training to generalise better.
FreeLB2019–20214