Training & optimization
Normalisation, initialisation, regularisation, distillation and optimiser fixes that make training work.
| Concept | Introduced by | Years | Papers |
|---|---|---|---|
| Batch normalisation Normalise activations per mini-batch to stabilise and speed up training. | BatchNorm | 2015–2020 | 10 |
| Knowledge distillation Train a small student to match a large teacher's softened predictions. | Knowledge Distillation | 2015–2022 | 5 |
| Label smoothing Soften one-hot targets to regularise the classifier. | Inception v3 | 2015 | 0 |
| PReLU & He initialisation Learnable leaky ReLUs plus an initialisation that lets very deep rectifier nets train from scratch. | PReLU / He init | 2015 | 0 |
| Decoupled weight decay (AdamW) Apply weight decay separately from Adam's adaptive step, fixing its regularisation. | AdamW | 2017–2023 | 7 |
| Large-batch training recipe Scale learning rate with batch size and warm up, so huge batches train like small ones. | Goyal large-batch SGD | 2017–2018 | 1 |
| Adversarial training in embedding space Perturb word or region embeddings adversarially during training to generalise better. | FreeLB | 2019–2021 | 4 |