optimizationIntroduced by Goyal large-batch SGD · 2017
Large-batch training recipe
Scale learning rate with batch size and warm up, so huge batches train like small ones.
Drafted by AI · not yet reviewed
How this idea evolved
Suggested from citations. No curator has recorded what this idea builds on yet. These ideas from the same theme come from papers that Goyal large-batch SGD cites, directly or one step removed. Citation is a fact; the connection between the ideas is not verified.
Normalise activations per mini-batch to stabilise and speed up training.
Learnable leaky ReLUs plus an initialisation that lets very deep rectifier nets train from scratch.
Soften one-hot targets to regularise the classifier.